Back to all articlesHow to Check the Evidence Behind an AI Analytics Answer
AI

How to Check the Evidence Behind an AI Analytics Answer

To check an AI analytics answer, connect each material claim to the metric definition, applicable business context, executed query or metric call, and returned result that supports it. Review the time period, access scope, and missing evidence before acting. SQL makes a calculation inspectable; an evidence trail makes its meaning and conclusions inspectable too.

Orion by Gravity supports this work through analysis notebooks, business-context pages, and documented metric and conversation tools. The worksheet below is a recommended review process. It is not a claim that Orion automatically captures every field in a complete, immutable audit record.

What is answer evidence in an AI context layer?

Answer evidence is the observable record that lets a person check how an agent connected a business question to a result. It includes the selected definition, relevant source material, calculation, returned values, and the relationship between those values and the explanation.

A context layer supplies business meaning beyond the calculation: which definition applies, whether a policy is current for the period being analyzed, and which exceptions matter. For the broader distinction, see Semantic Layer vs. Business Context.

Three kinds of evidence answer different questions:

  • A source citation points to a document or dataset. Does that source support this particular claim?
  • Data lineage describes relationships between data assets. Does the calculation use the intended model and grain?
  • An answer trace records observable activity for this answer, such as retrieved material, tool calls, query execution, and outputs. Did the agent actually use the appropriate information?

None of these alone proves that an answer is correct. A citation may point to the wrong policy period. A valid query may calculate the wrong business concept. A trace may record a successful run whose narrative overstates its findings.

Review recorded inputs and actions. Access to a model's private reasoning is unnecessary.

Why is showing SQL not enough?

SQL helps a reviewer inspect joins, filters, grouping, and date logic. It does not establish why the agent chose that definition, whether the query shown actually ran, or whether its output supports a causal explanation.

Suppose a query sums invoice values correctly, but the question asks for recognized revenue. The syntax and arithmetic can be valid while the answer remains wrong. Changing the prose around the result will not repair the definition.

Separate four checks: selected meaning, executed calculation, numerical reconciliation, and supported interpretation. A system that passes the calculation check can still fail the other three.

When your team governs metrics in dbt, preserve that authority. The dbt Semantic Layer documentation, reviewed October 6, 2026, describes central metric definitions powered by MetricFlow and access from downstream tools. A context page should refer to those approved calculations and explain their applicability. It should not quietly introduce a competing formula.

Verify the execution path. A dbt connection or model description is not, by itself, evidence that a particular answer invoked the governed metric service.

What should an answer-evidence worksheet contain?

Use the following worksheet for one important answer. The worksheet is an operating document your team can fill in manually or populate from supported tools. Record unavailable fields as unavailable; do not invent identifiers or infer retrieval from a source being enabled.

FieldWhat to recordReview question
Question and answerOriginal wording, material claims, conversation or artifact referenceIs this the actual answer under review?
ScopeProject, requesting identity or role, business unit, customer scope where relevantWas the answer produced within the intended access boundary?
Metric meaningApproved metric name, definition reference, applicable parametersDid it answer the business concept requested?
Context usedRelevant source passage and revision if exposed; effective dates and approval referenceDoes this source support the meaning for this period?
ExecutionQuery or metric call, parameters, completion status, execution reference if exposedIs this executed work rather than proposed code?
Numerical resultReturned aggregate or permitted supporting outputCan the reported number be reconciled?
TimeAnalysis window, time zone, data cutoff, execution time, relevant source refresh timeAre the compared values temporally comparable?
Claim supportLink each material sentence to its supporting result or sourceDoes the evidence justify the wording?
Review outcomeAccepted, corrected, or unresolved; reviewer, reason, next actionCan the next person tell what is safe to use?

Keep the worksheet proportional to the decision. A routine operational count may need a short record; a disputed executive explanation needs a fuller one. Agree on acceptable reconciliation tolerances before reviewing results.

Do not put access tokens, unrestricted query output, or sensitive row samples into a widely shared review document. Link to evidence through the access controls your team already uses. If a business reviewer cannot view raw rows, provide an approved aggregate or ask an authorized analyst to validate them.

How do you review one answer step by step?

  1. Capture the original answer before editing or rerunning it. Record the question, project, scope, period, and exact claims under review. A later successful run does not explain what happened in the earlier run.
  2. Resolve the business definition. Ask the metric owner which metric, dimensions, exclusions, and time basis should apply. If the question is ambiguous, obtain clarification before changing the calculation.
  3. Check the context actually used. Inspect available citations and source passages. Compare their effective period with the question. A current page in a knowledge base does not prove the original run read that version.
  4. Inspect executed work. Review the query or metric call and its parameters, then confirm completion. Check join grain, filters, grouping, and the data cutoff. Distinguish a displayed proposal from the work that returned the result.
  5. Reconcile the numbers. Compare the answer with the returned output and an independently approved reference for the same definition, scope, and period. Record differences and their explanation.
  6. Review the explanation sentence by sentence. Identify which statements are directly supported, which are hypotheses, and which require more investigation. Do not let a correct total validate an unsupported explanation.
  7. Record the disposition. Accept supported claims, correct misleading ones, or leave the answer unresolved with an owner and a specific evidence request. Preserve the original and corrected versions as separate records.

This is a review of a single answer. For vendor selection and a broader proof of concept, use the AI analytics evaluation scorecard. For context releases, reuse the reviewed question as a regression case after the underlying issue has been understood.

What does a claim-by-claim evidence check look like?

Consider this hypothetical example. The figures and names are invented solely to illustrate the review method; they are not customer results or an Orion benchmark.

A sales leader asks: “How did September recognized revenue compare with August, and why?”

The approved hypothetical metric returns August revenue of $120,000 and September revenue of $108,000 for the same business unit and currency, with closed periods and matching exclusions. The hypothetical agent writes: “Recognized revenue fell 10% because the new pricing policy reduced renewals.”

The numerical change is ($108,000 − $120,000) / $120,000 = −10%. That confirms the change in this example. It does not establish the cause.

Claim in the answerAvailable hypothetical evidenceReview decision
Recognized revenue fell 10%Executed approved metric result: $120,000 and $108,000; same scope and time basisAccept after reconciling definition and results
Renewal revenue explains the declineNo renewal-versus-new-business breakdown was returnedRequest the decomposition; leave unresolved
The pricing policy caused the declinePolicy document says the change began September 15; no analysis separates timing or other driversRewrite as a hypothesis requiring investigation

A supported corrected answer would say: “Recognized revenue declined 10% from August to September. The current evidence does not establish the cause. The pricing change is a hypothesis to investigate.”

The next step is to inspect renewal and new-business results, the relevant cohorts, and other plausible drivers. Even if renewal revenue accounts for the numerical decline, that still does not isolate pricing as its cause.

An answer can therefore contain one accepted claim and two unresolved claims. Review the claims separately instead of giving the whole paragraph a single trust label.

How do you tell what context the agent actually used?

Look for evidence attached to the run: a cited passage, observable retrieval output, a source revision where available, or an execution record linking context to the answer. A list of enabled documents establishes availability, not consumption.

A source URL also needs interpretation. Does it point to the exact content reviewed, or to a page that can change after the answer was produced? If the system does not retain the historical content used, state that limitation in the worksheet. Comparing today's page with yesterday's answer cannot reconstruct yesterday's input with certainty.

Separate business validity from technical freshness. A document imported today may contain last year's rule; an older document may remain the authority for a historical period. Store the relevant effective dates alongside available import or revision evidence.

For Orion's imported context, the external wiki documentation, reviewed October 6, 2026, distinguishes daily Sync Only updates from editable Read & Write imports with manual sync. Syncing replaces page content. These modes affect which copy you should inspect; they do not establish automatic write-back or preservation of every historical input.

What can you inspect in Orion by Gravity today?

The following capabilities are described in public Orion documentation reviewed October 6, 2026. Those pages do not state publication dates. Verify the exact deployed behavior in a demonstration before relying on it for a required control.

Orion's analysis documentation describes natural-language questions that produce charts, tables, and code. It also describes notebooks retained in chat history with analysis instructions, logic, and code. That provides a place to inspect and revisit analytical work. A rerun obtains a new result; it should not be treated as proof of the old result's inputs.

The Knowledge Base documentation describes wiki-page citations in reports, slide decks, and dashboards. Check that the citation supports the particular sentence being reviewed. This documented citation behavior does not establish exhaustive source attribution for every output.

The MCP tools reference describes metric detail with definitions, Python code, sources, and the latest result, plus historical metric results and conversation history. Clients using ask_orion must retrieve the completed response rather than treating the initial request as the final analysis. These are useful inspection surfaces; a complete answer-specific evidence record still needs verification.

Orion's dbt enrichment documentation describes adding model descriptions and lineage to an existing warehouse connection. This helps inspect the data foundation. It does not establish mandatory MetricFlow execution or automatic tracing from every business claim to an application pull request.

Ask Orion by Gravity to demonstrate one real answer from question to returned result and source citation, including what can be retained after a context edit. Confirm which evidence fields are exposed, who can access them, and which must be recorded outside the product.

When is a simpler approach enough?

If your existing BI tool already exposes the approved metric, filters, result, and supporting source for the decisions you make, a short manual worksheet may be sufficient. Adding a separate context platform is not necessary just to collect a few links.

Snowflake's Cortex Analyst documentation, reviewed October 6, 2026, describes semantic views for interpreting business questions and responses containing text and SQL. If your team already uses that stack, assess the surrounding application's execution records, results, and source presentation before adding another tool.

A notebook-oriented analytics product can be useful when analysts need to inspect executable work. A semantic layer is useful for governed calculations. A context workflow is useful when interpretation depends on policies, exceptions, or shared business knowledge. Match the evidence gap to the function you need.

Orion by Gravity is worth evaluating when your team wants natural-language analysis with reusable analytical work and managed business context. If your requirement is a complete immutable audit trail, automatic application-code provenance, or universal oversight of unrelated third-party agents, verify that requirement explicitly before choosing a product.

For a wider product comparison, see AI business intelligence tools and platforms, then test the relevant evidence path with Orion by Gravity.

Frequently asked questions

Can a business user review an answer without knowing SQL?

Yes. A business reviewer can check the meaning, period, exceptions, and whether the explanation fits the evidence. Pair that review with an authorized analyst who checks execution and reconciliation. Assign different checks to the people qualified to perform them.

Is a citation proof that an answer is correct?

No. A citation identifies a source to examine. Verify that the passage supports the claim, applies to the period and scope, and matches the information actually used where that evidence is available.

Should an evidence trail expose chain of thought?

No. Observable inputs, citations, tool calls, executable calculations, and results are the appropriate review material. The worksheet does not require private model reasoning.

Can I rerun the query to verify an old answer?

A rerun checks current execution against the data and settings available now. To verify a historical answer, retain or obtain the relevant historical result, definition, scope, and source state. Record anything that cannot be reconstructed.

What if the number is correct but the explanation is wrong?

Accept the reconciled number separately and correct or withhold the unsupported explanation. A correct metric does not prove a causal claim. Specify the additional evidence needed for the unresolved sentence.

What should happen when the evidence is incomplete?

Mark the relevant claim unresolved and request the missing evidence from its owner. Apply your team's decision policy for the intended use. Do not replace a missing record with a confident narrative or an invented source version.

Keep reading

View all