Back to all articlesHow to Test an AI Context Layer After dbt or Business Rules Change
AI

How to Test an AI Context Layer After dbt or Business Rules Change

To test an AI context layer after a dbt model or business rule changes, rerun a fixed set of business questions against the candidate configuration, compare each answer with an approved reference, and inspect the definitions, sources, queries, and permissions behind the result. Block the release when a material mismatch cannot be explained, evidence is missing, or access rules fail.

An answer that changed is not automatically wrong. An answer that stayed the same is not automatically right. The review must show whether the change was intended and whether the agent used the right version of the rule.

Source and product documentation reviewed October 2, 2026. The worksheets and examples below are recommended operating practices. All example counts are synthetic.

What does an AI context-layer evaluation test?

An AI context-layer evaluation is a repeatable check of whether an agent uses the right business meaning, evidence, and access scope to answer a question. Context includes metric references, terminology, policy dates, business exceptions, analytical instructions, and the sources an agent can consult.

The evaluation should test the whole observable path: question, selected definition, retrieved evidence, query or tool call, result, and explanation. Inspect recorded work and sources; access to a model's private reasoning is unnecessary.

For example, an agent might calculate a valid number using an obsolete definition of active customer. A second agent might use the current definition but describe the increase as customer growth when it actually came from a policy change. Both answers need correction.

Use this guide for maintaining an agent after deployment. For an initial buying decision, start with How to Evaluate an AI Analytics Agent Before You Buy.

Why are passing dbt tests not enough?

Keep the tests you already run in dbt. Data tests assert conditions on data resources, such as uniqueness, valid values, and relationships. Unit tests check SQL model logic against specified inputs and expected outputs. They address different failures and remain part of the release process. See the official data-test documentation and unit-test documentation.

An answer evaluation adds a separate check: did the agent select that tested model or metric, apply the approved interpretation, and support its explanation? A passing model test does not establish that an agent retrieved a current policy or interpreted a vague question correctly.

The dbt Semantic Layer centralizes metric definitions on existing models and uses MetricFlow to handle joins. Keep governed calculations there when that is your authoritative metric system. Additional context should point to the metric and explain its scope or business implications, rather than maintain a competing formula in prose.

For that boundary, see Semantic Layer vs. Business Context.

What should you record before changing context?

Create a change record before editing the rule. It should answer seven questions:

  1. What changed: a metric calculation, a business note, a document, a schema, permissions, or an agent configuration?
  2. Who owns the authoritative source and who can approve the intended behavior?
  3. Which projects, audiences, and consuming agents should receive the change?
  4. When does the rule take effect, and should it apply to historical questions?
  5. Which answers should change, and which must remain stable?
  6. How will reviewers identify the exact candidate and reference configuration?
  7. What is the recovery action if the evaluation fails?

Record the data cutoff, dbt commit or deployed metric version, context-page revisions, agent configuration, tool configuration, identity, project, reporting period, and time zone. Capture identifiers actually exposed by your systems. If a version cannot be pinned, record that limitation and reduce what you claim the comparison proves.

Use explicit dates in the evaluation. A question about “last month” can silently change between runs. If you cannot hold the data constant, independently recompute the approved reference for each run and identify the refresh.

How do you distinguish a context regression from a data refresh?

Use a four-run comparison when the environment lets you isolate candidate context from candidate data. Here, “context” includes the definition references and agent instructions under review; “data” means the governed dataset or model build used for the test.

RunContext and agent configurationData or model buildReview purpose
AApproved baselineFixed reference fixtureConfirm that the baseline still reproduces
BCandidateSame reference fixtureIsolate the effect of the context change
CApproved baselineCandidate build or refreshed dataIsolate the effect of the data change
DCandidateCandidate build or refreshed dataCheck the configuration intended for release

Keep the identity, question, dates, and relevant settings constant across the runs. Compare each result with an independently approved answer for that combination. The old agent answer is evidence of prior behavior, not the source of truth.

When a model change makes an old configuration incompatible, do not force a misleading comparison. Mark that combination unsupported, test the intended configuration against a reference, and retain the incompatibility in the review record.

If you cannot create separate configurations or fixtures, start with a manual before-and-after evaluation. State which inputs changed together. This can still find failures, but it cannot isolate their cause as confidently.

Which regression cases should a data team keep?

Start with questions people already rely on, plus cases that expose the changed rule. The following worksheet is a proposed starting set, not an industry benchmark or a claim about Orion's test coverage. Replace every expected behavior with one your business owner approves.

CaseTest question or actionExpected behavior after the changeEvidence to retain
1. Intended changeAsk for the metric in the new rule's effective periodApplies the new approved definition and explains the changeMetric reference, effective date, query and approved result
2. Historical boundaryAsk for the same metric before the effective dateUses the historical rule unless restatement was explicitly approvedHistorical rule and period-specific reference
3. Unaffected metricAsk a related question outside the change's scopePreserves the approved interpretationUnchanged definition and reconciled result
4. SynonymAsk with the business team's alternate nameMaps to the same scoped metric or clarifies ambiguityTerm mapping and selected metric
5. Conflicting documentProvide an obsolete and a current policySelects the applicable policy and identifies the conflictDocument revisions, effective dates and cited source
6. Retired columnAsk a question that previously used a removed or deprecated fieldUses the approved replacement or stops with an actionable limitationSchema or model reference and executed query
7. Restricted audienceRun as a user without access to part of the evidenceEnforces the restriction across results, explanation and cited materialTest identity, policies and observed output
8. Consuming agentRun through each supported client or channelUses the intended definitions and access scope on every tested pathClient, project, tool route and returned evidence
9. Saved analysisRerun a saved analysis and also ask the question afreshBoth follow their documented version policy; differences are explainedSaved logic, new query, context revisions and results
10. Missing or stale sourceWithhold a required refresh or sourceWarns, clarifies, or declines according to the approved ruleSource cutoff, expected warning and observed behavior

Include a boundary case, not just a typical example. A rule effective October 1 needs tests immediately before and after that date. A renamed field needs a test that would still return a plausible result from the old field.

Repeat important cases to reveal unstable behavior. Choose the repetitions and numerical tolerances before seeing the candidate results. Do not require identical prose or identical SQL text when equivalent logic is valid. Require consistent business meaning, reconciled values, permitted evidence, and supported conclusions.

What does a worked context-change test look like?

Consider a hypothetical business that changes its active-customer policy from activity within 30 days to activity within 60 days, effective October 1. Finance approves the definition and decides that earlier reporting periods will retain the old rule.

On a fixed synthetic fixture for an October reporting date, an independently checked reference produces 80 customers under the old rule and 95 under the new rule. These numbers are invented for the example; they are not customer results or a product benchmark.

A useful test record would contain:

  • Question: “How many active customers were there as of October 2?”
  • Scope: the approved account population, identity, project, and reporting time zone.
  • Authoritative metric: the governed definition, with the candidate version and effective date.
  • Expected result: 95 on this fixture, with an explanation of the 60-day window.
  • Historical case: a September question that still uses the approved 30-day rule.
  • Prohibited interpretation: describing the definition-driven difference as evidence of customer growth.
  • Required evidence: the selected definition, date logic, query or governed metric call, and returned result.
  • Approval: the metric owner confirms the reference; the technical owner confirms the execution path.

If the candidate returns 95 but cites the old policy, it has failed the evidence check. If it returns 80 without explaining that it used the old rule, it has failed the current-definition check. If it returns 95 and asserts that the business gained 15 customers, it has failed the interpretation check.

Test a saved analysis separately. A reusable notebook might preserve older calculation logic. Decide whether saved work should adopt new rules, retain historical rules, or require an explicit migration. Do not assume rerunning means every definition is automatically updated.

What should block a release?

Define release decisions by failure type and business impact, not a universal accuracy percentage.

Stop for unauthorized disclosure, fabricated evidence, an incorrect material metric, or an unresolved conflict in the rule being released. Also stop when reviewers cannot establish which configuration produced the answer for a consequential workflow.

An intentional change can pass when the business owner approves its expected effect and the evidence reconciles. A cosmetic wording change can usually pass when the claims remain equivalent. An unexplained change should enter review, even if the number appears reasonable.

Keep separate outcomes for passed, failed, needs review, and not run. A case skipped because its integration was unavailable is not a passing case. Preserve case-level results so a high average cannot hide a permission failure.

Route the failure to its source: the metric owner for calculation errors, the document owner for stale policy, the integration owner for incorrect routing, or the application owner for permissions. Change the appropriate source and rerun the affected cases plus a small set of unaffected controls.

How can this fit into CI without becoming another hidden process?

A practical sequence is to validate the dbt build, assemble the candidate context configuration, run the impacted question set, collect results and evidence, and present a visible report beside the change.

Start manually if the interfaces cannot produce reliable machine-readable evidence. Automate only the parts whose results you can verify. The report should link each failing question to its expected behavior, observed answer, relevant source versions, and owner.

For an automated integration, your team needs authenticated access to the candidate environment, a way to submit questions, reliable completion detection, observable results, and explicit failure handling. Timeouts and missing evidence should produce an incomplete or failed result, according to the policy you choose.

An LLM judge can help triage explanations. Use deterministic checks for exact values and permissions where possible, and have a qualified reviewer inspect material semantic differences. A model agreeing with another model does not establish that the business rule is correct.

CI gating, immutable context snapshots, and automatic recovery require implementation in your chosen stack. The workflow described here is a recommendation, not a claim that Orion provides all of those controls natively.

Where does Orion by Gravity fit?

Orion by Gravity provides documented building blocks for maintaining and reviewing business context.

Its Knowledge Base page documentation describes editable pages, revision history, project-selected pages, and optional tenant or group defaults. Locked pages support change requests, with direct-edit exceptions for creators and admins described in the page workflow. Verify the behavior of the specific editing path and roles in your deployment before treating a lock as a mandatory approval gate.

The analysis documentation describes notebooks that retain instructions, logic, and code and can be rerun. These provide reviewable analytical work. A notebook's existence alone does not prove that every context input is preserved in an immutable snapshot.

The data-source documentation describes enriching warehouse connections with dbt model descriptions and lineage. Confirm the actual metric execution path separately; metadata enrichment is not evidence that every answer must call MetricFlow.

Orion's MCP documentation describes access from compatible assistants to Orion projects, metrics, workflows, and knowledge. Its tools reference documents tools for asking questions and retrieving conversation history. A team can assess those interfaces when designing an evaluation integration. Test the path you deploy rather than assuming Orion controls every third-party agent or every route to the warehouse.

A bounded pilot should demonstrate one approved context change, inspectable answer evidence, a historical case, and one restricted-user case. Expand to additional consuming agents only after their paths are tested.

Frequently asked questions

Can business users contribute regression cases?

Yes. A business user's correction is a useful source of a new test. Have the appropriate owner approve the intended rule, scope, and expected behavior before adding it to the reference set. A correction to one answer should not silently become company-wide policy.

Should all clients return exactly the same answer?

They should apply the same approved definition, data cutoff, and permission scope when those inputs are equivalent. Presentation can differ. Test each client because its instructions, tools, caching, or identity mapping can change the execution path.

How do you prevent an agent from using stale documents?

Record ownership, applicable dates, revisions, and supersession rules, then test retrieval with both old and current documents present. “Most recently uploaded” and “currently applicable” can mean different things. Historical questions may correctly require the older policy.

How many questions are enough?

There is no universal number. Cover the changed rule, its boundaries, unaffected controls, permissions, and supported consuming paths. Use the ten cases above as a starting worksheet and add real failures from production. A large set of easy questions can still miss the important regression.

Can a rerun against fresh data prove that a context correction worked?

Only if the reference is recomputed for the same data and all other relevant inputs are controlled or recorded. Otherwise, a refreshed dataset may hide or mimic the effect of the correction. Use a fixed fixture where possible.

What should happen after a failed evaluation?

Hold the affected rollout, identify the authoritative source that needs correction, retain the failed evidence, and rerun after the fix. If reverting is appropriate, follow the recovery process your deployment actually supports. Reverting context alone will not undo a data migration or a permission change.

Keep reading

View all