

How to Test an AI Context Layer After dbt or Business Rules Change
To test an AI context layer after a dbt model or business rule changes, rerun a fixed set of business questions against the candidate configuration, compare each answer with an approved reference, and inspect the definitions, sources, queries, and permissions behind the result. Block the release when a material mismatch cannot be explained, evidence is missing, or access rules fail.
An answer that changed is not automatically wrong. An answer that stayed the same is not automatically right. The review must show whether the change was intended and whether the agent used the right version of the rule.
Source and product documentation reviewed October 2, 2026. The worksheets and examples below are recommended operating practices. All example counts are synthetic.
What does an AI context-layer evaluation test?
An AI context-layer evaluation is a repeatable check of whether an agent uses the right business meaning, evidence, and access scope to answer a question. Context includes metric references, terminology, policy dates, business exceptions, analytical instructions, and the sources an agent can consult.
The evaluation should test the whole observable path: question, selected definition, retrieved evidence, query or tool call, result, and explanation. Inspect recorded work and sources; access to a model's private reasoning is unnecessary.
For example, an agent might calculate a valid number using an obsolete definition of active customer. A second agent might use the current definition but describe the increase as customer growth when it actually came from a policy change. Both answers need correction.
Use this guide for maintaining an agent after deployment. For an initial buying decision, start with How to Evaluate an AI Analytics Agent Before You Buy.
Why are passing dbt tests not enough?
Keep the tests you already run in dbt. Data tests assert conditions on data resources, such as uniqueness, valid values, and relationships. Unit tests check SQL model logic against specified inputs and expected outputs. They address different failures and remain part of the release process. See the official data-test documentation and unit-test documentation.
An answer evaluation adds a separate check: did the agent select that tested model or metric, apply the approved interpretation, and support its explanation? A passing model test does not establish that an agent retrieved a current policy or interpreted a vague question correctly.
The dbt Semantic Layer centralizes metric definitions on existing models and uses MetricFlow to handle joins. Keep governed calculations there when that is your authoritative metric system. Additional context should point to the metric and explain its scope or business implications, rather than maintain a competing formula in prose.
For that boundary, see Semantic Layer vs. Business Context.
What should you record before changing context?
Create a change record before editing the rule. It should answer seven questions:
- What changed: a metric calculation, a business note, a document, a schema, permissions, or an agent configuration?
- Who owns the authoritative source and who can approve the intended behavior?
- Which projects, audiences, and consuming agents should receive the change?
- When does the rule take effect, and should it apply to historical questions?
- Which answers should change, and which must remain stable?
- How will reviewers identify the exact candidate and reference configuration?
- What is the recovery action if the evaluation fails?
Record the data cutoff, dbt commit or deployed metric version, context-page revisions, agent configuration, tool configuration, identity, project, reporting period, and time zone. Capture identifiers actually exposed by your systems. If a version cannot be pinned, record that limitation and reduce what you claim the comparison proves.
Use explicit dates in the evaluation. A question about “last month” can silently change between runs. If you cannot hold the data constant, independently recompute the approved reference for each run and identify the refresh.
How do you distinguish a context regression from a data refresh?
Use a four-run comparison when the environment lets you isolate candidate context from candidate data. Here, “context” includes the definition references and agent instructions under review; “data” means the governed dataset or model build used for the test.
| Run | Context and agent configuration | Data or model build | Review purpose |
|---|---|---|---|
| A | Approved baseline | Fixed reference fixture | Confirm that the baseline still reproduces |
| B | Candidate | Same reference fixture | Isolate the effect of the context change |
| C | Approved baseline | Candidate build or refreshed data | Isolate the effect of the data change |
| D | Candidate | Candidate build or refreshed data | Check the configuration intended for release |
Keep the identity, question, dates, and relevant settings constant across the runs. Compare each result with an independently approved answer for that combination. The old agent answer is evidence of prior behavior, not the source of truth.
When a model change makes an old configuration incompatible, do not force a misleading comparison. Mark that combination unsupported, test the intended configuration against a reference, and retain the incompatibility in the review record.
If you cannot create separate configurations or fixtures, start with a manual before-and-after evaluation. State which inputs changed together. This can still find failures, but it cannot isolate their cause as confidently.
Which regression cases should a data team keep?
Start with questions people already rely on, plus cases that expose the changed rule. The following worksheet is a proposed starting set, not an industry benchmark or a claim about Orion's test coverage. Replace every expected behavior with one your business owner approves.
| Case | Test question or action | Expected behavior after the change | Evidence to retain |
|---|---|---|---|
| 1. Intended change | Ask for the metric in the new rule's effective period | Applies the new approved definition and explains the change | Metric reference, effective date, query and approved result |
| 2. Historical boundary | Ask for the same metric before the effective date | Uses the historical rule unless restatement was explicitly approved | Historical rule and period-specific reference |
| 3. Unaffected metric | Ask a related question outside the change's scope | Preserves the approved interpretation | Unchanged definition and reconciled result |
| 4. Synonym | Ask with the business team's alternate name | Maps to the same scoped metric or clarifies ambiguity | Term mapping and selected metric |
| 5. Conflicting document | Provide an obsolete and a current policy | Selects the applicable policy and identifies the conflict | Document revisions, effective dates and cited source |
| 6. Retired column | Ask a question that previously used a removed or deprecated field | Uses the approved replacement or stops with an actionable limitation | Schema or model reference and executed query |
| 7. Restricted audience | Run as a user without access to part of the evidence | Enforces the restriction across results, explanation and cited material | Test identity, policies and observed output |
| 8. Consuming agent | Run through each supported client or channel | Uses the intended definitions and access scope on every tested path | Client, project, tool route and returned evidence |
| 9. Saved analysis | Rerun a saved analysis and also ask the question afresh | Both follow their documented version policy; differences are explained | Saved logic, new query, context revisions and results |
| 10. Missing or stale source | Withhold a required refresh or source | Warns, clarifies, or declines according to the approved rule | Source cutoff, expected warning and observed behavior |
Include a boundary case, not just a typical example. A rule effective October 1 needs tests immediately before and after that date. A renamed field needs a test that would still return a plausible result from the old field.
Repeat important cases to reveal unstable behavior. Choose the repetitions and numerical tolerances before seeing the candidate results. Do not require identical prose or identical SQL text when equivalent logic is valid. Require consistent business meaning, reconciled values, permitted evidence, and supported conclusions.
What does a worked context-change test look like?
Consider a hypothetical business that changes its active-customer policy from activity within 30 days to activity within 60 days, effective October 1. Finance approves the definition and decides that earlier reporting periods will retain the old rule.
On a fixed synthetic fixture for an October reporting date, an independently checked reference produces 80 customers under the old rule and 95 under the new rule. These numbers are invented for the example; they are not customer results or a product benchmark.
A useful test record would contain:
- Question: “How many active customers were there as of October 2?”
- Scope: the approved account population, identity, project, and reporting time zone.
- Authoritative metric: the governed definition, with the candidate version and effective date.
- Expected result: 95 on this fixture, with an explanation of the 60-day window.
- Historical case: a September question that still uses the approved 30-day rule.
- Prohibited interpretation: describing the definition-driven difference as evidence of customer growth.
- Required evidence: the selected definition, date logic, query or governed metric call, and returned result.
- Approval: the metric owner confirms the reference; the technical owner confirms the execution path.
If the candidate returns 95 but cites the old policy, it has failed the evidence check. If it returns 80 without explaining that it used the old rule, it has failed the current-definition check. If it returns 95 and asserts that the business gained 15 customers, it has failed the interpretation check.
Test a saved analysis separately. A reusable notebook might preserve older calculation logic. Decide whether saved work should adopt new rules, retain historical rules, or require an explicit migration. Do not assume rerunning means every definition is automatically updated.
What should block a release?
Define release decisions by failure type and business impact, not a universal accuracy percentage.
Stop for unauthorized disclosure, fabricated evidence, an incorrect material metric, or an unresolved conflict in the rule being released. Also stop when reviewers cannot establish which configuration produced the answer for a consequential workflow.
An intentional change can pass when the business owner approves its expected effect and the evidence reconciles. A cosmetic wording change can usually pass when the claims remain equivalent. An unexplained change should enter review, even if the number appears reasonable.
Keep separate outcomes for passed, failed, needs review, and not run. A case skipped because its integration was unavailable is not a passing case. Preserve case-level results so a high average cannot hide a permission failure.
Route the failure to its source: the metric owner for calculation errors, the document owner for stale policy, the integration owner for incorrect routing, or the application owner for permissions. Change the appropriate source and rerun the affected cases plus a small set of unaffected controls.
How can this fit into CI without becoming another hidden process?
A practical sequence is to validate the dbt build, assemble the candidate context configuration, run the impacted question set, collect results and evidence, and present a visible report beside the change.
Start manually if the interfaces cannot produce reliable machine-readable evidence. Automate only the parts whose results you can verify. The report should link each failing question to its expected behavior, observed answer, relevant source versions, and owner.
For an automated integration, your team needs authenticated access to the candidate environment, a way to submit questions, reliable completion detection, observable results, and explicit failure handling. Timeouts and missing evidence should produce an incomplete or failed result, according to the policy you choose.
An LLM judge can help triage explanations. Use deterministic checks for exact values and permissions where possible, and have a qualified reviewer inspect material semantic differences. A model agreeing with another model does not establish that the business rule is correct.
CI gating, immutable context snapshots, and automatic recovery require implementation in your chosen stack. The workflow described here is a recommendation, not a claim that Orion provides all of those controls natively.
Where does Orion by Gravity fit?
Orion by Gravity provides documented building blocks for maintaining and reviewing business context.
Its Knowledge Base page documentation describes editable pages, revision history, project-selected pages, and optional tenant or group defaults. Locked pages support change requests, with direct-edit exceptions for creators and admins described in the page workflow. Verify the behavior of the specific editing path and roles in your deployment before treating a lock as a mandatory approval gate.
The analysis documentation describes notebooks that retain instructions, logic, and code and can be rerun. These provide reviewable analytical work. A notebook's existence alone does not prove that every context input is preserved in an immutable snapshot.
The data-source documentation describes enriching warehouse connections with dbt model descriptions and lineage. Confirm the actual metric execution path separately; metadata enrichment is not evidence that every answer must call MetricFlow.
Orion's MCP documentation describes access from compatible assistants to Orion projects, metrics, workflows, and knowledge. Its tools reference documents tools for asking questions and retrieving conversation history. A team can assess those interfaces when designing an evaluation integration. Test the path you deploy rather than assuming Orion controls every third-party agent or every route to the warehouse.
A bounded pilot should demonstrate one approved context change, inspectable answer evidence, a historical case, and one restricted-user case. Expand to additional consuming agents only after their paths are tested.
Frequently asked questions
Can business users contribute regression cases?
Yes. A business user's correction is a useful source of a new test. Have the appropriate owner approve the intended rule, scope, and expected behavior before adding it to the reference set. A correction to one answer should not silently become company-wide policy.
Should all clients return exactly the same answer?
They should apply the same approved definition, data cutoff, and permission scope when those inputs are equivalent. Presentation can differ. Test each client because its instructions, tools, caching, or identity mapping can change the execution path.
How do you prevent an agent from using stale documents?
Record ownership, applicable dates, revisions, and supersession rules, then test retrieval with both old and current documents present. “Most recently uploaded” and “currently applicable” can mean different things. Historical questions may correctly require the older policy.
How many questions are enough?
There is no universal number. Cover the changed rule, its boundaries, unaffected controls, permissions, and supported consuming paths. Use the ten cases above as a starting worksheet and add real failures from production. A large set of easy questions can still miss the important regression.
Can a rerun against fresh data prove that a context correction worked?
Only if the reference is recomputed for the same data and all other relevant inputs are controlled or recorded. Otherwise, a refreshed dataset may hide or mimic the effect of the correction. Use a fixed fixture where possible.
What should happen after a failed evaluation?
Hold the affected rollout, identify the authoritative source that needs correction, retain the failed evidence, and rerun after the fix. If reverting is appropriate, follow the recovery process your deployment actually supports. Reverting context alone will not undo a data migration or a permission change.

