Context graph validation is the structured process of proving that a context graph can answer agreed business questions correctly, historically, explainably, securely, and within operational service levels using production-representative data and repeatable acceptance tests.
Replace schema sign-off with question-driven acceptance #
Teams often validate a graph the way they validate a database migration: confirm that expected labels and properties exist, count loaded records, run a few sample queries, and review a visualization. Those checks are necessary but insufficient. A context graph exists to assemble situations for decisions. The business question, not the node, is therefore the right unit of acceptance.
Build the validation program around the recurring questions people will ask on day one and the questions auditors, operators, and incident responders will ask on day 100. Each question should specify the expected answer, acceptable alternatives, evidence required, time perspective, permitted users, and the action the answer may trigger.
Stage 1: build the validation question bank #
Interview decision owners, frontline operators, risk, data governance, security, and platform teams. Capture questions in their natural language, then normalize them into testable forms. Include direct lookups, multi-hop relationships, aggregates, historical reconstruction, “why” explanations, and “what changed” queries.
Classify questions by criticality. Tier 1 questions gate go-live because a wrong answer can create regulatory, financial, safety, or customer harm. Tier 2 questions affect operational quality. Tier 3 questions are useful but can be deferred. This prevents a high average pass rate from hiding failure on the few questions that matter most.
Stage 2: trace each question to graph requirements #
For every question, list the required entity types, relationship types, source systems, temporal fields, policies, and freshness thresholds. This creates a traceability matrix from business intent to ontology, ingestion, storage, query, and reasoning. Missing coverage becomes visible before testing begins.
Define expected provenance. A correct answer without explainable evidence is not acceptable for regulated or high-impact decisions. The test should assert which source facts, transformations, and inferred edges must be visible and which users may inspect them.
Stage 3: test answers, explanations, and denial behavior #
Run the question bank against production-like data. Validate exact answers where deterministic, bounded sets where multiple answers are legitimate, and explicit uncertainty where evidence is incomplete. Test historical “as of” questions and late-arriving facts. Test whether identity merges and splits preserve past truth.
Negative tests matter equally. Confirm that the graph returns “unknown” rather than inventing an answer, denies access when purpose or role is insufficient, excludes expired facts, and keeps inferred relationships labeled. A secure refusal is a successful result, not a failed query.
Validation question bank sample #
| Question ID | Business question | Required proof | Pass condition |
|---|---|---|---|
| Q-01 | Who ultimately owns this customer as of 31 March? | Ownership chain, valid-time state, source lineage | Correct terminal owner; full chain replayable |
| Q-02 | Which products depend on Supplier X? | Multi-hop supplier-part-product traversal | All known affected products; no unrelated products |
| Q-03 | Why was this account flagged? | Evidence, rule or inference, policy decision | Every contributing fact and model output visible |
| Q-04 | What changed since yesterday? | Temporal delta and source-event lineage | Complete approved change set within freshness SLO |
| Q-05 | Can this analyst view the supporting evidence? | Purpose, role, region, consent policy | Correct allow or deny with policy explanation |
Stage 4: test operational failure modes #
Inject delayed streams, source outages, schema changes, duplicate identities, conflicting records, deletion requests, and policy changes. Observe whether quality scores degrade, stale facts are labeled, alerts fire, and affected decisions are gated. The graph must know when it does not know.
Load-test representative traversals through the same Context Harness controls used in production. Side-stepping policy for performance testing produces meaningless numbers. Measure p50, p95, and p99 latency by question class, along with timeout and partial-answer behavior.
The pattern across three industries #
Banking. Validate beneficial-ownership, exposure, and related-party questions using historical ownership states, source lineage, and role-based denial tests.
Healthcare. Validate patient, provider, encounter, authorization, and consent relationships, including identity ambiguity and purpose-specific access.
Manufacturing. Validate transitive supplier, part, plant, and product dependencies, then inject a supplier outage to test propagation and affected-product answers.
A realistic enterprise scenario #
A global bank is preparing a customer-risk context graph for go-live. The architecture team initially reports 98 percent source coverage and successful load tests. The validation lead instead builds a 52-question bank with compliance, operations, and audit.
Understand. Testing reveals that current beneficial ownership is correct, but historical ownership is wrong after entity merges. Several “why was this customer flagged?” queries return the risk score without source evidence. A regional analyst can also see documents outside the approved purpose.
Decide. The team blocks go-live on three Tier 1 failures: historical reconstruction, provenance, and policy denial. Node-count and average latency results remain green but are no longer treated as sufficient.
Execute. Storage is corrected to preserve valid-time history, lineage is attached to each ownership edge, and the Context Harness adds purpose-specific evidence controls. The complete Tier 1 bank then passes, and all 52 questions become automated regression tests for subsequent releases.
Common mistakes to avoid #
- Using node counts, edge counts, or schema completeness as the primary readiness measure.
- Testing only questions chosen by the graph team rather than questions used by decision owners.
- Ignoring historical, ambiguous, missing-data, and denied-access cases.
- Validating query speed without the production policy layer enabled.
- Accepting correct answers that cannot show provenance or distinguish inference from evidence.
- Treating validation as a one-time gate instead of a permanent regression suite.
Define go-live gates and ongoing regression #
Convert the question bank into automated regression tests that run on every ontology, pipeline, policy, and resolution change. Critical questions should have zero unresolved defects at launch. Lower tiers may have explicit waivers with owners and dates. Publish a readiness scorecard that separates correctness, coverage, freshness, explainability, policy, and performance rather than blending them into one opaque percentage.
After launch, production queries and decision outcomes should continuously expand the test bank. Every serious incident becomes a permanent regression case. In this way, context graph validation becomes an operating discipline rather than a one-time project phase.
How OpenKnowra approaches this #
OpenKnowra validates context graphs through a question-to-decision framework. The question bank traces each business need through the Context Graph Engine, Enterprise Digital Twin, Context Harness, Decision Layer, and Execution Grid. Tests cover answers, evidence, temporal replay, uncertainty, denial behavior, latency, and operational degradation. Production outcomes and incidents are then converted into ongoing regression cases so quality improves with use.