Context quality metrics are the measurable dimensions of whether an AI system's assembled context is fit for decisions: coverage (the share of needed entities, relationships, and facts present), freshness (fact currency against SLA tiers), lineage (traceability of every fact to source and version), and consistency (internal coherence: resolved identities, no contradictions, model-conformant relationships). They are computed per decision domain and enforced at decision time, not merely reported.
Context quality metrics: measuring the picture, not the pixels #
Context quality metrics answer a question data quality never asked. Data quality asks whether individual records are accurate, complete, and valid, whether the pixels are good. Context quality asks whether the assembled picture an AI system reasons over is good: whether the situation it sees, this customer, these relationships, this state, these constraints, is complete enough, current enough, traceable enough, and coherent enough to decide on. An enterprise can score superbly on the first and fail the second, because perfect records from four systems still assemble into a broken picture if identities are unresolved, relationships are missing, or two systems disagree and nothing arbitrates.
For a CDO, the practical consequence is that existing data quality programs, necessary as they are, do not certify AI readiness. The certification requires measuring the graph the AI consumes, and it decomposes into four dimensions that map onto how context fails: coverage fails as blind spots, freshness fails as confident staleness, lineage fails as unexplainable decisions, and consistency fails as contradiction and identity drift. Each is measurable, and Figure 1 sketches the scorecard that makes them managerial objects.
The four dimensions, precisely #
Coverage is measured against a decision template, not against the universe: for each decision type, the context model declares which entities, relationships, and facts a well-made decision requires, and coverage is the share actually present. This scoping is what makes the metric honest; "we ingested forty systems" is activity, while "94 percent of what credit review needs is in the graph" is coverage. The template comes from the decision-first scoping practice described in Context Engineering Explained. Freshness was treated at length in Context Freshness: Keeping the Graph Alive: per-fact age against tiered SLAs, summarized here as the share of facts within tier and the lag distribution per tier.
Lineage is the traceability of every fact to its source system, document version, and transformation path. It is what lets the Decision Layer attach real evidence to recommendations and what turns an audit from archaeology into replay: reconstruct the situation as the AI saw it, with every input's provenance. Lineage is binary per fact and should be near-total in regulated domains, because a decision resting on even one untraceable fact is an unexplainable decision. Consistency is the graph's internal coherence: each real-world thing resolved to one identity, no unquarantined contradictory facts, relationships conforming to the context model's types. Its enemies are identity drift, the slow accumulation of duplicates that resolution misses, and silent contradiction, where two sources disagree and the graph serves whichever a query happens to hit. The Context Graph Engine computes all four continuously; the discipline is deciding thresholds, owners, and what happens on breach.
Quality dimensions and their metrics #
| Dimension | Question answered | Core metrics | Failure it prevents | Natural owner |
|---|---|---|---|---|
| Coverage | Is what this decision needs present? | Share of template entities, relationships, and facts present; gap list by source | Blind-spot decisions; confident answers from partial pictures | Domain steward |
| Freshness | Is what is present current? | Share of facts within SLA tier; lag distribution per tier; breach count | Confidently stale answers; decisions on yesterday's business | Platform owner |
| Lineage | Can every fact be traced and replayed? | Share of decision-used facts with full provenance; audit replay time | Unexplainable decisions; failed audits; untrusted evidence | Policy owner |
| Consistency | Is the graph internally coherent? | Duplicate-identity rate; contradictory fact pairs open and aging; model-violation count | Wrong-entity answers; contradiction served as fact | Context engineer |
Two design rules make the table operational. Score per decision domain, never enterprise-wide, because a 92 percent global average hides the one domain at 60 that is quietly generating bad decisions. And wire the scores into the Context Harness: below threshold, dependent decisions are escalated to humans with the gap named in the evidence, which converts quality from reporting into behavior.
The dimensions at work in three industries #
Healthcare. A care-coordination twin scores high on coverage and freshness but 88 percent on consistency: patient identities drift across admission systems. The scorecard localizes the risk, duplicate patients, and the Harness responds proportionately: coordination decisions for unresolved identities route to humans while the resolution backlog burns down.
Manufacturing. A quality-investigation domain discovers its lineage score is the binding constraint: supplier batch facts arrive through a transformation nobody documented, so recall decisions cannot cite provenance. Fixing lineage, not adding data, is what makes automated recall scoping defensible to regulators.
Retail banking. A collections domain scores 97 across three dimensions and 71 on coverage: hardship indicators live in call transcripts not yet unified into the graph, the gap pattern addressed in Structured and Unstructured Data in One Context. The scorecard turns a vague fairness concern into a specific ingestion project with a number attached.
A realistic enterprise scenario #
A CDO at a multinational insurer faces a board asking a simple question about the new claims automation: how do we know what it knows? The existing data quality program reports 96 percent record accuracy, which answers a different question. The CDO commissions a context quality baseline for the claims domain.
Understand. The team writes the claims decision template, what a well-adjudicated claim must know, and scores the twin against it: coverage 89 percent (medical report facts under-extracted), freshness 99, lineage 84 (a legacy transformation strips provenance from payment history), consistency 93 (claimant identity drift across two acquired books of business).
Decide. Thresholds and owners are set per dimension, and the Context Harness begins gating: claims touching unresolved identities or unprovenanced payment facts route to adjusters with the specific gap named. Remediation is prioritized by which gap blocks the most decisions.
Execute. The Execution Grid continues settling in-threshold claims while scores climb on a standing dashboard. As an illustrative range, organizations report that the first baseline typically finds one dimension far below the others, and that closing that single gap recovers the majority of blocked automation within one to two quarters. The board question now has a standing answer: four numbers, refreshed continuously.
Common mistakes to avoid #
- Reporting data quality as context quality: clean records do not certify the assembled picture; the graph the AI consumes is the thing to measure.
- Enterprise-wide averages: quality is per decision domain; one global score hides exactly the domain that needs attention.
- Coverage without a template: measuring presence against "everything we could ingest" inflates the score; measure against what each decision requires.
- Dashboards without gates: quality that does not flow into the Context Harness changes reports, not decisions.
- Letting contradiction age silently: consistency conflicts need quarantine and an aging metric; a contradiction served as fact is worse than a gap.
- Measuring once: quality decays with business change; the scorecard is a standing instrument with owners, not a project deliverable.
The CDO's first scorecard #
The path to the first scorecard fits in a quarter. Choose the decision domain where AI is furthest along, write its decision template with the domain steward, and let the Context Graph Engine compute the four dimensions against it. Set thresholds deliberately conservative, assign the four natural owners, and wire the Harness gates. Then publish the scorecard where the AI program's sponsors already look. The cultural effect tends to exceed the technical one: "trust the AI" becomes "coverage 94, freshness 98, lineage 100, consistency 96, gated," which is a sentence a board can actually govern with.
Quality presupposes the anatomy being measured, laid out in The Anatomy of Enterprise Context, and the freshness dimension has its own deep treatment in Context Freshness: Keeping the Graph Alive. The lifecycle that keeps scores improving is the operate stage of Context Engineering Explained, all in the Fundamentals cluster.
How OpenKnowra approaches this #
The four-dimension scorecard is method, not product, and any team can compute a first baseline manually. OpenKnowra makes it continuous: the Context Graph Engine scores coverage against decision templates, tracks per-fact freshness in the Enterprise Digital Twin, maintains full lineage so every Decision Layer recommendation is replayable, and quarantines consistency conflicts with aging metrics. The Context Harness turns thresholds into gates, so low-quality regions of the graph produce escalations with the gap named, never confident errors, and the Execution Grid's write-backs keep the scores honest.
A first step we offer: a baseline scorecard for one decision domain, your data, our computation, four numbers with a gap list, which is typically the fastest way to move the AI-trust conversation from adjectives to metrics.