Lineage and Provenance in Context Graphs

Trust in a context graph depends on a simple question: where did this fact come from? Context graph lineage must answer more than source-system name. It must preserve the record, extraction, transformation, resolution, inference, policy, approval, and decision history that turned raw evidence into enterprise action.

The 60-second read

Context graph lineage is the chain of evidence and transformation behind every node, relationship, property, inference, and decision. A production design records source identity, source record and field, event or document anchor, ingestion time, valid time, transformation version, entity-resolution decision, ontology version, confidence, reviewer state, policy decision, and downstream consumers. Provenance must remain attached through merges, splits, corrections, and derived relationships. The Context Harness uses this metadata to govern access and explain permitted use; the Decision Layer uses it to justify recommendations; the Execution Grid records which evidence supported an action and what happened afterward. The objective is not to create a larger metadata catalog. It is to make every material answer replayable, challengeable, and auditable.

Key takeaways

Definition

Context graph lineage is the persistent, queryable record of the sources, transformations, identity decisions, temporal states, inferences, policies, approvals, and downstream uses that produced and consumed each material fact or relationship in an enterprise context graph.

Why source-system labels are not enough #

A property that says “source: CRM” does not explain which record, field, event, or version supplied the value; whether it was normalized; whether another source disagreed; how two identities were merged; or whether the relationship was inferred. That information is essential when a user challenges an answer, an auditor requests replay, a policy changes, or a source correction arrives.

Lineage should therefore be modeled as operational context that travels with the fact. It must be available to APIs, reviewers, policies, decisions, and audits without requiring a separate forensic project.

UNDERSTAND · GOVERN · DECIDE · EXECUTE1 · EvidenceSource identity2 · TransformPipeline + ontology3 · ResolveIdentity decisions4 · TemporalValid + record time5 · DecidePolicy + approval6 · OutcomeAction + resultOutcomes and corrections feed the next operating cycle
Figure 1. A governed operating workflow that connects source evidence to context, decisions, and outcomes.

Stage 1: assign durable identities to evidence #

Every source system, dataset, stream, document, message, record, field, and event needs a durable identifier. For unstructured sources, retain a stable content hash, version, page or message position, and character or token span. For streams, retain topic, partition, offset, event ID, producer, and event time.

Durable evidence identity allows corrections and deletions to propagate. When a contract version is replaced or a source record is deleted, the engine can find every graph assertion derived from it and determine whether to retract, expire, or recompute them.

Stage 2: record transformation and modeling lineage #

Capture the pipeline, code or prompt version, ontology version, mapping rule, normalization steps, and quality checks applied to each assertion. For entity resolution, retain candidate set, matching features, score, threshold, survivorship rule, and any human override. For relationship extraction, preserve the evidence span and whether the edge is stated, calculated, or inferred.

This metadata should be queryable but tiered. Most users need a concise explanation; data stewards and auditors need the full chain. The Context Harness determines which level may be exposed for each purpose.

Stage 3: preserve temporal and correction history #

Lineage and time are inseparable. Record when the source fact was valid, when the engine observed it, when it was transformed, and when it was superseded or retracted. Never overwrite a prior assertion without preserving its history and the reason for change.

Merges and splits require particular care. If two customer identities are merged and later separated, past decisions must still show the identity state and evidence available at the time. The lineage chain must survive the correction.

Provenance metadata fields #

Field groupRequired metadataWhy it matters
Source identitySystem, dataset, record, field, event, document version, source anchorReconstruct the original evidence
Temporal stateValid from/to, observed at, processed at, superseded atReplay what was true and known
TransformationPipeline, code/prompt version, mapping, ontology versionExplain how evidence became context
ResolutionCandidates, features, score, threshold, survivorship, overrideExplain identity merges and splits
InferenceMethod, confidence, supporting facts, reviewer statusDistinguish derived knowledge from evidence
Policy and usePurpose, access decision, consumer, decision, action, outcomeAudit how context influenced enterprise action

Stage 4: extend lineage through decisions and execution #

Data lineage ends too early if it stops at the graph. Decision lineage should connect the assembled situation, evidence set, quality state, policy evaluation, model or rule version, recommendation, human approval, executed action, and observed outcome. This is what makes the Understand-Decide-Execute loop auditable.

The same chain supports learning. When an outcome is poor, teams can determine whether the failure came from missing source data, stale context, incorrect resolution, a weak rule, an inappropriate policy, or execution failure. Without end-to-end lineage, every incident becomes argument instead of diagnosis.

The pattern across three industries #

Financial services. Trace a risk flag through transaction events, entity-resolution decisions, beneficial-ownership edges, rules, analyst approval, and filed action.

Pharmaceuticals. Trace a quality decision through batch records, laboratory results, equipment state, deviations, document versions, and release approval.

Energy. Trace a maintenance recommendation through sensor events, asset hierarchy, work history, model version, safety policy, operator approval, and work-order completion.

A realistic enterprise scenario #

Enterprise scenario

A pharmaceutical manufacturer must explain why a production batch was released and later recalled. The release decision used laboratory results, equipment readings, deviation records, supplier certificates, and an approved exception.

Understand. The Enterprise Digital Twin reconstructs the exact batch situation as known on the release date. Each property and relationship links to its source record, document version, extraction span, transformation version, and valid-time state.

Decide. Decision lineage shows that an exception policy allowed release after a human approver accepted a supplier certificate. A later source correction changed the certificate status, but the original decision remains replayable with the evidence then available.

Execute. The recall action links back to the corrected evidence, policy evaluation, approval, affected products, and downstream notifications. Investigators identify the failure as delayed source correction rather than missing graph data or model error, allowing the right control to be strengthened.

Common mistakes to avoid #

Watch out for
  1. Recording only the source-system name instead of record, field, event, or passage-level evidence.
  2. Storing lineage in a separate catalog that does not travel with graph answers.
  3. Overwriting assertions and losing prior values, identity states, or reasons for correction.
  4. Omitting entity-resolution and inference decisions from provenance.
  5. Stopping lineage at the graph and failing to link evidence to decisions, approvals, actions, and outcomes.
  6. Collecting exhaustive metadata without defining which fields are mandatory, queryable, and governed.

Operating lineage as a product capability #

Define mandatory lineage fields by assertion class and decision criticality. Monitor coverage, broken source links, orphaned assertions, unknown transformation versions, and decisions missing evidence. Include provenance in contract tests for every ingestion and modeling pipeline.

Provide human-readable explanations alongside machine-readable provenance. A lineage system succeeds when an executive, steward, architect, auditor, and API can each obtain the level of evidence appropriate to their role from the same underlying chain.

How OpenKnowra approaches this #

OpenKnowra models lineage as first-class context. The Context Graph Engine attaches source, transformation, temporal, resolution, and inference provenance to graph assertions. The Context Harness governs who can inspect which level of evidence. The Decision Layer records the evidence and policy behind recommendations, and the Execution Grid links approved action and observed outcome back to the original situation. This creates a complete Understand-Decide-Execute chain that can be replayed and challenged.

Frequently asked questions

What is context graph lineage?
It is the persistent, queryable chain showing which sources, records, fields, events, documents, transformations, identity decisions, inferences, policies, and approvals produced each graph assertion and how that assertion was later used.
How is provenance different from lineage?
Provenance describes origin and authorship of a specific fact or artifact. Lineage describes the broader chain of movement and transformation from source through graph, decision, and outcome. In practice, production context graphs need both and should expose them through one connected model.
What provenance metadata should every graph fact contain?
At minimum: source identity, source record or anchor, valid and observed time, ingestion and transformation version, ontology mapping, confidence or resolution status, and supersession state. High-impact facts also need policy, reviewer, consumer, and decision-use metadata.
How should lineage work for inferred relationships?
Inferred edges must be clearly labeled and linked to the supporting facts, inference method or model version, confidence, creation time, reviewer status, and any policy that limits their use. They should never be indistinguishable from directly evidenced relationships.
Why extend lineage to decisions and actions?
Because data can be correct while a decision or execution step is wrong. Linking evidence to policy, recommendation, approval, action, and outcome allows enterprises to diagnose failures accurately, satisfy audits, and improve the operating loop.

Keep exploring this cluster