Context graph lineage is the persistent, queryable record of the sources, transformations, identity decisions, temporal states, inferences, policies, approvals, and downstream uses that produced and consumed each material fact or relationship in an enterprise context graph.
Why source-system labels are not enough #
A property that says “source: CRM” does not explain which record, field, event, or version supplied the value; whether it was normalized; whether another source disagreed; how two identities were merged; or whether the relationship was inferred. That information is essential when a user challenges an answer, an auditor requests replay, a policy changes, or a source correction arrives.
Lineage should therefore be modeled as operational context that travels with the fact. It must be available to APIs, reviewers, policies, decisions, and audits without requiring a separate forensic project.
Stage 1: assign durable identities to evidence #
Every source system, dataset, stream, document, message, record, field, and event needs a durable identifier. For unstructured sources, retain a stable content hash, version, page or message position, and character or token span. For streams, retain topic, partition, offset, event ID, producer, and event time.
Durable evidence identity allows corrections and deletions to propagate. When a contract version is replaced or a source record is deleted, the engine can find every graph assertion derived from it and determine whether to retract, expire, or recompute them.
Stage 2: record transformation and modeling lineage #
Capture the pipeline, code or prompt version, ontology version, mapping rule, normalization steps, and quality checks applied to each assertion. For entity resolution, retain candidate set, matching features, score, threshold, survivorship rule, and any human override. For relationship extraction, preserve the evidence span and whether the edge is stated, calculated, or inferred.
This metadata should be queryable but tiered. Most users need a concise explanation; data stewards and auditors need the full chain. The Context Harness determines which level may be exposed for each purpose.
Stage 3: preserve temporal and correction history #
Lineage and time are inseparable. Record when the source fact was valid, when the engine observed it, when it was transformed, and when it was superseded or retracted. Never overwrite a prior assertion without preserving its history and the reason for change.
Merges and splits require particular care. If two customer identities are merged and later separated, past decisions must still show the identity state and evidence available at the time. The lineage chain must survive the correction.
Provenance metadata fields #
| Field group | Required metadata | Why it matters |
|---|---|---|
| Source identity | System, dataset, record, field, event, document version, source anchor | Reconstruct the original evidence |
| Temporal state | Valid from/to, observed at, processed at, superseded at | Replay what was true and known |
| Transformation | Pipeline, code/prompt version, mapping, ontology version | Explain how evidence became context |
| Resolution | Candidates, features, score, threshold, survivorship, override | Explain identity merges and splits |
| Inference | Method, confidence, supporting facts, reviewer status | Distinguish derived knowledge from evidence |
| Policy and use | Purpose, access decision, consumer, decision, action, outcome | Audit how context influenced enterprise action |
Stage 4: extend lineage through decisions and execution #
Data lineage ends too early if it stops at the graph. Decision lineage should connect the assembled situation, evidence set, quality state, policy evaluation, model or rule version, recommendation, human approval, executed action, and observed outcome. This is what makes the Understand-Decide-Execute loop auditable.
The same chain supports learning. When an outcome is poor, teams can determine whether the failure came from missing source data, stale context, incorrect resolution, a weak rule, an inappropriate policy, or execution failure. Without end-to-end lineage, every incident becomes argument instead of diagnosis.
The pattern across three industries #
Financial services. Trace a risk flag through transaction events, entity-resolution decisions, beneficial-ownership edges, rules, analyst approval, and filed action.
Pharmaceuticals. Trace a quality decision through batch records, laboratory results, equipment state, deviations, document versions, and release approval.
Energy. Trace a maintenance recommendation through sensor events, asset hierarchy, work history, model version, safety policy, operator approval, and work-order completion.
A realistic enterprise scenario #
A pharmaceutical manufacturer must explain why a production batch was released and later recalled. The release decision used laboratory results, equipment readings, deviation records, supplier certificates, and an approved exception.
Understand. The Enterprise Digital Twin reconstructs the exact batch situation as known on the release date. Each property and relationship links to its source record, document version, extraction span, transformation version, and valid-time state.
Decide. Decision lineage shows that an exception policy allowed release after a human approver accepted a supplier certificate. A later source correction changed the certificate status, but the original decision remains replayable with the evidence then available.
Execute. The recall action links back to the corrected evidence, policy evaluation, approval, affected products, and downstream notifications. Investigators identify the failure as delayed source correction rather than missing graph data or model error, allowing the right control to be strengthened.
Common mistakes to avoid #
- Recording only the source-system name instead of record, field, event, or passage-level evidence.
- Storing lineage in a separate catalog that does not travel with graph answers.
- Overwriting assertions and losing prior values, identity states, or reasons for correction.
- Omitting entity-resolution and inference decisions from provenance.
- Stopping lineage at the graph and failing to link evidence to decisions, approvals, actions, and outcomes.
- Collecting exhaustive metadata without defining which fields are mandatory, queryable, and governed.
Operating lineage as a product capability #
Define mandatory lineage fields by assertion class and decision criticality. Monitor coverage, broken source links, orphaned assertions, unknown transformation versions, and decisions missing evidence. Include provenance in contract tests for every ingestion and modeling pipeline.
Provide human-readable explanations alongside machine-readable provenance. A lineage system succeeds when an executive, steward, architect, auditor, and API can each obtain the level of evidence appropriate to their role from the same underlying chain.
How OpenKnowra approaches this #
OpenKnowra models lineage as first-class context. The Context Graph Engine attaches source, transformation, temporal, resolution, and inference provenance to graph assertions. The Context Harness governs who can inspect which level of evidence. The Decision Layer records the evidence and policy behind recommendations, and the Execution Grid links approved action and observed outcome back to the original situation. This creates a complete Understand-Decide-Execute chain that can be replayed and challenged.