What Belongs in Your Context Graph (and What Does Not)

A context graph should contain the minimum connected enterprise knowledge required to understand decisions and govern action. It should not become a copy of every source system. This guide provides a practical method for setting context graph scope.

The 60-second read

Context graph scope should be driven by decisions, not by source-system inventories. Begin with high-value entities, relationships, state, policies, and outcomes that several use cases can reuse. Keep large documents, raw telemetry, detailed transactions, and analytical history in their systems of record unless the graph needs selected references or summaries. A disciplined scope produces a smaller, fresher, more governable Enterprise Digital Twin and avoids turning the graph into another uncontrolled data platform.

Key takeaways

Definition

Context graph scope is the governed boundary that determines which entities, relationships, attributes, events, policies, and references must be represented to support defined enterprise decisions.

Context graph scoping decision tree Does a defined decision need it? Does it change interpretation or policy? Is it reusable and governable? Include Reference or exclude
Figure 1. A decision-led test keeps the graph focused on reusable operational context.

Stage 1: Start from decisions, not data sources #

List the decisions the organization wants to improve, automate, or make consistent. Examples include prioritizing a customer intervention, allocating constrained inventory, approving an exception, routing a case, or scheduling qualified work. For each decision, identify the situation the decision-maker must understand.

A source-first approach usually creates a graph that is broad but shallow. A decision-first approach creates a graph that is narrow but useful. The first artifact should be a decision-context matrix, not a catalog of available tables.

Stage 2: Identify anchor entities #

Anchor entities are the durable business objects around which decisions recur: customer, employee, supplier, product, contract, asset, location, order, project, account, policy, or case. Include an entity when it has a stable identity, appears in several decisions, and must be connected to other objects.

Do not create entities for every code, row, or report label. Some values are attributes, classifications, or references rather than independent nodes.

Candidate elementGraph treatmentReason
Customer, asset, employee, supplierInclude as entityDurable identity reused across decisions
Contract entitlement or certificationInclude as relationship or policy-bearing entityChanges what actions are permitted
Current capacity, eligibility, or statusInclude as stateMaterial to operational decisions
Raw sensor streamReference or summarizeHigh volume; historian remains source of detail
Full document bodyReference with metadataDocument store remains authoritative
Old report extractsUsually excludeDuplicative and weakly governed
Every transaction lineUsually reference or aggregateUse source platform for detail unless the decision traverses individual lines

Stage 3: Model only meaningful relationships #

Relationships belong in the graph when they change how a decision is interpreted. “Customer holds contract,” “component is used in product,” “employee is certified for asset,” and “order is allocated to facility” carry operational meaning. Generic links such as “related to” weaken the model because they hide business semantics.

Every relationship should have an owner, direction, time validity where relevant, and source provenance.

Stage 4: Separate state from history and payload #

Current status, eligibility, risk, capacity, ownership, and policy-relevant state often belong in the graph. Full transaction lines, document bodies, high-frequency sensor streams, and long analytical histories usually belong elsewhere. The graph can retain a reference, summary, embedding pointer, or latest material state.

This keeps the Context Graph Engine responsive and reduces synchronization burden without losing access to evidence.

Stage 5: Include policy and action context #

A graph becomes decision-grade when it contains or can resolve the policies, permissions, obligations, thresholds, and action interfaces that govern what may happen next. A customer entity alone is descriptive. A customer connected to contract terms, consent, service level, active incidents, and permitted remedies is operational context.

Stage 6: Apply the inclusion test #

Include an element when it is required by a defined decision, reused by more than one use case, changes interpretation, must be current, or is needed for policy and audit. Reference it when the detail is large, rarely traversed, or already served reliably elsewhere. Exclude it when no decision uses it, ownership is unclear, or it duplicates another representation without added meaning.

The pattern across three industries #

A bank may include customers, accounts, products, devices, beneficiaries, cases, permissions, and risk state, while leaving full ledger history in the transaction platform. A manufacturer may include products, components, suppliers, plants, equipment, orders, qualifications, and capacity state, while leaving raw sensor streams in the historian. A healthcare provider may include patients, clinicians, appointments, care plans, consent, facilities, and eligibility, while leaving full clinical documents in the electronic health record.

A realistic enterprise scenario #

Enterprise scenario

An energy company initially proposes copying its asset registry, maintenance history, telemetry, work orders, manuals, and geographic data into one graph. The architecture team instead starts with the outage-response decision. It models asset, site, feeder, customer segment, crew, skill, permit, work order, and safety policy. Current alarms and availability enter as state. Detailed telemetry remains in the historian, and manuals remain in document storage with governed links. The first graph is smaller than the source inventory but supports outage triage, crew assignment, customer prioritization, and restoration communication.

Common mistakes #

Watch out for

1. Equating the graph with a data lake.

2. Modelling every source field before identifying decision use.

3. Creating duplicate entities for each source system.

4. Loading raw event volume that should remain in a stream or historian.

5. Excluding policy, provenance, and action metadata because they are not “business data.”

Stage 7: Review scope as the decision portfolio grows #

Treat scope as a governed product backlog. New use cases should reuse existing entities before adding new ones. Every addition should document the decision served, source of truth, refresh requirement, access rule, and retirement condition. Review low-use nodes and relationships periodically. The objective is not maximum coverage; it is maximum decision reuse per governed element.

How OpenKnowra approaches this #

The explanation above is category-level guidance and applies regardless of platform. OpenKnowra implements the pattern through a Context Engine that synchronizes selected enterprise state, resolves it through the Context Graph Engine, evaluates choices in the Decision Layer, carries permitted actions through the Execution Grid, and governs the full loop through the Context Harness. The objective is a progressively richer Enterprise Digital Twin built around measurable Understand-Decide-Execute loops.

Frequently asked questions

What should be the first entities in a context graph?
Start with durable entities that anchor several high-value decisions, such as customer, product, supplier, employee, asset, contract, order, and location.
Should the graph contain all enterprise data?
No. It should contain decision-relevant entities, relationships, state, policy, and references. Large payloads and raw history can remain in systems designed for them.
How do we decide whether something is an entity or an attribute?
Use an entity when it has independent identity, lifecycle, relationships, or policy significance. Use an attribute when it simply describes another entity.
Can documents and telemetry still be used?
Yes. The graph can hold governed references, summaries, and retrieval metadata while the detailed content remains in document stores or historians.
How often should scope be reviewed?
Review it whenever new decision domains are added and periodically remove or consolidate elements that no longer support active use cases.

Keep exploring this cluster