Context graph scope is the governed boundary that determines which entities, relationships, attributes, events, policies, and references must be represented to support defined enterprise decisions.
Stage 1: Start from decisions, not data sources #
List the decisions the organization wants to improve, automate, or make consistent. Examples include prioritizing a customer intervention, allocating constrained inventory, approving an exception, routing a case, or scheduling qualified work. For each decision, identify the situation the decision-maker must understand.
A source-first approach usually creates a graph that is broad but shallow. A decision-first approach creates a graph that is narrow but useful. The first artifact should be a decision-context matrix, not a catalog of available tables.
Stage 2: Identify anchor entities #
Anchor entities are the durable business objects around which decisions recur: customer, employee, supplier, product, contract, asset, location, order, project, account, policy, or case. Include an entity when it has a stable identity, appears in several decisions, and must be connected to other objects.
Do not create entities for every code, row, or report label. Some values are attributes, classifications, or references rather than independent nodes.
| Candidate element | Graph treatment | Reason |
|---|---|---|
| Customer, asset, employee, supplier | Include as entity | Durable identity reused across decisions |
| Contract entitlement or certification | Include as relationship or policy-bearing entity | Changes what actions are permitted |
| Current capacity, eligibility, or status | Include as state | Material to operational decisions |
| Raw sensor stream | Reference or summarize | High volume; historian remains source of detail |
| Full document body | Reference with metadata | Document store remains authoritative |
| Old report extracts | Usually exclude | Duplicative and weakly governed |
| Every transaction line | Usually reference or aggregate | Use source platform for detail unless the decision traverses individual lines |
Stage 3: Model only meaningful relationships #
Relationships belong in the graph when they change how a decision is interpreted. “Customer holds contract,” “component is used in product,” “employee is certified for asset,” and “order is allocated to facility” carry operational meaning. Generic links such as “related to” weaken the model because they hide business semantics.
Every relationship should have an owner, direction, time validity where relevant, and source provenance.
Stage 4: Separate state from history and payload #
Current status, eligibility, risk, capacity, ownership, and policy-relevant state often belong in the graph. Full transaction lines, document bodies, high-frequency sensor streams, and long analytical histories usually belong elsewhere. The graph can retain a reference, summary, embedding pointer, or latest material state.
This keeps the Context Graph Engine responsive and reduces synchronization burden without losing access to evidence.
Stage 5: Include policy and action context #
A graph becomes decision-grade when it contains or can resolve the policies, permissions, obligations, thresholds, and action interfaces that govern what may happen next. A customer entity alone is descriptive. A customer connected to contract terms, consent, service level, active incidents, and permitted remedies is operational context.
Stage 6: Apply the inclusion test #
Include an element when it is required by a defined decision, reused by more than one use case, changes interpretation, must be current, or is needed for policy and audit. Reference it when the detail is large, rarely traversed, or already served reliably elsewhere. Exclude it when no decision uses it, ownership is unclear, or it duplicates another representation without added meaning.
The pattern across three industries #
A bank may include customers, accounts, products, devices, beneficiaries, cases, permissions, and risk state, while leaving full ledger history in the transaction platform. A manufacturer may include products, components, suppliers, plants, equipment, orders, qualifications, and capacity state, while leaving raw sensor streams in the historian. A healthcare provider may include patients, clinicians, appointments, care plans, consent, facilities, and eligibility, while leaving full clinical documents in the electronic health record.
A realistic enterprise scenario #
An energy company initially proposes copying its asset registry, maintenance history, telemetry, work orders, manuals, and geographic data into one graph. The architecture team instead starts with the outage-response decision. It models asset, site, feeder, customer segment, crew, skill, permit, work order, and safety policy. Current alarms and availability enter as state. Detailed telemetry remains in the historian, and manuals remain in document storage with governed links. The first graph is smaller than the source inventory but supports outage triage, crew assignment, customer prioritization, and restoration communication.
Common mistakes #
1. Equating the graph with a data lake.
2. Modelling every source field before identifying decision use.
3. Creating duplicate entities for each source system.
4. Loading raw event volume that should remain in a stream or historian.
5. Excluding policy, provenance, and action metadata because they are not “business data.”
Stage 7: Review scope as the decision portfolio grows #
Treat scope as a governed product backlog. New use cases should reuse existing entities before adding new ones. Every addition should document the decision served, source of truth, refresh requirement, access rule, and retirement condition. Review low-use nodes and relationships periodically. The objective is not maximum coverage; it is maximum decision reuse per governed element.
How OpenKnowra approaches this #
The explanation above is category-level guidance and applies regardless of platform. OpenKnowra implements the pattern through a Context Engine that synchronizes selected enterprise state, resolves it through the Context Graph Engine, evaluates choices in the Decision Layer, carries permitted actions through the Execution Grid, and governs the full loop through the Context Harness. The objective is a progressively richer Enterprise Digital Twin built around measurable Understand-Decide-Execute loops.