Context graph modeling is the practice of defining the entities, typed relationships, states, events, policies, and evidence required to represent a decision domain as a living, governed model. Its purpose is not to reproduce source systems, but to assemble the minimum connected context needed for the enterprise to understand situations, make decisions, and execute actions safely.
Stage 1: Begin with decisions, not data #
The first question in context graph modeling is not, “What data do we have?” It is, “What recurring decisions are important enough to deserve reliable context?” Source-first modeling almost always inherits the boundaries, naming, duplication, and omissions of the systems being integrated. Decision-first modeling reverses the sequence. It identifies the situation a person or agent must understand, the alternatives available, the policies that constrain the choice, the evidence required for approval, and the actions that follow.
For a supply disruption decision, the useful situation may include a supplier, its facilities, the components it supplies, the products that depend on those components, current inventory, open purchase orders, contractual alternatives, and customer commitments. The ERP schema is an input, but it is not the model. The model is the connected business reality required to decide whether to expedite, substitute, reallocate, or escalate.
Artifact: the decision-context canvas
Create a one-page canvas for each target decision. Record the trigger, decision owner, time horizon, required evidence, candidate actions, governing policies, expected outcome, and post-action feedback. This canvas becomes the acceptance test for the model. A concept that does not help populate the canvas is probably not in the first release.
Stage 2: Select the core entities #
An entity deserves first-class identity when it persists through time, participates in important relationships, carries decision-relevant state, or anchors evidence and policy. Customers, suppliers, employees, contracts, products, assets, facilities, accounts, cases, orders, and organizational units are common examples. Not every noun becomes an entity. “Priority,” “status,” and “temperature” are usually attributes. A shipment may be an entity if it is tracked, related to orders and carriers, and governed by service commitments. It may be only an event if the domain cares solely that dispatch occurred.
Define each entity with a business identity statement, not a database key. A Supplier is “a legal or operating party from which the enterprise procures goods or services.” Then define identity rules, lifecycle states, authoritative sources, and resolution evidence. This prevents three supplier IDs from three systems from becoming three suppliers in the twin.
Stable identity: does the thing remain recognizable across systems and time? Decision relevance: can its state change a decision? Relationship centrality: does it connect important parts of the domain? Lifecycle: do creation, change, and retirement matter? Governance: does policy attach to it? If fewer than two answers are yes, begin as an attribute or event and promote later only if needed.
Stage 3: Make relationships first-class #
Relationships are not decorative connectors. They encode the operational meaning that lets a graph answer multi-entity questions. A useful relationship has a business verb, direction, cardinality, valid time, provenance, and, where relevant, attributes and confidence. Supplier supplies Component is stronger than Supplier related_to Component. Employee manages Team differs from Employee belongs_to Team. Account owned_by Customer differs from Customer authorized_on Account.
Many relationships deserve their own lifecycle. A supplier relationship can carry tier, allocation share, approved-site scope, contract reference, effective dates, and alternate-source status. An employment relationship carries role, manager, location, cost center, start date, and end date. Treating these as foreign keys or undifferentiated links discards exactly the context the Decision Layer needs.
Artifact: the relationship catalog
For every edge type, document the source and target entity types, natural-language definition, inverse label, cardinality, temporal behavior, required provenance, confidence treatment, and policy implications. The catalog becomes both ontology documentation and a test specification for ingestion and reasoning.
Stage 4: Separate state, events, and evidence #
Current state answers, “What is true now?” Events answer, “What happened?” Evidence answers, “Why do we believe it?” Combining them creates ambiguity. A machine's operating status is state; a shutdown at 14:03 is an event; the sensor reading and maintenance ticket are evidence. A customer risk tier is state; a missed payment is an event; the ledger entry is evidence.
State should be time-bounded and reconstructable. Events should be immutable, ordered, and linked to affected entities. Evidence should preserve source, extraction method, document anchor, and confidence. This separation enables the Context Graph Engine to replay situations, reason over change, and show the Decision Layer both the current picture and the chain of facts behind it.
Stage 5: Model policy and permitted action #
A context graph becomes operational when it represents not only what exists, but what may be done. Policies can be modeled as entities connected to the subjects, objects, jurisdictions, thresholds, and actions they govern. A credit policy may constrain an Offer action based on customer segment, exposure, jurisdiction, and approval level. A workforce policy may constrain Assignment based on certification, location, labor rules, and conflict-of-interest relationships.
The Context Harness evaluates these connections at decision time. The model should therefore make policy scope explicit and versioned. Avoid burying important rules in application code or unstructured documents without anchors. The graph need not execute every rule itself, but it must be able to identify the relevant policy and provide the evidence needed for enforcement.
Modeling patterns versus anti-patterns #
| Modeling concern | Pattern | Anti-pattern | Why it matters |
|---|---|---|---|
| Starting point | Begin with a bounded decision and its situation | Copy all source schemas into a graph | Decision-first models stay purposeful and testable |
| Entity design | Use stable business identity and lifecycle | Promote every table row or noun to an entity | Prevents node explosion and duplicate concepts |
| Relationships | Use typed, directed, dated, evidenced edges | Use generic related_to links | Preserves business meaning and supports reasoning |
| Attributes | Keep simple descriptive values on entities or edges | Turn every field into a node | Reduces traversal cost and cognitive clutter |
| Time | Separate state, events, valid time, and record time | Overwrite current values without history | Enables replay, audit, and change reasoning |
| Inference | Label inferred facts with confidence and method | Mix inference with observed evidence | Allows policy to distinguish belief from fact |
| Scope | Expand domain by domain along proven decisions | Design the complete enterprise ontology first | Accelerates production value and learning |
Three industry patterns #
Banking. Core entities include Customer, Legal Entity, Account, Product, Transaction, Case, and Policy. High-value edges include owns, controls, authorized_on, transacted_with, and subject_to. Events such as payment, alert, review, and ownership change remain distinct from state such as risk tier and account status. This model supports KYC refresh, exposure analysis, fraud-ring detection, and governed next-best action.
Manufacturing. Core entities include Supplier, Facility, Component, Product, Asset, Work Order, Contract, and Inventory Position. Relationships such as supplies, manufactured_at, depends_on, substitutes_for, and covered_by expose transitive risk. Events capture failures, inspections, shipments, and maintenance. State captures capacity, health, availability, and quality status.
Healthcare. Core entities include Patient, Practitioner, Encounter, Condition, Medication, Facility, Care Plan, and Authorization. Relationships such as treated_by, prescribed, contraindicated_with, referred_to, and covered_by must be time-bounded and purpose-governed. Events record encounters, administrations, tests, and consent changes. The Context Harness restricts which relationships and evidence can be assembled for each care or administrative purpose.
A realistic enterprise scenario #
A global manufacturer wants to reduce the time required to respond when a critical supplier reports a disruption. Its first attempt created a graph by loading ERP tables, procurement records, and supplier-master data. The graph contained millions of nodes but could not answer the operational question: which customer commitments are at risk, what alternatives exist, and which actions are permitted?
Understand. The architecture team reframes the domain around the disruption decision. It defines Supplier, Facility, Component, Product, Inventory Position, Contract, Customer Commitment, and Alternate Source as core entities. It types relationships such as Supplier supplies Component at Facility, Component used_in Product, Product fulfills Commitment, and Contract permits_substitution_with conditions. Current inventory is modeled as dated state; shipment, failure, and quality-hold records are events; contracts and notices remain anchored evidence.
Decide. The Decision Layer assembles a situation for the affected facility. Multi-hop traversal identifies dependent products and commitments. Policy edges identify which substitutions are pre-approved, which require engineering review, and which customer tiers require notification. Inferred alternate-source edges remain confidence-labeled and cannot trigger automatic substitution without approval.
Execute. The Execution Grid opens engineering reviews, proposes purchase-order reallocations, and drafts customer notifications. Outcomes return through ingestion, updating allocation state and creating an auditable decision record. The model proves useful because every element was selected for this loop, not because it mirrored every source field.
Common mistakes to avoid #
- Copying schemas instead of modeling the business: source systems reflect application boundaries, not the connected reality required for decisions.
- Using nouns without identity rules: a named entity without resolution and lifecycle rules becomes a duplicate generator.
- Using vague relationships: generic edges make diagrams look connected while preventing reliable queries, policy, and reasoning.
- Over-modeling attributes as nodes: this increases graph size and traversal complexity without adding meaning.
- Overwriting state: without valid time, record time, and events, the enterprise cannot reconstruct what was known when a decision was made.
- Building the enterprise ontology before a production domain: breadth delays feedback and turns governance into abstract debate.
Stage 6: Validate, govern, and evolve the model #
Validate the model with questions, not diagrams. Write representative situation queries and expected answers. Can the model identify every product affected by a facility outage? Can it explain why an account is linked to a beneficial owner? Can it reconstruct the relationships that were valid when a past decision occurred? Can policy distinguish observed, asserted, and inferred facts?
Govern the model through named domain owners, versioned definitions, compatibility rules, and release tests. Changes should include impact analysis for ingestion mappings, stored graph data, context APIs, Decision Layer logic, Harness policies, and Execution Grid actions. A new edge type is not complete until its provenance, temporal behavior, permissions, and operational owner are defined.
Practical delivery sequence
Week 1 to 2, illustrative: define decision canvases and vocabulary. Week 3 to 4: agree core entities and identity rules. Week 5 to 6: build the relationship catalog and temporal model. Week 7 to 9: connect representative sources, resolve entities, and validate situations. Week 10 to 12: add policy, permitted actions, quality gates, and human approvals. The exact range depends on source access and governance, but the sequence remains useful: prove the model through a working Understand-Decide-Execute loop.
The modeling artifact pack #
A production model should leave behind a small set of concrete artifacts rather than only a diagram. The first is the domain glossary: each concept has one preferred name, a plain-language definition, examples, exclusions, owner, and lifecycle. The second is the entity profile: identity rules, source identifiers, resolution keys, survivorship policy, mandatory state, retention requirements, and sensitivity classification. The third is the relationship catalog: definitions, direction, inverse labels, cardinality, temporal rules, provenance, confidence, and permissions. The fourth is the event catalog: event names, triggers, affected entities, ordering requirements, payload, and replay behavior. The fifth is the decision trace: a representative situation showing exactly which nodes, edges, policies, and evidence support a decision and which action records return afterward.
These artifacts divide responsibility cleanly. Business owners validate meaning. Data and integration teams map source records into the model. Security and risk teams define purpose and access. Product and operations teams confirm that the assembled situation supports the real decision. Architecture owns coherence across domains, but it should not become the sole author of business semantics. A context model is durable only when the people who operate the domain recognize it as their reality.
Entity profile example
For a Supplier entity, the identity statement may define a legal or operating party from which goods or services are procured. Resolution evidence can include legal registration number, tax identifier, approved supplier number, bank-account ownership, address, and parent organization. The profile should distinguish legal entity from operating site, because combining them creates false relationships between contracts, risks, and facilities. Lifecycle states might include Prospective, Approved, Restricted, Suspended, and Retired. Decision-relevant state can include criticality, financial risk, compliance status, capacity, and last-assessed date. Every field should identify whether it is observed, asserted by a source, calculated, or inferred.
Relationship profile example
For Supplier supplies Component, define whether the relationship applies globally or only through a named facility, contract, product program, geography, or time period. Record effective and end dates, allocation share, approved status, qualification evidence, lead time, minimum order quantity, and the source that asserted the relationship. Define the inverse label, Component supplied_by Supplier, for human clarity, but preserve one canonical direction for queries. State whether multiple suppliers can supply the same component and whether the same supplier can use several facilities. These details turn an attractive diagram into an executable semantic contract.
How to test whether the model is useful #
Use three classes of tests. Coverage tests ask whether every required concept and relationship for the target decision can be represented. Integrity tests check cardinality, required provenance, allowed state transitions, temporal consistency, and identity constraints. Decision tests replay realistic situations and verify that the graph assembles the expected evidence, policy, alternatives, and actions. A model can pass coverage and integrity while still failing the decision test, which is why the final test must remain operational.
Include negative cases. What happens when two sources disagree on a supplier's parent? When a relationship is valid but stale? When an inferred beneficial-owner edge has only moderate confidence? When policy changed after the underlying event but before the decision? The Context Harness should be able to treat these conditions differently. The model must expose them rather than normalize them away.
Measure quality by decision domain, not globally. A customer graph may be complete enough for marketing segmentation but insufficient for credit approval. A supplier model may support spend analytics while lacking the facility-level dependencies needed for disruption response. Coverage, freshness, lineage, and consistency should therefore be scored against the situation the model is expected to serve. This prevents a high aggregate quality score from hiding a critical gap in one operational path.
How OpenKnowra approaches this #
The modeling practices above are vendor-neutral. OpenKnowra implements them through a Context Graph Engine that resolves identity, validates concepts against governed domain models, types and timestamps relationships, preserves lineage, and serves decision-ready situations as part of an Enterprise Digital Twin.
Assemblies provide reusable domain models, context APIs, and policy patterns. The Context Harness controls purpose, evidence, confidence, and permitted action. The Decision Layer consumes the resulting situations, while the Execution Grid records governed outcomes back into the graph. The objective is not the largest ontology. It is the smallest durable model that can support many high-value decisions and improve as the operating loop runs.