A context graph is a governed, continuously updated model of one business domain: its entities resolved across systems, the relationships between them, their live state and history, and the policies that constrain decisions. Connected domain graphs, kept current inside a Context Engine, grow into the Enterprise Digital Twin.
Before you build a context graph: what you are actually making #
A context graph is not a data warehouse with edges, and it is not the corporate ontology program wearing new clothes. It is the smallest connected model of a domain that lets software answer a decision-grade question: not "what data do we have about suppliers" but "which purchase orders are exposed if this supplier's audit fails, and what does policy let us do about it tonight".
That distinction drives everything in this guide. A knowledge graph connects entities and relationships; a context graph adds live state, time, policy, and decision memory, because its consumer is not a human browsing a visualization but a Decision Layer assembling a situation at the moment of action. Get this consumer right and the build decisions get easier: freshness matters where the graph will act, entity resolution matters where systems disagree, and policy belongs in the model because the Context Harness enforces it there once instead of every application enforcing it badly.
This is also why the work is increasingly called context engineering rather than data modeling. The discipline is not describing data; it is engineering what a decision needs to know into a shape software can consume: the right subgraph, at the right freshness, with the right policies attached, served at the moment of action. Hold that consumer in mind at every stage and most design debates resolve themselves.
The sequence below is the pattern that repeatedly works when teams build a context graph for the first time: five stages, each with a concrete exit artifact, each small enough to fail cheaply and fix quickly. Figure 1 shows the whole sequence; the rest of the guide takes it stage by stage, with worked examples from banking, manufacturing, and healthcare throughout.
Stage 1: Pick the domain and the decision #
The single most predictive choice in the whole build is made before any modeling: which decision will this graph serve? Not which domain has the most data, not which executive is loudest, but which recurring decision is expensive because context is fragmented. The decision defines the graph: its entities, its relationships, its freshness requirements, and its first three systems all fall out of one well-chosen decision.
Good first decisions share four properties. They recur weekly or daily, so the loop gets reps quickly. They are cross-system, so the graph adds something no source system has alone. They have a measurable outcome, cycle time, leakage, escalation rate, so value is provable. And they are consequential but survivable: important enough to fund, safe enough to run with human approval while trust builds. In banking, collections treatment for mid-risk accounts fits; a first-line chatbot answer does not. In manufacturing, supplier-quality disposition fits; annual network redesign does not. In healthcare, complex discharge coordination fits; diagnosis does not and should not.
Write the choice down as a one-page decision brief, the stage's exit artifact: the decision in one sentence, who makes it today and how long it takes, the systems consulted, the policies that constrain it, and the metric that will prove improvement. This page becomes the acceptance test for everything that follows; if a proposed entity or feed does not serve the brief, it waits.
- One decision, one sentence, signed by the domain owner
- Baseline measured: current cycle time, volume, and error or escalation rate
- The three candidate systems named, with their owners identified
- Constraining policies listed, with where they are written down today
- The Understand-Decide-Execute loop sketched for this decision on one whiteboard photo
Stage 2: Connect your first three systems #
Three systems is the proven starting shape, and they are not any three. You want the system of record for the decision's core entity, the system of state that carries what is happening right now, and the system of constraints that holds the rules, usually contracts, policy documents, or a policy engine. In the banking collections example that is core banking, the collections workflow platform, and the hardship-policy repository. In supplier quality it is the ERP vendor master, the quality management system, and the contract repository. In discharge coordination it is the EHR, the bed and staffing system, and payer authorization rules.
Connect each at the freshness the decision needs, not the freshness the platform brochure advertises. The system of state usually needs events or change-data-capture, because the graph will act on it. The system of record often tolerates hourly sync. The system of constraints changes slowly but must carry effective dates, because policy at decision time is what the Harness will enforce and what an auditor will ask about. Every feed gets an explicit staleness signal so the Decision Layer can tell fresh context from stale.
Of the three, the constraint system is the one teams most often skip, and skipping it quietly caps the graph's value. Without contracts and policy in the model, the graph can describe a situation but not what is permissible in it, which pushes every rule back into application code and out of the Harness's reach. If the constraints live in documents, budget for extraction in this stage: effective-dated clauses linked to the entities they govern is the target shape, and even a partial first pass, the ten clauses your decision actually touches, beats a complete pass that arrives next year.
The exit artifact is a source contract per system: one page listing the entities and fields consumed, the sync method and latency, the identifier the system uses for the core entity, known data-quality issues, and the named owner who will fix breaks. Teams that skip this page rediscover every one of these facts during an incident instead.
- Record, state, and constraint systems flowing into the Context Graph Engine
- A source contract signed per system, with a named break-fix owner
- Freshness matched to the decision: events where the graph will act, batch where it will inform
- Staleness signals emitted per feed and visible to the Decision Layer
- Access provisioned through the Context Harness, not through per-app credentials
Stage 3: Model entities, relationships, and time #
Now the graph itself. Start from the decision brief and list every noun the decision touches; a first domain model lands, as an illustrative range, at 8 to 15 entity types and 15 to 30 relationship types. More than that in week one is a sign the ontology is leading and the decision is following, which is the classic failure. The collections graph needs customer, account, arrangement, hardship flag, policy clause, and interaction; it does not need the full party model on day one.
Four modeling disciplines separate a decision-grade context graph from a diagram. Entity resolution first: decide how the core entity is identified across your three systems, define match rules, and keep every source identifier on the resolved node so lineage survives. Relationships carry meaning: name edges as verbs a domain expert would say aloud, "guarantees", "supplied by", "authorized under", and give every edge validity dates. Time is bitemporal where decisions are audited: record both when something became true and when the graph learned it, because "what did we know when we decided" is the question regulators actually ask. Policy lives in the model: represent constraining rules as first-class nodes with effective dates and link them to what they govern, so the Context Harness enforces from the graph rather than from application code.
Resist two temptations. Do not model for hypothetical future decisions; the second decision will extend the model more cheaply than speculation ever prices in. And do not let attribute completeness block edge coverage: a sparse graph with the right relationships beats a rich one with the wrong shape, because the Decision Layer traverses before it reads. The exit artifact is domain model v1: the entity and relationship catalog with definitions, identity rules, temporal treatment per entity, and the policy nodes mapped, all reviewed aloud with the domain expert until the model reads like their language.
- 8 to 15 entity types, 15 to 30 relationship types, all traceable to the decision brief
- Match rules written and tested for the core entity across all three systems
- Every edge dated; audited entities modeled bitemporally
- Policies represented as nodes with effective dates, enforceable by the Harness
- The domain expert can narrate the model without the architect translating
Stage 4: Validate with real questions #
A context graph is validated by interrogation, not by row counts. Before the graph touches any decision, build a question ledger: 20 to 30 real questions, sourced from the people who make the decision today, that no single source system can answer alone. Then answer them from the graph, in front of those people, and log hits, misses, and surprises.
Good validation questions have a shape: they traverse at least two systems, they involve time or policy, and the domain expert already knows roughly what the answer should be. From the three running examples: in banking, "which customers in an active hardship arrangement received a collections call last week, against policy?"; in manufacturing, "which open purchase orders depend on a supplier whose quality certification lapses within 60 days?"; in healthcare, "which planned discharges this week require a home-care skill we have no credentialed staff for on the scheduled day?". Each is unanswerable from one system, embarrassing to answer manually, and instantly convincing when the graph returns it in seconds.
Structure the ledger as a simple table your team maintains for the life of the graph: the question in the domain expert's words, the systems it spans, the expected answer or its shape, the graph's actual answer, a pass or fail, and the fix if it failed. Aim for a deliberate mix, as an illustrative guide: roughly half traversal questions that walk two or three relationships, a quarter temporal questions that ask about a past state or an approaching deadline, and a quarter policy questions that test whether constraints bind where they should. The ledger is not throwaway QA; it becomes the regression suite you re-run after every model change and every new feed, and the first artifact a new team member reads to understand what the graph is for.
Score the ledger honestly. Misses are gold: each one names a missing edge, a broken match rule, or a stale feed, and fixing them at this stage costs hours instead of incidents. The stage exits when, as an illustrative bar, the graph cleanly answers roughly 80 percent of the ledger and the domain expert has stopped checking its answers against the source systems, because that behavioral shift, not a dashboard, is what trust looks like.
Stage 5: Put the graph to work in the loop #
A validated graph that only answers questions is a successful knowledge graph and an unfinished context graph. Stage 5 closes the loop on the decision from your brief. In Understand, the Context Graph Engine assembles the situation subgraph: the entities, relationships, history, and applicable policy for one case. In Decide, the Decision Layer frames options against constraints, scores confidence, and attaches the evidence trail. In Execute, the Execution Grid carries the approved action through APIs, agents, and human tasks, and writes the outcome back into the graph as decision memory.
Wire the guardrails before the first action, not after. In the Harness, define who and what may see each part of the subgraph, the approval thresholds per action type, the mandate boundary beyond which the Execution Grid must route to a human, and the audit fields every decision must carry: inputs seen, options considered, policy version in force, approver, and outcome. None of this is bureaucracy; it is the evidence base that lets autonomy widen later without a leap of faith.
Run the first weeks with a human approving every action; the point is not autonomy but instrumentation. Track the decision KPIs from your brief alongside graph health, coverage, freshness, resolution accuracy, and let the Context Harness accumulate the audit evidence that later justifies widening autonomy on the low-risk slice of cases. This is also the moment your first domain graph becomes something larger: the identity spine, policy nodes, and decision memory you built are the first cells of the Enterprise Digital Twin, and the second domain will inherit them rather than rebuild them.
Architecture decisions you will face on the way #
Four design debates recur in every first build, and all four have decision-first answers. Storage: the temptation is to open with a graph-database selection; resist it. Specify the capabilities first, entity resolution, relationship and temporal modeling, policy attachment, subgraph APIs, and choose or accept storage after. Inside a Context Engine the storage layer is the platform's concern; in a self-build it is a consequence of the model, never the starting point.
Materialized vs virtual: federated, query-time stitching of source systems demos well and fails at decision time, because latency, source outages, and the absence of history all surface exactly when the Decision Layer needs an answer. Materialize the graph, keep lineage to sources, and let the source contracts govern freshness. Push vs pull: default to events and change-data-capture for the state system and scheduled sync elsewhere, and make every feed's staleness observable rather than assumed. Serving: expose the graph through situation-shaped context APIs that return a decision's subgraph in one call, not through a generic query endpoint that forces every consumer to relearn the model. Agents and the Decision Layer consume situations, not schemas.
One organizational decision belongs on the same list: where the Context Harness policies live. Decide in week 1 that access, masking, and audit are configured in the Harness and inherited by every consumer, and get the governance owner to sign that model early. Teams that defer this decision rediscover it as a blocking finding in the first security review, usually in week 9, when it is most expensive.
The week-by-week build plan #
The plan below compresses the five stages into a single quarter for a scoped domain with cooperative system owners. All durations are illustrative planning ranges: access delays and governance sign-off stretch timelines far more often than modeling does. The minimal team is one architect who owns the model, one data engineer who owns synchronization, one embedded domain expert who owns meaning and the question ledger, and a part-time governance owner for Harness policies.
| Weeks | Stage | Work | Exit artifact |
|---|---|---|---|
| 1 | 1: Pick | Decision selection, baseline measurement, system and owner mapping | Decision brief, signed |
| 2 to 4 | 2: Connect | Access via the Harness, feeds landed at decision-grade freshness, staleness signals wired | Three source contracts |
| 4 to 6 | 3: Model | Entity and relationship catalog, match rules, temporal and policy modeling, expert review | Domain model v1 |
| 6 to 8 | 4: Validate | Question ledger built and answered live; misses triaged into model and feed fixes | Ledger at roughly 80% clean |
| 9 to 12 | 5: Run | Decision Layer and Execution Grid wired for the one decision; humans approve every action; KPIs instrumented | First governed loop, audit trail included |
Weeks 2 to 4 and 9 to 12 carry the real schedule risk. Start access requests and Harness policy sign-off in week 1, in parallel with the decision brief, and the quarter usually holds.
A realistic enterprise scenario #
An industrial manufacturer, roughly 40 plants, chooses supplier-quality disposition as its first decision: when a deviation or audit finding lands, what happens to the affected supplier, parts, and open orders? Today that answer takes a war room and, illustratively, three to five days.
Weeks 1 to 4. The decision brief names the metric: time from finding to disposition. The three systems are the ERP vendor master (record), the quality management system (state), and the contract repository (constraints). The QMS lands as events because dispositions must react same-day; contracts land with effective-dated clauses.
Weeks 4 to 8. The model settles at 11 entity types and 22 relationships; the hard modeling problem is supplier identity, since plants onboarded vendors independently for a decade. Validation questions do their job: "which open POs depend on suppliers with findings open more than 30 days" exposes two broken match rules and one stale feed, all fixed before anything acts.
Weeks 9 to 12. The loop goes live on the narrow decision: the Decision Layer drafts dispositions with evidence, quality engineers approve every one, and the Execution Grid updates the ERP, notifies plants, and opens supplier tasks as one traceable act. By the end of the quarter the graph has decision memory from dozens of dispositions. Illustratively, teams report first-quarter results in the range of 40 to 60 percent faster disposition cycles on the scoped decision; treat that as directional, and note the quieter win: the supplier identity spine the graph built is now the inherited foundation for the second domain, procurement risk.
Common mistakes to avoid #
- Starting from the ontology instead of the decision: enterprise-wide entity models are where first graphs go to die. The decision brief is the scope control.
- Connecting five or more systems in the first build: every added source multiplies identity conflicts before the match rules are proven. Three, then earn the fourth.
- Skipping entity resolution because "we have an MDM": master data covers the record system; your state and constraint systems still disagree with it.
- Validating with demos instead of the question ledger: a graph that impresses in a walkthrough and fails the ledger is not ready to act.
- Leaving time and policy out of v1 to "add later": bitemporal history and policy nodes are structural; retrofitting them means remodeling, not extending.
- Stopping at Stage 4: a validated graph nobody runs decisions through becomes shelfware within quarters, because unused context silently goes stale.
After the first graph: scaling toward the twin #
Two disciplines keep the scale-out honest. First, measure every domain the same way: each new graph carries its own decision brief, baseline, question ledger, and decision KPIs, so the portfolio reports outcomes rather than node counts. A twin program that reports "12 million entities connected" is drifting; one that reports "nine decisions running governed loops across four domains" is compounding. Second, resist the platform team's urge to centralize modeling: domain experts own meaning, the central team owns the identity spine, the modeling patterns, and the Harness, and the boundary between those responsibilities is what lets domain three onboard in weeks rather than months.
Choose the second domain by adjacency, not ambition: the domain that shares the most entities with your first graph inherits its identity spine, its policy patterns, and its Harness setup, which is why second builds run materially faster than firsts. Repeat the five stages each time, keep every domain anchored to a decision, and the connected graphs grow, domain by domain, into the Enterprise Digital Twin, with the Understand-Decide-Execute loop compounding decision memory across all of them.
For the surrounding concepts in the Getting Started path: What is a Context Engine? defines the system this graph lives inside, The Enterprise Digital Twin, Explained shows the architecture your domains grow into, and Context Engine vs Semantic Layer settles how the graph sits beside your analytics stack.
How OpenKnowra approaches this #
The five-stage sequence above is method, not product, and it works on any credible platform. OpenKnowra's contribution is making the stages shorter: the Context Graph Engine ships entity resolution, bitemporal history, and policy nodes as built-in modeling primitives rather than design projects, the Decision Layer and Execution Grid are already wired for Stage 5, and the Context Harness carries access, policy, and audit from the first feed onward, so governance sign-off happens once instead of per application. Domain models start from industry templates that teams prune against their decision brief rather than blank pages.
A first engagement follows this guide literally: your decision brief, your three systems, and a validated graph running one governed loop inside the quarter, question ledger and audit trail included.