Structured unstructured context unification is the practice of representing tabular records and text-borne knowledge in one context graph: structured sources populate entities, relationships, and state directly, while a Context Graph Engine extracts entities and facts from documents, logs, and conversations, resolves them against the graph, and anchors each passage to the entities it concerns. The unified Enterprise Digital Twin lets AI reason over numbers and narrative together under one governance model.
Structured unstructured context: healing the oldest split in enterprise data #
The structured unstructured context problem is older than AI and was survivable until AI arrived. Enterprises organized their knowledge into two estates: the structured estate of tables, rows, and keys, where amounts, dates, and statuses live, and the unstructured estate of contracts, emails, reports, tickets, and calls, where the reasons live. Analytics grew up in the first estate; content management grew up in the second; and for decades the border between them was policed only by humans, who read documents and typed the relevant bits into systems.
AI raised the stakes because language models can finally read the second estate at scale, and enterprises rushed to let them, mostly by embedding documents into vector indexes. But retrieval over an isolated document estate reproduces the original split in new form: the model that reads the contract cannot see the account balance, and the model that queries the warehouse cannot see the clause that explains the balance. Real decisions cross the border constantly. Whether to release a shipment depends on an order status (table) and a penalty clause (document). Whether to escalate a customer depends on revenue (table) and the tone of the last three tickets (text). Unification is not a data hygiene nicety; it is the precondition for AI reasoning about situations rather than fragments. Figure 1 shows how the two estates converge into one graph.
How unification actually works #
Structured sources are the easy half, at least conceptually. Rows map to graph structure: master records become entities, foreign keys become relationships, transactional attributes become state with freshness stamps, as the Context Graph Engine resolves identities across systems so the customer in billing and the customer in CRM become one node. This is the graph's skeleton, and its own difficulties, resolution above all, are treated in Entity Resolution Across Enterprise Systems.
The unstructured half requires two distinct operations, and conflating them is the commonest design error. Extraction pulls graph-shaped knowledge out of text: the parties, dates, amounts, and obligations in a contract; the entities and complaints in a ticket; the commitments in an email thread. Extracted facts are then resolved against the existing graph, so the "Meridian Group" in the clause becomes the same node as customer 48291 in billing. Anchoring preserves what extraction cannot: the passage itself, attached to the entities it concerns, with document version and effective dates. Extraction gives the graph facts; anchoring gives every fact its evidence and every entity its narrative. A payment-term fact in the twin points to the clause that set it; when the Decision Layer cites the term, the clause travels with it. Conversations and logs follow the same pattern at higher volume and lower ceremony: entities and signals extracted, threads anchored, sentiment and intent attached as state.
Source type handling matrix #
| Source type | What it contributes | Handling | Graph representation | Watch out for |
|---|---|---|---|---|
| Transactional tables | Amounts, statuses, events | Direct mapping via CDC | State on entities, with freshness stamps | Keys that encode identity assumptions the graph must not inherit |
| Master data | Core entities and hierarchies | Direct mapping plus resolution | Entities and structural relationships | Competing masters; survivorship rules needed |
| Contracts and policies | Obligations, terms, permissions | Extraction + anchoring, version-aware | Facts and policy objects linked to source clauses | Superseded versions polluting the graph as current |
| Emails and correspondence | Commitments, escalations, intent | Extraction + thread anchoring | Events and signals on entity relationships | Privacy boundaries; Harness rules before ingestion |
| Tickets and cases | Issues, resolutions, sentiment | Extraction + anchoring | Events and state on customer and asset entities | Free-text entity mentions needing resolution |
| Call transcripts | Verbal commitments, tone, churn signals | Transcription, extraction, anchoring | Signals as state; passages as evidence | Consent and retention policy per jurisdiction |
| Operational logs | What happened, in order | Streaming, pattern extraction | Events and history on entities | Volume; extract signals, anchor selectively |
The final column is where CDO attention earns its pay: the recurring risks are versioning, identity, and policy, not parsing. Text parsing is a solved commodity; deciding which version is in force, which entity a mention means, and which passages the Harness may show whom is the governed work.
Unification in three industries #
Insurance. Claims sit in tables; the reasons claims succeed or fail sit in adjuster notes, medical reports, and policy wordings. Unified, the twin links every claim entity to its coverage clauses and its correspondence, so an adjudication agent can check the numeric reserve against the verbal commitment an adjuster made last month, and cite both.
Telecommunications. Network inventory and billing are structured; the enterprise customer's real state lives in account team emails and support tickets. Extraction turns complaint patterns into churn-signal state on the customer entity; anchoring keeps the threads attached, so a retention decision reads the revenue and the frustration in one situation.
Energy. Asset registries and sensor telemetry are structured; decades of inspection reports and incident narratives are not. Extracting findings into the graph and anchoring the reports gives a maintenance agent the pattern history the tables never captured, that this valve type fails after cold snaps, with the source paragraphs one click away.
A realistic enterprise scenario #
A CDO at a commercial lender inherits two proud, separate estates: a governed warehouse the analytics team trusts, and a document platform holding half a million credit agreements, amendments, and covenant letters. The AI program straddles them badly: the covenant-monitoring pilot reads documents but not balances, the early-warning model reads balances but not amendment letters, and each misses breaches the other would catch.
Understand. The team scopes unification to the lending domain. Structured sources map borrowers, facilities, and exposures into the twin; the Context Graph Engine extracts parties, covenants, thresholds, and effective dates from agreements and amendments, resolves them to the same borrower entities, and anchors every clause with its version state.
Decide. Covenant monitoring becomes one Decision Layer task instead of two half-blind ones: current financials from state, thresholds from extracted covenant facts, and the governing clause attached as evidence, with the Context Harness applying the same visibility rules to clauses as to balances.
Execute. The Execution Grid raises breach alerts and drafts notices, writing outcomes back as history. As an illustrative range, lenders unifying the two estates for covenant work typically report catching breach conditions weeks earlier than either estate caught alone, with review time per alert falling because the evidence arrives pre-assembled.
Common mistakes to avoid #
- Embedding instead of unifying: a vector index over documents is search, not unification; without extraction and anchoring, text knowledge never meets table knowledge, the limitation detailed in Context vs Vector Database.
- Extracting without resolving: facts pulled from text but not matched to graph entities create a parallel shadow graph of unresolved mentions.
- Anchoring without versions: a clause anchored to an entity is misleading unless the graph knows whether it is in force, superseded, or draft.
- Ingesting correspondence before policy: emails and transcripts carry personal data; Harness rules on visibility and retention come before connectors, not after.
- Treating logs like documents: anchoring every log line drowns the graph; extract signals and events, anchor selectively.
- Governing the halves differently: if documents escape the access rules that columns obey, the unified graph's weakest policy becomes its effective policy.
The CDO's unification sequence #
Unify by decision domain, not by estate. Pick one domain where decisions visibly straddle the border, lending, claims, retention, and build its slice: structured skeleton first, then extraction and anchoring for the two or three document types that matter most, with version handling and Harness policy from day one. Measure the unification by what it catches that either estate alone missed, because that delta is the business case in its purest form. Then repeat, domain by domain, letting the resolved entity backbone compound across them.
For the components the two estates populate, see The Anatomy of Enterprise Context; for where the unified layer sits among warehouses and lakehouses, see The Context Layer in the Modern Data Stack; and for the machinery performing extraction and anchoring, The Context Graph Engine, Explained, all in the Fundamentals cluster.
How OpenKnowra approaches this #
The unification pattern above, direct mapping for structure, extraction plus anchoring for text, one governance model over both, is implementable on any modern stack. OpenKnowra ships it as the default: the Context Graph Engine maps structured sources, extracts and resolves entities and facts from documents, tickets, and transcripts, and anchors every passage with version state into the Enterprise Digital Twin. The Context Harness governs clauses and columns identically, and the Decision Layer cites numeric facts and their textual evidence in the same trail before the Execution Grid acts.
A concrete starting exercise: bring one decision that requires both a table and a document to make correctly, and we will build its unified slice, so your team sees the anchored evidence pattern on your own data before committing to a domain.