A document extraction context graph pipeline is a governed process that parses enterprise files, extracts candidate entities and relationships, resolves them to graph identities, validates them against ontology and policy, and stores accepted facts with confidence, temporal validity, and evidence-level lineage.
Begin with document classes and decisions #
Do not begin by sending every enterprise file to a model. Start with a decision that lacks structured context and identify the document classes that contain the missing facts. A supplier-risk decision may need termination clauses, liability caps, renewal dates, service obligations, and named parties from contracts. A quality decision may need tolerances and escalation rules from SOPs. Each class has different layouts, vocabulary, and risk.
Create a document-class specification: source repository, access policy, expected layouts, target entities and relationships, mandatory evidence anchors, confidence thresholds, review owner, retention rule, and downstream decisions. This artifact prevents the pipeline from becoming an uncontrolled summarization service.
Stage 1: parse structure before meaning #
Reliable extraction begins with document structure. Detect pages, headings, paragraphs, tables, lists, signatures, footnotes, and repeated headers. For scans, OCR must preserve coordinates and confidence. Segment the document into stable units that can be cited later. A clause anchor should survive reprocessing and allow a reviewer to open the exact region that supported a graph fact.
Store the raw file and a normalized representation separately. Hash the source, record version and repository metadata, and detect duplicates or superseded versions before extraction. A graph should not treat an obsolete draft and an executed contract as equally valid evidence.
Stage 2: extract against the ontology #
Define the target schema before prompting or training extraction. Candidate outputs should use ontology concepts such as Party, Agreement, Obligation, Asset, Product, Policy, Location, Event, EffectiveDate, and GoverningJurisdiction. Relationship extraction is often more valuable than isolated entities: supplier is obligated to meet a service level; policy applies to a process; report identifies a control failure.
Use a combination of deterministic parsers, dictionaries, pattern rules, classifiers, and language models. The best method varies by field. Dates and identifiers may be rule-friendly; ambiguous obligations require semantic extraction. Keep model output as candidate assertions until validation and resolution succeed.
Stage 3: validate, resolve, and score evidence #
Resolve extracted parties, assets, products, and locations to existing graph entities. Validate type, cardinality, date logic, and domain constraints. A contract end date before its start date, an obligation without an obligated party, or a currency amount without currency should fail validation. Confidence should combine extraction probability, OCR quality, ontology fit, resolution certainty, and rule checks.
| Evaluation slice | What to measure | Why aggregate accuracy is insufficient | Example acceptance path |
|---|---|---|---|
| Entity extraction | Precision and recall by entity type | Party names and product codes have different difficulty | Auto-accept known IDs; review ambiguous names |
| Relationship extraction | Correct subject, predicate, object | All entities can be right while the relationship is wrong | Require ontology and direction validation |
| Dates and amounts | Value, unit, qualifier, context | A correct number with the wrong clause meaning is unsafe | Use deterministic checks plus evidence review |
| Obligations and exceptions | Completeness and scope | Missing an exception may reverse the decision | Human review for high-impact clauses |
| Evidence grounding | Anchor correctness | Ungrounded facts cannot be audited | Reject facts without stable source anchors |
Stage 4: write facts, not summaries #
The graph should receive typed facts and relationships with evidence, confidence, valid time, and source lineage. A summary can remain useful for human orientation, but it should not substitute for explicit assertions. Store whether a fact was extracted, inferred, reviewed, or sourced directly. Policy can then treat each class differently.
Three industry patterns
Insurance. Policies and endorsements yield coverage, exclusions, limits, and effective dates connected to customer and asset context. Manufacturing. SOPs and quality reports yield process requirements, control limits, deviations, and responsible roles. Pharmaceuticals. study reports, procedures, and regulatory correspondence yield obligations, evidence, sites, products, and deadlines under strict review.
A realistic enterprise scenario #
A global manufacturer needs to understand supplier exposure hidden in thousands of contracts, quality reports, and corrective-action documents. Procurement systems identify suppliers and spend but do not encode termination rights, alternate-source clauses, quality obligations, or site-specific exceptions.
Understand. The pipeline classifies documents, parses clauses and tables, extracts suppliers, facilities, products, obligations, dates, and exceptions, and resolves them to the existing supplier and product graph. High-impact obligations require review; every accepted assertion retains a clause anchor and model version.
Decide. The Decision Layer combines extracted obligations with live inventory, product dependencies, incidents, and geographic risk. It can explain that a product is exposed because one supplier serves two plants, the alternate-source clause is absent, and a corrective action is overdue.
Execute. Approved remediation actions flow through the Execution Grid to procurement and quality systems. When a reviewer corrects an extracted clause, the graph preserves the prior assertion, updates current context, and adds the correction to the evaluation set for future model improvement.
Common mistakes to avoid #
- Treating document summarization as equivalent to extracting governed facts.
- Failing to preserve page, clause, table cell, or image-region evidence anchors.
- Using one aggregate accuracy score across all fields and document classes.
- Writing model output directly to the graph before ontology validation and entity resolution.
- Ignoring version, execution status, and supersession of source documents.
- Using human review as an untracked manual step instead of a versioned correction workflow.
How OpenKnowra approaches this #
The pipeline above is platform-neutral. OpenKnowra implements document intelligence inside the Context Graph Engine, combining layout-aware ingestion, evidence anchoring, ontology-constrained extraction, entity resolution, confidence, temporal lineage, and review workflows. The Context Harness controls which documents and extracted facts may be accessed or used; the Decision Layer combines them with structured context; and the Execution Grid records approved outcomes.
The recommended first deployment uses one document class, one decision, a limited target ontology, and explicit field-level acceptance criteria. That boundary produces a trustworthy loop before the organization expands to more files.