All field notes
Knowledge graphs10 min read

How to Build a Knowledge Graph That Preserves Evidence

Build an evidence-preserving knowledge graph by treating each claim as a first-class record, not an unquestioned edge. Keep the proposition separate from the assertion that it holds, then attach the source, exact evidence location, extraction or review activity, responsible agent, time, scope, version, and status. Conflicting assertions can coexist; applications decide which ones are eligible for a particular question.

By Sahil Maheshwari

Two sources making scoped claims about the same relationship, with derivation, version, and review paths preserved in a knowledge graphTRACE A CLAIMSOURCE 03WORKING CLAIMContext is more usefulwhen its origin survives.3 sources · 2 relationships · 1 open question

The short answer: model claims before facts

A simple knowledge graph might store an edge such as `Policy A — excludes → Flood`. That is easy to query, but it hides the questions a reviewer will eventually ask. Which version of the policy? Was the edge extracted from an issued document or inferred from a summary? Does the exclusion apply in every jurisdiction? Was it later superseded?

Promote that relationship into an addressable claim. The claim points to its subject, predicate, and object, while also carrying links to evidence and context. The graph may still expose a convenient direct edge for ordinary queries, but that edge should be generated from accepted claims rather than becoming the only record of truth.

This distinction is useful outside RDF. A property graph can use a Claim node, a relational system can use an assertions table, and a JSON document can give each assertion a stable ID. The important design choice is that provenance belongs to the claim at the same granularity at which the system answers questions.

Begin with the questions the graph must support: “Why do we believe this?”, “Who or what asserted it?”, “Which source span supports it?”, “When and where does it apply?”, and “What changed?” If the schema cannot answer those questions without opening an unrelated log system, provenance will disappear from the workflow.

An edge is a useful view. The assertion, evidence, and derivation behind it are the durable record.

Separate the proposition from the assertion

A proposition describes a possible relationship: Policy A excludes flood. An assertion says that a particular source, person, or process claims that the proposition holds. This separation matters when the proposition is uncertain, disputed, historical, or merely being evaluated. Recording a proposition should not automatically promote it to accepted fact.

RDF 1.2 Concepts introduces triple terms and reifiers for this pattern. A triple term can denote a proposition without asserting that it is true. A reifier can denote a statement or belief related to that proposition and carry further description. As of September 2026, the document is a W3C Candidate Recommendation, so teams adopting its syntax should track implementation support rather than treating it as a settled deployment requirement.

You can implement the same logic with a small claim record: claim ID, subject ID, relation type, object or literal value, asserting agent, assertion time, status, and links to evidence. Keep entity identity separate from source wording. “Professional indemnity policy,” “PI cover,” and a product code may resolve to one entity while the original phrase remains attached to the evidence span.

Give each assertion its own identity even when two sources state the same proposition. Their dates, authority, methods, and scopes may differ. Deduplicate source documents and entities where appropriate, but do not collapse independent acts of assertion before you have compared them.

Attach evidence at the claim level

A document-level citation is often too broad. Store the smallest stable locator that lets a person inspect support for the claim: page and bounding box, section path and paragraph, table and cell range, transcript timestamp, database record and field, or API object and version. Preserve the exact quoted span where licensing and retention rules allow it, plus a checksum so later changes are detectable.

The source record should identify the work itself, not only the URL used during ingestion. Capture a canonical URI or document ID, title, publisher or owner, publication time, effective interval, retrieved time, version, and content hash. A redirecting web address is not a durable source identity; neither is the day a PDF entered the index.

Wikidata offers a practical public example. Its statement model adds qualifiers, references, and ranks to property-value claims. Its source guidance ties references to the specific statement they support. The model is not a universal schema, but it demonstrates why a source belongs beside the claim rather than in a general bibliography attached to an entity.

One source can support several claims, and one claim can have several evidence records. Keep those as explicit many-to-many relationships. Label whether a source directly states the claim, provides a primary observation, supports an inference, or merely repeats another source. Ten mirrors of one press release are not ten independent pieces of evidence.

  • Source identity: canonical ID, version, owner, dates, and content hash.
  • Evidence locator: exact page, section, span, cell, timestamp, or record field.
  • Support role: direct statement, primary observation, inference input, or quotation.
  • Access context: permission scope, retention rule, and whether content may be displayed.

Record how each claim entered the graph

Evidence says where a claim came from. Derivation says what happened between the source and the graph. Was the relationship copied from a structured field, extracted by a model, inferred by a rule, normalised by an entity resolver, or accepted by a human reviewer? Store that activity as its own record with inputs, outputs, time, method or software version, and responsible agent.

The W3C PROV Data Model defines provenance as a record of the entities, activities, and agents involved in producing or influencing a thing. PROV-O maps that model into an ontology with relationships such as `wasDerivedFrom`, `wasGeneratedBy`, `used`, and `wasAttributedTo`. Its qualified relations allow more detail when a binary link is not enough.

A practical chain might say: evidence span E came from document version D; extraction run X used E and model configuration M; X generated assertion A; reviewer R accepted A under review policy P. If the parser, model, or policy changes, the graph can identify which assertions need reprocessing instead of rebuilding everything blindly.

Keep source authorship separate from graph authorship. The author of a policy did not create your extracted assertion, and the model that extracted it did not author the policy. This distinction prevents a common attribution error and makes responsibility legible when a claim is challenged.

Store the source of the claim and the process that interpreted it as two different provenance paths.

Give claims scope, versions, and status

Many apparent contradictions are really differences in time, jurisdiction, product, customer segment, or definition. Model those qualifiers on the assertion. For a policy claim, that may include effective-from and effective-to dates, territory, form version, endorsement, coverage section, and applicable conditions. Unknown scope should remain unknown; a missing value must not mean global applicability.

Never overwrite a changed assertion in place. Create a new assertion and link it to the earlier one as a revision, correction, or replacement. PROV-DM treats revision as a kind of derivation from a preceding entity. Keep both records so a historical question can use the evidence that was valid at that time.

Use a small lifecycle vocabulary such as proposed, extracted, reviewed, accepted, disputed, superseded, and rejected. Status describes how your system handles an assertion; it does not prove the proposition. Confidence should also be typed. Extraction confidence, source authority, reviewer confidence, and evidential strength answer different questions and should not be averaged into one impressive-looking number.

Decide which claims produce the convenient direct-edge view. A production query may use only accepted, in-scope, current assertions from permitted sources. A research view may include disputed and historical assertions. Both should be derivable from the same underlying claim records.

  • Scope: time, place, jurisdiction, product, population, and stated conditions.
  • Version: stable assertion ID plus explicit revision or replacement links.
  • Status: workflow state with actor, timestamp, and reason for every transition.
  • Confidence: separate extraction quality, source quality, and review judgement.

Preserve disagreement instead of forcing consensus

When two sources disagree, store two assertions connected to their respective evidence. First check whether their propositions are truly incompatible after scope and time are considered. If they remain in conflict, add an explicit conflict relation or conflict set and record the rule used by each consuming application.

Do not resolve conflict by counting sources. Several articles may repeat one original claim, while one primary record provides stronger evidence. Resolution policy can consider source type, authority, effective date, directness, independence, and review status. It should be visible and domain-specific; the graph itself should not quietly delete the losing assertion.

The nanopublication approach offers another useful pattern. Nanopublications: A Growing Resource of Provenance-Centric Scientific Linked Data describes atomic assertions packaged with assertion provenance and publication information. Its scientific Linked Data setting will not fit every organisation, but the separation between what is claimed, how the claim arose, and how the record was published transfers well.

An answer layer can now say, “Source A states X for the 2025 form; Source B states Y for the revised 2026 form,” and cite both paths. That is more useful than flattening them into one timeless edge or asking a language model to decide without an explicit policy.

Validate the graph and test the evidence path

Define machine-checkable constraints for every claim type. A reviewed assertion may require a subject, predicate, object, evidence locator, source version, derivation activity, asserting agent, timestamp, status, and applicable scope fields. Reject or quarantine incomplete records rather than filling them with plausible defaults.

SHACL is a W3C Recommendation for expressing constraints over RDF graphs and producing validation reports. A property-graph or relational implementation can enforce equivalent rules with application schemas and database constraints. The technology matters less than making provenance completeness testable at ingestion and after migrations.

Build evaluation questions that follow the evidence path: show the source for a claim, retrieve the claim valid on a past date, compare two versions, identify every assertion produced by a retired extractor, disclose unresolved conflicts, and remove everything derived from a deleted source. Test permissions at the evidence, assertion, and generated-edge layers; a safe answer must not reveal a restricted claim through a public relationship.

Track provenance coverage, broken locators, orphaned assertions, claims without derivation, conflicting current assertions, revision-chain integrity, review latency, and deletion completeness. Sample the evidence itself. A graph can be structurally perfect while pointing to passages that do not support the relationship.

Do not capture unlimited lineage by default. Full token-level traces and every intermediate model thought are expensive, unstable, and often inappropriate to retain. Store the minimum record needed to reproduce, review, update, and govern the claim. Add finer detail only for a concrete decision or risk.

If your knowledge graph contains useful relationships but cannot explain why anyone should trust them, share one real claim and its current evidence trail at sahil@granveo.com. The missing provenance fields will usually reveal the next schema decision.

  • Can a reviewer open the exact evidence for any answerable claim?
  • Can the graph distinguish source authorship from extraction and review responsibility?
  • Can it reproduce the accepted view for a historical date and scope?
  • Can it retain disagreement without treating every assertion as equally authoritative?
  • Can it retract all derived views when a source or extraction run is invalidated?

Sources and further reading

  1. 1
    PROV-O: The PROV Ontology

    W3C — Defines an RDF vocabulary for entities, activities, agents, derivation, generation, attribution, and qualified provenance relationships.

  2. 2
    PROV-DM: The PROV Data Model

    W3C — Provides the conceptual model for provenance, including derivations, agents, activities, bundles, revisions, and responsibility.

  3. 3
    RDF 1.2 Concepts and Abstract Data Model

    W3C — Introduces triple terms and reifiers that can describe asserted or unasserted propositions; currently a Candidate Recommendation.

  4. 4
    Help:Statements

    Wikidata — Documents Wikidata's practical statement model with property-value claims, qualifiers, references, and ranks.

  5. 5
    Help:Sources

    Wikidata — Explains how references connect verifiable sources to individual Wikidata statements.

  6. 6
    Nanopublications: A Growing Resource of Provenance-Centric Scientific Linked Data

    Kuhn et al., PeerJ Computer Science 2018 — Describes atomic assertions packaged separately with assertion provenance and publication information.

  7. 7
    Shapes Constraint Language (SHACL)

    W3C — Defines machine-readable shapes, constraint components, validation results, and conformance reports for RDF graphs.

Continue the conversation

Where does context get lost in your work?

I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.