How to Resolve Entities Without Corrupting a Knowledge Graph
Resolve entities by keeping source records intact and treating identity as a supported, reversible decision. Use reliable identifiers first, generate a manageable set of candidates, compare names and context, record positive and negative evidence, and allow three outcomes: same, different, or unresolved. A canonical entity should connect records; it should not erase them.
By Sahil Maheshwari
The short answer: identity is a decision, not a string
Entity resolution asks whether two records refer to the same real-world person, organisation, product, place, or event. It is also called record linkage, deduplication, or entity matching. The hard part is not finding similar text. It is deciding when similarity means identity and when it does not.
Two records can describe the same entity with different names: a company may rebrand, a person may use initials, and a policy may appear under a marketing name and a regulatory product code. The reverse is equally common. Two people can share a name, a group can share a phone number, and a parent company can resemble its subsidiary. Normalising case and punctuation helps comparison, but it cannot settle identity.
This matters more in a graph than in a flat search index. Merge two records incorrectly and their addresses, roles, documents, claims, and relationships appear to belong to one thing. If that merged node feeds retrieval, an answer can combine evidence from different people or products while every individual statement still looks plausible.
Start with an explicit policy for each entity type. Define what counts as the same entity, which identifiers are authoritative, which contradictions rule a match out, who reviews ambiguous cases, and what the cost of a false merge is. A legal entity, a trading brand, and an office branch may be related, but they are not interchangeable identities.
A false merge does not create one bad field. It creates a bridge that can carry unrelated evidence across the graph.
Preserve records before creating canonical entities
Give every source record a stable identity based on its source system and source key. Keep its original values, source, version, retrieval time, and effective dates. Then create a separate canonical entity and connect the record to it with a resolution decision. This keeps the raw observation distinct from your interpretation of what it represents.
RDF 1.1 Concepts defines IRIs as identifiers used in RDF triples, but an identifier only names a resource within a chosen data model. It does not prove that two IRIs identify the same thing. OWL 2 makes the consequence explicit: `SameIndividual` means the named individuals are equal and can substitute for each other in expressions. That is much stronger than “these records look related.”
Avoid asserting `owl:sameAs` for fuzzy or provisional matches. Use a scoped relationship such as `possiblyRepresents`, `sameAsInSource`, or a dedicated resolution record until the evidence satisfies your policy. Keep subsidiaries, previous organisations, product variants, and household members linked with the correct relationship rather than forcing them into one node.
External identifiers are often the strongest starting point. Wikidata's external identifier guidance treats them as links to records in authority files and other databases. Store the issuing system and identifier type with the value. An identical-looking number from two namespaces is not automatically the same key, and even authoritative systems can recycle, retire, or correct identifiers.
- Source record: immutable source key, original fields, source version, and dates.
- Canonical entity: a workspace identity with type and lifecycle, not a copied row.
- Resolution decision: record pair, outcome, evidence, method, reviewer, and time.
- Relationship: a precise link such as subsidiary of, renamed from, branch of, or possible match.
Generate candidates without comparing everything
Comparing every record with every other record grows quadratically. Candidate generation, often called blocking, narrows the work to pairs that have a plausible reason to match. The aim is not to make the final decision. It is to retain nearly all true matches while removing obviously unrelated pairs.
Begin with deterministic routes: exact identifiers from the same issuing authority, verified email addresses, registration numbers, product codes, or source-maintained crosswalks. Add broader blocks for records without dependable keys: normalised name plus country, phonetic surname plus birth year, domain plus organisation type, or product family plus issuer. Use several overlapping routes because any single field can be missing or wrong.
Candidate generation should respect time and type. A company name reused twenty years later may not identify the same legal entity. A broker branch should not compete with an insurer product. Transliteration, aliases, reordered names, and changed addresses need separate handling rather than one universal similarity function.
Log why each candidate pair was created. That makes recall failures diagnosable: if a true match never reached the matcher, tuning the scoring model cannot recover it. The Swoosh work on generic entity resolution also shows why comparison and merging cannot be treated as unrelated operations; a merged record may expose new evidence and create new possible matches. In production, that feedback loop needs limits and an audit trail.
Compare evidence, including reasons not to merge
For each candidate, calculate interpretable features: exact identifier agreement, name similarity, compatible addresses, overlapping active periods, shared domains, common relationships, and source authority. Keep contradictions beside supporting signals. Different verified registration numbers, incompatible dates of birth, mutually exclusive jurisdictions, or impossible timelines may outweigh a very similar name.
Separate hard rules from learned scores. A hard rule can declare two records different when a trusted unique identifier conflicts. A model can rank cases where evidence is incomplete. Ditto casts entity matching as classification over pairs of records using pretrained language models and allows domain knowledge to highlight important fields. That is useful for messy descriptions, but the model still learns from labelled examples and can inherit their blind spots.
Do not turn one similarity score into an unexplained verdict. Store the feature values, model and version, threshold, source fields used, and a concise reason. Calibrate thresholds by entity type and consequence. A research reading list may tolerate duplicate authors; a claims workflow should be far more cautious about joining people, policies, or beneficiaries.
Graph context can help. Shared directors, addresses, citations, or product families may support a match when names differ. It can also create circular confidence: two uncertain records look similar because they are both connected to a third uncertain merge. Distinguish independently observed relationships from relationships inherited through resolution, and prevent inferred context from counting as fresh evidence for itself.
Positive evidence asks why two records may match. Negative evidence asks what would make that match impossible. Keep both.
Use three outcomes and make merges reversible
A reliable resolver supports same, different, and unresolved. Forcing every candidate into a binary answer converts uncertainty into graph structure. Marking a reviewed non-match also has value: it prevents the same tempting false pair from returning after every reindex.
When a match is accepted, add both records to a canonical cluster; do not overwrite one with the other. Choose display values through a separate survivorship policy based on source authority, recency, completeness, and scope. Preserve alternative names and historical values with their provenance. The canonical view can change without rewriting the records that justified it.
Make every decision append-only or versioned. A split should restore the earlier clusters, relationships, and downstream indexes without reconstructing history from logs. Record which derived claims, embeddings, summaries, and permissions depended on the old cluster so they can be invalidated. High-impact merges should enter a review queue, especially when evidence barely clears a threshold or when the resulting cluster becomes unusually large.
Be careful with transitivity. If A matches B and B matches C, a graph may infer that A matches C. OWL equality has exactly that kind of substitutive force. In operational data, pairwise model scores are not necessarily transitive. Before closing a cluster, check every member against cluster-level constraints and surface the weakest or most contradictory links.
Evaluate entity resolution at pair and cluster level
Build a labelled test set from real edge cases, not only easy exact matches. Include common names, shared contact details, renamed organisations, subsidiaries, branches, transliteration, identifier corrections, sparse records, and entities that change over time. Split training and evaluation so near-duplicate examples from one cluster do not leak across both sets.
Measure candidate recall first: what share of true matches entered the candidate set? Then measure match precision and recall. Precision matters when false merges are expensive; recall matters when duplicate entities hide connected evidence. Pair metrics are not enough because one bad bridge can combine two large clusters. Add cluster-level measures, the number and size of false merges, false splits, unresolved review volume, and the cost of undoing decisions.
The OAEI Knowledge Graph Track evaluates instance alignments with precision, recall, and F-measure, and its results show substantial trade-offs between matchers. Its gold standard also relies on stated assumptions, including one representation per concept within each graph. Your evaluation should document comparable assumptions instead of presenting one score as universal truth.
Monitor drift after launch. New sources, naming conventions, countries, or product lines change candidate distributions. Track decisions by rule and model version, reviewer overturn rates, cluster growth, identifier conflicts, and downstream incidents. Review samples from confident matches as well as uncertain ones; a miscalibrated system can be confidently wrong.
You may not need probabilistic resolution at all when one authoritative system already assigns durable IDs and every source carries them correctly. In that case, preserve the mapping and test identifier quality. Use complex matching only where records are genuinely incomplete, inconsistent, or disconnected.
- Candidate recall: did the blocking stage retain the matches you needed?
- Pair quality: how many accepted and missed pairs were correct?
- Cluster quality: did one false edge combine distinct real entities?
- Operational quality: can reviewers explain, reverse, and propagate a correction?
An operating checklist for a trustworthy resolver
Treat entity resolution as a maintained knowledge process. Owners need a queue for ambiguous pairs, a way to inspect evidence in context, and a clear route to correct the graph. Data pipelines must re-evaluate affected candidates when identifiers, source records, or policies change. Consumers should know whether they are viewing original records, a canonical entity, or a provisional match.
Before a merge reaches the graph, ask whether the records have compatible types and timelines, which independent evidence supports identity, which evidence contradicts it, whether a stronger identifier is available, and what downstream material will inherit the decision. Afterward, retain the decision record and schedule re-evaluation when its inputs change.
If you have a real entity-resolution problem—such as duplicated organisations, policy products, experts, or research authors—share an anonymised example and the decision that currently gets stuck at sahil@granveo.com. The most useful design usually starts with the false merge you cannot afford.
Sources and further reading
- 1RDF 1.1 Concepts and Abstract Syntax
W3C — Defines RDF identifiers, triples, graphs, and the distinction between naming a resource and proving identity across records.
- 2OWL 2 Web Ontology Language Structural Specification and Functional-Style Syntax
W3C — Defines SameIndividual and DifferentIndividuals semantics, including the substitutive consequences of asserting identity.
- 3Wikidata: External identifiers
Wikidata — Documents how external identifiers link Wikidata items to authority files and records in other databases.
- 4Swoosh: A Generic Approach to Entity Resolution
Stanford University — Formalises entity resolution around comparison and merge functions and explains how merges can expose further candidate matches.
- 5Deep Entity Matching with Pre-Trained Language Models
Li et al., VLDB 2020 — Introduces Ditto, which classifies record pairs with pretrained language models and domain-aware representation techniques.
- 6Results for OAEI 2024 – Knowledge Graph Track
Ontology Alignment Evaluation Initiative — Reports precision, recall, F-measure, assumptions, and matcher trade-offs for schema and instance alignment across knowledge graphs.
Continue the conversation
Where does context get lost in your work?
I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.