Back to blog
Deduplicate Entities in an Agent Memory Knowledge Graph

This blog is written by AI for SEO

Deduplicate Entities in an Agent Memory Knowledge Graph

HelixDB10 min read

Every developer building memory for autonomous agents runs into the duplicate entity trap. An LLM ingests an email mentioning Sam Lee, parses a Slack thread citing Sam from engineering, and indexes a calendar invite with samuel@acme.com. Left unchecked, the database creates three separate nodes for one human being. When the agent later tries to answer a simple question about Sam's projects, retrieval splits across three disconnected islands of context.

Prompting the LLM to check for duplicates before writing won't fix this. Large language models are nondeterministic, and nobody can afford full-scan context comparisons on every conversational turn. You need a deterministic entity resolution pipeline. This guide walks through how to deduplicate entities in an agent memory knowledge graph across six steps and a testing phase. By the end, you'll have a pipeline that blocks candidates cheaply, scores them across multiple signals, rewires edges safely, preserves provenance, and executes merges atomically.

Step 1: Block candidates before you compare anything (normalized names, exact emails via BM25, type filters)

Pairwise entity comparison across an entire graph is an O(n^2) disaster. If your agent memory contains 50,000 entities, evaluating every candidate pair takes over 1.2 billion comparisons. Blocking cuts that search space to a few dozen likely candidates before any expensive scorer runs.

Start with strict entity type filtering. A company node labeled Acme Corp should never be compared against an individual human labeled Acme Corp. Type guards kill cross-domain noise immediately.

Next, generate blocking keys using string normalization and inverted indexes. For person and organization entities, lowercase the raw display name, strip punctuation, remove corporate suffixes like Inc or LLC, and sort tokens alphabetically. The string Lee, Sam produces the blocking token sam lee, which catches simple inverted mentions. For high-entropy identifiers like email addresses, domain handles, or GitHub usernames, run exact BM25 text index queries. If an incoming memory payload references samuel@acme.com, query the BM25 index on the identity property and pull every entity that shares that identifier or domain.

In our pipeline, any candidate that shares an exact email address, a common blocking token, or an identical normalized slug qualifies for the candidate pool. Everything else is ignored. Blocking keeps the candidate list under 50 records per incoming entity, which turns a multi-second bottleneck into a sub-millisecond property scan.

Step 2: Embed the name plus its context, not the name alone (and scope the vector search to the entity's graph neighbourhood with HelixDB pre-filtering)

Embedding a bare entity name creates false matches. The vector for the string Alex carries almost no distinguishing information. It clusters near every other Alex in the embedding space, whether that person is your lead backend engineer or a fictional character in a support ticket. High-precision resolution means embedding identity plus local context.

When your agent mints an entity mention, generate an entity context passage before embedding. A typical format looks like this: Alex: Staff Platform Engineer at Acme Corp; works on Kafka pipelines; mentioned in thread #infra-deployments. Embedding this summary produces a vector that clusters near infrastructure professionals, not marketing coordinators who also happen to be named Alex.

Candidate search fails when you run a top-k vector search against the whole database. A global search compares the incoming node against hundreds of thousands of unrelated entities across irrelevant tenants or subgraphs. Scope the similarity query to the entity's graph neighborhood instead, using vector search on graph edges.

HelixDB provides native vector pre-filtering in the same engine. You don't search a standalone vector database and intersect the results with a graph database in application glue code. HelixDB runs approximate nearest neighbor search directly over candidates connected by existing relationships. Vectors in HelixDB normalize to float32 and live directly on nodes and edges. You can traverse outward two hops from the relevant project node, filter to person entities, and run cosine distance ranking only across those related entities.

Step 3: Score candidates with several signals and pick a merge threshold (name, email, type, shared neighbours, embedding similarity; merge / review / keep separate)

Never merge two entities on vector similarity alone. Embeddings propose candidates, and deterministic structural evidence decides identity. A production deduplication pipeline uses an ensemble score built from five signals.

First, calculate string similarity on normalized names using Jaro-Winkler or Levenshtein distance. Jaro-Winkler works well for person names because it penalizes prefixes less heavily, so Sam can match Samuel. Second, apply an exact match weight to high-cardinality keys like emails, user IDs, or tax numbers. Third, verify type parity. Fourth, compute neighborhood overlap using the Jaccard similarity of shared 1-hop neighbors. If Candidate A and Candidate B both report to the same manager and belong to the same Slack channel, their neighborhood score spikes. Fifth, include the cosine similarity of the contextual vector embeddings from Step 2.

Combine these weights into a composite confidence score between 0.0 and 1.0. An effective heuristic configuration looks like this:

Composite Score = (0.25 * NameSim) + (0.35 * ExactKeyMatch) + (0.20 * JaccardNeighbors) + (0.20 * VectorSim)

Define three operational ranges based on this score:

Auto-Merge (Score >= 0.88): The evidence is clear. The entities are the same concept, and the background worker can fuse them immediately.

Human or Agent Review (0.65 <= Score < 0.88): Ambiguous matches, like Alex Rivera and Alex Ruiz sharing an office location. Flag these pairs for an asynchronous resolution queue where an agent can ask the user for clarification during conversation.

Keep Separate (Score < 0.65): Reject the candidate match. The incoming mention becomes a distinct new entity in your agent memory knowledge graph. This three-range threshold stops false positives from collapsing distinct concepts into corrupted supernodes.

Step 4: Write a merge rule: choose a canonical node, fold aliases and properties, and rewire edges

Once an entity pair passes your merge threshold, run the merge rule. Merging is not just deleting one node. You need a deterministic procedure to select a canonical target, preserve historical identity strings, reconcile property conflicts, and rewire graph connections without losing semantic directionality.

First, choose the canonical entity. Pick the node with the highest degree centrality, the oldest creation timestamp, or the most complete metadata. The surviving entity keeps its primary UUID.

Second, fold aliases into an array property on the canonical node. If node_1 has the primary name Samuel Lee and node_2 is Sam Lee, append Sam Lee to the aliases property of node_1. Add all unique identifiers, such as secondary email addresses or alternate usernames, into the canonical node's identity arrays. When property conflicts arise, like conflicting job titles, keep the value with the most recent timestamp or store historical values in an audit map. If you're building a people search knowledge graph, keeping aliases inside the canonical node means downstream BM25 full-text search still finds the record whichever variant a user searches for.

Third, rewire all incoming and outgoing edges from the discarded entity to the canonical entity. If node_2 was connected to Project Alpha via an ASSIGNED_TO edge, point that edge directly to node_1. If both nodes already share an identical relationship to the same target node, merge the edges: keep the earliest created_at property, take the higher edge weight, and discard the duplicate edge. Finally, soft-delete or remove the obsolete node so traversal queries never hit ghost entities.

Step 5: Keep provenance edges to every source mention so merges are auditable and reversible

Entity resolution algorithms make mistakes. If your agent merges two people with the same name and later discovers they work at different companies, an irreversible merge permanently corrupts your knowledge graph. Every production agent memory system needs auditability and reversibility.

Don't treat an entity node as raw source text. Separate mentions from entities. When an LLM ingests a Slack thread or a document chunk, create a Mention node for that raw observation with its verbatim string, source URL, and timestamp. Then add an explicit MENTIONED_IN or EXTRACTED_FROM edge connecting the canonical Entity node to the Mention node.

When node_2 merges into node_1, rewire node_2's provenance edges to point to node_1, and record the merge event itself as a structured graph mutation. Add properties to the rewired edge or store a MergeEvent record with the original node ID, the calculated confidence score, the algorithm version, and the feature weights that triggered the merge.

Say downstream observations contradict the merge, like the two entities turning out to have different parent companies. Reversing it is straightforward. Read the provenance edges, filter mentions by their original source metadata, recreate the split entity, and restore the edges to their original configuration. Explicit provenance edges turn entity resolution from a dangerous, destructive write into an auditable, self-healing memory layer.

Step 6: Commit the merge atomically in HelixDB so graph and vector writes succeed or roll back together

The worst state for an agent memory system is a partial failure during a merge. Suppose your pipeline rewires half of an entity's graph edges in a graph database, deletes the candidate node, and then crashes before updating the vector index or deleting the old embedding from a separate vector database. Your agent memory is now desynchronized. Traversal paths point to ghost IDs, and similarity queries recall dead entities.

HelixDB removes this distributed failure mode by putting knowledge graphs, vector search, and BM25 full-text indexing inside a single Rust-built engine. Graph topology, vector properties, and text search live in the same operational store, and graph mutations and vector updates commit together in atomic transactions.

Applications submit operations as JSON queries via the native TypeScript, Python, Go, or Rust SDKs directly to the POST /v2/query HTTP endpoint. Here is how a merge operation looks in the TypeScript SDK (@helix-db/helix-db):

import { writeBatch } from "@helix-db/helix-db";

// Execute the canonical update, edge rewires, and node deletion atomically
const result = await writeBatch([
  {
    action: "update_node",
    id: canonicalId,
    properties: {
      aliases: updatedAliases,
      title: resolvedTitle,
      embedding: updatedContextVector,
    },
  },
  {
    action: "rewire_edges",
    from_node: duplicateId,
    to_node: canonicalId,
  },
  {
    action: "delete_node",
    id: duplicateId,
  },
]);

HelixDB processes this mutation in one atomic transaction. Either the property updates, edge rewires, and vector modifications all succeed, or the whole operation rolls back. HelixDB Cloud runs serializable snapshot-isolation ACID transactions, so simultaneous memory updates from parallel agents never leave the knowledge graph inconsistent. For multi-tenant setups, check out our guide on multi-tenant agent memory to keep merges strictly partitioned by tenant ID.

Step 7: Reject merges below the threshold and test with labeled pairs (troubleshooting false merges and what to do next)

When a candidate pair falls below your match threshold, leave the entities distinct. Keeping two separate nodes that might later be unified is far safer than prematurely collapsing two different entities. When the confidence score lands in the review bracket, create an UNRESOLVED_SIMILARITY edge between the two nodes with the calculated score. That edge prevents duplicate resolution runs and flags the pair so background agents can look for corroborating evidence in future interactions.

To keep your deduplication pipeline accurate over time, build an evaluation benchmark of 200 to 500 labeled entity pairs taken directly from production logs. Include obvious duplicates (nicknames, minor spelling variations), near-miss duplicates (two employees named John Smith in different departments), and unrelated entities that share common names.

Evaluate the pipeline on precision and recall. If you see false merges (precision drop), your vector similarity weight or name similarity threshold is too loose. Increase the required shared-neighbor threshold or raise the auto-merge cutoff. If you see graph fragmentation (recall drop), your blocking step is discarding valid candidates before scoring. Widen your BM25 query expansion or allow looser token containment on initial retrieval.

Run this benchmark in your continuous integration pipeline whenever you update prompt extraction templates, swap embedding models, or adjust scoring weights. Deterministic evaluation catches regressions as your agent memory graph scales to millions of nodes.

Conclusion

Entity deduplication is not an LLM prompting problem. It's a systems problem that needs candidate blocking, contextual embedding, multi-signal scoring, auditable provenance, and atomic transactional writes. Gluing together a graph database, a separate vector database, and full-text indexes introduces multi-database synchronization failures that routinely break agent memory under load.

HelixDB replaces that brittle stack with a single graph-vector engine written in Rust. You get native graph traversal, float32 vector pre-filtering, and BM25 search in one database, with atomic transactions that keep your knowledge graph clean and consistent. Star HelixDB on GitHub, pull the open-source Docker container, and start building deterministic memory graphs that don't corrupt over time.

Build with HelixDB

Give your coding agent the setup prompt, or sign up and deploy a database.

Sign up