Back to blog
Multi-Tenant Agent Memory: How to Isolate per Tenant

This blog is written by AI for SEO

Multi-Tenant Agent Memory: How to Isolate per Tenant

HelixDB9 min read

Picture two enterprise customers sharing one autonomous agent platform. Tenant A asks their agent to summarize recent vendor negotiations, and the agent recalls a contract clause that belongs to Tenant B. Your security model is broken and your product is compromised.

Most teams start with a patchwork. They spin up a shared vector database, inject a tenant_id tag into chunk metadata, and run a naive filter after retrieving nearest neighbors. Then queries start dropping valid results, latencies climb, and relational context across entities disappears. You don't need a separate database cluster for every customer to fix this. In this guide, you'll configure multi-tenant agent memory inside a single graph-vector database engine, combine relational graph structure with tenant-partitioned vector indexes, and run pre-filtered searches that prevent cross-tenant data leaks.

Step 1: Pick an isolation model (database per tenant, shared index with post-filter, or tenant-partitioned index)

Engineering teams typically consider three architectural patterns for multi-tenant agent memory. Each one shifts the trade-off between operational overhead, query latency, and resource isolation.

The first is database-per-tenant. You spin up a distinct database instance or namespace for each customer. The boundaries are clean and physical. If an attacker injects a prompt that tries to traverse unauthorized nodes, there is no other customer data inside the process memory or on the disk volume. But the operational cost is brutal. Managing schema migrations, connection pooling, and baseline memory footprints across 500 separate database instances slows developers down and burns cloud spend.

The second is a shared index with a tenant_id post-filter. All customer data goes into a single global vector index. Queries run approximate nearest neighbor (ANN) retrieval across the entire dataset first, then discard matches where tenant_id doesn't equal the requester's tenant. It's operationally simple. It also breaks down under load and drops recall badly.

The third is a tenant-partitioned index within a unified engine. The database keeps separate, isolated index structures per tenant on disk and in memory, but shares the underlying storage and runtime engine. HelixDB runs multi-tenancy on shared infrastructure. You get bounded index scans, fast retrieval, and low operational overhead. Tenant partitioning is strictly an organizational and indexing boundary, not access control. We'll come back to that distinction later.

Step 2: Understand why a tenant_id post-filter on ANN under-returns and can leak ranking

A post-filter on an approximate nearest neighbor search is one of the most common pitfalls in agent memory design. To see why, trace what the retrieval algorithm actually does.

When your application sends an embedding vector to an ANN index with top_k set to 10, the index traverses its graph structure (such as HNSW) to find the 10 closest vectors in the global space. Only after those 10 candidates are selected does your application or database evaluate the metadata filter (tenant_id == 'tenant_123'). If tenant_456 owns 80 percent of the items near that query vector, eight of the 10 returned vectors belong to tenant_456 and get discarded. Your agent receives two results instead of 10. As we cover in why does hybrid search return fewer results than you asked for, post-filtering produces silent recall starvation.

Post-filtering also creates a subtle ranking leak. Say your agent requests the top five memories related to a sensitive incident. If a competitor tenant has dozens of highly similar embeddings nearby in the vector space, those embeddings push your customer's memories further down the candidate list, past the exploration boundary (efSearch). You strip out the competitor text before returning the payload, but the competitor's data still changed which memories your tenant got to see.

Pre-filtering or true index partitioning fixes this. When you restrict the search to the tenant partition before calculating vector distance, the search budget is spent only on that customer's vectors. Top_k returns k genuine results every time.

Step 3: Run HelixDB locally (helix init, helix start dev) and model tenant-scoped memory as a graph

To see how partitioned indexing works alongside graph relationships, run HelixDB locally. Install the CLI and initialize a fresh working directory:

helix init agent-memory-demo
cd agent-memory-demo
helix start dev

This starts the open-source engine, ready to receive HTTP queries. HelixDB combines graph traversal, approximate vector search, and full-text search in one Rust engine.

Agent memory is not a flat list of strings. Real agent state is entities, conversations, and observations connected by explicit relationships. We go deeper on this pattern in how to give AI agents persistent memory in one database.

Model your tenant data as a property graph. Every node representing an agent memory must carry a tenant_id property. A simple schema includes:

  • Tenant node: contains customer metadata (tenant_id, organization_name).

  • Agent node: represents the AI persona or assistant scoped to that tenant.

  • MemoryItem node: contains textual memory, timestamp, and a top-level vector property (embedding).

  • OBSERVED edge: connects Agent to MemoryItem, carrying metadata such as confidence or emotional valence.

In HelixDB, vector properties must be declared as top-level properties on nodes or edges. Writing tenant_id directly on the nodes anchors your relationship graph to the tenant boundary.

Step 4: Create a tenant-partitioned vector index and write memories with tenant_id

HelixDB lets vector indexes be global or tenant-partitioned. When an index is tenant-partitioned, vectors are indexed inside segregated internal structures keyed by tenant_id. Vector values normalise to float32, and the engine supports cosine, euclidean, and manhattan distance metrics.

Write your data with the official TypeScript SDK. Import named exports directly from '@helix-db/helix-db' (v3.0.4):

import {
  g,
  writeBatch,
  defineParams,
  param,
  VectorDistanceMetric
} from "@helix-db/helix-db";

// Write an agent memory node with a top-level vector
const memoryPayload = {
  tenant_id: "cust_corp_a",
  content: "User prefers quarterly summaries in markdown tables.",
  memory_type: "preference",
  created_at: 1714521600,
  embedding: [0.012, -0.045, 0.089, /* ... float32 values */]
};

const query = writeBatch((b) => {
  const memoryNode = b.addNode("MemoryItem", {
    tenant_id: memoryPayload.tenant_id,
    content: memoryPayload.content,
    memory_type: memoryPayload.memory_type,
    created_at: memoryPayload.created_at,
    embedding: memoryPayload.embedding
  });
  return memoryNode;
});

When HelixDB ingests this node, the atomic transaction engine commits the graph attributes and vector index entries at the same time. Because the index partitions by tenant_id, the embedding goes directly into the index segment allocated for 'cust_corp_a'. It shares no centroid or neighbor list with other tenants' data.

Step 5: Search within one tenant's partition via POST /v2/query

Once memories are partitioned, queries must run strictly within the requesting tenant's context. The SDK generates JSON queries that are posted to HelixDB's HTTP endpoint at POST /v2/query. The database parses this JSON directly and turns it into execution code.

Here is how to query vector similarity within a tenant-partitioned index using the TypeScript SDK:

import {
  g,
  readBatch,
  defineParams,
  param,
  Predicate,
  VectorDistanceMetric
} from "@helix-db/helix-db";

const queryEmbedding = [0.011, -0.042, 0.091, /* ... 1536 dims */];
const activeTenant = "cust_corp_a";

const searchQuery = readBatch((b) => {
  return b.nodes("MemoryItem")
    .where(Predicate.eq("tenant_id", activeTenant))
    .findSimilar({
      vectorProperty: "embedding",
      target: queryEmbedding,
      metric: VectorDistanceMetric.Cosine,
      limit: 10
    });
});

The JSON sent to POST /v2/query tells the engine to skip all vectors belonging to other partitions. The ANN search runs entirely against the candidate space of 'cust_corp_a'. An unrestricted vector search in HelixDB caps at 800 effective results. Because the search evaluates strictly within the tenant index structure, you avoid under-returns and keep true rank proximity.

Step 6: Pre-filter similarity on graph-scoped candidates, including vectors on edges

Multi-tenant agent memory gets much stronger when you combine graph topology with vector search. Often you don't want to search every memory a tenant has ever accumulated. You want only the memories connected to a specific active project, conversation thread, or user session.

HelixDB natively supports vector pre-filtering scoped to related entities. Instead of running vector similarity across the whole tenant partition and filtering by relationships afterward, the engine traverses the graph first to establish a candidate set, then ranks vectors only across those candidates. To see how embeddings attach to relationships, read our tutorial on pre-filtering vector search on graph edges.

Vectors can live directly on edges as well as nodes. For example, if an edge represents a 'DISCUSSED_TOPIC' relationship between a User and a Project, the edge itself can store an embedding of the semantic gist of that discussion:

import {
  g,
  readBatch,
  Predicate,
  VectorDistanceMetric
} from "@helix-db/helix-db";

const agentQuery = readBatch((b) => {
  return b.nodes("Tenant")
    .where(Predicate.eq("tenant_id", "cust_corp_a"))
    .outEdges("ACTIVE_PROJECT")
    .toNodes("Project")
    .where(Predicate.eq("project_id", "proj_compliance_2026"))
    .inEdges("REFERENCES")
    .whereSimilar({
      vectorProperty: "context_vector",
      target: [0.034, 0.12, -0.05, /* ... */],
      metric: VectorDistanceMetric.Cosine,
      limit: 5
    });
});

HelixDB applies the 800-result ceiling after candidate intersection, and it rejects the query if the candidate stream exceeds 1,000,000 unique entities. This graph-filtered approach constrains semantic similarity by both tenant ownership and structural relevance.

Step 7: Enforce access in your application, because partitioning is not authorization

Tenant partitioning is not access control. It's an index optimization and data organization mechanism. Partitioning keeps queries fast and search candidate sets isolated, but the database engine doesn't check whether an incoming HTTP request has legitimate authority to query that partition.

Never accept a tenant_id directly from untrusted client input. If your agent UI or MCP client passes a tenant parameter in the request body, an attacker can change that value and query another customer's partition. Validate authorization in your application layer before you construct queries.

Build an enforcement layer in your backend service:

  1. Validate Session Authentication: Extract identity from a cryptographically signed token (such as a verified JWT or mTLS session).

  2. Resolve Tenant Identity: Map the authenticated user or agent identity to their verified tenant_id on the server side.

  3. Bind Tenant Parameters: Hardcode the tenant_id inside your query builder using bound parameters. Don't let client payloads override it.

  4. Apply Scoped RBAC: If your tenant has internal roles (such as admin versus member), enforce those entity-level restrictions in your query predicates.

Treat tenant partitioning as physical isolation and your application middleware as the security gate. If you do, accidental data bleeding across customer boundaries stops.

What to do next: open-source core vs HelixDB Cloud, and troubleshooting empty results

When you move from local development to production, pick the deployment model that fits your operational needs. HelixDB's open-source core is Apache-2.0 licensed, so you can self-host it embedded in your application, as a Docker server on disk, or backed by S3-compatible object storage. For managed infrastructure, HelixDB Cloud is a fully managed, object-storage-backed service. A gateway routes traffic, a single writer serializes mutations, and reader nodes auto-scale horizontally with tiered caching.

If your tenant-partitioned queries return zero results during initial setup, check these common failure modes:

  • Vector Normalization: Make sure your embedding dimensions match the index configuration. Mismatched vector dimensions fail silently or reject insertions.

  • Top-Level Property Check: Confirm that vector fields are declared as top-level properties on nodes or edges. HelixDB doesn't index vectors nested inside generic JSON blob fields.

  • Tenant ID Exact Match: Verify that the tenant_id filter string exactly matches what was written during ingestion. Check for trailing whitespace or case mismatches in customer slugs.

  • Candidate Stream Limits: Queries with candidate streams over 1,000,000 unique entities trigger an error instead of a partial scan.

Conclusion

Isolating agent memory across hundreds of enterprise tenants doesn't require managing hundreds of separate database instances or gambling on leaky post-filters. HelixDB combines graph structure with tenant-partitioned vector indexing, so you store entities, relationships, and embeddings in one Rust engine while keeping each customer's memory strictly separated.

Clone the open-source repository at github.com/HelixDB/helix-db, start a local instance with helix start dev, and test your first tenant-partitioned agent memory schema today.

Build with HelixDB

Give your coding agent the setup prompt, or sign up and deploy a database.

Sign up