
This blog is written by AI for SEO
How to Query Agent Memory by Time Range Without a Full Scan
Ask an LLM agent what a user said about their billing refactor last Tuesday, and the usual retrieval pipeline falls apart. The standard recipe is predictable: embed the question, pull the top k nearest neighbors from a vector index, and post-filter in application code to drop anything outside Tuesday's window.
That pipeline fails silently. If the user has talked about billing many times over the past months, older high-similarity chunks take every one of the k slots. The memories from last Tuesday never enter the candidate set, the post-filter discards everything, and the agent either invents an answer or says it has no memory of the conversation.
Raising k and hoping is not the fix. The fix is to bound the candidates by owner and by time first, then rank only that set by similarity. This guide models agent memory as events with a timestamp, indexes that timestamp, and queries a time window with a traversal and prefiltered vector search in HelixDB.
To follow this guide, you need a Node.js or TypeScript development environment, the @helix-db/helix-db package installed, and access to a running HelixDB instance.
Step 1: Model Events as Nodes With a Top-Level Timestamp Property
Many vector stores encourage dumping operational metadata into a nested JSON payload. A timestamp buried in a nested field is awkward to index, so range filtering over it tends to mean inspecting every record.
In HelixDB, you model time yourself: store it as a top-level property, the same way a vector property has to be top-level. Model every conversational turn, tool call or agent reflection as an Event node connected to the user it belongs to.
An event node needs four properties here: an integer timestamp in epoch milliseconds, an eventType string, a text content field, and an embedding vector. HelixDB assigns the node ID, which comes back as $id.
interface EventNode {
timestamp: number; // Unix epoch in milliseconds
eventType: "user_message" | "agent_thought" | "tool_output";
content: string;
embedding: number[]; // Float32 vector
}Relationships carry the scoping. Attach each Event node to its User node with a LOGGED edge. A traversal from that user then decides which memories are eligible at all, before any timestamp comparison or similarity ranking runs.
Step 2: Create a Range Index on the Timestamp and a Vector Index on the Embedding
You model time with your own properties. What HelixDB provides is efficient time-range indexing, through ordered secondary indexes.
Without an index, checking whether a timestamp falls between two values means inspecting every node that carries the label. An ordered index lets the engine seek to the start of the window and read forward to its end, so the work tracks the size of the window rather than the size of the whole history. The secondary indexes guide lists four families: equality, unique equality, range ascending and range descending. The range families are the ones that serve time windows.
Index creation is a write query. It runs asynchronously and backfills existing data, so poll the returned receipt until each index is active before depending on it. Prefiltered vector search also needs an active vector index, and this example uses a global cosine index with a three-dimensional toy embedding; use your embedding model's real dimension:
import { IndexSpec, VectorDistanceMetric, g, writeBatch } from "@helix-db/helix-db";
const createIndexes = writeBatch()
.varAs(
"timestampIndex",
g().createIndexIfNotExists(IndexSpec.nodeRange("Event", "timestamp")),
)
.varAs(
"embeddingIndex",
g().createIndexIfNotExists(
IndexSpec.nodeVector("Event", "embedding", 3, VectorDistanceMetric.Cosine),
),
)
.returning(["timestampIndex", "embeddingIndex"]);Vector indexes can be global or tenant-partitioned. A tenant-partitioned index needs the same tenant value that built the candidate stream, and tenant partitioning keeps data isolated on shared infrastructure; it is not access control.
Step 3: Write Events Into the Graph With writeBatch()
Writing agent memory needs transactional integrity. If the node, its embedding and its edge lived in different stores, a failure halfway through would leave memory in a state no query can trust.
In HelixDB, a write batch commits or rolls back as a unit, so the event node, its vector property and the edge to its user land together. The writing data guide covers the full set of write operations.
Install the TypeScript SDK:
npm install @helix-db/helix-dbThis batch looks up the user by an indexed username, creates the Event node, and connects the two:
import { NodeRef, SourcePredicate, g, writeBatch } from "@helix-db/helix-db";
const logEvent = writeBatch()
.varAs(
"user",
g().nWithLabelWhere("User", SourcePredicate.eq("username", "alice")),
)
.varAs(
"event",
g().addN("Event", {
timestamp: 1790672400000,
eventType: "user_message",
content: "Let's push the billing refactor to next sprint.",
embedding: [0.12, -0.03, 0.88],
}),
)
.varAs(
"logged",
g().n(NodeRef.var("user")).addE("LOGGED", NodeRef.var("event")),
)
.returning(["event"]);Every write persists the node properties and the vector property in SlateDB, HelixDB's storage engine. Because the batch is atomic, a query never sees an event whose timestamp exists but whose embedding or edge is missing.
Step 4: Bound the Time Window With a Traversal
There are two ways to apply the window. A source predicate such as nWithLabelWhere("Event", SourcePredicate.between("timestamp", start, end)) is eligible for index push-down, so it reads only the events inside the window, but across every user. That suits a global view, such as everything logged in the last hour.
For one user's memory, start at the User node, follow LOGGED, and filter the reached events with .where(...). Here the user's edges bound the stream before the timestamp check runs.
The filtering guide draws the same line: source predicates select initial candidates and can use an index, while .where(...) filters the stream later in a traversal. For a very active user with a long history, the traversal stream grows with that history, which is the trade-off to watch.
import { Predicate, SourcePredicate, g } from "@helix-db/helix-db";
const start = 1790640000000; // Tuesday 2026-09-29 00:00 UTC
const end = 1790726399999; // Tuesday 2026-09-29 23:59:59.999 UTC
const tuesdayEvents = g()
.nWithLabelWhere("User", SourcePredicate.eq("username", "alice"))
.out("LOGGED")
.where(Predicate.between("timestamp", start, end))
.valueMap(["$id", "timestamp", "content"]);The engine only reads events reachable from that user, so other users' memories never enter the stream.
Step 5: Add Prefiltered Vector Search Inside the Time Window
Once the traversal bounds the candidates to one user and one window, rank them by similarity. This is vector pre-filtering: distance is computed only over the nodes the traversal reaches, not over the whole index.
Post-filtering asks the vector index for nearest neighbors across the entire database, then discards candidates that fail the time filter. If recent entries share lower semantic scores than older entries, the required memories never appear in the top-k results. Pre-filtering guarantees that every candidate evaluated against the query embedding already satisfies the time window and relational boundary.
The prefiltered search guide describes the order every prefiltered search runs in: traversal, exact candidate membership, ranking, then top k. A result outside the candidate set is never returned. For scoping similarity across relationships in more depth, see our guide on vector search on graph edges.
Keep the documented limits in mind. An unrestricted vector search caps at 800 effective results. A traversal-scoped search applies that ceiling after candidate intersection and rejects the query, rather than silently clamping it, when the smaller of k and the unique candidate count exceeds 800. A candidate stream above 1,000,000 unique entities is a query error.
Step 6: The Full Query: Traversal, Time Window and Semantic Recall
Put the pieces in one read query. The SDK builds the JSON request sent to POST /v2/query; HelixDB parses it and turns it into the code that executes the query, so there is no separate query language and nothing compiled ahead of time. Typed parameters keep the operation tree stable while the values change, as the parameters guide explains.
import {
Predicate,
SourcePredicate,
defineParams,
g,
param,
readBatch,
} from "@helix-db/helix-db";
const params = defineParams({
username: param.string(),
start: param.i64(),
end: param.i64(),
query_vector: param.array(param.f32()),
limit: param.i64(),
});
const recallWindow = readBatch()
.varAs(
"memories",
g()
.nWithLabelWhere("User", SourcePredicate.eq("username", params.username))
.out("LOGGED")
.where(Predicate.between("timestamp", params.start, params.end))
.vectorSearchWith("Event", "embedding", params.query_vector, params.limit)
.valueMap(["$id", "timestamp", "content", "$distance"]),
)
.returning(["memories"]);
const request = recallWindow.toQueryRequest(params, {
username: "alice",
start: 1790640000000,
end: 1790726399999,
query_vector: [0.1, -0.02, 0.9],
limit: 5,
});This one query roots at the user, keeps only events inside the window, and ranks what is left by cosine distance, returning each memory with its $distance. There is no second store to keep in sync and no post-filter to starve the result.
What to Do Next (and What to Watch Out For)
To implement time-range memory retrieval in production, keep these operational realities in mind.
First, mind the candidate limits. An agent that logs every tool call can put a lot of events inside a loose window. Keep k well under the 800 ceiling, and narrow the window or add a traversal step, such as a session or topic, when one user's history gets long.
Second, tenancy is not authorization. HelixDB's tenant partitioning keeps tenant data isolated on shared infrastructure, but it is not per-user permissions or row-level access control. Enforce who may read which memories in your application before you build the traversal root.
Third, create the indexes before bulk ingestion and wait for them to report active, so your first range queries do not fall back to scanning.
Run it locally with helix init then helix start dev, as the quickstart shows. The same engine also runs embedded in your application, as a self-hosted Docker server in memory, on disk or against S3-compatible object storage, or on HelixDB Cloud.
Conclusion
Querying agent memory by time range should not mean syncing a vector store with a graph store, or trusting a post-filter that drops the context you needed. Store events as nodes with a top-level timestamp, index it, and scope the vector search to one user's window in a single query.
HelixDB is open source under Apache-2.0, and HelixDB Cloud starts at $5 a month for 200K reads and 100K writes. Star HelixDB on GitHub, run helix init and helix start dev, and try a time-bounded recall query against your own agent's memory.