Back to blog
Why Does Hybrid Search Return Fewer Results Than You Asked For?

This blog is written by AI for SEO

Why Does Hybrid Search Return Fewer Results Than You Asked For?

HelixDB13 min read

You asked for twenty documents. Your retriever dispatched ten candidates to the vector index, ten to BM25, ran Reciprocal Rank Fusion, and delivered seven chunks to your prompt. Nothing failed. The database returned a 200 OK and the call was fast, but your context window is half-empty.

Engineers debugging why hybrid search returns fewer results than requested usually assume their fusion math broke. They tweak RRF smoothing constants, relax score thresholds, or blame tokenizers. That diagnosis is almost always wrong. Rank fusion merges and orders candidates; it rarely deletes them unless your pipeline suffered an architectural breakdown earlier in the call.

In production retrieval pipelines, a truncated result set points to one of two structural realities. Either an unindexed post-filter is silently dropping valid matches from your candidate pool, or a scoped graph traversal did its job and hit the natural boundary of a small candidate universe. One is a bug in your execution pipeline. The other is a mathematical correctness guarantee.

The Short Answer: Two Causes, One Fix Each

When a hybrid search query asks for k items and returns fewer than k, the root cause is never random. It falls into two distinct operational categories: an execution ordering failure or an intentional boundary constraint.

The first scenario is a post-filtering defect. Your retrieval engine pulls a fixed candidate pool (say, twenty items from dense vector search and twenty from sparse BM25), merges them, and then passes the combined list through metadata filters in application memory or a secondary database query. If fifteen of those forty candidates belong to an inactive tenant, an expired permission group, or a non-matching document category, your pipeline discards them after the top-k cutoff already occurred. Your final count collapses from twenty to five. The fix is moving predicates into the retrieval step via true pre-filtering or unified index evaluation.

The second scenario is your query running against a strictly scoped candidate universe. Consider a system querying a specific sub-graph, such as chunks linked directly to an active incident ticket or a single legal agreement. If that incident ticket only contains four linked notes, a query specifying k=20 cannot yield twenty notes without violating query isolation. The retriever returns four notes because four notes exist. The fix here is changing your downstream expectations: do not pad the result set with irrelevant global matches when the query requested an isolated entity scope.

Nearly every truncated result set is one of these two. An over-aggressive score threshold is a third and rarer possibility, covered further down. Conflating them leads developers to write workarounds like speculative over-fetching, which add latency and pollute the context window with irrelevant chunks.

Cause 1: Post-Filtering Is Consuming Your Top-K Slots

Post-filtering is the most common reason hybrid search under-delivers on candidate counts. To see why, examine the lifecycle of a standard hybrid retrieval query across dense and sparse indexes.

Assume your retriever targets top_k=20. The dense index uses an Approximate Nearest Neighbor (ANN) algorithm like HNSW to explore its vector graph. It halts traversal once it finds the twenty closest vectors according to cosine distance. Simultaneously, your sparse BM25 engine scans its inverted index, evaluates term frequencies, and isolates its top twenty matching lexical records. Next, Reciprocal Rank Fusion, introduced by Cormack, Clarke and Buettcher in 2009, combines these ranked lists by rank rather than by raw score into a unified candidate pool.

Up to this point, the pipeline holds enough candidates to satisfy your limit. The breakdown occurs when your application applies metadata conditions: tenant_id == 'tenant_beta', visibility == 'public', or created_at >= 2026-01-01.

If the ANN traversal and BM25 scans were blind to these metadata conditions, they selected candidates from your entire multi-tenant corpus. When your application layer loops over the fused top-k array and evaluates the predicates, it drops non-matching elements. If thirty-two of the forty candidates belong to other tenants or fail the date filter, only eight candidates survive.

Your retriever returns eight items, not because the corpus lacked twenty valid records, but because your candidate budget was consumed by documents that were immediately disqualified. The vector index selected the closest points in embedding space without knowing they were ineligible, and the fusion step locked in those wasted slots before the filter had a chance to run.

Why Is This Anti-Pattern So Easy to Write?

Engineers do not write post-filtering logic because they misunderstand set theory. They write it because their database architecture forces them to.

When a team builds RAG by duct-taping independent systems together (Pinecone for vector embeddings, Elasticsearch for BM25 lexical search, and Postgres for business metadata), no single component owns the full schema. Pinecone holds vector IDs and payload blobs. Elasticsearch stores tokenized text.

Because these systems operate as isolated network hops, running a combined query requires application glue code. A Python backend queries the vector store, queries the search cluster, merges the IDs in memory, and then runs a SQL query: SELECT id FROM documents WHERE id IN (...) AND tenant_id = $1. Writing this pipeline takes twenty lines of Python, making it deceptively fast to prototype.

In local testing with five hundred documents, every document belongs to the test tenant, so every top-k slot survives. In production with ten thousand tenants, nearly all of the global vector space belongs to organizations other than the caller. The isolated vector engine returns twenty vectors from across the entire database, and the SQL check discards nineteen of them.

Overcoming this failure requires unified execution. HelixDB is an open-source graph-vector database written in Rust, and it keeps graph structure, vector properties and BM25 full-text indexes in one storage engine. When the relational boundaries sit alongside the vector properties, the engine evaluates predicates during traversal instead of in client memory afterwards. The mechanics of scoping a similarity search to a relationship are in our guide on vector search over graph edges.

Cause 2: A Scoped Search That Is Working Correctly

Not every missing result is an execution defect. In advanced retrieval architectures, such as GraphRAG or entity-scoped agent memory, returning fewer than k results is the correct behavior.

Consider an agent designed to answer questions about a specific customer support ticket. The prompt targets the ticket entity, traverses outgoing graph edges to discover associated chat transcripts, customer profile notes, and billing entries, and then runs hybrid search across the text chunks tied to those nodes. This pattern prevents cross-entity hallucination by strictly scoping retrieval to a verified subgraph.

If the customer ticket contains only three conversation turns and one billing entry, exactly four relevant text chunks exist in that sub-graph. When the agent issues a query with top_k=10, the engine should return four records.

If the engine returned ten records, six of them would have to come from outside the requested relational boundary: either from different support tickets or unrelated customer accounts. That would fill the prompt with material the question never asked about.

One thing worth being exact about, because it is easy to over-read. Scoping a retrieval to a subgraph is a correctness property of the query, not an authorization layer. It governs what a query can return, not who is allowed to run it. Access control stays where it already lives in your stack.

When teams transition from naive RAG to relational systems, developers often misinterpret this behavior as a retrieval bug. They see results.length < requested_k in their monitoring dashboards and assume the index failed. The query engine enforced an absolute boundary condition. When you ask a database to search within a set of size N, the maximum possible result count is N. If N is smaller than k, receiving N results means your scoping logic functioned precisely as requested. For a practical walkthrough of setting up entity-bounded retrieval, consult our guide on how to build a GraphRAG pipeline.

What Execution Order Actually Guarantees a Full Result Set?

Whether your short result count is a bug or an intended constraint depends entirely on pipeline execution order. Retrieval systems follow one of three architectural pipelines, and each produces different mathematical guarantees.

Pipeline 1: Post-Filtering (Non-Deterministic Count) Retrieve(Global Index, k) -> Rank Fusion -> Filter(Metadata) -> Return (Count <= k) In this order, candidate collection happens globally. Filters run on the tail end. Result counts vary unpredictably based on how heavily your target metadata is represented in the top-k candidates of the global vector space.

Pipeline 2: Attribute-Based Pre-Filtering (Guaranteed min(k, Total Matching)) Filter(Metadata Index) -> Retrieve(Filtered Candidates, k) -> Rank Fusion -> Return (Count = min(k, Total Matching)) Here, the index resolves the metadata predicate first, constructing a bitset or candidate ID set. The vector and lexical indexes constrain their search space to matching nodes only. If twenty matching records exist anywhere in the corpus, the pipeline is guaranteed to return twenty.

Pipeline 3: Structural Traversal with Hybrid Scoring (Guaranteed min(k, Scoped Universe)) Traverse(Graph Entity, Relationships) -> Bound Candidate Set -> Hybrid Score(Dense + Sparse) -> Return (Count = min(k, Scoped Universe)) In this pipeline, graph relationships define the candidate universe before any scoring begins. The retriever does not search a global vector space; it scores only the nodes reachable via the specified relationship paths.

HelixDB runs queries over an HTTP API at its /v2/query endpoint, as JSON built by native SDKs in TypeScript, Python, Go and Rust. Because vectors are properties attached directly to graph nodes and edges, the traversal and the ranking are one execution rather than two systems meeting in your application.

The documented order is graph traversal, then exact candidate membership, then vector ranking, then top k, and the traversal membership is authoritative: a result outside the candidate set cannot come back. The honest caveat, which is the part worth stating plainly, is that exact membership does not mean the engine compares every candidate embedding one by one. Approximate structures still do the ranking. What is guaranteed is the boundary, not an exhaustive scan.

Two limits are worth knowing before you pick a k. The server caps unrestricted vector search at 800 effective results. A traversal-scoped search applies that ceiling after candidate intersection and rejects the query when min(k, unique candidates) goes over 800, rather than quietly clamping it. The rejection is the useful half: a silent clamp is precisely what makes a short result set inexplicable elsewhere, because the engine hands you less than you asked for and never mentions it. Above a million unique entities in the candidate stream, the query is an error rather than a slow success.

The full-text half behaves identically, which is the part a two-system stack cannot reproduce. textSearchWith chains onto a traversal exactly as vectorSearchWith does, takes the same four arguments, carries a BM25 $score where the vector form carries $distance, and follows the same documented order. A hybrid query is therefore scoped once and both halves obey that scope. You are not intersecting two independently retrieved lists and hoping enough of them overlap.

How Do You Tell Which Problem You Actually Have?

When your hybrid retriever returns fewer items than requested, three diagnostic signals will tell you which cause you are looking at.

First, check candidate pool sizes before and after filtering. Log the count of raw matches immediately following dense and sparse retrieval, before any application-level filtering or rank fusion deduplication. If your dense retriever returns twenty IDs, your sparse retriever returns twenty IDs, but your final output array drops to four items, you have a post-filtering defect. Your candidate slots were occupied by disqualified documents.

Second, inspect partition cardinality. Run an explicit count query against your datastore using the exact metadata filters from your request. Check COUNT(*) where organization_id = 'org_abc'. If the database reports only four records match that organization in the entire database, your query returned four records because that is the total population size. You are observing Cause 2. Do not attempt to fix this by modifying retrieval thresholds.

Third, verify score distribution and threshold clipping. A common subtle variant of post-filtering occurs when engineers apply hard score cutoffs (dropping items with cosine similarity below 0.70 or RRF reciprocal rank scores below 0.02). If five documents pass the metadata check but only two exceed the similarity cutoff, your threshold is too aggressive for your embedding model. As discussed in our analysis of why semantic search over internal documents falls short, dense similarity scores drift between domains and between chunk lengths. Discarding candidates via arbitrary score floors produces the exact same symptom as broken post-filtering.

How Do You Fix It Without Breaking the Scoping Guarantee?

If diagnostics reveal that post-filtering is consuming your top-k slots, you must eliminate post-filtering without sacrificing relational guarantees.

Many teams attempt a temporary patch: speculative over-fetching. If you need twenty results, you retrieve two hundred candidates, run the filter, and slice the first twenty survivors. This masks the defect in staging environments but creates real production problems. On queries where metadata filters are selective, two hundred candidates may still yield only two survivors. On queries where filters are permissive, retrieving two hundred candidates wastes memory, increases serialization overhead, and introduces tail latency spikes during peak traffic.

The real solution requires shifting predicate evaluation into the index traversal itself. Both dense vector search and sparse BM25 have to restrict their search space before they rank anything. If your relationships already live in a graph database and it is the retrieval half you want to change, the mechanics of that move are in our guide on moving agent memory off Memgraph.

Start with the shape to avoid. This is the anti-pattern the docs call out by name: search the whole label, then filter what came back.

// Anti-pattern. Do not ship this.
// The text search takes the top k across EVERY Document, and the filter
// then discards the ones outside your workspace. Those discarded high
// scorers have already spent your k, so you get back fewer than you
// asked for even though enough matching documents exist.
g()
  .textSearchWith("Document", "body", params.query_text, params.limit)
  .where(/* workspace predicate, applied far too late */);

The scoped version builds the candidate set from the traversal first, so the ranking only ever sees documents that belong to the workspace.

import {
  SourcePredicate,
  defineParams,
  g,
  param,
  readBatch,
} from "@helix-db/helix-db";

const params = defineParams({
  workspace: param.string(),
  query_vector: param.array(param.f32()),
  limit: param.i64(),
});

const query = readBatch()
  .varAs(
    "matches",
    g()
      .nWithLabelWhere("Workspace", SourcePredicate.eq("slug", params.workspace))
      .out("CONTAINS_DOC")
      .vectorSearchWith("Document", "embedding", params.query_vector, params.limit)
      .valueMap(["$id", "title", "$distance"]),
  )
  .returning(["matches"]);

const request = query.toQueryRequest(
  params,
  { workspace: "acme", query_vector: queryVector, limit: 20n },
  { queryName: "workspace_doc_matches" },
);

Note the 20n: the limit is an i64, so it wants a BigInt literal. The keyword half is the same query with one operator swapped, and it reads a $score instead of a $distance.

// Same imports as above.
const params = defineParams({
  workspace: param.string(),
  query_text: param.string(),
  limit: param.i64(),
});

const query = readBatch()
  .varAs(
    "matches",
    g()
      .nWithLabelWhere("Workspace", SourcePredicate.eq("slug", params.workspace))
      .out("CONTAINS_DOC")
      .textSearchWith("Document", "body", params.query_text, params.limit)
      .valueMap(["$id", "title", "$score"]),
  )
  .returning(["matches"]);

Both halves start from the same traversal, so both are bounded by the same candidate set. That is what makes the count predictable: you get min(k, matching documents in the workspace), and the only way to come back with fewer is for the workspace to hold fewer. Ask for more than the ceiling described above and you get an error rather than a quietly shortened list, which is the behaviour you want when the alternative is debugging a number nobody told you was capped.

The same query shape runs whichever way you have deployed the engine, because running fully in memory, on disk, or against S3-compatible object storage is a startup flag rather than a different product.

Conclusion

When your hybrid search returns fewer results than you asked for, do not reach for fusion weights or speculative over-fetching. Check your execution order. If your pipeline retrieves globally and discards locally, your top-k slots are being wasted on records that never had a chance to qualify.

HelixDB puts graph traversal, vector similarity and BM25 full-text search in one engine written in Rust, so a hybrid query is scoped once and both halves respect that scope. The execution order, the 800-result ceiling and the exact query shapes above are all in the filtering guide. The engine is on GitHub, and a star helps if this is the shape of bug you keep chasing.

Build with HelixDB

Give your coding agent the setup prompt, or sign up and deploy a database.

Sign up