---
title: "RAG Retrieval Strategies on a 50k-Document Dataset: BM25 vs Dense vs Hybrid"
description: "How BM25, dense embeddings and hybrid search (RRF) behave on a 50k-document RAG corpus: working code, honest trade-offs, and when each one actually wins."
author: "Syed Ahmer Shah"
date: 2026-10-05
url: https://ahmershah.dev/blogs/rag-retrieval-strategies-bm25-vs-dense-vs-hybrid
tags: ["rag", "ai", "information-retrieval", "bm25", "embeddings", "python", "search", "engineering-logs"]
series: "The Engineering Logs" (part 13)
---

# RAG Retrieval Strategies on a 50k-Document Dataset: BM25 vs Dense vs Hybrid

_At 50k documents you can skip the vector database drama. What matters is how you retrieve, how you fuse, and whether you measured._

![RAG Retrieval Strategies on a 50k-Document Dataset — Ahmer behind three chess pieces labelled BM25, Dense and Hybrid, each orbited by documents](https://ahmershah.dev/blog/rag-retrieval-strategies-bm25-vs-dense-vs-hybrid/cover-ff949b93.webp)


> **TL;DR**
> - On 50k documents, retrieval quality matters more than your choice of vector database. At this size you can run exact search and skip the infrastructure.
> - **BM25** is great at exact terms (error codes, names, IDs) and weak at paraphrases. **Dense** retrieval is the reverse.
> - **Hybrid** (BM25 + dense, merged with Reciprocal Rank Fusion) is the best default. It's cheap to build and covers both failure modes.
> - Add a reranker only after you've measured that hybrid isn't enough. Build a small eval set first, because without one you're guessing.

## The setup nobody tells you about

Say you've built a RAG app over 50,000 documents: support tickets, internal docs, maybe PDFs. The demo worked. Then a real user asked a real question, and the model confidently answered from the wrong paragraph.

Nine times out of ten, the LLM isn't the problem. **Retrieval is.** If the right chunk never reaches the prompt, no model can save you.

One clarification before we start. "50k documents" usually means far more *chunks*. If each document splits into about 10 chunks, you're searching roughly 500k pieces of text. That number matters for memory and latency, so keep it in mind below.

I haven't run a fresh benchmark on your data, and I won't pretend to. This article uses published research and documented numbers to show what each approach does, where it breaks, and how to decide. If you want the background on what embeddings actually are, my post on [the anatomy of AI](/blogs/syedahmershah-the-anatomy-of-ai-deconstructing-the-brain-into-vectors-and-math) covers vectors from the ground up.

## The three contenders at a glance

| | BM25 | Dense | Hybrid (RRF) |
|---|---|---|---|
| Matches on | Exact words | Meaning | Both |
| Strong at | IDs, codes, names, jargon | Paraphrases, natural questions | Mixed real-world queries |
| Weak at | Synonyms, vocabulary mismatch | Rare tokens, unseen domains | Slightly more moving parts |
| Needs a model | No | Yes (embeddings) | Yes |
| Build effort | Low | Low to medium | Low, once the other two exist |

## BM25: the librarian with perfect index cards

**BM25** is a keyword-ranking function from the 1990s. It scores a chunk higher when:

1. it contains your query words (term frequency),
2. those words are rare across the corpus (inverse document frequency), and
3. the chunk isn't just long and rambling (length normalization).

*Analogy:* imagine a librarian who never reads the books. She only keeps meticulous index cards of which words appear where. Ask for "ERR-4012" and she hands you the exact page in seconds. Ask "why does my app crash on startup" and she's stuck unless the page literally says "crash" and "startup."

Here's a minimal version using `rank_bm25`:

```python
import re
from rank_bm25 import BM25Okapi

def tokenize(text: str) -> list[str]:
    # keep dots, dashes and underscores *inside* a token (ERR-4012, user_id,
    # v2.4.1) but not at its edges, so "startup." still matches "startup"
    return re.findall(r"[a-z0-9]+(?:[._\-][a-z0-9]+)*", text.lower())

tokenized_chunks = [tokenize(c) for c in chunks]
bm25 = BM25Okapi(tokenized_chunks)

def bm25_search(query: str, k: int = 50):
    scores = bm25.get_scores(tokenize(query))
    top = scores.argsort()[::-1][:k]
    return [(int(i), float(scores[i])) for i in top]
```

Two honest notes. First, `rank_bm25` is fine for prototypes, but it scores every document per query. For production, use something with an inverted index like Elasticsearch, OpenSearch, Tantivy, or Postgres full-text search. Second, your *tokenizer* is half the battle. The regex above keeps `ERR-4012` as one token on purpose. A default tokenizer that splits on punctuation can quietly wreck exact-match queries.

**Where BM25 wins:** product codes, function names, legal clause numbers, rare jargon, anything where the user types the exact words.

**Where it loses:** vocabulary mismatch. The user says "refund," the doc says "reimbursement." BM25 sees two unrelated strings.

## Dense retrieval: the friend who gets what you mean

**Dense retrieval** turns text into vectors (embeddings) using a neural model. Chunks with similar *meaning* land close together in vector space, and you search by finding the nearest neighbors to your query vector.

*Analogy:* instead of index cards, you have a well-read friend. You describe a vague idea ("that thing where the server keeps retrying and makes it worse") and they say "oh, you want the doc on retry storms." They understand meaning, but they might blank on the exact serial number you asked for.

```python
import numpy as np
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-small-en-v1.5")

# 50k chunks x 384 float32 is about 77 MB. 500k chunks is about 770 MB.
doc_vecs = model.encode(chunks, normalize_embeddings=True, batch_size=64)

def dense_search(query: str, k: int = 50):
    q = model.encode([query], normalize_embeddings=True)[0]
    scores = doc_vecs @ q          # cosine similarity (vectors are normalized)
    top = np.argpartition(-scores, k)[:k]
    top = top[np.argsort(-scores[top])]
    return [(int(i), float(scores[i])) for i in top]
```

Notice what's missing: no vector database, no ANN index. That's deliberate.

### You probably don't need an ANN index yet

Approximate nearest neighbor indexes like HNSW trade a bit of recall for speed. At small scale, you may be paying that trade for nothing. pgvector, for example, does exact nearest neighbor search by default, which gives perfect recall. Plenty of teams run tens of thousands of vectors on a plain sequential scan, and skipping the index avoids build time, maintenance cost, and the recall loss of approximate search.

If your 50k documents become 500k chunks, revisit that, but measure first. A brute-force matrix multiply over a few hundred thousand vectors is still quick on a decent machine; it's the memory (see the comment in the code above) you'll feel first. When you do need HNSW, here's the standard pgvector shape:

```sql
CREATE INDEX ON chunks
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);

-- higher ef_search = better recall, slower queries
SET hnsw.ef_search = 100;
```

Always compare ANN results against exact results on your own queries before trusting the index.

**Where dense wins:** paraphrases, natural-language questions, cross-vocabulary matches, multilingual queries.

**Where it loses:** exact identifiers, rare tokens, and anything where the embedding model has never seen your domain's vocabulary.

## What the research says (and what it doesn't)

The most cited evidence here is **BEIR**, a benchmark covering zero-shot retrieval across a wide range of settings. Its headline finding surprised people at the time. Dense models beat BM25 by a clear margin on MS MARCO, the data they were trained on, yet BM25 held up as a robust baseline almost everywhere else. The authors found dense models do well when training and target data overlap heavily, but fall short on datasets with large domain shift.

That's the key insight for your 50k-document corpus. Your data is probably *not* MS MARCO. It has your company's acronyms, your product names, your weird internal vocabulary. A dense model trained on generic web text is, in effect, out-of-domain. That's exactly where BM25 refuses to embarrass itself.

Two honest caveats:

1. **BEIR is from 2021.** Embedding models have improved a lot since. Don't read it as "BM25 beats dense." Read it as "dense models can fail unpredictably on unfamiliar domains, and lexical search is a safety net."
2. **Benchmark averages hide your queries.** A model can win on average and still lose on the 20% of queries your users care about most.

BEIR also showed the cost side. BM25 answers in milliseconds on a CPU, while a cross-encoder reranker on top of it was orders of magnitude slower per query in the paper's setup. Reranking helps, but it isn't free, and we'll come back to that.

## Hybrid: two witnesses are better than one

*Analogy:* picture two witnesses to a crime. One remembers exact details (the license plate), and the other remembers the overall impression (a dark sedan, driving erratically). Either alone can mislead you. Together, they converge on the right car.

That's **hybrid retrieval**: run BM25 and dense search in parallel, then merge the two ranked lists.

The tricky part is the merge. BM25 scores and cosine similarities live on completely different scales, so you can't just add them. You *can* normalize and weight them, but the weights drift as your data changes.

### Reciprocal Rank Fusion (RRF)

RRF sidesteps the scale problem by ignoring scores entirely and using only **rank positions**. It comes from a 2009 SIGIR paper by Cormack, Clarke, and Büttcher. The formula is one line:

```
RRF(d) = Σ over rankers  1 / (k + rank(d))
```

with k = 60 as the usual default. The paper found that value near-optimal in a pilot study, and it has stuck around ever since. Because RRF is score-independent, using only ranks, one retriever with weird score ranges can't dominate the result.

```python
from collections import defaultdict

def rrf(ranked_lists: list[list[int]], k: int = 60) -> list[tuple[int, float]]:
    fused = defaultdict(float)
    for ranking in ranked_lists:
        for rank, doc_id in enumerate(ranking, start=1):
            fused[doc_id] += 1.0 / (k + rank)
    return sorted(fused.items(), key=lambda x: x[1], reverse=True)

def hybrid_search(query: str, k: int = 10, pool: int = 50):
    bm25_ids  = [i for i, _ in bm25_search(query, pool)]
    dense_ids = [i for i, _ in dense_search(query, pool)]
    return rrf([bm25_ids, dense_ids])[:k]
```

*Analogy:* RRF is like a talent show with two judges, where each judge only gives rankings, never scores. A contestant who is 2nd on both lists beats one who is 1st on one list and 40th on the other. **Consensus beats a single loud opinion.**

Two practical tips: pull a generous candidate pool (50 or so) from each retriever before fusing, and treat `k = 60` as a starting point rather than a law.

## A real-world data point on hybrid

Anthropic published numbers on this in their *Contextual Retrieval* write-up. The headline result: combining contextual embeddings with contextual BM25 cut the top-20-chunk retrieval failure rate by 49%, from 5.7% to 2.9%. They also explain *why* BM25 earns its place: embedding models can miss exact-match queries like unique identifiers, which BM25 handles well.

A fair caution comes from critics of that post. The best number, a 67% reduction, comes from stacking *several* techniques, including reranking, not from any single trick. One analysis pointed out that the headline figures measure only top-k retrieval failure rate, which may not match what *you* care about. So don't copy their pipeline wholesale. Take the lesson: **hybrid beats either alone, and each added layer needs its own measurement.**

## How to actually decide: build a tiny eval

This is the part most tutorials skip, and it's the part that matters. Before you pick a strategy, write down 50 to 100 real questions with the chunk IDs that *should* answer them. Then measure **recall@k**: what fraction of the time the right chunk shows up in the top k results.

```python
def recall_at_k(search_fn, eval_set, k=10):
    hits = 0
    for query, relevant_ids in eval_set:
        retrieved = {i for i, _ in search_fn(query)[:k]}
        hits += bool(retrieved & set(relevant_ids))
    return hits / len(eval_set)

for name, fn in [("bm25", bm25_search),
                 ("dense", dense_search),
                 ("hybrid", hybrid_search)]:
    print(name, recall_at_k(fn, eval_set, k=10))
```

Pull your eval questions from real user logs if you have them, and include the ugly ones: typos, error codes, vague questions. Hand-label them yourself. It's boring, and it's the highest-leverage hour you'll spend on this project.

**Look at the failures, not just the number.** If BM25 misses are paraphrase problems and dense misses are ID problems, you've just proven hybrid is justified for *your* data, not someone else's.

## What I'd do on a 50k-document corpus

If I were handed this project tomorrow, here's the order I'd work in:

1. **Start with BM25 alone.** It's fast to build and a strong baseline. If it already hits your recall target, you may be done.
2. **Add dense retrieval with a decent open embedding model**, brute-force, no vector DB. Compare the two on your eval set.
3. **Fuse with RRF.** In most real corpora with mixed vocabulary, this is where recall jumps.
4. **Fix chunking before adding fancier models.** Chunks that lose their context (like a paragraph saying "it increased 12%" with no idea what "it" is) hurt *both* retrievers. This is the problem Anthropic's contextual approach targets.
5. **Add a reranker last**, and only if hybrid recall@50 is good but precision@5 is poor. That pattern means the right chunk is in the pool but ranked too low, which is exactly the reranker's job. Budget for the latency, since reranking can cost far more per query than the retrieval itself.
6. **Reach for HNSW only when measurements say so**, such as high query volume or a corpus that grew far past your current size.

## Common mistakes I see

- **Skipping evaluation** and trusting vibes from five hand-picked demo queries.
- **Tokenizing BM25 carelessly**, so identifiers and code get shredded.
- **Normalizing and summing raw scores** from different retrievers instead of fusing ranks.
- **Over-engineering infrastructure** (a managed vector DB, three services) for a dataset whose vectors fit in under a gigabyte of RAM.
- **Blaming the LLM** for answers that were doomed by what retrieval handed it.
- **Ignoring metadata filters.** If users usually ask about one product or one date range, filtering before search often beats any clever ranking.

## Final thoughts

There's no universally best retriever. BM25 is the reliable, literal-minded librarian. Dense retrieval is the intuitive friend who sometimes gets overconfident. Hybrid puts them in the same room and lets them check each other's work.

At 50k documents, you have the luxury of simplicity: brute-force vectors, a plain keyword index, and fusion code you can read in 15 seconds. Spend your effort on a good eval set and good chunking. Those two things will move your answer quality more than any model swap.

The snippets above are illustrative, so test them on your own data before relying on them.

* * *

## References

1. Thakur et al. (2021). [BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models](https://arxiv.org/pdf/2104.08663). NeurIPS Datasets and Benchmarks.
2. Cormack, Clarke, Büttcher (2009). [Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods](https://dl.acm.org/doi/10.1145/1571941.1572114). SIGIR '09.
3. Anthropic. [Introducing Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval).
4. Almond AI. [Responding to Anthropic's Contextual Retrieval: Why Context is NOT All You Need](https://medium.com/almond-ai/responding-to-anthropics-contextual-retrieval-for-rag-apps-why-context-is-not-all-you-need-af503985aa55).
5. pgvector. [pgvector on GitHub](https://github.com/pgvector/pgvector).


---

*Originally published at https://ahmershah.dev/blogs/rag-retrieval-strategies-bm25-vs-dense-vs-hybrid — © Syed Ahmer Shah*
