Hybrid search and reranking in RAG: BM25 plus embeddings, reciprocal rank fusion, cross-encoders, and why this pair fixes most bad answers

Why vector search alone misses exact terms and keyword search alone misses meaning, how to fuse them with reciprocal rank fusion in a dozen lines, what a cross-encoder reranker does and why it beats retrieving more, how to tune k at each stage, how to measure the gain with recall, and how the hybrid path works in a production RAG platform.

On the mic
On the mic

Key takeaways

  • Embeddings find meaning and miss exact strings; BM25 finds exact strings and misses paraphrase. Running both and fusing fixes both failure modes.
  • Reciprocal rank fusion merges ranked lists with no score normalisation and no tuning; it is a dozen lines and a 2009 paper.
  • A cross-encoder reranker reads the question and each candidate together, which is far more accurate than comparing vectors, so retrieve wide and rerank narrow.
  • Tune k per stage: retrieve 30 to 50 from each retriever, fuse, rerank to 3 to 8, and measure recall at each boundary.
  • Hybrid plus reranking is the single largest quality jump available to most RAG systems, and it costs milliseconds, not GPUs.

What is hybrid search in RAG?

Hybrid search runs two retrievers on every query, a keyword retriever such as BM25 and a dense embedding retriever, and merges their ranked results into one list. Keyword search catches exact terms, product codes, names, error numbers and rare words that embeddings blur together. Embedding search catches paraphrases and meaning that share no words with the document. Fusing them gives higher recall than either alone, and the fusion itself can be a dozen lines of code with no tuning.

Also asked as: what is hybrid search in rag · hybrid search rag · hybrid retrieval · hybrid search vs vector search · bm25 vs embeddings · dense vs sparse retrieval · hybrid search explained · keyword search vs semantic search rag · lexical and semantic search · why hybrid search

Hybrid retrieval with reciprocal rank fusion and a cross-encoder reranker is the core of RAG.NextUpgrad's query path, and it is the part of the platform that moved retrieval recall most on the labelled question set [13]. My shorter notes on each half are on the portfolio [11][12]. This page is the pair together, with the papers and the code.

Users ask for "error E-4021" as often as they ask "why did my upload fail". Only one of those is a semantic query. Ship both retrievers or you will fail half your users. Pranjul Rathour, from building RAG.NextUpgrad

Why does vector search alone miss results?

Because embeddings compress meaning and lose exact strings. A product code, a person's name, an error number, an acronym or a version string carries little meaning to the model, so its vector lands near other codes and names rather than near the exact match. Numbers and negation blur too. Dense retrieval was shown to beat keywords for open-domain questions on average [3], and the average hides exactly the queries that matter most in support, legal and technical corpora.

Also asked as: limitations of vector search · why semantic search fails · vector search exact match · embeddings miss keywords · semantic search product codes · vector search names · when vector search fails · dense retrieval weaknesses

What is BM25, and why is it still used?

BM25 is a ranking function from classical information retrieval that scores a document for a query by how often the query's terms appear in it, weighted by how rare each term is across the collection and normalised for document length [1]. It needs no training, runs in microseconds over an inverted index, and finds exact terms perfectly. It cannot see that "reset my password" and "credential recovery" are the same request, which is exactly what embeddings are for. Together they cover each other's blind spots.

Also asked as: what is bm25 · bm25 explained · bm25 algorithm · bm25 vs tf idf · bm25 in rag · how bm25 works · bm25 python · okapi bm25 · sparse retrieval bm25 · keyword search algorithm

Learned sparse models such as SPLADE bring some semantic expansion to the keyword side [6], and late-interaction models such as ColBERT keep per-token vectors to recover exact-term precision on the dense side [5]. Both are worth knowing; BM25 plus a dense model plus fusion remains the practical default.

What is reciprocal rank fusion?

Reciprocal rank fusion merges several ranked lists into one by giving each document a score equal to the sum, over every list it appears in, of one divided by a constant plus its rank in that list. Documents ranked high in either list rise; documents ranked high in both rise most. The constant, usually 60, dampens the influence of top ranks. It needs no score normalisation, because it uses ranks not scores, and Cormack, Clarke and Buettcher showed it beating more complex fusion methods [2].

Also asked as: reciprocal rank fusion · rrf rag · what is rrf · reciprocal rank fusion explained · rrf formula · how to combine bm25 and vector search · rank fusion · hybrid search fusion methods · rrf k parameter · weighted rrf

def reciprocal_rank_fusion(*ranked_lists: list[str], k: int = 60) -> list[str]:
    scores: dict[str, float] = {}
    for ranked in ranked_lists:
        for rank, doc_id in enumerate(ranked, start=1):
            scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

fused = reciprocal_rank_fusion(dense_top_40, bm25_top_40)[:30]   # then rerank these

That function is in RAG.NextUpgrad almost exactly as written, because the paper's simplicity is the point: no learned weights, no score scales to reconcile, nothing to retune when you swap an embedding model [13].

What is reranking, and why not just retrieve more?

A reranker is a model that reads the question and a candidate passage together and outputs a relevance score. Retrievers compare a query vector with passage vectors that were computed separately, for speed; a cross-encoder looks at the pair jointly, which is far more accurate but too slow to run over the whole corpus. So you retrieve wide with the fast methods, thirty to fifty candidates, and rerank narrow, to the best three to eight. Retrieving more without reranking sends noise to the model, and models use long contexts unevenly [8].

Also asked as: what is reranking in rag · reranking explained · cross encoder reranker · bi encoder vs cross encoder · why rerank in rag · reranker vs retriever · how does a reranker work · best reranker for rag · does rag need a reranker · two stage retrieval

Passage reranking with BERT-style cross-encoders was introduced by Nogueira and Cho and remains the recipe [4]; Sentence Transformers publishes pretrained rerankers you can run on a CPU for small candidate sets [10].

How do I choose k at each stage?

Measure recall at each boundary on a labelled question set. Retrieve enough from each retriever that recall at the fusion output is above ninety percent; forty each is a common starting point. Fuse and keep enough that the reranker sees every plausible passage, twenty to fifty. Rerank down to what the model needs, three to eight, and check recall at that k is still above eighty-five. If recall drops sharply at the rerank step, the reranker is wrong for your domain; if it drops at fusion, one retriever is starving the other.

Also asked as: how many chunks to retrieve rag · top k rag · how many documents to send to llm · rerank top k · retrieval k value · how many results for reranking · tuning k in rag · recall at k tuning

On the mic
On the mic

How much does hybrid search and reranking improve results?

On the labelled sets I have worked with, fusion raised recall at thirty over either retriever alone, and the reranker raised recall at five without touching recall at thirty, which is exactly what each is supposed to do. Public benchmarks such as BEIR show dense models varying widely across domains and BM25 remaining a hard baseline on several [7], which is the research version of the same lesson. Your gain depends on your queries: corpora heavy in codes, names and numbers gain most from BM25; corpora of prose gain most from the reranker.

Also asked as: hybrid search performance · does hybrid search improve rag · reranking improvement · hybrid search benchmark · bm25 vs dense benchmark · beir hybrid · rag accuracy improvement reranking · recall improvement hybrid search

Keep a keyword index alongside the vector index: Postgres full-text search or a BM25 library over the same chunks [9]; Elasticsearch, OpenSearch, Qdrant, Weaviate and others also offer sparse or hybrid support. At query time, run both, fuse with the function above, rerank with a cross-encoder from Sentence Transformers or a hosted reranking API [10], and send the top few to the model. The whole path is a few hundred lines; the design decisions are the k values and the reranker choice.

Also asked as: how to implement hybrid search · hybrid search python · bm25 and embeddings together · postgres hybrid search · elasticsearch hybrid search · qdrant hybrid search · weaviate hybrid search · langchain hybrid retriever · llamaindex hybrid search · hybrid search implementation

When is hybrid search not worth it?

Rarely, but there are cases. A tiny corpus where everything fits in the prompt does not need retrieval at all. A corpus of pure prose with no codes or names, where the reranker alone recovers what dense misses. A latency budget so tight that even microseconds matter, which is unusual. In every other case I have seen, the cost of a second retriever is small and the recall gain on exact-term queries is large. Default to hybrid and let your measurements argue you out of it.

Also asked as: when not to use hybrid search · is hybrid search always better · hybrid search overhead · hybrid search latency · simple rag without bm25 · rag without reranking

Hybrid search and reranking interview questions

Explain why embeddings miss exact terms and BM25 misses paraphrase. Explain reciprocal rank fusion and why it needs no normalisation. Explain the difference between a bi-encoder and a cross-encoder and why you cannot rerank the whole corpus. Describe how you would choose k at each stage and measure it. Say when you would use learned sparse or late interaction models. A retrieval system you built and ablated is the best evidence.

Also asked as: hybrid search interview questions · reranking interview questions · retrieval interview questions rag · bm25 interview question · rrf interview · information retrieval interview

Where should I start?

Add BM25 next to your vector search over the same chunks, fuse with the twelve-line function, and measure recall at thirty on twenty real questions before and after. Then add a cross-encoder and measure recall at five. Both changes take an afternoon and usually deliver the largest quality jump the system will ever see. For a hands-on retrieval session at your college or team, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Robertson & Zaragoza, The Probabilistic Relevance Framework: BM25 and Beyond (2009)staff.city.ac.uk
  2. Cormack, Clarke & Buettcher, Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods (SIGIR 2009)plg.uwaterloo.ca
  3. Karpukhin et al., Dense Passage Retrieval for Open-Domain Question Answering (2020)arxiv.org
  4. Nogueira & Cho, Passage Re-ranking with BERT (2019)arxiv.org
  5. Khattab & Zaharia, ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT (2020)arxiv.org
  6. Formal et al., SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking (2021)arxiv.org
  7. Thakur et al., BEIR benchmark (2021)arxiv.org
  8. Liu et al., Lost in the Middle (2023)arxiv.org
  9. rank_bm25: BM25 implementations in Pythongithub.com
  10. Sentence Transformers: cross-encoder rerankerssbert.net
  11. Hybrid retrieval: dense, BM25 and RRF, Pranjul Rathourpranjulrathour.scult.in
  12. Reranking in RAG: why a cross-encoder second pass fixes most bad answers, Pranjul Rathourpranjulrathour.scult.in
  13. RAG.NextUpgrad source code, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur