Cross-Encoder Reranking
A secondary retrieval pipeline step that scores document-query pairs using deep attention classification.
Last reviewed: July 25, 2026
Cross-encoder reranking is the second-stage relevance scoring step in retrieval-augmented generation (RAG) pipelines, using a model architecture that processes a query and a candidate document together — rather than separately — to produce a much more accurate relevance score than initial retrieval alone.
Cross-Encoders vs. Bi-Encoders
The embedding models used for initial retrieval are bi-encoders: they encode the query and each document independently into separate vectors, which are then compared using a fast similarity calculation like cosine similarity. This independence is exactly what makes bi-encoders fast enough to search millions of documents — every document’s embedding can be precomputed once, and a new query only needs to be embedded and compared against that existing index.
A cross-encoder instead feeds the query and a candidate document into the model together, as a single combined input, letting the model’s attention layers directly compare specific terms and phrases between the two texts. This produces substantially more accurate relevance judgments because the model can reason about the specific interaction between query and document, rather than relying on how well two independently computed vectors happen to align. The cost is that this comparison can’t be precomputed — it has to be run fresh for every query-document pair, making cross-encoders far too slow to run against an entire corpus.
How They’re Used Together
This is why cross-encoder reranking is always a second stage: a fast bi-encoder (or keyword search) first narrows a large corpus down to a shortlist of perhaps 20-100 candidates, and only that shortlist is rescored with the more expensive but more accurate cross-encoder, with the top few results after reranking being what’s actually passed into the LLM’s context window. Popular cross-encoder rerankers include Cohere Rerank, BGE-reranker, and Jina Reranker.
Reranking Isn’t Always Necessary
Not every RAG pipeline needs a reranking stage — for smaller corpora, or queries where the initial retrieval step already returns highly relevant results with little noise, adding a reranker introduces latency without a proportional quality improvement. Reranking tends to matter most when the initial retrieval candidate pool is large and noisy (broad vector similarity search across a large, heterogeneous corpus tends to surface some marginally relevant or off-topic candidates alongside the genuinely relevant ones), which is precisely the scenario where a cross-encoder’s more precise, context-aware scoring provides the clearest benefit over the faster but coarser initial ranking.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.