Reranker Model
A specialized model that re-evaluates the relevance of document chunks retrieved by search stages.
Last reviewed: July 25, 2026
A reranker is a second-stage model in a retrieval pipeline that re-scores a shortlist of candidate documents for relevance to a query, sitting between an initial fast retrieval step and the final context that gets passed to an LLM. Rerankers exist because the retrieval methods that can search millions of documents quickly — vector similarity search or BM25 keyword matching — trade off some accuracy for speed, and a reranker recovers that accuracy on a much smaller candidate set where speed is less of a constraint.
How Rerankers Work
Most rerankers are cross-encoders: unlike the bi-encoder architecture used for initial retrieval (which embeds the query and each document independently, then compares vectors), a cross-encoder processes the query and a candidate document together in a single forward pass, letting the model’s attention mechanism directly compare query terms against document terms. This produces substantially more accurate relevance judgments, but it’s too slow to run against an entire corpus — hence using it only on the top 20-100 candidates that a faster retriever has already narrowed down.
Where It Fits in RAG
A typical retrieval-augmented generation pipeline looks like: embed the query, retrieve the top 50-100 candidate chunks via vector search, rerank those candidates down to the top 3-10 most relevant, then pass only those to the LLM’s context window. This two-stage design — cheap broad retrieval followed by expensive precise reranking — is one of the most reliable ways to improve RAG answer quality, since irrelevant chunks that make it into the context window are a common source of hallucinated or off-topic answers. Popular reranker models include Cohere’s Rerank API and open-weight options like BGE-reranker and Jina Reranker.
Latency Considerations
Because cross-encoder reranking requires a full forward pass per query-document pair rather than a precomputed comparison, it adds meaningful latency to a RAG pipeline compared to retrieval alone — typically tens to low hundreds of milliseconds depending on how many candidates are being reranked and the reranker model’s size. Production systems balance this by tuning how many candidates get passed to the reranker (reranking the top 20 rather than the top 100 candidates trades some potential recall for lower latency) and by choosing smaller, distilled reranker models when latency budgets are tight, accepting a modest accuracy tradeoff in exchange for faster response times.
Some newer reranker offerings address this with distilled or specifically optimized small cross-encoders designed to keep reranking overhead in the tens-of-milliseconds range rather than requiring a full large-model forward pass per candidate.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.