Reranking — more accurate RAG
The highest-ROI upgrade for RAG: a second layer that re-ranks results and pushes the truly relevant ones to the top. A dramatic accuracy boost for little effort.
Why vector retrieval alone isn't enough
In basic RAG, you search a Vector DB for the passages whose embedding is closest to the query. That's fast and great for "coarse filtering," but not always accurate: the embedding captures general meaning, and sometimes ranks a passage that "sounds similar" but doesn't really answer the question highly. The result — irrelevant passages in the context, which lead to worse answers and even hallucinations.
Reranking is usually the highest-ROI improvement you can make to RAG — a noticeable accuracy gain in a few lines of code.
Two-stage retrieval
The idea: combine two complementary stages —
- Broad retrieval: the Vector DB returns many candidates (say 50) — fast, narrowing to a relevant subset.
- Precise ranking (rerank): a reranker model examines each candidate against the query and returns the truly best N (say 5) — slower, but far more accurate.
You feed the LLM only the top 5 after reranking — a concise, precise context (see also Context Engineering).
How a reranker works — cross-encoder vs bi-encoder
The technical difference that explains the accuracy:
- Bi-encoder (embeddings): encodes the query and the passage separately into vectors, and compares. Very fast (you can precompute), but less accurate because there's no interaction between the query and the passage.
- Cross-encoder (reranker): feeds the query and the passage together into a model that returns a relevance score. Much more accurate because it "reads" both together — but slow, so you run it only on a few candidates.
Hence the logic of two stages: a fast bi-encoder for coarse filtering, an accurate cross-encoder for the polish.
Hybrid Search — a bonus
A complementary improvement: combine semantic search (embeddings) with keyword search (BM25). Each has an advantage — semantic captures meaning, keyword captures exact terms, names and codes. You merge both candidate lists, then run a reranker over the union. This gives the best of both worlds, and the gain is largest for technical terms and for languages the embedding model saw less of during training.
Code example — a reranker in service
Reranking services (like Cohere Rerank) take a query and a list of documents and return them ranked. Open models also exist (BGE-reranker) via Hugging Face.
# Stage 1: broad retrieval from the Vector DB
candidates = vector_search(query, top_k=50) # 50 candidates
# Stage 2: re-rank (example with a reranker library)
scored = reranker.rank(query=query,
documents=[c.text for c in candidates])
top = sorted(scored, key=lambda x: x.score, reverse=True)[:5]
# Stage 3: feed the LLM only the top 5
context = "\n\n".join(t.text for t in top)
answer = llm_answer(query, context)
What retrieval actually gets wrong
"Embeddings are less accurate" is true and not useful. The failures have a shape, and knowing it tells you when a reranker will help and when something else is broken.
An embedding compresses a passage into a fixed vector — a summary of what it is broadly about. Two things follow. It is excellent at topic and poor at specificity, and it has no notion of what the question was asking for.
So the characteristic mistakes are:
- Same topic, wrong question. A passage about your refund policy scores well for "how long does delivery take" because both live in the same region of meaning.
- Negation and opposites. "Supported platforms" and "unsupported platforms" embed close together, because they are about the same thing.
- Identifiers smeared. Model numbers, error codes, version strings and people's names lose their distinctness in a vector. SKU-4471 and SKU-4417 look nearly identical to an embedding and are entirely different to a customer.
- Long chunks diluted. A passage covering five subjects has a vector that is the average of five subjects, which is close to nothing in particular.
A cross-encoder fixes the first two, because it reads the query and the passage together and can tell that this passage is about the right thing but does not answer the question. It does not fix the third — that is what keyword search is for — and it does not fix the fourth, which is a chunking problem.
The reranker can only reorder what you gave it
The most important operational fact on this page, and the one that causes teams to conclude reranking "did not work".
Stage two is a reordering. If the passage that answers the question was not among the fifty candidates that came back from the vector search, no reranker can retrieve it. Your accuracy is capped by stage one's recall, permanently.
So measure that first. Take your evaluation questions, run only the retrieval stage, and ask: what fraction of the time is a correct passage anywhere in the top 50? That number is your ceiling.
- If recall at 50 is high — say most questions have the answer in there somewhere — reranking is exactly the right investment, and it will convert that latent recall into actual precision.
- If recall at 50 is poor, reranking is polishing the wrong candidates. The problem is upstream: chunking, embedding model, missing content, or a query that needs rewriting before it is embedded.
This one measurement decides where your next week goes, and it takes an afternoon.
Why not just cross-encode everything?
The obvious question, and the answer explains the whole two-stage architecture.
A bi-encoder lets you precompute. Every passage in your corpus gets embedded once, in advance, and stored. At query time you embed one short query and do a nearest-neighbour lookup — an operation that stays fast whether the corpus holds ten thousand passages or ten million.
A cross-encoder cannot precompute anything, because the score depends on the query and the passage together. Scoring a corpus means one forward pass per passage, per query. That is fine for fifty candidates and impossible for a million.
Hence: a cheap approximate filter that scales, then an expensive accurate ranking on a bounded set. The same pattern appears throughout search systems, and once you see it the parameter choices stop being arbitrary — stage one is sized by your recall requirement, stage two by your latency budget.
Hybrid search, and how to merge two lists
Combining semantic search with keyword search (BM25) covers the identifier problem embeddings have. Keyword search is exact where vectors are fuzzy: product codes, error numbers, proper nouns, rare technical terms, and anything a user copied and pasted out of an error message.
The practical question is how to combine two ranked lists whose scores are not comparable — a cosine similarity and a BM25 score live on different scales, so adding them is meaningless.
The standard answer is reciprocal rank fusion: ignore the scores and use the positions. Each document gets a contribution of roughly one over its rank in each list, and the contributions are summed. A document ranked third by both methods beats one ranked first by one and fiftieth by the other. It needs no tuning, no score normalisation, and it is robust — which is why it is the default in most search stacks that offer hybrid retrieval.
Then rerank the fused list. Retrieval casts a wide net in two different ways; the cross-encoder decides what actually answers the question.
This matters more in languages other than English, and for any corpus full of names and codes — the two cases where pure semantic search disappoints most reliably.
Chunking decides how well ranking can work
Reranking gets blamed for problems that are really chunking problems, so it is worth being explicit about the interaction.
Chunks that are too large dilute both the embedding and the rerank score. A two-thousand-word section containing one relevant paragraph scores as mostly-irrelevant, and if it wins anyway you have spent context on 1,900 useless words.
Chunks that are too small score well and then fail to answer, because the sentence that matched has lost the context that made it meaningful — the pronoun whose referent was two paragraphs up, the number whose units were in the heading.
Two techniques address this directly. Retrieve small, return large: index small chunks for precise matching, but feed the model the surrounding section once a small chunk wins. And keep the heading path in the chunk text — prefixing each chunk with its document title and section headings gives both the embedder and the reranker the context a human would get from the page layout, and it is close to free.
Measuring it properly
"It seems better" is how teams ship a reranker that costs latency and adds nothing. The measurement is not difficult.
Build a set of questions with known good answers — fifty is enough to be useful, drawn from real user queries rather than invented ones. For each, mark which passages in your corpus genuinely answer it. That marking is the tedious part and it is the whole value.
Then track three numbers:
- Recall at k for stage one — your ceiling, as above.
- Where the correct passage lands after reranking. If it is reliably first or second, the reranker is working. Mean reciprocal rank captures this in one figure if you want a single number.
- The end-to-end answer quality, which is what you actually care about and does not always move with the ranking metrics.
Keep the set and re-run it on every change — a new embedding model, a different chunk size, a reranker version. Retrieval systems are full of changes that improve one query visibly and degrade twenty invisibly, and the set is the only defence against that. See evals for the wider practice.
Latency, and where the time goes
A reranker adds a step, and whether that matters depends on where it sits in the user's experience.
In a system that then calls an LLM to generate an answer, the generation usually dominates — so the reranking step is often a small fraction of total response time, and the improved context can even make generation faster by shortening the prompt. In a search interface that returns links directly, the reranker is the latency, and the trade is real.
Three ways to buy it back. Reduce stage-one k — going from 100 candidates to 40 roughly halves the reranking work, and if your recall measurement says 40 is enough, the accuracy cost is nothing. Batch the candidates in one call rather than scoring them one at a time, which most APIs and libraries support and which is a large difference. And cache — repeated queries are common in support and documentation search, and the ranking for an identical query over unchanged content is reusable.
The cost is usually negative
Worth stating plainly because it is counter-intuitive: adding a reranker frequently makes the system cheaper.
Reranking is priced per document scored and is cheap — these are small models. The LLM call is the expensive part, and it is priced per token. If reranking lets you send five well-chosen passages instead of twenty mediocre ones, you have cut the largest cost in the pipeline by a large fraction, and paid a small amount for the privilege.
The secondary saving is answer quality: fewer irrelevant passages means less for the model to be distracted by, which reduces both wrong answers and the retries that follow them.
Choosing a reranker
The options divide into hosted APIs and open models you run yourself, and the decision is the same one that appears everywhere else in this stack.
A hosted reranking API is a single call, needs no infrastructure, and is priced per document scored. For most teams this is the right starting point: the quality is good, the integration is an afternoon, and the cost is small next to the LLM call it is saving you money on.
An open cross-encoder run on your own hardware costs nothing per query, keeps your documents inside your network, and can be fine-tuned on your own domain — which is where it beats the hosted option outright, because a reranker trained on your queries and your content understands what relevance means in your context. The price is a GPU and someone to look after it.
Whichever you pick, three things to check before committing:
- The maximum input length. Rerankers truncate, and a model with a short limit silently cuts your chunks in half — scoring only the beginning while you believe it read the whole thing.
- Language coverage, tested on your own content rather than taken from a model card.
- Whether scores are comparable between queries. Some rerankers return calibrated scores you can threshold globally; others return values that only rank within one query. This decides whether you can implement the "answer nothing if nothing clears the floor" rule.
The whole pipeline, in order
Putting the pieces together, a retrieval stack that works tends to look like this — and the order matters, because each step narrows what the next one has to handle.
- Rewrite the query into a standalone question, resolving pronouns and conversation context.
- Filter by metadata — product, language, date, tenant. Cheap, exact, and it removes whole categories of wrong answer.
- Retrieve twice, semantically and by keyword, and fuse the two lists by rank.
- Rerank the fused candidates with a cross-encoder.
- Apply a score floor. Nothing above it means the honest answer is that you do not know.
- Expand the winners to their surrounding context if you indexed small chunks.
- Generate, with instructions to answer only from the supplied passages and to cite which one each claim came from.
Most systems that disappoint are missing steps one, two and five — the cheap ones. Reranking at step four gets the attention because it is the interesting part, but a pipeline with steps one to three done properly and no reranker usually beats one with a reranker and nothing else.
Two things to try before reranking
Reranking is the highest-return change for a RAG system that is basically working. If yours is not, these are usually the problem.
Metadata filtering. If your corpus spans products, dates, languages or customers, filtering to the right subset before the vector search beats any amount of clever ranking. Searching one product's documentation instead of all twelve removes the wrong-answer class entirely rather than ranking it lower. This is frequently a bigger win than reranking and almost nobody does it first.
Query rewriting. Real user queries are short, ambiguous, and full of pronouns referring to earlier conversation. Expanding "does it support that?" into a standalone question before embedding it fixes retrieval failures that no reranker can recover from, because the wrong candidates were fetched.
When reranking will not help
- A small corpus. With a few hundred passages, retrieval already returns almost everything relevant, and a larger context window can simply take more of it.
- When recall is the bottleneck — measured, not assumed. Fix retrieval first.
- When the content is not there. No ranking step invents an answer your documents do not contain, and this is a surprisingly common root cause.
- Strict low-latency search-as-you-type. The budget may simply not exist.
Tips
- Retrieve broad, return narrow. Something like 30–50 candidates in, 3–8 out, tuned by your own recall measurement rather than by these numbers.
- Measure with evals, per stage, not by eye.
- Check multilingual support before choosing a reranker if your content is not English — quality varies far more across languages than the headline claims suggest, and hybrid search helps most exactly there.
- Prefix chunks with their heading path before indexing. Cheapest accuracy improvement available.
- Keep the score the reranker returns and set a floor — if nothing clears it, answering "I do not have that information" is better than feeding the best of a bad set.
Next step
Reranking is part of advanced RAG. Go deeper, and measure the improvement with evals.