Skip to main content
Level: Advanced Updated: September 2026

Reranking — more accurate RAG

The highest-ROI upgrade for RAG: a second layer that re-ranks results and pushes the truly relevant ones to the top. A dramatic accuracy boost for little effort.

Why vector retrieval alone isn't enough

In basic RAG, you search a Vector DB for the passages whose embedding is closest to the query. That's fast and great for "coarse filtering," but not always accurate: the embedding captures general meaning, and sometimes ranks a passage that "sounds similar" but doesn't really answer the question highly. The result — irrelevant passages in the context, which lead to worse answers and even hallucinations.

The insight

Reranking is usually the highest-ROI improvement you can make to RAG — a noticeable accuracy gain in a few lines of code.

Two-stage retrieval

The idea: combine two complementary stages —

  1. Broad retrieval: the Vector DB returns many candidates (say 50) — fast, narrowing to a relevant subset.
  2. Precise ranking (rerank): a reranker model examines each candidate against the query and returns the truly best N (say 5) — slower, but far more accurate.

You feed the LLM only the top 5 after reranking — a concise, precise context (see also Context Engineering).

How a reranker works — cross-encoder vs bi-encoder

The technical difference that explains the accuracy:

Hence the logic of two stages: a fast bi-encoder for coarse filtering, an accurate cross-encoder for the polish.

Hybrid Search — a bonus

A complementary improvement: combine semantic search (embeddings) with keyword search (BM25). Each has an advantage — semantic captures meaning, keyword captures exact terms, names and codes. You merge both candidate lists, then run a reranker over the union. This gives the best of both worlds, and the gain is largest for technical terms and for languages the embedding model saw less of during training.

Code example — a reranker in service

Reranking services (like Cohere Rerank) take a query and a list of documents and return them ranked. Open models also exist (BGE-reranker) via Hugging Face.

# Stage 1: broad retrieval from the Vector DB
candidates = vector_search(query, top_k=50)   # 50 candidates

# Stage 2: re-rank (example with a reranker library)
scored = reranker.rank(query=query,
                       documents=[c.text for c in candidates])
top = sorted(scored, key=lambda x: x.score, reverse=True)[:5]

# Stage 3: feed the LLM only the top 5
context = "\n\n".join(t.text for t in top)
answer = llm_answer(query, context)

What retrieval actually gets wrong

"Embeddings are less accurate" is true and not useful. The failures have a shape, and knowing it tells you when a reranker will help and when something else is broken.

An embedding compresses a passage into a fixed vector — a summary of what it is broadly about. Two things follow. It is excellent at topic and poor at specificity, and it has no notion of what the question was asking for.

So the characteristic mistakes are:

A cross-encoder fixes the first two, because it reads the query and the passage together and can tell that this passage is about the right thing but does not answer the question. It does not fix the third — that is what keyword search is for — and it does not fix the fourth, which is a chunking problem.

The reranker can only reorder what you gave it

The most important operational fact on this page, and the one that causes teams to conclude reranking "did not work".

Stage two is a reordering. If the passage that answers the question was not among the fifty candidates that came back from the vector search, no reranker can retrieve it. Your accuracy is capped by stage one's recall, permanently.

So measure that first. Take your evaluation questions, run only the retrieval stage, and ask: what fraction of the time is a correct passage anywhere in the top 50? That number is your ceiling.

This one measurement decides where your next week goes, and it takes an afternoon.

Why not just cross-encode everything?

The obvious question, and the answer explains the whole two-stage architecture.

A bi-encoder lets you precompute. Every passage in your corpus gets embedded once, in advance, and stored. At query time you embed one short query and do a nearest-neighbour lookup — an operation that stays fast whether the corpus holds ten thousand passages or ten million.

A cross-encoder cannot precompute anything, because the score depends on the query and the passage together. Scoring a corpus means one forward pass per passage, per query. That is fine for fifty candidates and impossible for a million.

Hence: a cheap approximate filter that scales, then an expensive accurate ranking on a bounded set. The same pattern appears throughout search systems, and once you see it the parameter choices stop being arbitrary — stage one is sized by your recall requirement, stage two by your latency budget.

Hybrid search, and how to merge two lists

Combining semantic search with keyword search (BM25) covers the identifier problem embeddings have. Keyword search is exact where vectors are fuzzy: product codes, error numbers, proper nouns, rare technical terms, and anything a user copied and pasted out of an error message.

The practical question is how to combine two ranked lists whose scores are not comparable — a cosine similarity and a BM25 score live on different scales, so adding them is meaningless.

The standard answer is reciprocal rank fusion: ignore the scores and use the positions. Each document gets a contribution of roughly one over its rank in each list, and the contributions are summed. A document ranked third by both methods beats one ranked first by one and fiftieth by the other. It needs no tuning, no score normalisation, and it is robust — which is why it is the default in most search stacks that offer hybrid retrieval.

Then rerank the fused list. Retrieval casts a wide net in two different ways; the cross-encoder decides what actually answers the question.

This matters more in languages other than English, and for any corpus full of names and codes — the two cases where pure semantic search disappoints most reliably.

Chunking decides how well ranking can work

Reranking gets blamed for problems that are really chunking problems, so it is worth being explicit about the interaction.

Chunks that are too large dilute both the embedding and the rerank score. A two-thousand-word section containing one relevant paragraph scores as mostly-irrelevant, and if it wins anyway you have spent context on 1,900 useless words.

Chunks that are too small score well and then fail to answer, because the sentence that matched has lost the context that made it meaningful — the pronoun whose referent was two paragraphs up, the number whose units were in the heading.

Two techniques address this directly. Retrieve small, return large: index small chunks for precise matching, but feed the model the surrounding section once a small chunk wins. And keep the heading path in the chunk text — prefixing each chunk with its document title and section headings gives both the embedder and the reranker the context a human would get from the page layout, and it is close to free.

Measuring it properly

"It seems better" is how teams ship a reranker that costs latency and adds nothing. The measurement is not difficult.

Build a set of questions with known good answers — fifty is enough to be useful, drawn from real user queries rather than invented ones. For each, mark which passages in your corpus genuinely answer it. That marking is the tedious part and it is the whole value.

Then track three numbers:

Keep the set and re-run it on every change — a new embedding model, a different chunk size, a reranker version. Retrieval systems are full of changes that improve one query visibly and degrade twenty invisibly, and the set is the only defence against that. See evals for the wider practice.

Latency, and where the time goes

A reranker adds a step, and whether that matters depends on where it sits in the user's experience.

In a system that then calls an LLM to generate an answer, the generation usually dominates — so the reranking step is often a small fraction of total response time, and the improved context can even make generation faster by shortening the prompt. In a search interface that returns links directly, the reranker is the latency, and the trade is real.

Three ways to buy it back. Reduce stage-one k — going from 100 candidates to 40 roughly halves the reranking work, and if your recall measurement says 40 is enough, the accuracy cost is nothing. Batch the candidates in one call rather than scoring them one at a time, which most APIs and libraries support and which is a large difference. And cache — repeated queries are common in support and documentation search, and the ranking for an identical query over unchanged content is reusable.

The cost is usually negative

Worth stating plainly because it is counter-intuitive: adding a reranker frequently makes the system cheaper.

Reranking is priced per document scored and is cheap — these are small models. The LLM call is the expensive part, and it is priced per token. If reranking lets you send five well-chosen passages instead of twenty mediocre ones, you have cut the largest cost in the pipeline by a large fraction, and paid a small amount for the privilege.

The secondary saving is answer quality: fewer irrelevant passages means less for the model to be distracted by, which reduces both wrong answers and the retries that follow them.

Choosing a reranker

The options divide into hosted APIs and open models you run yourself, and the decision is the same one that appears everywhere else in this stack.

A hosted reranking API is a single call, needs no infrastructure, and is priced per document scored. For most teams this is the right starting point: the quality is good, the integration is an afternoon, and the cost is small next to the LLM call it is saving you money on.

An open cross-encoder run on your own hardware costs nothing per query, keeps your documents inside your network, and can be fine-tuned on your own domain — which is where it beats the hosted option outright, because a reranker trained on your queries and your content understands what relevance means in your context. The price is a GPU and someone to look after it.

Whichever you pick, three things to check before committing:

The whole pipeline, in order

Putting the pieces together, a retrieval stack that works tends to look like this — and the order matters, because each step narrows what the next one has to handle.

  1. Rewrite the query into a standalone question, resolving pronouns and conversation context.
  2. Filter by metadata — product, language, date, tenant. Cheap, exact, and it removes whole categories of wrong answer.
  3. Retrieve twice, semantically and by keyword, and fuse the two lists by rank.
  4. Rerank the fused candidates with a cross-encoder.
  5. Apply a score floor. Nothing above it means the honest answer is that you do not know.
  6. Expand the winners to their surrounding context if you indexed small chunks.
  7. Generate, with instructions to answer only from the supplied passages and to cite which one each claim came from.

Most systems that disappoint are missing steps one, two and five — the cheap ones. Reranking at step four gets the attention because it is the interesting part, but a pipeline with steps one to three done properly and no reranker usually beats one with a reranker and nothing else.

Two things to try before reranking

Reranking is the highest-return change for a RAG system that is basically working. If yours is not, these are usually the problem.

Metadata filtering. If your corpus spans products, dates, languages or customers, filtering to the right subset before the vector search beats any amount of clever ranking. Searching one product's documentation instead of all twelve removes the wrong-answer class entirely rather than ranking it lower. This is frequently a bigger win than reranking and almost nobody does it first.

Query rewriting. Real user queries are short, ambiguous, and full of pronouns referring to earlier conversation. Expanding "does it support that?" into a standalone question before embedding it fixes retrieval failures that no reranker can recover from, because the wrong candidates were fetched.

When reranking will not help

Tips

Next step

Reranking is part of advanced RAG. Go deeper, and measure the improvement with evals.