Insight · October 2, 2026
The right passage
was already there.
Retrieval finds the right passage and then buries it in the middle of the pile. Reranking is the second pass that reads each candidate against the query and lifts the best one to the top.
01 · What it is
A first stage finds candidates. A second stage puts them in order.
Reranking is a second pass in a retrieval system. A first stage pulls a wide set of candidate passages fast, then a slower, more accurate model scores each one against the query and reorders them, so the passages most likely to answer sit at the top before anything reaches the model.
It exists because the fast first stage and a precise final order are two different jobs. A vector search is built to scan millions of passages in a few milliseconds. It is not built to tell the fourth best passage from the first. Reranking takes the short list the first stage returns and spends real compute to get the order right where it matters most, at the very top.
02 · Why the first pass misses
It compares the query and the document separately. The reranker reads them together.
The first stage is almost always a bi encoder. It turns the query into one vector and every document into its own vector, each computed on its own, ahead of time. At query time it just compares those finished vectors and returns the closest. That is what makes it fast, and it is also the ceiling. The two texts are never read side by side, so the match is a rough estimate of overlap rather than a real judgment of whether this passage answers this question.
A reranker is a cross encoder. The query and one candidate are passed to the model together as a single input, and it returns one relevance score. Because it runs attention across the query and the document at once, it can weigh how the two actually relate, not just whether their separate vectors happen to point the same way. That is why the cross encoder is the more accurate judge, and why it is worth running on the short list even though it is far too slow to run on the whole index.
The gap shows up most on the hard queries, the ones with a negation, a specific constraint, or two passages that share words but mean different things. Those are exactly the queries where the retrieved order is unreliable and the right answer is often sitting a few rows down. Reranking is the step that reaches down and pulls it up. The same discipline runs through semantic search and the vector database that feed it.
The one line to keep
“The first stage decides what the model could read. The reranker decides what it reads first.”
03 · How it runs
Four steps, from a wide net to a tight order.
01
Cast a wide net
A fast first stage, a vector search or a keyword index, pulls a large candidate set, often around a hundred passages. The goal here is recall, not order. Get the right passage somewhere in the pile, even if it lands near the bottom.
02
Score each pair
The reranker reads the query and one candidate together and returns a single relevance score. It does this for every candidate in the set, one pair at a time, which is why it is reserved for the short list rather than the whole corpus.
03
Reorder
Sort the candidates by that score. The passage the first stage left in the middle rises to the top, and the near misses that happened to share a few words sink.
04
Keep the top few
Pass only the best handful to the model. A tighter, better ordered context means the model reads less and answers from the passages most likely to hold the facts.
04 · The cost is real
You buy accuracy with throughput, one pair at a time.
A bi encoder can embed a document once and reuse that vector forever. A cross encoder cannot. It computes a fresh score for each query and document pair, so its work grows with the size of the candidate set, and it is slower than a bi encoder for the same reason it is more accurate: it reads the pair rather than two cached vectors. Scoring millions of pairs this way would be far too slow, which is the whole reason reranking runs on a short list and never the full corpus.
The size of the reranker is a dial, not a fixed choice. On the public MS MARCO benchmark the sentence transformers cross encoder models trace a clean line from fast to accurate. A tiny two layer model scores about 9,000 documents a second at an NDCG at 10 of 69.84. The six layer MiniLM reaches 74.30, the better order, at around 1,800 documents a second. A base size model lands near 340 a second. Same job, an order of magnitude of speed traded for a few points of ranking quality.
That dial is the real decision. Rerank too many candidates with too large a model and you add latency a user feels. Rerank a sensible short list with a model sized to your traffic and the added delay stays small while the order gets markedly better. The number of candidates and the model size are the two knobs, and both cost time linearly.
05 · One result, measured
Retrieve a hundred and fifty. Rerank. Keep twenty.
Anthropic published a retrieval study in September 2024 that put a number on the step. The pipeline retrieved 150 chunks for each query, reranked them, and passed only the top 20 to the model. Adding the reranking pass on top of an already strong retrieval setup reduced the top 20 failure rate by 67%, from 5.7% down to 1.9%. The reranker did not find anything new. It reordered what was already retrieved so the right chunks landed inside the top twenty.
The idea is older than the current wave. The cross encoder as a reranker was set out in 2019, when a BERT model was used to reorder the top 1,000 passages that BM25 had retrieved for each query, lifting the MS MARCO development score to an MRR at 10 of 36.53 from the high twenties of the prior leaders. Retrieve wide with something cheap, reorder with something precise. The models have changed. The shape has not. It is the same pattern that retrieval augmented generation leans on today.
cross encoder reranker · MS MARCO · NDCG@10 · docs/sec ms-marco-TinyBERT-L2-v2 69.84 ~9000 fastest ms-marco-MiniLM-L6-v2 74.30 ~1800 common default ms-marco-electra-base 71.99 ~340 heaviest one dial · more layers buy a better order and cost throughput
Public sentence transformers cross encoder models · accuracy and speed move in opposite directions
06 · When to reach for it
It earns its place where the order at the top decides the answer.
Reranking pays when a model reads only the first few passages and the quality of the answer depends on getting those few right. A support assistant, a search over a large knowledge base, a question with a precise constraint buried in a long document. Anywhere the first stage returns a plausible but loosely ordered list, a reranker turns a wide recall into a clean order.
It is wasted effort in the other cases. If the corpus is small and the first stage already returns the right passage at the top, the reorder changes nothing. If the query is an exact lookup by id or by title, there is nothing to judge. And if the latency budget cannot absorb another model call, the honest answer is to skip it or shrink it rather than ship a slower product for a gain users will not feel.
The sound way to decide is to measure, not to assume. Turn reranking on over a sample of real queries, look at how far the right passage moved and what the extra step cost in time, and keep it only where the numbers reward it. The step that precisely orders a messy pile is only worth paying for when the pile is actually messy. Getting the pieces right before the cut, through sound chunking, decides how much work is left for the reranker to do.
Closing
Finding it was never the hard part.
Take one query your system answers weakly and read the passages it retrieved. The right one is usually in the list, a few rows from the top. Rerank that list and watch it rise. The answer was in reach the whole time. It was sitting in the wrong order.

Written by · October 2, 2026
Donny Smith · ECD, Founder, Bttr.
Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.
LinkedInShare this perspective
More insights
Adjacent perspectives.
Bttr. Field Brief
The brief Bttr. writes for senior buyers.
Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.