Skip to main content

Insight · October 2, 2026

The right passage
was already there.

Retrieval finds the right passage and then buries it in the middle of the pile. Reranking is the second pass that reads each candidate against the query and lifts the best one to the top.

01 · What it is

A first stage finds candidates. A second stage puts them in order.

Reranking is a second pass in a retrieval system. A first stage pulls a wide set of candidate passages fast, then a slower, more accurate model scores each one against the query and reorders them, so the passages most likely to answer sit at the top before anything reaches the model.

It exists because the fast first stage and a precise final order are two different jobs. A vector search is built to scan millions of passages in a few milliseconds. It is not built to tell the fourth best passage from the first. Reranking takes the short list the first stage returns and spends real compute to get the order right where it matters most, at the very top.

02 · Why the first pass misses

It compares the query and the document separately. The reranker reads them together.

The first stage is almost always a bi encoder. It turns the query into one vector and every document into its own vector, each computed on its own, ahead of time. At query time it just compares those finished vectors and returns the closest. That is what makes it fast, and it is also the ceiling. The two texts are never read side by side, so the match is a rough estimate of overlap rather than a real judgment of whether this passage answers this question.

A reranker is a cross encoder. The query and one candidate are passed to the model together as a single input, and it returns one relevance score. Because it runs attention across the query and the document at once, it can weigh how the two actually relate, not just whether their separate vectors happen to point the same way. That is why the cross encoder is the more accurate judge, and why it is worth running on the short list even though it is far too slow to run on the whole index.

The gap shows up most on the hard queries, the ones with a negation, a specific constraint, or two passages that share words but mean different things. Those are exactly the queries where the retrieved order is unreliable and the right answer is often sitting a few rows down. Reranking is the step that reaches down and pulls it up. The same discipline runs through semantic search and the vector database that feed it.

The one line to keep

“The first stage decides what the model could read. The reranker decides what it reads first.”

03 · How it runs

Four steps, from a wide net to a tight order.

01

Cast a wide net

A fast first stage, a vector search or a keyword index, pulls a large candidate set, often around a hundred passages. The goal here is recall, not order. Get the right passage somewhere in the pile, even if it lands near the bottom.

02

Score each pair

The reranker reads the query and one candidate together and returns a single relevance score. It does this for every candidate in the set, one pair at a time, which is why it is reserved for the short list rather than the whole corpus.

03

Reorder

Sort the candidates by that score. The passage the first stage left in the middle rises to the top, and the near misses that happened to share a few words sink.

04

Keep the top few

Pass only the best handful to the model. A tighter, better ordered context means the model reads less and answers from the passages most likely to hold the facts.

04 · The cost is real

You buy accuracy with throughput, one pair at a time.

A bi encoder can embed a document once and reuse that vector forever. A cross encoder cannot. It computes a fresh score for each query and document pair, so its work grows with the size of the candidate set, and it is slower than a bi encoder for the same reason it is more accurate: it reads the pair rather than two cached vectors. Scoring millions of pairs this way would be far too slow, which is the whole reason reranking runs on a short list and never the full corpus.

The size of the reranker is a dial, not a fixed choice. On the public MS MARCO benchmark the sentence transformers cross encoder models trace a clean line from fast to accurate. A tiny two layer model scores about 9,000 documents a second at an NDCG at 10 of 69.84. The six layer MiniLM reaches 74.30, the better order, at around 1,800 documents a second. A base size model lands near 340 a second. Same job, an order of magnitude of speed traded for a few points of ranking quality.

That dial is the real decision. Rerank too many candidates with too large a model and you add latency a user feels. Rerank a sensible short list with a model sized to your traffic and the added delay stays small while the order gets markedly better. The number of candidates and the model size are the two knobs, and both cost time linearly.

05 · One result, measured

Retrieve a hundred and fifty. Rerank. Keep twenty.

Anthropic published a retrieval study in September 2024 that put a number on the step. The pipeline retrieved 150 chunks for each query, reranked them, and passed only the top 20 to the model. Adding the reranking pass on top of an already strong retrieval setup reduced the top 20 failure rate by 67%, from 5.7% down to 1.9%. The reranker did not find anything new. It reordered what was already retrieved so the right chunks landed inside the top twenty.

The idea is older than the current wave. The cross encoder as a reranker was set out in 2019, when a BERT model was used to reorder the top 1,000 passages that BM25 had retrieved for each query, lifting the MS MARCO development score to an MRR at 10 of 36.53 from the high twenties of the prior leaders. Retrieve wide with something cheap, reorder with something precise. The models have changed. The shape has not. It is the same pattern that retrieval augmented generation leans on today.

cross encoder reranker · MS MARCO · NDCG@10 · docs/sec
  ms-marco-TinyBERT-L2-v2    69.84    ~9000   fastest
  ms-marco-MiniLM-L6-v2      74.30    ~1800   common default
  ms-marco-electra-base      71.99     ~340   heaviest

one dial · more layers buy a better order and cost throughput

Public sentence transformers cross encoder models · accuracy and speed move in opposite directions

06 · When to reach for it

It earns its place where the order at the top decides the answer.

Reranking pays when a model reads only the first few passages and the quality of the answer depends on getting those few right. A support assistant, a search over a large knowledge base, a question with a precise constraint buried in a long document. Anywhere the first stage returns a plausible but loosely ordered list, a reranker turns a wide recall into a clean order.

It is wasted effort in the other cases. If the corpus is small and the first stage already returns the right passage at the top, the reorder changes nothing. If the query is an exact lookup by id or by title, there is nothing to judge. And if the latency budget cannot absorb another model call, the honest answer is to skip it or shrink it rather than ship a slower product for a gain users will not feel.

The sound way to decide is to measure, not to assume. Turn reranking on over a sample of real queries, look at how far the right passage moved and what the extra step cost in time, and keep it only where the numbers reward it. The step that precisely orders a messy pile is only worth paying for when the pile is actually messy. Getting the pieces right before the cut, through sound chunking, decides how much work is left for the reranker to do.

Closing

Finding it was never the hard part.

Take one query your system answers weakly and read the passages it retrieved. The right one is usually in the list, a few rows from the top. Rerank that list and watch it rise. The answer was in reach the whole time. It was sitting in the wrong order.

Donny Smith

Written by · October 2, 2026

· ECD, Founder, Bttr.

Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.

LinkedIn

Share this perspective

More insights

Adjacent perspectives.

What Is Chunking in RAG

8 min read

What Is Chunking in RAG

A retrieval system never reads your whole document. It reads the passages you cut it into, and chunking is that cut. Split too coarse and one chunk carries three ideas, so its embedding blurs and the right query misses. Split too fine and a passage loses the context that made it mean anything, the way a line reading the company revenue grew by 3% no longer says which company or which quarter. There is no universal size. LangChain base text splitter defaults to 4,000 characters, LlamaIndex sentence splitter to 1,024 tokens, because the right cut depends on your documents and your queries. Anthropic reported that adding a short generated context to each chunk before embedding cut the failure rate for the top 20 retrieved chunks by 35%, and by 67% once reranking was stacked on top. Retrieval quality does not start at the model. It starts at the cut.

What Is Semantic Search

8 min read

What Is Semantic Search

A search box that matches words cannot tell that refund and get my money back mean the same thing, because the two share no letters. Semantic search closes that gap. It turns every document and every query into an embedding, a list of numbers that places meaning at a point in a large space and puts similar meanings nearby, then returns the documents whose points sit closest, measured by cosine similarity, the angle between two vectors on a scale from minus one to one. The idea runs from word2vec in 2013 through Sentence BERT in 2019, and Google folded it into Search in October 2019 to better understand one in ten queries. Across billions of vectors it stays fast through approximate nearest neighbor search, most often the hierarchical navigable small world graph. It does not replace keyword search, it joins it, and the pairing most teams ship is hybrid search.

What Is Spec-Driven Development

8 min read

What Is Spec-Driven Development

A model writes plausible code from a loose prompt in seconds, then the code drifts from what you meant because the only record of intent was a chat that scrolled away. Spec-driven development is the discipline that grew up around that failure. You write a precise specification first, and you treat the code as something the spec generates rather than the thing you argue about. GitHub frames its Spec Kit toolkit around one line, define what and why before deciding how to build it, and states the inversion plainly: specifications do not serve code, code serves specifications. The workflow is four moves, specify, plan, tasks, implement, with the spec kept as a plain file in the repo rather than a throwaway chat. OpenSpec keeps its specs in Markdown and reports working with more than thirty AI assistants, and Amazon Kiro writes requirements in EARS notation so every line is one condition and one behavior a machine can test. Review moves to the small artifact. You read a page instead of a pull request.

Bttr. Field Brief

The brief Bttr. writes for senior buyers.

Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.

Industries We Serve

Aerospace & DefenseBiotechnologyMedical & HealthcareManufacturingFinancial ServicesConsumer ProductsEnterprise Software

New Business

Start a project

Headquarters

North America

© 2026 Bttr. All rights reserved.