Insight · September 30, 2026
The cut
decides the answer.
A retrieval system never reads your whole document. It reads the pieces you cut it into. Chunking is that cut, and it sets the ceiling on what the model can find.
01 · What it is
A model retrieves passages, not documents.
Chunking is the step in a retrieval system that splits a long document into smaller passages before each one is turned into an embedding and stored. The system retrieves passages, not whole files, so where you cut the text sets the ceiling on what it can find. Size and boundaries are a real tuning decision.
It sits quietly in front of everything else. Before a query is ever embedded, before a vector database is ever searched, the source has already been carved into pieces. Every later stage can only work with the pieces it was handed. Get the cut wrong and no embedding model, no reranker, and no larger context window fully recovers it.
02 · Why it matters
Cut too coarse and the meaning blurs. Cut too fine and it leaks out.
An embedding is one point that stands for the whole passage. Pack three separate ideas into a single chunk and that point lands in the blurred average of all three, close to nothing in particular, so the query that wanted one of them retrieves a weak match or misses entirely. The chunk was too coarse.
Cut the other way and the opposite failure appears. A passage trimmed to one sentence can lose the context that made it mean anything. Anthropic gives the cleanest example. A chunk pulled from a filing that reads, the company's revenue grew by 3% over the previous quarter, no longer says which company or which quarter. It reads fine and retrieves badly, because the words a searcher would use are the very words the cut removed.
This is why practitioners often say the splitter matters more than the embedding model. Both failures are decided before the model is ever called. Overlap, letting each chunk repeat the last stretch of the one before it, softens the second failure but does not remove the underlying choice of where to draw the line.
The one line to keep
“The best embedding model cannot retrieve a passage the chunker split in half.”
03 · The methods
Four ways to cut, from blunt to aware.
01
Fixed size
Cut every so many tokens or characters, with a small overlap so a sentence split across the seam still appears whole in one piece. Fast, cheap, and blind to where meaning actually ends.
02
Recursive
Try to break on the largest natural unit first, a paragraph, then a line, then a space, and only fall to a raw cut when nothing else fits. The default in most tooling because it respects structure at almost no cost.
03
Document structure
Split on the markup the document already carries. Headings, sections, table rows, and code blocks become the boundaries, so each piece is a unit the author meant to stand together.
04
Semantic
Embed the sentences first and cut where the meaning shifts, not at a fixed count. Higher retrieval quality in the right corpus, at the cost of running a model just to decide the boundaries.
04 · There is no default that fits everything
Two mature tools ship two different defaults.
If there were one correct chunk size, the popular libraries would agree on it. They do not. LangChain's base text splitter defaults to 4,000 characters with an overlap of 200, and its recursive splitter tries the separators paragraph, line, space, then a raw character split, in that order. LlamaIndex's sentence splitter defaults to 1,024 tokens with an overlap of 200. Different units, different sizes, same job.
The gap is not a bug. It reflects a real truth: the right cut depends on the shape of your documents and the shape of your queries. Dense reference text with self contained paragraphs wants smaller pieces. Narrative or legal text where meaning carries across pages wants larger ones, or a structure aware cut. The defaults are a starting line, not an answer.
Chroma's open retrieval study made the size question measurable. Evaluating recall at the level of individual tokens, it found the boundary choice alone moved recall by close to nine points between two common strategies. The lesson was not pick this number. It was measure your own.
05 · One fix, measured
Give each piece back the context the cut removed.
In September 2024 Anthropic published a method that treats the lost context problem head on. Before a chunk is embedded, a model writes one short line that situates it inside the whole document, which company, which period, which section, and prepends that line to the chunk. The passage that used to read only, revenue grew by 3%, now carries the name and the quarter with it.
The results were specific. Adding this generated context before embedding reduced the failure rate for the top 20 retrieved chunks by 35%, from 5.7% to 3.7%. Pairing it with a keyword index took the reduction to 49%. Adding a reranking pass on top took it to 67%, down to a 1.9% failure rate. With prompt caching, the one time cost to generate the context came to $1.02 per million document tokens.
raw chunk "The company's revenue grew by 3% over the previous quarter." · retrieves badly · no company, no period contextual chunk "From Acme's Q2 2023 filing, the section on quarterly results. The company's revenue grew by 3% over the previous quarter." · retrieves on the terms a searcher would type
Pattern from Anthropic's contextual retrieval · the added line carries the context the cut removed
06 · The smarter cut costs more
Semantic chunking can win on recall. It also runs a model to decide.
The appeal of semantic chunking is obvious. Instead of cutting at a fixed count, it embeds the sentences and draws the line where the topic shifts, so each piece is about one thing. In Chroma's open evaluation, a cluster based semantic chunker held to a small maximum reached a mean recall of 0.897, against 0.809 for a fixed 1,200 character cut, close to nine points of recall from the boundary choice alone.
The cost is that you now run an embedding model just to place the cuts, on top of embedding the chunks themselves, and you add a tuning knob that can misfire on text without clear topic seams. That is why recursive character splitting stays the common default even in 2026. It captures most of the structure for almost none of the cost, and it fails in predictable ways.
The honest read is that no single method wins everywhere. The right cut is the one your own retrieval numbers reward, found by testing on your documents rather than adopted from a blog post.
07 · Where to start
Start plain, measure, then earn the complexity.
Begin with a recursive splitter at a moderate size with a small overlap. It is the cheapest thing that respects structure, and it gives you a baseline to beat. Then look at the failures. When a retrieved chunk is right but missing its subject, add context to each piece before embedding. When chunks carry several ideas at once, cut smaller or cut on structure.
Reach for semantic chunking only after the simpler moves stop paying, and only where your evaluation shows it helps. Every added method is another model call, another parameter, another way to be wrong at scale. The discipline is not finding the cleverest cut. It is measuring retrieval on your own content and letting the numbers decide how much cleverness the cut has earned.
Closing
Retrieval was always decided at the cut.
Take one query your system gets wrong and read the chunk it pulled. Nine times in ten the fault is in the boundary, not the model. Fix where the text was cut, and watch the right passage rise to the top.

Written by · September 30, 2026
Donny Smith · ECD, Founder, Bttr.
Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.
LinkedInShare this perspective
More insights
Adjacent perspectives.
Bttr. Field Brief
The brief Bttr. writes for senior buyers.
Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.