Skip to main content

Insight · September 30, 2026

The cut
decides the answer.

A retrieval system never reads your whole document. It reads the pieces you cut it into. Chunking is that cut, and it sets the ceiling on what the model can find.

01 · What it is

A model retrieves passages, not documents.

Chunking is the step in a retrieval system that splits a long document into smaller passages before each one is turned into an embedding and stored. The system retrieves passages, not whole files, so where you cut the text sets the ceiling on what it can find. Size and boundaries are a real tuning decision.

It sits quietly in front of everything else. Before a query is ever embedded, before a vector database is ever searched, the source has already been carved into pieces. Every later stage can only work with the pieces it was handed. Get the cut wrong and no embedding model, no reranker, and no larger context window fully recovers it.

02 · Why it matters

Cut too coarse and the meaning blurs. Cut too fine and it leaks out.

An embedding is one point that stands for the whole passage. Pack three separate ideas into a single chunk and that point lands in the blurred average of all three, close to nothing in particular, so the query that wanted one of them retrieves a weak match or misses entirely. The chunk was too coarse.

Cut the other way and the opposite failure appears. A passage trimmed to one sentence can lose the context that made it mean anything. Anthropic gives the cleanest example. A chunk pulled from a filing that reads, the company's revenue grew by 3% over the previous quarter, no longer says which company or which quarter. It reads fine and retrieves badly, because the words a searcher would use are the very words the cut removed.

This is why practitioners often say the splitter matters more than the embedding model. Both failures are decided before the model is ever called. Overlap, letting each chunk repeat the last stretch of the one before it, softens the second failure but does not remove the underlying choice of where to draw the line.

The one line to keep

“The best embedding model cannot retrieve a passage the chunker split in half.”

03 · The methods

Four ways to cut, from blunt to aware.

01

Fixed size

Cut every so many tokens or characters, with a small overlap so a sentence split across the seam still appears whole in one piece. Fast, cheap, and blind to where meaning actually ends.

02

Recursive

Try to break on the largest natural unit first, a paragraph, then a line, then a space, and only fall to a raw cut when nothing else fits. The default in most tooling because it respects structure at almost no cost.

03

Document structure

Split on the markup the document already carries. Headings, sections, table rows, and code blocks become the boundaries, so each piece is a unit the author meant to stand together.

04

Semantic

Embed the sentences first and cut where the meaning shifts, not at a fixed count. Higher retrieval quality in the right corpus, at the cost of running a model just to decide the boundaries.

04 · There is no default that fits everything

Two mature tools ship two different defaults.

If there were one correct chunk size, the popular libraries would agree on it. They do not. LangChain's base text splitter defaults to 4,000 characters with an overlap of 200, and its recursive splitter tries the separators paragraph, line, space, then a raw character split, in that order. LlamaIndex's sentence splitter defaults to 1,024 tokens with an overlap of 200. Different units, different sizes, same job.

The gap is not a bug. It reflects a real truth: the right cut depends on the shape of your documents and the shape of your queries. Dense reference text with self contained paragraphs wants smaller pieces. Narrative or legal text where meaning carries across pages wants larger ones, or a structure aware cut. The defaults are a starting line, not an answer.

Chroma's open retrieval study made the size question measurable. Evaluating recall at the level of individual tokens, it found the boundary choice alone moved recall by close to nine points between two common strategies. The lesson was not pick this number. It was measure your own.

05 · One fix, measured

Give each piece back the context the cut removed.

In September 2024 Anthropic published a method that treats the lost context problem head on. Before a chunk is embedded, a model writes one short line that situates it inside the whole document, which company, which period, which section, and prepends that line to the chunk. The passage that used to read only, revenue grew by 3%, now carries the name and the quarter with it.

The results were specific. Adding this generated context before embedding reduced the failure rate for the top 20 retrieved chunks by 35%, from 5.7% to 3.7%. Pairing it with a keyword index took the reduction to 49%. Adding a reranking pass on top took it to 67%, down to a 1.9% failure rate. With prompt caching, the one time cost to generate the context came to $1.02 per million document tokens.

raw chunk
  "The company's revenue grew by 3% over the
   previous quarter."
  · retrieves badly · no company, no period

contextual chunk
  "From Acme's Q2 2023 filing, the section on
   quarterly results. The company's revenue grew
   by 3% over the previous quarter."
  · retrieves on the terms a searcher would type

Pattern from Anthropic's contextual retrieval · the added line carries the context the cut removed

06 · The smarter cut costs more

Semantic chunking can win on recall. It also runs a model to decide.

The appeal of semantic chunking is obvious. Instead of cutting at a fixed count, it embeds the sentences and draws the line where the topic shifts, so each piece is about one thing. In Chroma's open evaluation, a cluster based semantic chunker held to a small maximum reached a mean recall of 0.897, against 0.809 for a fixed 1,200 character cut, close to nine points of recall from the boundary choice alone.

The cost is that you now run an embedding model just to place the cuts, on top of embedding the chunks themselves, and you add a tuning knob that can misfire on text without clear topic seams. That is why recursive character splitting stays the common default even in 2026. It captures most of the structure for almost none of the cost, and it fails in predictable ways.

The honest read is that no single method wins everywhere. The right cut is the one your own retrieval numbers reward, found by testing on your documents rather than adopted from a blog post.

07 · Where to start

Start plain, measure, then earn the complexity.

Begin with a recursive splitter at a moderate size with a small overlap. It is the cheapest thing that respects structure, and it gives you a baseline to beat. Then look at the failures. When a retrieved chunk is right but missing its subject, add context to each piece before embedding. When chunks carry several ideas at once, cut smaller or cut on structure.

Reach for semantic chunking only after the simpler moves stop paying, and only where your evaluation shows it helps. Every added method is another model call, another parameter, another way to be wrong at scale. The discipline is not finding the cleverest cut. It is measuring retrieval on your own content and letting the numbers decide how much cleverness the cut has earned.

Closing

Retrieval was always decided at the cut.

Take one query your system gets wrong and read the chunk it pulled. Nine times in ten the fault is in the boundary, not the model. Fix where the text was cut, and watch the right passage rise to the top.

Donny Smith

Written by · September 30, 2026

· ECD, Founder, Bttr.

Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.

LinkedIn

Share this perspective

More insights

Adjacent perspectives.

What Is Semantic Search

8 min read

What Is Semantic Search

A search box that matches words cannot tell that refund and get my money back mean the same thing, because the two share no letters. Semantic search closes that gap. It turns every document and every query into an embedding, a list of numbers that places meaning at a point in a large space and puts similar meanings nearby, then returns the documents whose points sit closest, measured by cosine similarity, the angle between two vectors on a scale from minus one to one. The idea runs from word2vec in 2013 through Sentence BERT in 2019, and Google folded it into Search in October 2019 to better understand one in ten queries. Across billions of vectors it stays fast through approximate nearest neighbor search, most often the hierarchical navigable small world graph. It does not replace keyword search, it joins it, and the pairing most teams ship is hybrid search.

What Is Spec-Driven Development

8 min read

What Is Spec-Driven Development

A model writes plausible code from a loose prompt in seconds, then the code drifts from what you meant because the only record of intent was a chat that scrolled away. Spec-driven development is the discipline that grew up around that failure. You write a precise specification first, and you treat the code as something the spec generates rather than the thing you argue about. GitHub frames its Spec Kit toolkit around one line, define what and why before deciding how to build it, and states the inversion plainly: specifications do not serve code, code serves specifications. The workflow is four moves, specify, plan, tasks, implement, with the spec kept as a plain file in the repo rather than a throwaway chat. OpenSpec keeps its specs in Markdown and reports working with more than thirty AI assistants, and Amazon Kiro writes requirements in EARS notation so every line is one condition and one behavior a machine can test. Review moves to the small artifact. You read a page instead of a pull request.

What Are Embeddings

7 min read

What Are Embeddings

A computer cannot read. It matches characters, so refund and get my money back look like strangers even though they mean the same thing. An embedding closes that gap. It turns a piece of text into a list of numbers that places its meaning at a single point in a large space, and puts text that means something similar nearby. OpenAI’s text-embedding-3-small returns 1,536 numbers per input and its larger model returns 3,072, and no one sets them by hand. The idea is older than the chatbot: Google published word2vec in 2013 and Stanford followed with GloVe in 2014, both trained on raw text with nothing labeled. Once meaning is a location, closeness becomes a number you can compute with cosine similarity, and that one move is the layer under semantic search, retrieval augmented generation, and recommendations. A token is not an embedding. Tokenizing tells you which pieces you have, embedding tells you what they mean.

Bttr. Field Brief

The brief Bttr. writes for senior buyers.

Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.

Industries We Serve

Aerospace & DefenseBiotechnologyMedical & HealthcareManufacturingFinancial ServicesConsumer ProductsEnterprise Software

New Business

Start a project

Headquarters

North America

© 2026 Bttr. All rights reserved.