Skip to main content

Insight · October 4, 2026

Look it up,
or bake it in.

RAG gives a model facts it can look up. Fine-tuning changes how it behaves. The decision buyers keep getting wrong, and the rule that gets it right.

01 · The distinction

Two methods, two different problems.

RAG and fine-tuning solve different problems. RAG gives a model facts it can look up at answer time, retrieved from your own documents. Fine-tuning changes how the model behaves, its tone, its format, its grip on a task. Use RAG for knowledge that changes. Use fine-tuning for behavior that should hold. Most production systems use both.

The two get argued about as if you must choose one. You rarely do. They sit at different layers of the system. One decides what the model can know. The other decides how the model acts. Pick before you name the problem and you will spend a training budget solving something retrieval would have fixed for free. A team that fine-tunes to patch a wrong fact ships a model that is confidently wrong in a fresh way, and pays to retrain it the next time that fact moves.

02 · What RAG does

RAG hands the model notes at answer time.

Retrieval augmented generation keeps your knowledge outside the model. At answer time it searches that knowledge, pulls the few passages that match the question, and appends them to the prompt. The model reads them and answers. The documents are cut into chunks, turned into embeddings, and stored in a vector database so the search matches by meaning rather than exact words.

Anthropic notes that when the whole knowledge base fits inside roughly 200,000 tokens you can skip retrieval and put everything in the prompt. Past that, retrieval is what lets a model draw on a corpus far larger than any single prompt could hold. The fact never enters the weights. It is fetched, read, and discarded on the next question. For the mechanics end to end, see what RAG is.

The one line to keep

“RAG for facts. Fine-tuning for behavior. In production, usually both.”

03 · What fine-tuning does

Fine-tuning changes the model itself.

Fine-tuning trains the base model further on your own examples. It does not add a fact you can point to. It adjusts the weights so the model leans toward a behavior: a house tone, a strict output shape, a way of reasoning through a task it kept getting wrong. You are not giving it something to read. You are changing what it does by reflex.

That power has a price. The OpenAI Cookbook notes that fine-tuning to maximize performance on a task can take thousands of worked examples, and that it suits teaching a skill or a style rather than a fact. Build the dataset, run the training, evaluate the result. When the base model or the data moves, you run it again. The knowledge is now set in the weights, which is exactly why it is the wrong tool for a fact that will change next week.

The behaviors it is good at are the ones a prompt keeps failing to pin down. A strict output shape the model drifts away from after a few turns. A house voice that a style instruction only ever approximates. A narrow classification the base model fumbles because it has met the task too rarely. In each case you are not handing the model something to read. You are teaching it a reflex, and a reflex is cheaper to train once than to describe again in every prompt forever.

04 · The example that explains it

An exam explains the whole choice.

The OpenAI Cookbook frames it as two ways to pass an exam. Fine-tuning is studying for an exam a week away. The model learns the material in advance, and like a student it can forget a detail or misremember a fact by the time the question comes. Search, the retrieval behind RAG, is taking the exam with open notes. The answer sits in front of the model while it works, so it is far more likely to get the fact right.

Their guidance is blunt. Search is the better way to recall facts. Fine-tuning is the better way to teach a skill or a format. The two map cleanly onto the two problems, and once you hear the analogy you stop reaching for the wrong one.

Need the model to KNOW a thing that changes?        → RAG
Need it to BEHAVE a certain way, every time?        → Fine-tune
Need a fresh fact AND a fixed behavior at once?      → Both
Still failing on a plain, well written prompt?       → Neither yet

Name the problem · the tool picks itself

05 · The decision

A table you can hold in one hand.

Knowledge changes often

RAG

Prices, docs, and policies live outside the model and update the moment you change the document. No retraining.

Answers must cite a source

RAG

The retrieved passage is the citation. Every claim traces back to a page you can open.

Output must hold a format or tone

Fine-tuning

The behavior is trained into the weights, so it holds without the instruction being repeated in every prompt.

A narrow task the model handles poorly

Fine-tuning

Thousands of worked examples teach the pattern more reliably than a long prompt ever will.

A private corpus larger than the context window

RAG

Retrieval draws on a body of documents no single prompt could ever hold at once.

A fresh fact and a fixed behavior at once

Both

RAG supplies the fact that changes. Fine-tuning holds the behavior that should not.

06 · Cost, freshness, the paper trail

Three questions settle most of it.

Freshness. RAG updates the moment you change a document. A fine-tuned model cannot learn a new fact without a new training run, so a fact that moves is baked in stale until you retrain. If the answer depends on this quarter, retrieval is the only honest option.

Cost. Fine-tuning is an up front training cost plus maintenance every time the model or the data shifts. RAG moves the cost to answer time, in the retrieval step and a longer prompt. Retrieval quality is itself tunable. Anthropic reported that adding a short generated context to each chunk cut the share of failed retrievals by 35%, and by 67% once a reranking pass was stacked on top.

Trust. RAG leaves a paper trail. The passage it retrieved is the citation, so a reviewer can open the source and check the claim. A fine-tuned model answers with no pointer to where the answer came from. For regulated work, that paper trail often decides the question on its own.

07 · In production

The answer is usually both, in order.

A mature system fine-tunes for the house voice and the output shape, then layers RAG for the facts. The two do not compete. One sets how the model speaks, the other feeds it what to speak about. The model sounds like you and cites a real source in the same answer.

Reach for them in order of cost. Before either, exhaust the prompt. The OpenAI Cookbook showed a single prompting change, asking the model to reason step by step, lifting accuracy on a math benchmark from 18% to 79%. Prompting is free. Add RAG when the model lacks facts. Add fine-tuning when it lacks the right behavior. Running the training job first, on a problem a better prompt or a retrieval step would have solved, is how budgets disappear.

Closing

The question was never which one.

Name the problem first. A fact that changes is a retrieval problem. A behavior that drifts is a training problem. Once the problem has a name, the tool picks itself, and the teams that get this wrong are almost always answering a question nobody asked.

Donny Smith

Written by · October 4, 2026

· ECD, Founder, Bttr.

Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.

LinkedIn

Share this perspective

More insights

Adjacent perspectives.

What Is Reranking

8 min read

What Is Reranking

A retrieval system usually finds the right passage and then leaves it in the wrong place. The first stage is a bi encoder that turns the query and every document into separate vectors and compares them, which is fast but never reads the two together, so the best passage often lands in the middle of the list. Reranking is the second pass that fixes the order. A cross encoder reads the query and one candidate as a single input, scores how they actually relate, and reorders the short list so the most relevant passages rise to the top. It is slower because it computes a fresh score for every query and document pair, so it runs on about a hundred candidates, never the whole corpus. On the public MS MARCO benchmark the sentence transformers cross encoder models trade speed for accuracy across sizes, from a tiny model near 9,000 documents a second to larger ones a fraction of that. Anthropic measured the payoff on a real pipeline: retrieve 150 chunks, rerank, keep the top 20, and the failure rate fell 67%. Retrieval finds the passage. Reranking makes sure the model reads it first.

What Is Chunking in RAG

8 min read

What Is Chunking in RAG

A retrieval system never reads your whole document. It reads the passages you cut it into, and chunking is that cut. Split too coarse and one chunk carries three ideas, so its embedding blurs and the right query misses. Split too fine and a passage loses the context that made it mean anything, the way a line reading the company revenue grew by 3% no longer says which company or which quarter. There is no universal size. LangChain base text splitter defaults to 4,000 characters, LlamaIndex sentence splitter to 1,024 tokens, because the right cut depends on your documents and your queries. Anthropic reported that adding a short generated context to each chunk before embedding cut the failure rate for the top 20 retrieved chunks by 35%, and by 67% once reranking was stacked on top. Retrieval quality does not start at the model. It starts at the cut.

What Is Semantic Search

8 min read

What Is Semantic Search

A search box that matches words cannot tell that refund and get my money back mean the same thing, because the two share no letters. Semantic search closes that gap. It turns every document and every query into an embedding, a list of numbers that places meaning at a point in a large space and puts similar meanings nearby, then returns the documents whose points sit closest, measured by cosine similarity, the angle between two vectors on a scale from minus one to one. The idea runs from word2vec in 2013 through Sentence BERT in 2019, and Google folded it into Search in October 2019 to better understand one in ten queries. Across billions of vectors it stays fast through approximate nearest neighbor search, most often the hierarchical navigable small world graph. It does not replace keyword search, it joins it, and the pairing most teams ship is hybrid search.

Bttr. Field Brief

The brief Bttr. writes for senior buyers.

Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.

Industries We Serve

Aerospace & DefenseBiotechnologyMedical & HealthcareManufacturingFinancial ServicesConsumer ProductsEnterprise Software

New Business

Start a project

Headquarters

North America

© 2026 Bttr. All rights reserved.