Skip to main content

Insight · September 28, 2026

Search by
meaning.

Keyword search matches the letters you typed. Semantic search matches what you meant. The difference is a number that places meaning in space.

01 · The problem

Keyword search cannot read. It matches characters.

A classic search box runs on an inverted index, a lookup table that maps every word to the list of documents that contain it. The engine answers a query by scoring the documents that share your exact terms. The ranking function behind almost every one of these boxes is BM25, short for Best Match, a formula refined by Stephen Robertson and Karen Spärck Jones and still the standard inside Lucene, Elasticsearch, and OpenSearch. It is fast, it is precise, and it powers most of the search you use every day.

It has one blind spot. It matches words, not meaning. Type refund and a document that only ever says get my money back looks like a stranger, even though the two mean the same thing. This is the vocabulary problem, and no amount of tuning the keyword formula solves it, because the words never overlap.

02 · What changed

Meaning became a location you can measure.

Semantic search closes that gap by changing what it compares. Instead of matching words, it turns each piece of text into an embedding, a list of numbers that places the text at a single point in a large space, and puts text that means something similar nearby. The idea is older than the chatbot. Google published word2vec in 2013 and Stanford followed with GloVe in 2014, both learning these positions from raw text with nothing labeled by hand.

Once meaning is a position, similar things sit close together. Cheap lands near affordable. A question about canceling a plan lands near a document about ending a subscription, even with no shared words. The search stops asking which documents contain these words and starts asking which documents mean this.

The one line to keep

“Keyword search finds the words you typed. Semantic search finds the thing you meant.”

03 · How it works

Four moves, whatever the stack.

01

Embed the corpus

Every document runs through an embedding model once and becomes a vector. The heavy work happens ahead of time, before anyone searches.

02

Store the vectors

The vectors go into an index built for similarity rather than exact match. This is the job a vector database exists to do.

03

Embed the query

At search time the query runs through the same model and lands as a point in the same space as the documents.

04

Rank by closeness

The engine returns the stored vectors sitting nearest the query, measured by the angle between them, and hands back the documents they came from.

04 · Why it scales

Closeness is an angle, and you never check every point.

Two choices make this work at scale. The first is how closeness gets measured. Most systems use cosine similarity, which compares the direction two vectors point rather than their length. It runs from minus one to one, where one means the same direction and zero means unrelated. Direction carries the meaning, and length tends to carry noise like how much text went in, so measuring only the angle keeps the comparison clean.

The second is that you never compare the query against every vector. A real corpus holds millions or billions of them, and checking each one is too slow. Engines use approximate nearest neighbor search instead, which trades a sliver of accuracy for a large amount of speed. The common method is the hierarchical navigable small world graph, published by Malkov and Yashunin in 2016, which walks a layered graph of points to reach the closest ones in a few hops rather than a full scan.

05 · One example you already use

The search box in your pocket already runs on this.

You do not have to look far for a real one. In October 2019 Google brought a model called BERT into Search and said it would help understand one in ten searches in United States English, with a rollout to more than seventy languages by December. The gain showed up most on long, conversational queries where the small words carry the meaning, the exact case that trips a keyword match.

Sentence embeddings made this practical at scale. Sentence BERT, published by Reimers and Gurevych in 2019, reshaped the model so a whole sentence could become one vector meant for comparison, which is what turns understanding a single query into searching a corpus of millions. Understanding the question rather than matching its words is now simply how modern search behaves.

query   how do I get my money back

semantic search · ranked by meaning
  ·  Refunds and returns policy
  ·  Cancel an order before it ships
  ·  Dispute a charge on your account

keyword search · ranked by the word money
  ·  no match · not one document says money

An illustration · the query and the right documents share almost no words

06 · Not a replacement

It does not replace keyword search. It joins it.

Semantic search has its own blind spot, the mirror image of the first. It is built to look past exact words, so it can miss them when they matter, a part number, a name, an error code, a term of art that has to match precisely. The honest answer most teams reach is not one method or the other. It is both.

That combination is called hybrid search. It runs the keyword ranker and the vector ranker side by side, then merges the two result lists. The common way to merge them is reciprocal rank fusion, which scores each result by its position in each list rather than by a raw score. That matters because the two scores sit on different scales, cosine similarity runs zero to one and BM25 has no ceiling, and working from ranks sidesteps the mismatch.

Hybrid search is now the default retrieval strategy for systems that feed context to a model, and every major vector engine ships it. When people say a product has semantic search, they usually mean this.

07 · When it earns its cost

Reach for it when words fail, not by default.

Semantic search is not free. You run an embedding model over the whole corpus, you keep a second index in sync, and you pay to embed every query. On a small catalog where people search by exact names and codes, a keyword box is faster and cheaper and does the job.

It earns its cost when the words in the question do not match the words in the answer. Support content, long documents, natural questions, anything a model will read and reason over. When meaning matters more than spelling, semantic search is the layer that finds the right thing even when it shares no words with the query.

Closing

Search was always about meaning.

Pick one search your users run in their own words, the kind where they never type your exact terms. Put a semantic layer behind it and watch what surfaces. The first time it returns the right document that shares not one word with the query, you will stop counting on the words to match.

Google Search BERT announcement, October 2019 · Sentence BERT, Reimers and Gurevych, 2019 · Hierarchical Navigable Small World graphs, Malkov and Yashunin, 2016 · BM25, Robertson and Spärck Jones · retrieval practice current as of September 2026

Share this perspective

More insights

Adjacent perspectives.

What Is Spec-Driven Development

8 min read

What Is Spec-Driven Development

A model writes plausible code from a loose prompt in seconds, then the code drifts from what you meant because the only record of intent was a chat that scrolled away. Spec-driven development is the discipline that grew up around that failure. You write a precise specification first, and you treat the code as something the spec generates rather than the thing you argue about. GitHub frames its Spec Kit toolkit around one line, define what and why before deciding how to build it, and states the inversion plainly: specifications do not serve code, code serves specifications. The workflow is four moves, specify, plan, tasks, implement, with the spec kept as a plain file in the repo rather than a throwaway chat. OpenSpec keeps its specs in Markdown and reports working with more than thirty AI assistants, and Amazon Kiro writes requirements in EARS notation so every line is one condition and one behavior a machine can test. Review moves to the small artifact. You read a page instead of a pull request.

What Are Embeddings

7 min read

What Are Embeddings

A computer cannot read. It matches characters, so refund and get my money back look like strangers even though they mean the same thing. An embedding closes that gap. It turns a piece of text into a list of numbers that places its meaning at a single point in a large space, and puts text that means something similar nearby. OpenAI’s text-embedding-3-small returns 1,536 numbers per input and its larger model returns 3,072, and no one sets them by hand. The idea is older than the chatbot: Google published word2vec in 2013 and Stanford followed with GloVe in 2014, both trained on raw text with nothing labeled. Once meaning is a location, closeness becomes a number you can compute with cosine similarity, and that one move is the layer under semantic search, retrieval augmented generation, and recommendations. A token is not an embedding. Tokenizing tells you which pieces you have, embedding tells you what they mean.

What Is a Vector Database

8 min read

What Is a Vector Database

A relational database finds a row by matching it exactly. That breaks the moment the match is about meaning, because a refund article and a customer asking to get their money back share almost no words. A vector database closes that gap. It stores each document as an embedding, a list of numbers that places its meaning at a point in space, and answers a query by returning the records whose points sit closest, measured by cosine similarity, Euclidean distance, or dot product. To stay fast across billions of vectors it searches for the approximate nearest neighbor rather than the exact one, most often through the hierarchical navigable small world graph introduced by Malkov and Yashunin in 2016. This is the infrastructure under retrieval augmented generation and agent memory, and for most teams in 2026 pgvector on the Postgres they already run is the honest place to start.

Bttr. Field Brief

The brief Bttr. writes for senior buyers.

Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.

Industries We Serve

Aerospace & DefenseBiotechnologyMedical & HealthcareManufacturingFinancial ServicesConsumer ProductsEnterprise Software

New Business

Start a project

Headquarters

North America

© 2026 Bttr. All rights reserved.