Skip to main content

January 2026

The Simulation
Worldview

Every Silicon Valley startup pitches a simulated, modeled future. The map has replaced the territory.

The Pitch

Every pitch deck tells the same story. We will model your business. Predict your customers. Optimize your operations. Simulate every outcome before it happens.

94%of enterprise software promises "predictive insights"
78%of AI startups claim to "model reality"

The assumption runs deep. If we can measure it, we can model it. If we can model it, we can predict it. If we can predict it, we can control it.

The map is not the territory. The model is not the system.

The Failures

0major recessions predicted by economic models
7 daysmaximum reliable weather prediction window
0cultural shifts predicted by social models
$2.3Tlost to model-based financial strategies in 2008

The world resists modeling. It resists prediction. It resists the clean lines of a simulation. The best weather models fail a week out. Economic models miss every recession. Social models fail to anticipate any cultural shift that matters.

And yet we keep building products that assume the model IS the reality. That treat the dashboard as truth. That mistake correlation for causation. That confuse precision with accuracy.

The Problem

The simulation worldview creates products that work brilliantly in demos and fail catastrophically in reality. They optimize for the modeled world, not the actual one.

Every model is a simplification. Every simplification is a loss. Every loss matters in ways we cannot predict.

The hubris is not in building models. Models are useful. The hubris is in forgetting they are incomplete. In treating them as oracles rather than approximations.

Data visualization dashboard

The Pattern

The same story plays out across industries. A model works in testing. It gets deployed. Reality diverges from the model. The system fails. Everyone is surprised.

Financial Models

Black-Scholes assumed markets were rational. They built trillion-dollar positions on this assumption. When markets stopped being rational, they collapsed.

Recommendation Engines

Optimized for engagement metrics. Created filter bubbles, radicalization pipelines, and mental health crises. The model worked. The outcome was disaster.

Self-Driving Cars

Trained on millions of miles. Still fail at scenarios no simulation anticipated. An edge case to a model is just Tuesday to reality.

Pandemic Models

Predicted curves and peaks with precision. Missed human behavior, political will, supply chains, and everything else that mattered.

The simulation worldview isn't wrong because simulations are useless. It's wrong because it forgets they're incomplete.

Build for variance.
Not just means.

What Good Looks Like

The best products know what they do not know. They build for the unexpected. They design for failure modes, not just success paths. They respect the complexity they cannot capture.

They treat models as tools, not truths. As starting points, not destinations. As maps that are always, necessarily, incomplete.

We believe:

  • 1Models are hypotheses. They should be tested, not trusted. Validated, not venerated.
  • 2Edge cases are not edge cases. They are the moments when your product meets reality.
  • 3Graceful degradation is required. When the model fails, the product should not.
  • 4Human override is non-negotiable. No system should be trusted more than the people using it.

What We Build

Robust Systems

Products that work when assumptions fail. Designed for the unexpected, not just the modeled.

Transparent Predictions

AI that shows its uncertainty. Models that communicate their limitations.

Human-Centered Automation

Systems that augment human judgment, not replace it. Tools, not oracles.

Reality-Tested Products

Built for actual conditions, not simulated ones. Validated in the wild.

Is your product designed for the model?

Or for reality?

Donny Smith

Written by · January 6, 2026

· ECD, Founder, Bttr.

Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.

LinkedIn

Share this perspective

More Insights

Perspectives on product design, technology, and the forces shaping how people interact with the things we build.

RAG vs Fine-Tuning
9 min read

RAG vs Fine-Tuning

The teams that waste the most on AI pick RAG or fine-tuning before they name the problem. The two are not answers to the same question. RAG gives a model facts it can look up at answer time, retrieved from your own documents, cut into chunks, turned into embeddings, and searched by meaning. The fact never enters the weights. Anthropic notes that once a knowledge base passes roughly 200,000 tokens, retrieval is what lets a model draw on a corpus no prompt could hold. Fine-tuning does the opposite. It trains the base model on your examples until a tone, a format, or a task is set into the weights, which the OpenAI Cookbook notes can take thousands of worked examples and suits teaching a skill rather than a fact. The clean rule is RAG for facts, fine-tuning for behavior, and in production usually both. The Cookbook frames it as an exam: fine-tuning is studying a week ahead and risking a forgotten detail, RAG is the open notes exam with the answer in front of the model while it works. Freshness, cost, and the paper trail settle the rest. A fact that changes is a retrieval problem. A behavior that drifts is a training problem. Name the problem first and the tool picks itself.

What Is Reranking
8 min read

What Is Reranking

A retrieval system usually finds the right passage and then leaves it in the wrong place. The first stage is a bi encoder that turns the query and every document into separate vectors and compares them, which is fast but never reads the two together, so the best passage often lands in the middle of the list. Reranking is the second pass that fixes the order. A cross encoder reads the query and one candidate as a single input, scores how they actually relate, and reorders the short list so the most relevant passages rise to the top. It is slower because it computes a fresh score for every query and document pair, so it runs on about a hundred candidates, never the whole corpus. On the public MS MARCO benchmark the sentence transformers cross encoder models trade speed for accuracy across sizes, from a tiny model near 9,000 documents a second to larger ones a fraction of that. Anthropic measured the payoff on a real pipeline: retrieve 150 chunks, rerank, keep the top 20, and the failure rate fell 67%. Retrieval finds the passage. Reranking makes sure the model reads it first.

What Is Chunking in RAG
8 min read

What Is Chunking in RAG

A retrieval system never reads your whole document. It reads the passages you cut it into, and chunking is that cut. Split too coarse and one chunk carries three ideas, so its embedding blurs and the right query misses. Split too fine and a passage loses the context that made it mean anything, the way a line reading the company revenue grew by 3% no longer says which company or which quarter. There is no universal size. LangChain base text splitter defaults to 4,000 characters, LlamaIndex sentence splitter to 1,024 tokens, because the right cut depends on your documents and your queries. Anthropic reported that adding a short generated context to each chunk before embedding cut the failure rate for the top 20 retrieved chunks by 35%, and by 67% once reranking was stacked on top. Retrieval quality does not start at the model. It starts at the cut.

Industries We Serve

Aerospace & DefenseBiotechnologyMedical & HealthcareManufacturingFinancial ServicesConsumer ProductsEnterprise Software

New Business

Start a project

Headquarters

North America

© 2026 Bttr. All rights reserved.