Skip to main content
Code on a dark screen

January 12, 2026

The Dumbest Smart Way
to Ship Code

How a bash while loop named after a Simpsons character is letting developers ship $50K projects for $300 in API costs.

The Technique

It's called the Ralph Loop. Named after Ralph Wiggum from The Simpsons· the character known for being simple, earnest, and persistently trying until something works.

$297shipped a $50K project
5,000+GitHub stars on Ryan Carson's implementation

In mid-2025, Australian developer Geoffrey Huntley was tending to his goats in the countryside when he had a realization about AI coding tools. Every time AI-generated code errored, the entire session would stall. Manual intervention was mandatory.

What if AI could act like a human·see the results of its own failed run and automatically retry based on that feedback?

The Code

Huntley's solution was almost disappointingly simple. In its purest form, Ralph Loop is just this:

while :; do cat PROMPT.md | claude ; done

That's it. An infinite bash loop that feeds the same prompt to an AI coding agent, over and over. But here's the clever part: progress doesn't persist in the AI's context window·it lives in your files and git history.

When the context fills up, you get a fresh agent with fresh context, picking up where the last one left off. The codebase becomes the memory.

"Deterministically bad in an undeterministic world."

Each individual iteration might produce garbage. But run enough iterations with clear success criteria, and the model eventually converges on something that works. Failures are predictable, and predictable failures can be fixed by tuning the prompt.

What Makes It Work

  • 1Objective success criteria. Progress isn't measured by the AI saying 'this looks good.' It's measured by tests passing, builds succeeding, and health checks completing.
  • 2State lives in files. The codebase is the memory. Git history is the log. Context window resets don't lose progress.
  • 3Failures are expected. The system assumes most iterations will fail. That's fine. Predictable failures inform prompt improvements.
  • 4Humans set strategy, AI executes. PRDs and acceptance criteria come from humans. The AI handles implementation.
Developer working late at night

The bottleneck was never the AI.
It was the human in the loop.

The Results

$297to ship a $50K project in API costs
6 reposbuilt by a YC team while they slept
3 monthsof autonomous iteration created CURSED language
Case Study: The CURSED Language

Geoffrey Huntley ran Claude in a Ralph Loop for three months with a simple prompt: create a programming language like Go, but with all lexical keywords swapped to Gen Z slang.

The result: a fully functional compiler with interpreted and compiled modes, producing binaries for macOS, Linux, and Windows via LLVM. Completely autonomous. No human code review during execution.

Case Study: Y Combinator Hackathon

A team at a YC hackathon set up Ralph loops before going to sleep. By morning, they had 6 complete repositories to demo·each with functional code, tests, and documentation. The AI ran overnight while they rested.

The Evolution

Ryan Carson, founder of Treehouse (which taught over a million people to code), took Huntley's minimal approach and added structure. His 3-file system has earned over 5,000 GitHub stars because it gives developers a process to follow.

Carson's 3-File System

1PRD (Product Requirements Document)
The human-written specification. What the product should do, acceptance criteria, constraints.
2Generate Tasks
Guides the AI to create a step-by-step task list from the PRD. Parent tasks and subtasks.
3Process Task List
Controls agent execution. Iterative code-test-commit workflow until all tasks complete.

The key insight: you're not "vibe coding" one giant prompt. You're giving the agent testable, bite-sized tickets and letting it execute like an engineering team. The AI handles implementation. Humans handle strategy and acceptance criteria.

The big idea is this: give the agent clear acceptance criteria and let it iterate until the tests pass. Not until the AI thinks it's done.

AI doesn't need to be smart.
It needs to be persistent.

Why This Matters

The Ralph Loop reveals something important about the current moment in software development. The constraint isn't AI capability·it's human bandwidth.

Before Ralph

Every AI error required human intervention. Context resets lost progress. Developers became bottlenecks in their own workflows.

After Ralph

Errors are expected and handled automatically. Progress persists across sessions. Developers can sleep while code ships.

This isn't about replacing developers. It's about changing what developers do. Less time babysitting AI outputs. More time on architecture, strategy, and the parts that require human judgment.

What would you build if the AI could run all night?

What could you ship if you weren't the bottleneck?

Share this perspective

Donny Smith

Written by · January 12, 2026

· ECD, Founder, Bttr.

Over the past 15 years, he has led creative teams and contributed to products used by millions of people worldwide, working with companies including Apple, Opendoor, JP Morgan, GE Aerospace, Pepsi, and Alterra Mountain Company.

LinkedIn

Sources: Geoffrey Huntley on the CURSED language · Ryan Carson's 3-File System Tutorial · Ralph Loop GitHub Implementation

More Insights

Perspectives on product design, technology, and the forces shaping how people interact with the things we build.

RAG vs Fine-Tuning
9 min read

RAG vs Fine-Tuning

The teams that waste the most on AI pick RAG or fine-tuning before they name the problem. The two are not answers to the same question. RAG gives a model facts it can look up at answer time, retrieved from your own documents, cut into chunks, turned into embeddings, and searched by meaning. The fact never enters the weights. Anthropic notes that once a knowledge base passes roughly 200,000 tokens, retrieval is what lets a model draw on a corpus no prompt could hold. Fine-tuning does the opposite. It trains the base model on your examples until a tone, a format, or a task is set into the weights, which the OpenAI Cookbook notes can take thousands of worked examples and suits teaching a skill rather than a fact. The clean rule is RAG for facts, fine-tuning for behavior, and in production usually both. The Cookbook frames it as an exam: fine-tuning is studying a week ahead and risking a forgotten detail, RAG is the open notes exam with the answer in front of the model while it works. Freshness, cost, and the paper trail settle the rest. A fact that changes is a retrieval problem. A behavior that drifts is a training problem. Name the problem first and the tool picks itself.

What Is Reranking
8 min read

What Is Reranking

A retrieval system usually finds the right passage and then leaves it in the wrong place. The first stage is a bi encoder that turns the query and every document into separate vectors and compares them, which is fast but never reads the two together, so the best passage often lands in the middle of the list. Reranking is the second pass that fixes the order. A cross encoder reads the query and one candidate as a single input, scores how they actually relate, and reorders the short list so the most relevant passages rise to the top. It is slower because it computes a fresh score for every query and document pair, so it runs on about a hundred candidates, never the whole corpus. On the public MS MARCO benchmark the sentence transformers cross encoder models trade speed for accuracy across sizes, from a tiny model near 9,000 documents a second to larger ones a fraction of that. Anthropic measured the payoff on a real pipeline: retrieve 150 chunks, rerank, keep the top 20, and the failure rate fell 67%. Retrieval finds the passage. Reranking makes sure the model reads it first.

Industries We Serve

Aerospace & DefenseBiotechnologyMedical & HealthcareManufacturingFinancial ServicesConsumer ProductsEnterprise Software

New Business

Start a project

Headquarters

North America

© 2026 Bttr. All rights reserved.