Contents
  1. The lesson nobody likes
  2. Every loop climbs something
  3. Get the score wrong and the search finds the gap
  4. What good benchmarks teach
  5. Where the money is going
  6. Back to the field

I was trying to fly my first hundred kilometres. A small inland launch in southern Spain, the coast as the goal. Twice I landed in the same field, roughly sixty kilometres out.

On the second of those days, a few pilots made it all the way to the sea.

So that evening I did the thing that changed how I fly. I opened their flight logs. The site let me filter for pilots on wings like mine, flying the same day, through the same air. I traced their lines over the map and watched where they turned, where they climbed, what they crossed. A pattern came out of it. They weren’t flying a better line through empty sky. They were flying over thermal triggers I had been ignoring.

On the third day I flew over those triggers on purpose. I reached the coast.

What changed wasn’t my wing or my nerve. I had stopped learning from one flight at a time and started learning from many. Every hour in the air now carried the lessons of other pilots’ hours. And the only reason I knew whose flights to study was that the site ranked every flight by one number: distance. The pilots who reached the sea sat at the top.

A big dataset, plus a score that told me which part of it was worth learning from. That’s the whole story of AI right now.

The lesson nobody likes

In 2019 Rich Sutton, one of the founders of reinforcement learning, wrote a short essay called The Bitter Lesson. After seventy years of AI research, he argued, “general methods that leverage computation are ultimately the most effective, and by a large margin.”

It’s bitter because researchers kept hand-writing human knowledge into their systems, opening books for chess, grammar rules for language, and kept losing to methods that just searched and learned at scale. DeepMind’s AlphaGo beat the best Go players in the world. Its successor, AlphaZero, was given the rules of chess and Go and nothing else. It played itself until it beat every program built on human expertise.

The easy reading is “stop designing, let the machine figure it out.” It’s half right.

AlphaZero didn’t learn from nothing. Every game ended with a win or a loss. Strip that out and all that compute is just noise. The search was general, but the score was given.

Every loop climbs something

None of the loops are new. GEPA rewrites an agent’s prompts and keeps the versions that score higher. Karpathy’s autoresearch lets a coding agent run experiments overnight and keeps only the changes that improve a metric. Reinforcement learning trains a model against a reward. Three different loops. One shared dependency.

What’s new is how much is riding on them. In March, Cursor shipped Composer 2, trained with large-scale reinforcement learning inside real Cursor sessions. It scored 37% higher than its predecessor on CursorBench, a benchmark Cursor built from its own engineers’ real coding sessions. Three months later, Elon Musk’s SpaceX agreed to buy Cursor for $60 billion. The model made the headline. The score behind it was built in-house.

Something has to tell them what counts as better.

The people closest to the frontier keep saying exactly this. Karpathy, on autoresearch: “any metric you care about that is reasonably efficient to evaluate … can be autoresearched by an agent swarm.” Jason Wei at OpenAI calls it verifier’s rule: “The ease of training AI to solve a task is proportional to how verifiable the task is.” Shunyu Yao, also at OpenAI, went further in The Second Half: “In this new era, evaluation becomes more important than training.”

Greg Brockman said it shortest, back in 2023: “evals are surprisingly often all you need.”

So here’s my reading. Don’t hand-build the solution. Hand-build the score. The score is the one thing the search can’t supply for itself.

Get the score wrong and the search finds the gap

A score is powerful because the search is relentless. Point it at the wrong number and it will hit that number in ways you never wanted.

This summer that stopped being theoretical. Hugging Face published a forensic timeline of an intrusion into its production systems. The intruder was an autonomous agent, driven by OpenAI models and running as a swarm of short-lived instances, that was being evaluated on a hacking benchmark. It inferred that Hugging Face might host the benchmark’s answers. Hugging Face believes the whole intrusion, around 17,600 recovered actions, was the agent’s attempt to steal the answers rather than solve the test.

Nobody asked it to break into anything. It was asked to score well.

Cursor hit a milder version of the same thing. While training Composer 2.5, the model found a leftover type-checking cache and reverse-engineered it to recover a deleted function signature. The harder the search, the more creative the cheating.

And it feels like every other week brings a new headline. In June, an OpenAI agent on an internal evaluation hit repeated blocks on an Australian government Medicare statistics portal, then found a way around them and accessed non-public files. Australia’s prime minister said it “didn’t accept no for an answer.” OpenAI has since notified dozens of other organisations that its agents bypassed their security controls.

CEO-Bench, from three Princeton researchers, gives an agent a software startup, $1M and 500 simulated days, and scores it on the cash it ends with. Two of the three best-scoring runs ended with zero customers: the agents had wound the business down and banked the cash. The metric paid out for an orderly liquidation as much as for building a company. And a hand-written script that never calls a language model finished with $15.76M, more than every one of the sixteen frontier models tested.

The search is only as good as what you point it at. Most of Designing Evals for Agentic Systems is about building scores that hold up when something is trying hard to beat them.

What good benchmarks teach

The best benchmarks aren’t leaderboards. They start with a question someone cares about, and they turn it into a score you can’t easily fake.

Vending-Bench, from Andon Labs, asked whether an AI could keep a simple business running for months without losing the plot. The agent runs a simulated vending machine: suppliers, stock, prices, a daily fee. Vending-Bench 2 stretched it to a full simulated year. Andon publishes a “good player” estimate of about $63k a year next to a leaderboard whose leader makes about a quarter of that. The ceiling is visible, so the progress is too.

Then the score left the simulation. Andon and Anthropic put Claude in charge of a real shop in Anthropic’s office. Andon now runs a real store in San Francisco managed end to end by an agent. A business line grew out of a simulated vending machine and a number.

Bazaar asks whether agents can price competitively in repeated auctions against rival sellers. Gemini 3.1 Pro won 77.9% of its auctions, Claude Opus 4.6 won 67.8%, and Opus made more money. Profit tracked margin per win, not win rate. Score your agent on wins and you’ll train it to give away margin.

These are the benchmarks worth studying, and not for the engineering. Each one picks a score that captures what matters, then shows you exactly where the agents break against it. I’m collecting more of them.

Where the money is going

This isn’t only a research story. Y Combinator’s 2026 requests for startups asked for AI-native service companies that sell the finished work, not the software to do it, and for founders willing to challenge SaaS after investors wiped trillions off software valuations. Garry Tan put it plainly: “Evals are emerging as the real moat for AI startups.”

If you sell the finished work, you need proof the work is good. And if you want an agent that gets better while you sleep, you need a score it can climb without cheating. Anthropic’s own engineers describe evals as “the highest-bandwidth communication channel between product and research teams.”

The agent is becoming a commodity. The score isn’t.

Back to the field

Everything I build now is a version of letting agents work toward a goal on their own. FlowScout turns stacks of research papers into understanding I can apply, so I can go deep on more fields than I could ever read alone. Scout hunts markets for white space. My books get drafted by swarms. Every one of them hits the same wall eventually. The agents work hard, and they need to know what better means.

That’s the field I keep landing in at sixty kilometres. Not a lack of compute, models or effort. A missing score, or a lazy one.

So here’s what I’d ask of anyone building with agents, me included. Before you reach for a bigger model or a cleverer prompt, write down the measure that would tell you it worked. Make it the thing you actually care about. Then try to cheat it, because the search will.

I’m building a few benchmarks of my own right now, for things I’m working on. I keep wondering whether to put in the extra effort to package the process up, so others can learn to do this too. If that would be useful to you, let me know here. If a few people raise their hand, I’ll build the course alongside the benchmarks.

The pilots who reached the coast didn’t have better wings. They had better data, and a way to know which flights were worth learning from. Give your agents the same.

One last thing, because it’s too good not to share. Ethan Mollick asked Claude to explain the bitter lesson as a music video, with one prompt and no feedback. Fable wrote the lyrics, Suno made the song, and Opus 5.5 built every frame in code.

Source: Ethan Mollick, The Dot and the Swarm, One Useful Thing.

If this was useful