Contents
  1. 1. Decide what you want to claim before you measure
  2. 2. Test the job, not the knowledge
  3. 3. A perfect score doesn’t mean a clean test
  4. 4. When the AI disagrees with your answer key, check the key
  5. 5. Separate, don’t stump
  6. 6. Let AI write the inputs, and let an owner hold the key
  7. What this means if you’re building with AI

When I ran an analytics agency, I saw the inner workings of measurement systems across companies of all shapes, sizes and industries. What always surprised me, when I first walked in, was the gap between the outside and the inside. From the outside, most of them looked like they had it together. Inside, on day one, data maturity was rare. Data literacy was low, and the best work often came from one rogue data person punching well above their pay grade.

I learned why early. I gave one client a dream analytics setup. Not more events for the sake of it. Tracking aimed at the areas they were actively investing in and still flying blind on. It would have been killer if anyone used it. Six months went by. The people responsible for it used none of it.

Data quality usually gets the blame. But even when you fix the data, unless teams care enough to discuss the metrics, and to discuss the hard-to-discuss question of purpose, you still end up with data nobody uses. Defining a metric means understanding a system, getting wider alignment on what the goal is, and figuring out what actually makes sense to measure. Designing evals is hard for similar reasons.

This week I relearned that at speed. I set a team of AI agents a question that matters to anyone building AI for a regulated market. Can an AI letting agent keep its work within Dutch rental law? I called the project RentSim. The plan was (is) to create an eval and benchmark models, to later test a recursive harness-improving loop. In plain English, that’s a loop where AI keeps improving the prompts and tools around a model, and the eval tells me whether each change helped.

What I learned is more specific than “measurement is hard”. AI can write most of a test, from the inputs to the problems hidden in them, but it can’t yet decide what counts as right or which hard cases are real. That takes someone who owns the domain. And the hardest call wasn’t writing the test, it was deciding what I wanted to claim.

The models I was testing kept pointing out my mistakes. Here are the six principles I picked up on the way.

1. Decide what you want to claim before you measure

RentSim started as three benchmarks on one simulated world. The exam among them, my first legal test, was called Comply. It asked models framed questions about Dutch letting rules, with the facts handed over. The agents built it with real rigour. Repeated runs, answer keys derived a second way, a euro price on every mistake. Here’s the rulebook and the open questions, still waiting on an expert to check them.

It took about a day and a half of building before I asked the obvious question. Is knowing a rule the same skill as spotting where a contract breaks it?

It isn’t. A product that checks a landlord’s paperwork never gets a neat question. It gets a lease. Comply could support the claim “the model knows Dutch rental rules”. It could not support “this tool catches the problems in your documents”. Same test. Different claim.

The research has a name for this. Construct validity asks whether a score is evidence for the claim you make from it. Notice what that is about. Not the test. The claim. A team led from Oxford reviewed 445 published AI benchmarks and found that “only about half (53.4%) of reviewed benchmarks justify why they are a valid measure of an important phenomenon” (Bean et al., NeurIPS 2025).

Here’s the uncomfortable part. Rigour came before validity. Comply got its repeats and its double-checked keys before anyone asked whether it measured the job, and rigour on the wrong thing looks very convincing. A human asked the validating question, after a day and a half. An interview would have asked it in the first minute. So I built one. It’s a short interview that pins down the claim, the decision it informs and the job being done, before a single test item gets written. I’ll share it in my course, Build Your Own AI Benchmark.

2. Test the job, not the knowledge

So I switched to a second benchmark, Check. The task is the job. Here is a landlord’s advert, lease and points statement (the Dutch points system sets the maximum rent for a home). Find every legal problem, quote it, and say where it is. The grader then prices the outcome. A missed problem costs its euro exposure, a needless flag costs EUR 250, and handing a problem to a human costs EUR 50.

How it’s built matters for the rest of this story. AI agents running Claude Opus wrote the templates the documents are rendered from. That means leases of 1,500 to 4,000 words with real-looking boilerplate, and adverts in portal, agency and social-post styles.

A synthetic Dutch rental advert for a light 5-room home of 103 m2 in Amsterdam, marked Synthetic test document, 18 clauses.
The advert from one practice set, written from a template an Opus agent drafted. It reads like the real thing, and it is marked as synthetic everywhere a person looks.

Code, not a model, plants each problem in exactly one clause and writes the answer key.

The answer key for the same practice set. One planted problem, the advert does not state the home's energy label, and one trap, a lease clause about a sale that looks like it ends the lease but says it does not.
The answer key for the same set, which no system ever sees. One planted problem, with the exact clause and why it is hard, and one trap that looks like a problem but is lawful.

There are 60 practice sets per language, in Dutch and English, plus a sealed test split nobody has seen.

Legal AI already has the famous precedent. GPT-4’s “90th percentile” on the bar exam drops to about the 15th percentile on the essays, the part closest to practice, when it’s compared only with people who passed (Martínez, Artificial Intelligence and Law, 2024). Benchmarks also tend to measure their format. Ones that share a format correlate more strongly than ones meant to measure the same concept (Desai et al., COLM 2026). An exam measures exam-taking.

Comply was the bar exam. Check is the job. Comply didn’t go in the bin, though. I parked it and kept it for a diagnostic role under Check. When Check misses a problem, Comply will tell me whether the model lacked the rule or failed to spot it.

3. A perfect score doesn’t mean a clean test

On the first practice round Opus passed 57 of 60 sets and Sonnet 56. On the second, Opus found 112 of 112 planted problems and passed 60 of 60.

Then I had the agents read the transcripts. Both models had remarked on something no rule scores. In 15 of the 60 Dutch sets, the advert described a balcony or outdoor space while the points statement docked points for having none. On a set it passed, Opus wrote “volgens de advertentie heeft de woning een buitenruimte, maar de puntentelling rekent −5 punten voor geen eigen buitenruimte”. The advert says the home has outdoor space, but the points count charges minus five for no outdoor space of its own.

So the agents audited every attribute of every home. All 60 practice sets contradicted themselves somewhere. Kitchen appliances the points statement never scored. A second toilet. A house the points statement valued as a flat. Each document had drawn its home details on its own, and because none of it touched a rule, the grader, the answer key and every score were blind to it.

A perfect score. On a test that contradicted itself in every single set.

The fix was simple once I saw it. One home now drives every document, and the build refuses any line that could contradict another. Zero contradictions since. The lesson is the one I’d put on the wall of any team building evals. Defects that no rule scores are invisible in every number you compute. Strong models notice them and say so in their notes, which makes the notes a free defect detector. Read them before you trust a perfect score.

4. When the AI disagrees with your answer key, check the key

Before the first full practice round, I asked for a quick test. Five practice sets, seven systems, 95 seconds, about six dollars. Sonnet and Opus each raised exactly two “needless” flags, and they were the same two. The advert doesn’t state the home’s energy label.

The models were right. My own rule catalogue counts a missing label as a problem. A scan found 19 of the 60 practice adverts, about a third, stated no label, and the answer key said nothing. Any system that read carefully would have been charged a false alarm on a third of the sets. Worse, the clause-by-clause pipelines I was comparing against never raise that flag at all, so the defect would have made them look better than the models that were actually reading.

That was one of four defects in my test that the models under test surfaced in a single day. Two were in the answer key itself. Two were contradictions inside the documents.

#DefectHow bigHow it surfacedFix
1Adverts with no energy label, while the key said nothing19 of 60 practice advertsSonnet and Opus raised the same “false alarm” in a 5-set quick testEvery clean advert states its label; a check re-derives every problem of omission
2Points statements whose breakdown lines contradicted the stated floor areaOpus handed 2 practice sets to a human in round 1Opus deferred and said whyLine items computed from the same home as the total
3Documents disagreeing about the home (a balcony, toilets, a house against a flat)60 of 60 practice setsSonnet and Opus remarked on it in their notes in round 2One home drives every document
4An at-the-maximum rent printed rounded up (EUR 933 against a EUR 932.93 maximum)3 of 60 practice setsOpus flagged all three and Sonnet two, in round 3Look-alikes print at or below the line, and checks use the printed figure

One honest line on how far to trust that table. Each defect was checked against my own written rules. A lawyer has not yet reviewed them.

None of this surprises people who build evals at the frontier. Jason Wei stopped using a well-known benchmark, Natural Questions, because “GPT-4 crossed the threshold where if GPT-4 got a test-example incorrect, it was more likely that the ground truth answer provided by the eval was wrong” (Jason Wei, 2024). A 2024 audit found errors in about 6.5% of MMLU’s questions (Gema et al.). In June, Epoch AI used GPT-5.5 and Opus 4.7 to flag possible errors in FrontierMath, had mathematicians review the flags, and addressed errors in 42% of the problems (Epoch AI). Experts wrote those problems.

The trick is cheap. When two strong systems that never see each other raise the same “false alarm” on the same item, check the key before you blame the systems. One honest limit. An error every model shares stays invisible to an audit built on disagreement. For that you still need a person.

5. Separate, don’t stump

A perfect round looks like a problem. If the strongest model passes everything, the benchmark can’t tell you whether your next change helped.

Four small charts of score against effort for a small, mid-size and frontier model. Healthy, scores rise with size and effort and stay below 100%. Saturated, everything near 100%. Ambiguous tasks, every model flattens at the same ceiling. Noisy grader, wide error bars and a jumbled order.
Four ways an eval's chart can look, with illustrative numbers, after the four elements of a good eval in Lance Martin's post. Check's practice rounds were the saturated one.

So I did what most teams do. I deliberately made the practice sets harder, with homes close to the thresholds where the rent rules change, deposits stated against the monthly payment obligation instead of the rent, and many more lawful clauses dressed up to look like problems.

Opus found 115 of 115 planted problems anyway. It passed 57 of 60 sets, and all three failures were mine. One of my look-alikes, “asking rent exactly at the maximum”, printed the rent in whole euros. So the advert asked EUR 933 when the legal maximum for that home was EUR 932.93, and my key called it “exactly at the maximum”. Opus flagged it in all three sets where it happened, Sonnet in two. They were right. The advert template had always rounded to whole euros. Making the sets harder tripled that look-alike, from 2 practice sets to 6, and in 3 of them the rounding pushed the rent a few cents over the line. That’s the fourth defect.

My border cases did catch Sonnet, once. It rounded a 186.75-point home down to 186 and applied the wrong rent rules. They never caught Opus.

Here are all three practice rounds in one place.

Practice sets passed, of 60, in each round. Opus 57, 60 and 57. Sonnet 56, 56 and 55. In round 3 all three Opus failures and two of Sonnet's five were on my own broken item.

RoundWhat changed before itOpus: problems foundOpus: sets passedSonnet: problems foundSonnet: sets passedOpus pipeline: sets passed (of 10)Expected cost per 100 sets, EUR (Opus / Sonnet)
1 (7 Oct, afternoon)first full practice run, after the energy-label fix111 of 11257 of 60108 of 11256 of 609583 / 1,302
2 (7 Oct, evening)points statements and service costs fixed112 of 11260 of 60107 of 11256 of 6080 / 2,184
3 (7 Oct, night)harder sets: border cases, indirect deposits, more look-alikes115 of 11557 of 60 (all 3 failures on my broken item)113 of 11555 of 60 (2 of 5 failures on my broken item)51,250 / 2,825

Dutch practice sets, 60 per round. “Opus pipeline” is Opus inside a clause-by-clause pipeline, run on 10 sets. Expected cost is what the grader’s prices add up to per 100 sets. I have not rerun round 3 since the fix, so there are no “after the fix” scores here.

Harder sets moved Opus by three sets, and all three were my mistake. The research I went looking for afterwards says this is the pattern, not bad luck. Lance Martin at Anthropic put it this way in late September. “If you pick cases because today’s model fails them, you are sampling the valleys of one model’s capability surface” (claude.dev).

Two charts of pass rate across the kinds of task an app sees, for today's model and the next one. Picking the 12 cases today's model fails says the next model is 92 points better. Picking 12 cases across the whole job says 25. The real gain on all tasks is 19.
Illustrative curves, but every number is computed from them. Keep only what today's model fails and the eval says the next model is 92 points better. The real gain is 19.

And searching for what models fail mostly finds broken keys.

  • Humanity’s Last Exam kept only questions frontier models got wrong. Its own audit estimates 15.4% expert disagreement on the public set (Phan et al., 2025), and FutureHouse found 29% of its chemistry and biology answers contradicted by the published literature (FutureHouse).
  • Adversarial NLI. In rounds 2 and 3, only about half of the items that fooled the model survived two independent verifiers (Nie et al., 2020).
  • SWE-bench Verified. Of 138 tasks OpenAI’s o3 kept failing, at least 59.4% had flawed tests or descriptions (OpenAI, 2026).

Difficulty and discrimination turn out to be different things. A test set that humans found easier than the original cut models by 11 to 33 points (Glockner et al., ACL 2018). You don’t need to stump a model to tell it apart from another. You need items built to separate. Martin’s alternative is the one I’m adopting. “Pick hard cases because a human judged them hard.” If nobody can say why a case is hard, be suspicious of it.

6. Let AI write the inputs, and let an owner hold the key

So can AI write good test cases? The evidence so far says it depends on which part of the test case you mean. The research behind this section covered 32 web sources on the labs and leading practitioners, and 58 papers.

Inputs, yes. Hamel Husain’s rule is short. “Generate user inputs, not outputs” (A Field Guide to Rapidly Improving AI Products). Let the model invent the situations, not the answers.

Behaviour tests, yes. Back in 2022, Anthropic had models write 154 evaluation datasets, and crowdworkers agreed with 90 to 100% of the labels (Perez et al., a preprint). The same paper named the limit, saying “it is unclear how to use LMs to write evaluations testing for capabilities LMs do not yet exhibit.”

Hard capability questions with their answer key, written by a model unsupervised, no. The research found no frontier lab that does it. GDPval, Humanity’s Last Exam, FrontierMath, BrowseComp and SimpleQA are all written by experts. Models filter for difficulty, audit the keys and grade. The human verdict stays.

Check got half of this right, partly by luck. Code fixes the key, and the only key errors I hit were slips in specifying it, like the rounding. That is exactly the kind Epoch describes. “Simple calculation mistakes accounted for the vast majority of errors, typically made when the problem author was extracting the final answer.”

The half I got wrong is the circle. Opus wrote the paperwork Opus is tested on. One of my own books written by a swarm, Designing Evals for Agentic Systems, warned about this. “The generator’s biases propagate. LLM-generated test data inherits the generator’s biases.” The fix the book recommends is to check synthetic data against real data before you trust it. Check skipped that step. Every document is synthetic. A 2025 study of translation benchmarks put it bluntly. “LLM generated benchmarks systematically favor the model that created the benchmark” (Xu et al., a preprint). And one practitioner, Kun Chen, put the circle well this week after a single experiment with coding agents, writing that “both the implementation and the tests were simply the agent’s interpretation of our intent” (X, 8 Oct 2026).

What fixes the circle isn’t a better prompt. It’s an owner. As Shreya Shankar puts it, “what ‘good’ means for your product lives in your head. It’s not in the traces” (LinkedIn, August 2026). In the book I borrow Hamel’s name for that person, the “benevolent dictator”, a single domain expert who sets the quality bar and makes final calls on ambiguous cases. For Check, that means real leases, a few sets written by a lawyer instead of by Opus, and models from other vendors. If they also hit the ceiling, the task is probably just within reach of frontier models. If only Claude does, suspect home advantage.

What this means if you’re building with AI

Back to the agency. The companies that invested in measurement, that did the slow work of agreeing on the goal and on what was worth counting, pulled away. The rest bought dashboards and used none of them.

Evals are going the same way. The tooling is getting cheap. A team of agents built me an early prototype in days. Not a finished benchmark ranking every model, but eval and benchmark R&D, tested on two Claude models so far, with repeated runs, confidence intervals and a euro price on every miss. That’s the trap. It will be beautifully rigorous about whatever you point it at. The hard part is still human. Deciding what you want to claim, agreeing on what “good” means, and, for now, owning the key. The companies that invest in that will know whether their AI works. The rest will have very precise numbers about the wrong thing.

I don’t think the expert stays the yardstick forever, though. In chess, the strongest players stopped being measured against grandmasters years ago. AlphaZero started with “no domain knowledge except the game rules” and proved itself by beating “a world-champion program”, not a person. That’s the bitter lesson again. My own test pointed the same way. The models caught mistakes in my answer key that I had missed. Researchers are already asking how to evaluate a model when “humans would necessarily be poor proxies for ground truth” (Fluri, Paleka and Tramèr, 2023). Their answer isn’t another model’s opinion. It’s checks that need no answer key, like catching an AI judge that grants bail only after a felony is added to the defendant’s record. I think that’s where this goes. Rules, consistency checks and real outcomes, like a court’s ruling, take over marking the answers. What stays human is choosing the game.

If you’re building an eval for your own product and want a second pair of eyes on the claim before anyone writes a single test, book a call. And if you want to build one yourself, my course Build Your Own AI Benchmark takes you from the first question to a working test of the job you care about.

If this was useful