No. 27 · JUL 2026 · 5 Min Read
Carry Your Own Clock
Abstract
The control group quit and the benchmarks broke. There is no shared instrument for what AI does to engineering work, so carry your own.
Latitude is written in the sky. Measure the sun at noon, do a little arithmetic, and you know how far north you are. Longitude is not up there at all. It is a difference between local noon and the time back at your home port, so without a clock that survives the voyage you can stare at a perfect sky all night and remain lost. Ships guessed instead, and some of them guessed onto rocks. The eventual fix was not a better way to read the sky. It was Harrison’s sea clock: a reference you carried with you.
We are trying to work out what AI does to engineering work, and we have no clock.
The Control Group Quit
Start with the fact that ought to be the loudest one in the field. METR ran the study everyone cites, the randomized trial where experienced developers were nineteen percent slower with AI tools while believing they were twenty percent faster. Then they tried to run it again. In February they published a note explaining that they were changing the design, because between a third and a half of participating developers were declining to submit tasks they did not want to do without AI. The control arm would not hold. Their updated cohort came in around four percent, with a confidence interval wide enough to drive a fleet through, and METR themselves now describe the signal as weak.
Read that as a measurement problem rather than a result. A profession has stopped being able to produce a group of practitioners willing to work the old way for long enough to be measured. There is no unmodified baseline left to subtract from. Whatever your position on AI, the delta you want to compute requires a world that no longer exists, and every study published from here forward inherits that hole.
The instruments went the same way. Benchmark contamination stopped being a footnote once models started reproducing known fixes from memory, including details the problem statement never contained. Identical weights score wildly differently across harnesses, which means the number describes the scaffolding as much as the model. And the word that took over this year is “jagged,” which is doing an enormous amount of unearned work. Jagged lets you claim you are eighty percent of the way to AGI and explain away a model that cannot read an analog clock, using the same syllable, in the same paragraph. A term that absorbs every possible outcome is a hedge with good branding.
Vacuums Do Not Produce Humility
Here is what I did not expect. Losing the instruments made the discourse more confident, not less.
An admitted measurement vacuum should have softened people. It freed them. When there is no shared number to be wrong about, whatever you believed in 2023 turns out to be exactly what the evidence now shows. The optimist cites throughput and pull request counts, which measure typing. The skeptic cites the nineteen percent, quietly dropping the part where the same lab walked it back and said the design broke. Both sides are arguing from anecdote and calling it data, and neither one can be checked, which is the appeal.
Meanwhile the honest reports coming out of large organizations say something less dramatic and more interesting: individual effectiveness up, organizational delivery roughly flat. That pattern is what you would expect if the work is moving rather than disappearing. Generation got cheap. Verification did not, and we only automate what we can verify. The queue of things waiting to be checked by a human is where the gains are going, and that queue shows up in almost no one’s reporting because it was never a line item before.
Keep Your Own
You’re not going to get a shared instrument. There is no forthcoming study that will tell you what agents did to your shop, because your shop is not in anyone’s sample and the sample is broken anyway. The number is not coming. Keep your own.
A usable clock has three properties. It measures an outcome you actually sell, not an activity. It is cheap enough that you take the reading every week without ceremony. And it is read the same way every time, by the same definition, even when the definition becomes unflattering.
What that looks like in practice is boring. Elapsed time from a request being accepted to the change being in production and still there thirty days later. Percentage of changes reverted or hotfixed. Count of open items sitting in review older than a week, trended. Escaped defects per release. Time from an incident starting to someone understanding it, which is the number that quietly degrades when comprehension of the codebase thins out. None of these are AI metrics. That is the point. They are business metrics that AI either moves or does not.
The timing is the part people miss. A difference needs both readings, and the first one has to be taken before you leave. Most organizations rolled out agents first and would now like a study to tell them what happened, which is asking for a subtraction with one number in it. If you have not adopted yet, take six weeks of readings before you do anything. It is the cheapest item on this list and the only one that expires.
Guessing Is Allowed
If you took no baseline, you are guessing, and guessing is a legitimate way to run a shop. Most people who guess are fine.
The damage starts when a guess gets reported as a measurement. An estimate with an error bar you never wrote down, stated in the register of an observation, and then acted on when the stakes are highest. That failure mode is running through the whole AI conversation right now, in both directions. People with no instrument speaking like people with one. If you are working off judgment and vibes, say so out loud, and hold the conclusion loosely enough to drop it when reality shows up somewhere it should not be. The alternative is a demo you have decided to trust.
A real instrument comes with the size of its own error attached. The productivity numbers being sold to you do not. The one worth having is small, local, ugly, and honest about how soft its reading is.