AI Is Not Magic
Sixteen experienced developers predicted AI would make them 24 percent faster. It made them 19 percent slower. Afterward they believed they had been 20 percent faster.
There is a study I think about often, because it is the most useful correction available to anybody about to make a claim about AI.
METR ran a randomized controlled trial with sixteen experienced open-source developers across 246 real tasks in codebases they already knew well. Before starting, those developers estimated AI tooling would cut their completion time by about 24 percent. The measured result was the opposite: completion time went up by about 19 percent. And afterward, having just been measurably slower, they estimated they had been roughly 20 percent faster.
Read that last part again, because it is the interesting one. It is not that the tools failed. It is that people using them could not tell.
I bring this up in a series about a platform I built with AI assistance because I would rather set the frame honestly than sell something. If your model of this technology is that it makes work go faster, that study says your model is wrong, and it says so with real developers doing real work rather than with a survey.
So what actually happened in my build, where the effect was clearly enormous?
The reconciliation is one sentence, and once you see it the METR result stops being surprising.
The value of this technology is inversely proportional to how much you already know how to do.
Those developers were experts operating inside code they had written. For them, describing what they wanted, reading what came back, and correcting what was subtly wrong cost more than doing it themselves. The overhead was real and the ceiling was low, because the task was already near the limit of how fast a competent person can go.
I was the opposite case in every respect. I did not know how to do any of it. My alternative was not doing it more slowly. My alternative was not doing it at all, or spending a year learning first, or raising money to hire people who already knew. Against that alternative the effect is not a percentage. It is the difference between a thing existing and not existing.
That is the honest shape of the gain, and it is not the shape the marketing suggests.
Here is what it actually did across this build, in the order the value showed up.
It held the structure while I worked. More than anything else early on, it kept the plan visible. What was open, what depended on what, what I had decided last week and why. Not doing the work, holding the shape of the work while I did it, at a point when my own capacity was one task at a time.
It compressed the learn-and-do loop into a single motion. Normally you learn a thing, then apply it, and the gap between those is where most people quit. Here they happened together. I learned what a migration was by needing one, and the explanation arrived at the same moment as the task did.
It surfaced options faster than I could have found them. Not better options necessarily, but more of them, quickly, which meant the decisions I made were made against a wider field. That matters more than speed. A fast decision from two choices is worse than a slower one from six.
It made me more accurate, but only through verification. This is the part people get backwards. The tool did not make the output correct. It produced plausible output constantly, and plausible is a category that includes wrong. Accuracy came from checking, from stating what I expected to find before I looked, and from knowing what the answer should be because I know the domain. Without the checking, more output at higher speed is a worse outcome, not a better one.
And it changed what one person could attempt. That is the part with commercial consequence. A platform serving multiple categories, with a scoring engine, an audit layer, tenant isolation, and payments, is not something one person builds in the ordinary course. The constraint that used to make that impossible was capital and headcount. That constraint moved.
Now the things it did not do, because this is where the magic framing does the damage.
It did not supply judgment. Every decision that mattered in this build was a judgment about how buyers actually behave, where implementations actually fail, and what vendors actually overstate. None of that came from a model. It came from thirty years of watching it happen, and a model with no such experience would have produced something reasonable-looking and wrong.
It did not supply accountability. When something breaks, there is nobody to point at. The gates, the reviews, the refusal to accept a change I cannot explain, all of that exists because the responsibility does not distribute. It concentrates.
It did not supply memory. It will solve the same class of problem a different correct way each time unless somebody holds the line, which means consistency across a codebase is a human job.
And it did not tell me how I was doing. That is the real lesson in the METR numbers. The tool has no view on whether you are being effective. It will help you enthusiastically in the wrong direction, and it will feel exactly like helping you in the right one.
Which brings me to the framing I would offer instead of magic.
AI is not a solution. It is an enabler, and the distinction is not semantic. A solution does the thing. An enabler removes the constraint that was stopping somebody from doing the thing. Everything after that depends on whether the person had something worth doing and enough judgment to know when the output is wrong.
For an operator, that is the whole opportunity and the whole risk in one sentence. If you know your domain deeply and the only thing standing between you and building the tool was that you could not build tools, that constraint is gone. If you do not know your domain, this will help you produce a great deal of confident, well-structured, useless work, faster than you could have produced it before.
The technology is genuinely remarkable and it is not doing what most people think it is doing. It is not thinking for you. It is not deciding for you. It removed a wall, and what you do on the other side of it is entirely a function of what you brought with you.
Selecting supply chain technology?
PreShiftIQ matches buyers to vendors on measured fit across TMS, dock scheduling, ELD, carrier vetting, and fleet management. Vendors pay for outcomes, never for influence, and the buyer match is free.
See how matching works at preshiftiq.com/how-it-works.
Sources
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: a randomized controlled trial with 16 developers across 246 real tasks.
Next in the series: learning when to take the answer and when to argue with it.

