Every couple of weeks, someone asks me what the AI tooling actually buys. Sometimes it is a client, working out whether my rate still makes sense now that everyone has heard a model can write code.

Sometimes it is another developer deciding whether to bother. Either way the question is fair, and for ten months my honest answer has been a shrug and a story.

I do not love answering with a story. Stories are how people talk themselves into things.

So when a proper study landed, I read it the same day. On the 1st of July, three Microsoft researchers published a paper on what happened when the company rolled Claude Code and GitHub Copilot CLI out to tens of thousands of its own engineers.

The number in the abstract is 24%. Adopters merged roughly 24% more pull requests than they otherwise would have. Then I read the rest of it, and the rest of it is more interesting than the number.

Somebody finally did the work properly

Most of what gets said about AI coding tools is a screenshot, a benchmark, or a man on a podcast. This is not that.

Emerson Murphy-Hill, Jenna Butler and Alexandra Savelieva took a pre-period running from October 2025, a rollout date of the 5th of January 2026, and an observation window closing on the 29th of April.

They restricted the analysis to active engineers, defined as anyone with at least two merged pull requests in a four-week window before the rollout, so the result is not carried by people who were barely committing anyway.

They built a synthetic control out of the engineers who did not adopt, and ran it through a Bayesian structural time-series model.

The headline comes out as a 24.0% lift in pull requests per engineer per day, with a 95% confidence interval of 14.5% to 33.7%.

Then they did the thing that separates real analysis from motivated analysis. They re-ran the whole model with a fake intervention date in October, months before anything actually happened, to see whether their method would invent an effect out of nothing.

It returned minus 1.1%, with an interval comfortably straddling zero.

That is a good paper. If your instinct is that developer productivity cannot be measured and anyone who tries is a consultant, this is the work that should make you pause.

The alternative to measuring badly is not measuring wisely. It is arguing from vibes forever, and I have sat in enough of those meetings.

The line under the headline

The paper also breaks the result down by tool, and this is the part I have not seen quoted anywhere. Copilot CLI adopters came in at plus 24.9%. Claude Code adopters came in at plus 11.4%. The authors put it plainly: Copilot CLI adopters saw about 2.2 times the lift of Claude Code adopters, at p below 0.0001.

Sit with that for a second. The gap between the two tools is larger than the effect everybody is arguing about.

Two products, both correctly described as command-line AI coding agents, both rolled out to the same population in the same building in the same quarter, and one of them moved the metric more than twice as far as the other.

Which means "AI coding agent" is not a unit of measurement. It is a category, and the variance inside the category swamps the average.

I am not going to pretend I know why the split came out that way. Nobody does, from the outside. It could be the tools. It could be who reached for which. It could be that the two agents encourage different sized pull requests, which would make the metric itself the explanation.

That last possibility is not a stray thought of mine, and I will come back to it, because the authors got there first.

What a merged pull request is, and what it is not

Here is the sentence from the abstract that almost nobody carried forward. The authors write that they use merged pull requests as their proxy for output, acknowledging that a merged PR is not the same as the value it delivers.

In the body they are blunter. Merged PRs, they say, are an imperfect proxy for throughput and reward small, frequent PRs, and they may miss quality costs such as added complexity.

They disclose that they are Microsoft employees, that Microsoft sells AI tools and owns GitHub, and that their proximity may have shaped their questions, their design and their interpretation.

Then they say the pressing open question is now about quality, whether the added throughput yields better software, and that the field still lacks agreed-upon measures to answer it.

I want to be precise about what happened here. The researchers were the most careful people in this story. They named their conflict of interest, the weakness of their metric, and the question their own paper could not answer.

Every one of those caveats is in the paper, in plain language, near the top.

And then the number travelled without them.

None of this is new ground. The SPACE framework said in 2021 that developer productivity cannot be captured by a single metric or dimension.

Kent Beck and Gergely Orosz spent a long piece in 2023 explaining that measuring effort and output is easy precisely because it sits early in the chain, and that being early in the chain is exactly what makes it easy to game. We know this. We knew it before agents existed.

A number that is easy to count is not thereby the right number. It is just the one that survives being turned into a headline.

The decision that happened down the hall

Now for the part that changed how I read the whole paper. Tucked into the background section, the authors explain why they closed their window on the 29th of April.

Shortly afterwards, they write, an internal announcement indicated that Claude Code licenses would be discontinued for most engineers in about a month, with affected engineers directed to transition to Copilot CLI.

They ended the analysis early because engineers being moved off a tool would have looked like organic adopters of the other one.

That is a researcher protecting their data. It is also, read a second time, the most consequential sentence in the paper.

The Verge broke the underlying story in May. Microsoft's Experiences and Devices group, the division behind Windows, Microsoft 365, Outlook, Teams and Surface, would be off Claude Code by the end of June.

The internal note from the executive in charge framed it as convergence. Claude Code had been an important part of the learning, and Copilot CLI offered something especially important: a product Microsoft can help shape directly with GitHub for its own repos, workflows and security expectations.

Forbes covered it on the 1st of June and put it more flatly, saying the tool was priced by the token, the engineers used it heavily, the costs ran past the annual budget early, and cost pulled the trigger.

I genuinely do not know which of those is the real reason. Neither does anyone outside that building. It might be both. It might be a third thing nobody wrote down.

But look at the shape of it. We have the effect measured to one decimal place, with a confidence interval, a significance test and a placebo control.

And we cannot say with any confidence at all why the decision went the way it did. The precision is entirely on the side of the thing that did not decide anything.

The 24% appears nowhere in the reasoning, on either side. Not as a defence of the tool, not as a cost that was worth paying, not as a number that failed to justify itself.

In every account I have read of that decision, the one rigorous measurement anyone had produced about those exact tools, on those exact engineers, in that exact window, simply does not come up.

What the number cannot see

The gap is not a scandal. It is the normal condition, and pretending otherwise is how people end up disappointed by good research.

The number measures throughput. The decision was about control, or spend, or both, plus a dozen things that never get written down. Those are different questions and the number was never going to answer the second one.

The trouble is that the things a throughput metric cannot see are not small. DORA's own work this year found that higher AI adoption is associated with an increase in software delivery throughput and an increase in software delivery instability, at the same time.

They also found that time saved in creation is frequently reallocated to auditing and verification. That is not a footnote. That is the entire experience of working this way, described in one sentence.

Peer-reviewed research from CodeScene, published in January, found that AI coding assistants increase defect risk by at least 30% when applied to unhealthy code.

Adam Tornhill's framing of it stuck with me: in unhealthy code, AI acts as a technical debt multiplier rather than an accelerator.

None of that shows up as fewer merged pull requests. Some of it shows up as more of them.

What I actually measure

My own output did not get 24% bigger over the last ten months. I have said this before and I will keep saying it, because the alternative is letting people assume a multiplier that is not there. The work moved. Different work, same hours.

The things I watch now are unglamorous and none of them have a confidence interval. How often I throw a branch away and start the plan again, because the plan was the problem. How long the review takes, which is the real bottleneck and has been for months.

Whether I can still explain, out loud, what shipped and why. Whether I am reading the diff properly or nodding at it. That last one is the one that scares me, and no metric anyone has proposed can see it.

So when I say I am not arguing against measurement, I mean it. I am arguing that you should know what a number is load-bearing for.

The Microsoft study is load-bearing for one claim: these tools move throughput, in a large organisation, and it is not a novelty effect that fades in month three. That claim is better supported today than it was a month ago, and I am glad someone did the work.

It is not load-bearing for whether you should adopt them, which tool to pick, what the code will be like to maintain in a year, or what happens to your bill. It was never going to be.

The number was never going to make the call

I still get asked what the AI tooling actually buys, and I still answer with a story.

The difference now is that the story has a footnote. There is a real, carefully constructed number in the world, produced by people who were honest about its limits in the abstract itself, and it says the throughput is real. Take it. It is more than we had.

Then notice that the same paper contains a quiet sentence about licenses being cancelled a month after the data stopped, and that the decision behind that sentence was made on grounds nobody has measured and probably could not.

That is not a failure of the research. It is a description of the job. The measurable part of engineering has never been the part that decides anything, which is why the work is still interesting and why it still needs someone doing the deciding.

If your answer to what the tooling buys is a story rather than a multiplier, you are not being unrigorous. You might just be the only person in the conversation being honest about what we can actually count.