AI coding costs land in review, not on the invoice
AI coding costs per engineer are a small share of pay. Review time decides whether delivery rises, and how to measure it at 20 to 150 engineers.
Your AI coding bill is paid in review
One agent-written change, from the invoice to the number worth dividing by.
What did your last merged change cost, counting the review?
KaaS Newsletterkaas.team/newsletter
Your AI coding costs keep climbing, and talk of usage pricing makes next year's number harder to guess. Engineering leaders keep asking two things: will this end up costing as much as an engineer, and what are we getting for it?
Faros AI's 2025 telemetry study found that teams with high AI adoption merged 98% more pull requests, while PR review time rose 91%1.
The AI coding bill is a small share of an engineer's pay; what decides the return is review, and without strong tests and fast feedback, more code turns into instability. To show where the money goes, we follow one change at Acme Inc: an agent writes "retry failed card payments from a queue" in an afternoon.
The AI coding bill is still a small share of an engineer's pay
Per engineer, the AI coding bill is small, though the invoice is where most of the worry sits. Anthropic's figures for enterprise use of its coding agent2 and DX's all-in estimate of seat plus token spend3 both come to a modest monthly sum per engineer.
Levels.fyi's 2025 report puts a US software engineer's total yearly pay at about 226 thousand dollars4. Against that, a year at the top of the vendor's average range is a little over one percent, and a year at the top of DX's range about three percent.
- Anthropic average, upper endjust over one cent
- DX all-in estimate, upper endabout three cents
- The engineerthe whole dollar
- One square is one hundredth of the whole.
That is our arithmetic on total pay rather than loaded cost, so the true share is lower still. The vendor figure is an aggregate with no sample size, so treat both shares as rough; even rough, the bill is nowhere near the cost of the person.
Right now, we're not sweating the costs because we're trying to evolve best practices for the tools, but that has resulted in some devs really blowing through budget, so we may start instituting caps on spending.
Caps may come, and a few heavy users can surprise you. At Acme, though, the token line for the payment-retry change is a rounding error on the month's invoice. On our arithmetic over a full working year, a whole active day of agent use at the vendor's average2 costs less than one hour of the engineer's pay at Levels.fyi's figure4.
More code arrives, but company delivery often does not move
The bill buys more code, but the evidence that it moves company delivery is thinner than the dashboards suggest. In a telemetry study of over ten thousand developers, Faros AI found that gains in team behavior do not scale when aggregated to the company1. Across more than 400 organizations, DX found a median PR throughput gain of 7.76%3: a real gain, and a modest one.
LinearB's 2026 benchmarks point to wasted output: AI-generated pull requests are accepted far less often than manual ones6. Code that never merges costs review time and ships nothing.
32.7% of AI-generated pull requests are accepted
84.4% of manual pull requests are
It's hard to keep our CFO supportive about investing in these tools because the productivity benefits have proven difficult to conclusively prove.
Not every study agrees. DORA's 2025 survey finds a positive link between AI adoption and delivery throughput7; the section on tests below covers what DORA says the return depends on.
Back at Acme, the agent opens the payment-retry PR the same afternoon, and the dashboard counts a pull request, though nothing has shipped yet.
The extra output piles up in review, on your senior engineers
The extra output stalls in review. In Faros AI's data, PRs grew 154% larger on average1, and review time rose with them.
- Pull requests merged+98%
- PR review time+91%
In LinearB's separate data, AI pull requests wait 4.6x longer before anyone picks them up, though they are reviewed 2x faster once someone does6. Addy Osmani, who spent over 14 years leading developer experience at Google and became a director at Google Cloud AI8, named the limit in a January 2026 essay.
When output increases faster than verification capacity, review becomes the rate limiter.
PR volume rises at companies your size too: Synthesia's 118 engineers went all in on AI coding tools in November 202510. Its CTO, Peter Hill, told IEEE Spectrum that pull requests had risen 120 percent year over year10. He gave the volume entering review but not the hours it took.
Faros AI's head of product marketing, Naomi Lurie, argues that senior engineers absorb the review load because they catch what AI gets subtly wrong11. Cui and colleagues recorded seniority in the Microsoft experiment. They found the AI assistant significantly raised task completion for recent hires and juniors, but not for senior, long-tenured developers12. If those PRs go to seniors for review, the juniors' extra output becomes the seniors' extra review; that step is our reading, not something either study measured.
None of this is on the invoice, and no study we found prices it. Price it yourself: one hour of your staff engineer against the agent's whole token line. At Acme, the payment-retry PR is far larger than the change needed. It waits days, then lands on the one staff engineer who knows the payment flow, the same person whose juniors fill her queue.
More code merges unreviewed, and adoption tracks with less stable delivery
Under high AI adoption, more pull requests merge with nobody looking, and DORA links adoption to less stable delivery. Faros AI's 2026 report found 31% more PRs merged with no review at all under high AI adoption11. Faros does not say whether that is a count of merges or a share of them. On our reading, each unreviewed merge can move cost from today's reviewer to whoever handles the next incident.
more pull requests merged with no review at all
DORA's 2025 survey finds AI adoption still has a negative relationship with delivery stability7. DORA reports the link, not its cause.
GitClear's analysis of changed lines shows a longer trend: refactoring shrank over the years AI tools spread13. The timing matches adoption, but GitClear does not show that AI caused it.
25% of changed lines were refactoring in 2021
Under 10% were by 2024
The evidence is mixed. Jellyfish, looking at bug tickets and reverts, found no significant relationship with AI adoption14. The studies measure different things in different ways, and we do not know why their results differ.
At Acme, under deadline, the second review on the payment-retry PR is waved through. A retry loop double-charges a handful of customers on a Friday. The rework and the incident are part of what the change cost.
Without tests and fast feedback, more code turns into instability
DORA's 2025 report says that without strong automated testing, mature version control and fast feedback loops, more change volume leads to instability7. DORA draws that from its survey, not from a controlled trial.
DORA adds that the greatest return comes from internal platforms, clear workflows and aligned teams rather than from the tools themselves7.
Jellyfish's Nicholas Arcolano runs research at a vendor that sells engineering measurement. In a talk on his firm's data, he compared teams by how their code is spread across repos.
What’s really interesting is this highly distributed case. There’s essentially no correlation here between AI adoption and PR throughput. And actually the, the weak trend that does exist is actually slightly negative.
Treat the size of any gain with care. Pooled field experiments at three companies found a 26.08% rise in completed tasks, with a standard error of 10.3%, so the true gain could be small12.
- SpecThe engineer writes what the change must and must not do.Engineer
- Agent writesThe agent drafts the change and its tests against the spec.Agent
- Tests gateNothing reaches a person until the tests pass.Tests gate
- Size checkOversized changes are split before review.Size gate
- Review on a targetA reviewer picks it up within the agreed pickup target.Staff engineer
- MergeThe change merges with its review hours logged.Tech lead
- Survives a monthNot reverted or rewritten, so it counts as delivered.Churn gate
So measure the return where it is decided. One vendor defines the unit as total fully loaded AI cost over a period, divided by the PRs that merged and survived. Survived means not reverted or substantially rewritten inside a churn window15. Put review hours inside that cost; our next post takes the metric apart.
Acme puts a size limit and a test gate in front of review, and sets a pickup target for agent PRs. It starts dividing all AI cost, review hours included, by changes that merged and stayed merged. The payment-retry change now reaches its staff engineer small, tested and on time.
Code was never the scarce part of your team
In a team of 20 to 150 engineers, the scarce part is the judgment of the few people who can say a change is safe. Our reading of the evidence is that every gain from cheaper code is capped by how fast that judgment can be applied. At Acme, the staff engineer who knows the payment flow was the limit before any agent arrived; the agent made the queue in front of her longer.
AI made writing cheap and left verifying slow to reach, so the cost moved to the step nobody bills for. DORA's word for it is amplifier: AI magnifies the strengths of high performers and the dysfunctions of struggling teams16. The evidence suggests that a team without good tests and small changes gets busier and less stable, and no faster.
So the budget question is about how you design review and testing more than about seats and tokens. The invoice is the easiest part to read and the least useful part to manage.
What to do on Monday
- Write down the share. Put last quarter's AI spend for each engineer beside your loaded pay. It is your baseline, and it is probably small.
- Split the review clock. Pull review pickup time and time in review for the last month, AI-assisted PRs apart from the rest.
- Gate before a person. Set a size limit and a pickup target for agent PRs, and make tests pass before anyone reviews.
- Count reviews by name. Count the reviews each senior engineer did last month. If one name carries the load, spread it.
- Start cost per merged change. All AI cost plus review hours, divided by changes that merged and were not reverted within a month.
The companion, The AI-ready engineering team playbook, is free with your work email.
What did your last merged change cost, counting the review?
Param
Sources
- Faros AI, The AI Productivity Paradox Report, 2025
- Anthropic, Claude Code docs: Manage costs effectively, 2026
- DX, AI coding assistant pricing and ROI guide (2026), Taylor Bruneaux, 2026
- Levels.fyi, 2025 Annual Report, 2025
- The Pragmatic Engineer, 'The impact of AI on software engineers in 2026' (Gergely Orosz, Elin Nilsson), 2026
- LinearB, 2026 Software Engineering Benchmarks Report, 2026
- Google Cloud blog: Announcing the 2025 DORA State of AI-assisted Software Development report, 2025
- Addy Osmani, Biography (addyosmani.com), 2026
- Addy Osmani, Code Review in the Age of AI (Substack), 2026
- IEEE Spectrum, 'AI Slop Is Changing How Engineers Review Code' (Aaron Mok), 2026
- Faros AI, How AI-Generated Code Is Increasing Code Review Burden, 2026
- Cui et al.: The Effects of Generative AI on High-Skilled Work (three field experiments), 2025
- GitClear, AI Copilot Code Quality: 2025 Research, 2025
- Jellyfish: What Data from 20 Million Pull Requests Reveal About AI Transformation, 2025
- Unblocked, Cost per Merged PR: The Unit Economics of AI Coding, 2026
- Google Research: DORA 2025 State of AI-assisted Software Development Report, 2025
