<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel>
<title>KaaS Newsletter</title><link>https://kaas.team/newsletter</link>
<atom:link href="https://kaas.team/newsletter/rss.xml" rel="self" type="application/rss+xml"/>
<description>Research on what AI changes in engineering delivery, for CTOs and VPs of Engineering at startups and scaleups. Every number cited.</description><language>en</language>
<item><title>AI coding costs land in review, not on the invoice</title><link>https://kaas.team/newsletter/p/ai-coding-costs</link><guid>https://kaas.team/newsletter/p/ai-coding-costs</guid><pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate><description>&#60;p&#62;AI coding costs land in review, not on the invoice&#60;/p&#62;&#60;p&#62;AI coding costs per engineer are a small share of pay. Review time decides whether delivery rises, and how to measure it at 20 to 150 engineers.&#60;/p&#62;&#60;p&#62;Your AI coding bill is paid in review&#60;/p&#62;&#60;p&#62;One agent-written change, from the invoice to the number worth dividing by.&#60;/p&#62;&#60;p&#62;Spend&#60;/p&#62;&#60;p&#62;A small share of the engineer&#38;#39;s pay.&#60;/p&#62;&#60;p&#62;Generate&#60;/p&#62;&#60;p&#62;More and larger pull requests, the same day.&#60;/p&#62;&#60;p&#62;Wait&#60;/p&#62;&#60;p&#62;The queue fills on your senior engineers.&#60;/p&#62;&#60;p&#62;Ship or slip&#60;/p&#62;&#60;p&#62;More merges unreviewed; adoption tracks with lower stability.&#60;/p&#62;&#60;p&#62;Measure&#60;/p&#62;&#60;p&#62;Cost per merged change, review included.&#60;/p&#62;&#60;p&#62;What did your last merged change cost, counting the review?&#60;/p&#62;&#60;p&#62;Your AI coding costs keep climbing, and talk of usage pricing makes next year&#38;#39;s number harder to guess. Engineering leaders keep asking two things: will this end up costing as much as an engineer, and what are we getting for it?&#60;/p&#62;&#60;p&#62;Faros AI&#38;#39;s 2025 telemetry study found that teams with high AI adoption merged 98% more pull requests, while PR review time rose 91%.&#60;/p&#62;&#60;p&#62;The AI coding bill is a small share of an engineer&#38;#39;s pay; what decides the return is review, and without strong tests and fast feedback, more code turns into instability. To show where the money goes, we follow one change at Acme Inc: an agent writes &#38;#34;retry failed card payments from a queue&#38;#34; in an afternoon.&#60;/p&#62;&#60;p&#62;Per engineer, the AI coding bill is small, though the invoice is where most of the worry sits. Anthropic&#38;#39;s figures for enterprise use of its coding agent  and DX&#38;#39;s all-in estimate of seat plus token spend  both come to a modest monthly sum per engineer.&#60;/p&#62;&#60;p&#62;Levels.fyi&#38;#39;s 2025 report puts a US software engineer&#38;#39;s total yearly pay at about 226 thousand dollars. Against that, a year at the top of the vendor&#38;#39;s average range is a little over one percent, and a year at the top of DX&#38;#39;s range about three percent.&#60;/p&#62;&#60;p&#62;A year of AI coding spend per engineer, out of each dollar a US engineer is paid&#60;/p&#62;&#60;p&#62;Anthropic average, upper end&#60;/p&#62;&#60;p&#62;DX all-in estimate, upper end&#60;/p&#62;&#60;p&#62;The engineer&#60;/p&#62;&#60;p&#62;That is our arithmetic on total pay rather than loaded cost, so the true share is lower still. The vendor figure is an aggregate with no sample size, so treat both shares as rough; even rough, the bill is nowhere near the cost of the person.&#60;/p&#62;&#60;p&#62;Caps may come, and a few heavy users can surprise you. At Acme, though, the token line for the payment-retry change is a rounding error on the month&#38;#39;s invoice. On our arithmetic over a full working year, a whole active day of agent use at the vendor&#38;#39;s average  costs less than one hour of the engineer&#38;#39;s pay at Levels.fyi&#38;#39;s figure.&#60;/p&#62;&#60;p&#62;The bill buys more code, but the evidence that it moves company delivery is thinner than the dashboards suggest. In a telemetry study of over ten thousand developers, Faros AI found that gains in team behavior do not scale when aggregated to the company. Across more than 400 organizations, DX found a median PR throughput gain of 7.76%: a real gain, and a modest one.&#60;/p&#62;&#60;p&#62;LinearB&#38;#39;s 2026 benchmarks point to wasted output: AI-generated pull requests are accepted far less often than manual ones. Code that never merges costs review time and ships nothing.&#60;/p&#62;&#60;p&#62;of AI-generated pull requests are accepted&#60;/p&#62;&#60;p&#62;LinearB benchmarks&#60;/p&#62;&#60;p&#62;Not every study agrees. DORA&#38;#39;s 2025 survey finds a positive link between AI adoption and delivery throughput; the section on tests below covers what DORA says the return depends on.&#60;/p&#62;&#60;p&#62;Back at Acme, the agent opens the payment-retry PR the same afternoon, and the dashboard counts a pull request, though nothing has shipped yet.&#60;/p&#62;&#60;p&#62;The extra output stalls in review. In Faros AI&#38;#39;s data, PRs grew 154% larger on average, and review time rose with them.&#60;/p&#62;&#60;p&#62;High-adoption teams against low, Faros AI&#60;/p&#62;&#60;p&#62;Pull requests merged&#60;/p&#62;&#60;p&#62;PR review time&#60;/p&#62;&#60;p&#62;In LinearB&#38;#39;s separate data, AI pull requests wait 4.6x longer before anyone picks them up, though they are reviewed 2x faster once someone does. Addy Osmani, who spent over 14 years leading developer experience at Google and became a director at Google Cloud AI, named the limit in a January 2026 essay.&#60;/p&#62;&#60;p&#62;PR volume rises at companies your size too: Synthesia&#38;#39;s 118 engineers went all in on AI coding tools in November 2025. Its CTO, Peter Hill, told IEEE Spectrum that pull requests had risen 120 percent year over year. He gave the volume entering review but not the hours it took.&#60;/p&#62;&#60;p&#62;Faros AI&#38;#39;s head of product marketing, Naomi Lurie, argues that senior engineers absorb the review load because they catch what AI gets subtly wrong. Cui and colleagues recorded seniority in the Microsoft experiment. They found the AI assistant significantly raised task completion for recent hires and juniors, but not for senior, long-tenured developers. If those PRs go to seniors for review, the juniors&#38;#39; extra output becomes the seniors&#38;#39; extra review; that step is our reading, not something either study measured.&#60;/p&#62;&#60;p&#62;None of this is on the invoice, and no study we found prices it. Price it yourself: one hour of your staff engineer against the agent&#38;#39;s whole token line. At Acme, the payment-retry PR is far larger than the change needed. It waits days, then lands on the one staff engineer who knows the payment flow, the same person whose juniors fill her queue.&#60;/p&#62;&#60;p&#62;Under high AI adoption, more pull requests merge with nobody looking, and DORA links adoption to less stable delivery. Faros AI&#38;#39;s 2026 report found 31% more PRs merged with no review at all under high AI adoption. Faros does not say whether that is a count of merges or a share of them. On our reading, each unreviewed merge can move cost from today&#38;#39;s reviewer to whoever handles the next incident.&#60;/p&#62;&#60;p&#62;more pull requests merged with no review at all&#60;/p&#62;&#60;p&#62;Faros AI telemetry, high vs low AI adoption&#60;/p&#62;&#60;p&#62;DORA&#38;#39;s 2025 survey finds AI adoption still has a negative relationship with delivery stability. DORA reports the link, not its cause.&#60;/p&#62;&#60;p&#62;GitClear&#38;#39;s analysis of changed lines shows a longer trend: refactoring shrank over the years AI tools spread. The timing matches adoption, but GitClear does not show that AI caused it.&#60;/p&#62;&#60;p&#62;of changed lines were refactoring in 2021&#60;/p&#62;&#60;p&#62;GitClear, changed lines across many repos&#60;/p&#62;&#60;p&#62;The evidence is mixed. Jellyfish, looking at bug tickets and reverts, found no significant relationship with AI adoption. The studies measure different things in different ways, and we do not know why their results differ.&#60;/p&#62;&#60;p&#62;At Acme, under deadline, the second review on the payment-retry PR is waved through. A retry loop double-charges a handful of customers on a Friday. The rework and the incident are part of what the change cost.&#60;/p&#62;&#60;p&#62;DORA&#38;#39;s 2025 report says that without strong automated testing, mature version control and fast feedback loops, more change volume leads to instability. DORA draws that from its survey, not from a controlled trial.&#60;/p&#62;&#60;p&#62;DORA adds that the greatest return comes from internal platforms, clear workflows and aligned teams rather than from the tools themselves.&#60;/p&#62;&#60;p&#62;Jellyfish&#38;#39;s Nicholas Arcolano runs research at a vendor that sells engineering measurement. In a talk on his firm&#38;#39;s data, he compared teams by how their code is spread across repos.&#60;/p&#62;&#60;p&#62;Treat the size of any gain with care. Pooled field experiments at three companies found a 26.08% rise in completed tasks, with a standard error of 10.3%, so the true gain could be small.&#60;/p&#62;&#60;p&#62;Acme&#38;#39;s payment-retry change, with review grown to fit&#60;/p&#62;&#60;p&#62;Spec&#60;/p&#62;&#60;p&#62;The engineer writes what the change must and must not do.&#60;/p&#62;&#60;p&#62;Agent writes&#60;/p&#62;&#60;p&#62;The agent drafts the change and its tests against the spec.&#60;/p&#62;&#60;p&#62;Tests gate&#60;/p&#62;&#60;p&#62;Nothing reaches a person until the tests pass.&#60;/p&#62;&#60;p&#62;Size check&#60;/p&#62;&#60;p&#62;Oversized changes are split before review.&#60;/p&#62;&#60;p&#62;Review on a target&#60;/p&#62;&#60;p&#62;A reviewer picks it up within the agreed pickup target.&#60;/p&#62;&#60;p&#62;Merge&#60;/p&#62;&#60;p&#62;The change merges with its review hours logged.&#60;/p&#62;&#60;p&#62;Survives a month&#60;/p&#62;&#60;p&#62;Not reverted or rewritten, so it counts as delivered.&#60;/p&#62;&#60;p&#62;So measure the return where it is decided. One vendor defines the unit as total fully loaded AI cost over a period, divided by the PRs that merged and survived. Survived means not reverted or substantially rewritten inside a churn window. Put review hours inside that cost; our next post takes the metric apart.&#60;/p&#62;&#60;p&#62;Acme puts a size limit and a test gate in front of review, and sets a pickup target for agent PRs. It starts dividing all AI cost, review hours included, by changes that merged and stayed merged. The payment-retry change now reaches its staff engineer small, tested and on time.&#60;/p&#62;&#60;p&#62;In a team of 20 to 150 engineers, the scarce part is the judgment of the few people who can say a change is safe. Our reading of the evidence is that every gain from cheaper code is capped by how fast that judgment can be applied. At Acme, the staff engineer who knows the payment flow was the limit before any agent arrived; the agent made the queue in front of her longer.&#60;/p&#62;&#60;p&#62;AI made writing cheap and left verifying slow to reach, so the cost moved to the step nobody bills for. DORA&#38;#39;s word for it is amplifier: AI magnifies the strengths of high performers and the dysfunctions of struggling teams. The evidence suggests that a team without good tests and small changes gets busier and less stable, and no faster.&#60;/p&#62;&#60;p&#62;So the budget question is about how you design review and testing more than about seats and tokens. The invoice is the easiest part to read and the least useful part to manage.&#60;/p&#62;&#60;p&#62;We read vendor telemetry studies, DORA&#38;#39;s survey, three field experiments and practitioner surveys, kept only figures we could quote word for word from the source, and name each vendor that sells measurement or review tools.&#60;/p&#62;&#60;p&#62;The AI coding bill is still a small share of an engineer&#38;#39;s pay&#60;/p&#62;&#60;p&#62;More code arrives, but company delivery often does not move&#60;/p&#62;&#60;p&#62;The extra output piles up in review, on your senior engineers&#60;/p&#62;&#60;p&#62;More code merges unreviewed, and adoption tracks with less stable delivery&#60;/p&#62;&#60;p&#62;Without tests and fast feedback, more code turns into instability&#60;/p&#62;&#60;p&#62;Code was never the scarce part of your team&#60;/p&#62;&#60;p&#62;Write down the share&#60;/p&#62;&#60;p&#62;Put last quarter&#38;#39;s AI spend for each engineer beside your loaded pay. It is your baseline, and it is probably small.&#60;/p&#62;&#60;p&#62;Split the review clock&#60;/p&#62;&#60;p&#62;Pull review pickup time and time in review for the last month, AI-assisted PRs apart from the rest.&#60;/p&#62;&#60;p&#62;Gate before a person&#60;/p&#62;&#60;p&#62;Set a size limit and a pickup target for agent PRs, and make tests pass before anyone reviews.&#60;/p&#62;&#60;p&#62;Count reviews by name&#60;/p&#62;&#60;p&#62;Count the reviews each senior engineer did last month. If one name carries the load, spread it.&#60;/p&#62;&#60;p&#62;Start cost per merged change&#60;/p&#62;&#60;p&#62;All AI cost plus review hours, divided by changes that merged and were not reverted within a month.&#60;/p&#62;&#60;p&#62;What did your last merged change cost, counting the review?&#60;/p&#62;&#60;p&#62;Right now, we&#38;#39;re not sweating the costs because we&#38;#39;re trying to evolve best practices for the tools, but that has resulted in some devs really blowing through budget, so we may start instituting caps on spending.&#60;/p&#62;&#60;p&#62;It&#38;#39;s hard to keep our CFO supportive about investing in these tools because the productivity benefits have proven difficult to conclusively prove.&#60;/p&#62;&#60;p&#62;When output increases faster than verification capacity, review becomes the rate limiter.&#60;/p&#62;&#60;p&#62;What’s really interesting is this highly distributed case. There’s essentially no correlation here between AI adoption and PR throughput. And actually the, the weak trend that does exist is actually slightly negative.&#60;/p&#62;</description></item>
</channel></rss>
