成果重于活动、代码背后的成本,以及各团队如何治理其 AI 花出的资金。
In LinearB's 2026 benchmark, autonomous AI agents got 79% of their pull requests merged at elite teams and just 37% at average ones. Same agents, less than half the yield. The number that moves is not the tool. It is everything the tool cannot see. Here is why AI coding output is not a fixed quantity, and what actually decides how much of it ships.
Salesforce said it would hire zero new engineers this year because AI made the team more productive. A year on, Forrester finds 55% of employers regret their AI-driven cuts and many are quietly rehiring. The gap between the two is a productivity claim that was asserted at headcount scale and never actually measured. Here is what measuring it would have required.
Microsoft is pulling Claude Code from thousands of engineers in its Experiences and Devices division and moving them to GitHub Copilot, after token bills reportedly hit around $2,000 per engineer per month. The decision was framed as a benchmark. A benchmark on cost alone is half a benchmark, and the half that was missing is the one that actually decides whether the switch was right.
Ask your engineers how much time AI saves them and someone will say two hours a day. It feels like proof the spend is working. A randomized trial found developers felt 20% faster with AI while actually finishing 19% slower, a perception gap of almost 40 points. Here is why self-reported time savings cannot go in a budget, and what number to use instead.
Uber put its engineering teams on internal leaderboards ranked by how much they used Claude Code and Cursor, then burned its entire 2026 AI budget in four months. Everyone points to the missing spend cap as the lesson. The real lesson is the metric at the center of the incentive. Here is the difference, and how to rank teams by outcome instead of activity.
Engineering teams pick AI coding models on price per token. The invoice does not arrive in tokens. It arrives in finished tasks. A model that looks 60% cheaper per token can cost more per merged PR once retries, context reloading, and correction time are counted. Here is how to measure the number that actually hits your budget.
Most AI coding ROI frameworks assume the result will be positive. Some pilots end at 90 days with flat or negative numbers. Here is how to detect the signal before it costs you a full quarter, and what to do when the number does not work.
The EU AI Act's full provisions apply from August 2026. Most engineering teams using AI coding tools have no documentation of which code was AI-generated, no PR-level attribution, and no audit trail. Here is what is required and how to close the gap.
AWS CloudWatch Coding Agent Insights tracks token spend, PR velocity, and cost-to-output ratio. The category is now validated at the enterprise level. Here is what the product covers and what it cannot cover by design.
Engineering approved the seat licenses. Finance signed the line item. Nobody set a token budget. Here is the governance layer that prevents the invoice from arriving as a surprise.
获取新文章和产品笔记,大约每月一次。
The 30-day checkpoint is where most AI coding pilots get their budget renewed. It is also the point most likely to show you a number that will not survive contact with reality.
Most AI coding tool comparisons measure the wrong thing. Here is what the numbers look like when you compare GitHub Copilot, Cursor, and Claude Code on cost per merged PR at 90 days.
Most engineering teams cannot answer the CFO question: what value are we actually getting from Claude Code, Cursor, and GitHub Copilot? Here is a measurement framework that produces numbers finance can verify.