Blog
AI FinOps

Microsoft Cut Claude Code for Thousands of Engineers. The Bill Decided It, Not the Output.

Microsoft is pulling Claude Code from thousands of engineers in its Experiences and Devices division and moving them to GitHub Copilot, after token bills reportedly hit around $2,000 per engineer per month. The decision was framed as a benchmark. A benchmark on cost alone is half a benchmark, and the half that was missing is the one that actually decides whether the switch was right.

Sep 2, 2026·8 min read
Microsoft Cut Claude Code for Thousands of Engineers. The Bill Decided It, Not the Output.
Microsoft moved thousands of engineers off Claude Code after token bills reportedly reached around $2,000 per engineer per month. The number that broke the pilot was the input. The number that decides whether the switch was correct never appears in the story.

Microsoft is pulling Claude Code from thousands of engineers in its Experiences and Devices division, the group behind Windows, Microsoft 365, Teams, Outlook, and Surface, and moving them to GitHub Copilot by the end of June.

The reason was the bill. Token-based billing reportedly pushed spend to around $2,000 per engineer per month and exposed a consumption pattern the pilot could not contain, with the fiscal year closing on June 30. So the seats got cancelled.

The decision was framed as a benchmark that led to convergence: try the tool, learn from it, standardize on the in-house platform. That is a reasonable way to describe it. It is also a decision made almost entirely on one number, the cost, and the number that would actually tell you whether the switch was right is not anywhere in the story.

Large office building with engineers working at rows of desks

What Actually Happened

The facts, as reported across several outlets, are straightforward. Microsoft ran a Claude Code pilot in its Experiences and Devices division starting around December 2025. It lasted roughly six months. By June 30, 2026, the division was cancelling the Claude Code seats for most engineers and steering thousands of them to GitHub Copilot CLI instead.

The driver named in the reporting was cost. Claude Code enterprise billing is a seat fee plus actual token consumption, so spend scales with usage, and heavy usage on a large engineering org produced monthly per-engineer numbers in the range of $2,000. That is the kind of figure that gets attention when the fiscal year is ending and someone is looking at external software spend.

An executive framed the move as benchmarking followed by convergence, letting Microsoft shape security review, repository integration, and workflow tooling directly with GitHub. The timing lined up with the fiscal-year close on June 30 and with a broader shift in how these tools are priced.

This Is Not a Story About Claude Code Being Too Expensive

It is tempting to read this as the cheaper tool winning, and that reading misses what actually happened.

Start with the billing detail almost everyone skipped. On June 1, 2026, GitHub moved Copilot toward usage-based billing with AI Credits. So the tool Microsoft is standardizing on is itself moving in the same metered direction that made Claude Code's spend visible. The escape from token billing is not as clean as the headline suggests, because the whole category is converging on usage-based pricing. Flat rate is not where the market is going.

What flat or in-house billing does change is predictability and control. Bringing the spend back inside the GitHub relationship makes it easier to forecast, easier to govern, and easier to reduce external software cost heading into a new fiscal year. Those are real and legitimate reasons. None of them are the same as knowing which tool produced more shipped work per dollar.

That is the point. This is not Claude Code being bad and Copilot being good, or the reverse. It is a very large company making a very large tooling decision on the cost side of the ledger, because the cost side is the side that had a number on it.

The Number Microsoft Used, and the One It Did Not

The decision had one number at its center: spend per engineer. That number was real, it was large, and it was visible, because token billing puts consumption right in front of you whether you asked for it or not.

The number that was missing is output per dollar. What did the Claude Code spend actually ship, in merged PRs, resolved issues, and delivered work, compared to what Copilot ships for the same engineers? That is the number that tells you whether $2,000 a month was expensive or a bargain, and it does not appear anywhere in the public account of the decision.

Here is why that gap matters. A tool that costs more per engineer can still be cheaper per merged PR, if it produces proportionally more merged PRs. The scary invoice and the better value can be the same tool. Cost per engineer cannot tell those apart. Only cost per unit of output can, and that is precisely the number that was not on the table.

Pro Tip: When a tool decision is being justified by per-seat or per-engineer spend, ask for the same spend divided by merged PRs. If nobody has that number, the decision is being made on the input, and the input is the half of the picture that cannot tell you whether you are getting value.

Why Cost-Only Benchmarking Is Half a Benchmark

A benchmark is a comparison, and a comparison needs two sides: what something costs and what it produces. Drop either side and you are not benchmarking, you are just reading a price tag.

Cost-only benchmarking feels rigorous because it has hard numbers in it. The invoices are exact. But exactness is not the same as completeness. You can know the cost of two tools to the cent and still have no idea which one was the better buy, because you never measured what either one delivered. The precision on the cost side creates a false sense that the analysis is done.

What a Cost-Only Decision Cannot See

  1. Whether the expensive tool shipped more · Higher spend per engineer with proportionally higher output is a good deal wearing a scary invoice. Cost-only analysis reads the invoice and stops.
  2. Whether the cheaper tool shipped less · A lower per-seat number with lower throughput can be the worse value. Flat and predictable is not the same as efficient.
  3. Where the spend was actually going · Cost per engineer is an average. It hides the fact that a subset of engineers and task types may have driven most of the cost and most of the value, which points to routing, not cancellation.
  4. What the switch will actually cost in output · Moving thousands of engineers to a different tool has a productivity cost during the transition that no seat-price comparison captures, and that only shows up in delivery metrics.

None of this means Microsoft made the wrong call. It might have been exactly right. The point is that from the outside, and quite possibly from the inside, the decision cannot be shown to be right, because the output half of the benchmark was never measured.

Two people reviewing charts and figures on a screen in a meeting

How Would You Benchmark Two AI Coding Tools on Output?

You run both against the same work and measure what comes out, not just what goes in. The mechanics are not exotic.

  1. Pick a representative cohort and split it. Comparable engineers on comparable work, some on each tool, for a fixed window long enough to produce a stable number.
  2. Measure cost per merged PR on each side. Total spend, seat plus tokens, divided by merged PRs, for each tool. This is the number that makes the comparison fair, because it puts the expensive-but-productive tool and the cheap-but-slow tool on the same axis.
  3. Hold quality constant. Track correction rate and change failure rate so you are not rewarding a tool for shipping more PRs that break. Faster and worse is not cheaper.
  4. Then look at total spend. Once you know cost per unit of output, the absolute bill becomes a budgeting question rather than the whole decision. You might still cap or switch, but now you are doing it knowing what you are trading away.

Pro Tip: If you are choosing between a metered tool and a flat-rate tool, the flat-rate option will always look safer on the spreadsheet because its number is predictable. Predictable and low are not the same thing. Run cost per merged PR on both before you let predictability decide, because a predictable bill for less output is not a saving.

The Whole Category Is Converging on Metered Billing

There is a second lesson in the timing. Windsurf moved toward flat, predictable pricing earlier in the year. GitHub Copilot moved the other way, to usage-based AI Credits, in June. Two vendors, same market, opposite conclusions about how to charge, in the same year.

What that means for a buyer is that cost comparisons between tools are getting harder, not easier. When one tool bills flat and another bills by token, the same amount of delivered work shows up as two very different-looking invoices. Comparing the invoices tells you almost nothing. The only thing that stays comparable across billing models is output per dollar, because a merged PR is a merged PR regardless of how the vendor decided to meter it.

As metered pricing spreads, the teams that can measure cost per merged PR will be the only ones able to compare tools honestly. Everyone else will be comparing pricing pages, which increasingly measure different things.

Key Takeaways

  • Microsoft decided on the input, not the output. The Claude Code cancellation was driven by per-engineer spend near $2,000 a month and a fiscal-year deadline. The output half of the comparison, what the spend shipped, is absent from the story.
  • This is not about Claude Code being too expensive. GitHub Copilot itself moved to usage-based billing in June, so the category is converging on metered pricing. The switch buys predictability and in-house control, which is not the same as knowing which tool delivered more per dollar.
  • A tool with a scarier invoice can be the better value. Higher cost per engineer with proportionally higher output is a lower cost per merged PR. Cost per engineer cannot see that. Only cost per unit of output can.
  • Cost-only benchmarking is half a benchmark. Exact invoices create a false sense of rigor. A comparison needs both what a tool costs and what it produces, and dropping the output side means the decision cannot be shown to be right.
  • Output per dollar is the only cross-tool comparison that survives. As flat and metered billing diverge, invoices stop being comparable. Cost per merged PR stays comparable because a merged PR means the same thing on every billing model.

Why Even Great Companies Decide on the Bill

Microsoft is about as sophisticated a buyer of software as exists, and it still made this call on the cost side. That is not a knock on Microsoft. It is a sign of how strong the pull is, because the reason is structural, not a lapse in judgment.

The cost number shows up on its own. Token billing generates it automatically, finance sees it, and a fiscal deadline gives it urgency. The output number has to be built. Somebody has to join the spend to the shipped work and produce cost per merged PR, and if no one has done that, then when a decision has to be made, the only hard number in the room is the invoice. So the invoice decides. Not because anyone believes cost is the whole story, but because cost is the only part of the story that was quantified.

The pattern I keep running into is exactly this, at every size of company. The spend is measured because it measures itself. The output is not, because it does not. And so tool decisions, cap decisions, and renewal decisions all get made on the half of the picture that happened to have a number attached, and everyone involved half-knows the other half is missing but has nothing to put there.

The teams that will make better calls than this are not the ones with smaller bills. They are the ones who built the output number before they needed it, so that when the fiscal deadline comes and someone points at the invoice, there is a second number on the table that says what the invoice actually bought. With both numbers, cancelling might still be the right move. Without the second one, no one can honestly say.

· Vukasin

How Yardstick Puts the Output Number Next to the Bill

Yardstick joins token spend from your AI coding tools to merged outcomes from your Git provider and shows cost per merged PR by tool, updated daily. When a per-engineer bill lands on someone's desk, the delivery number that says what the spend produced is already sitting beside it.

Because Yardstick is vendor and cloud neutral, it produces the one comparison that survives diverging billing models: cost per merged PR on Claude Code next to cost per merged PR on Copilot or Cursor, on the same axis, regardless of whether each vendor bills flat or by token.

For a team weighing a tool switch or a spend cap, Yardstick turns a cost-only decision into a real benchmark, so the call is made on what each tool shipped per dollar rather than on which invoice looked scarier at fiscal close.

Yardstick - Measure the value AI creates

To see the platform or join the early access list, visit yardstick.fi/platform or review pricing for current plan options.

FAQ

Why did Microsoft cancel Claude Code for its engineers?

As reported across several outlets, Microsoft's Experiences and Devices division ran a Claude Code pilot from around December 2025 and moved to cancel the seats for most engineers by June 30, 2026, steering them to GitHub Copilot CLI. The stated driver was cost: token-based billing reportedly pushed per-engineer spend to around $2,000 a month, and the fiscal year closing on June 30 added urgency to reducing external software spend.

Does this mean GitHub Copilot is cheaper than Claude Code?

Not necessarily. Copilot offered more predictable and in-house billing, but GitHub itself moved Copilot toward usage-based AI Credits in June 2026, so the category is converging on metered pricing. Predictable cost is not the same as lower cost per unit of output. Without measuring cost per merged PR on each tool for the same work, it is not possible to say which was actually the better value.

What is the difference between cost per engineer and cost per merged PR?

Cost per engineer is total tool spend divided by headcount. It measures the input and is easy to produce from a bill. Cost per merged PR is total spend divided by shipped work, which measures value delivered per dollar. A tool can have a high cost per engineer and a low cost per merged PR at the same time, if it produces proportionally more output. Only the second number tells you whether the spend was worth it.

How do you benchmark two AI coding tools fairly?

Run both against comparable work with comparable engineers for a fixed window, measure cost per merged PR on each side, and hold quality constant by tracking correction rate and change failure rate so you are not rewarding a tool for shipping broken code faster. Only after you know cost per unit of output does the absolute bill become a budgeting question rather than the entire decision.

Why is output per dollar better than comparing tool prices?

Because billing models are diverging. Some tools bill flat, others bill by token, so the same amount of delivered work produces very different-looking invoices, and comparing pricing pages compares different things. Output per dollar, such as cost per merged PR, stays comparable across billing models because a merged PR means the same thing regardless of how the vendor meters usage.

Recommended

Vukašin Kitanović · Co-founder

Leads commercial and go-to-market. Founder of Emberwood, an AI-driven cold-email agency for B2B, with a background in outbound, lead generation, and sales.

Partager

Mesurez la valeur créée par votre IA.

Recevez les nouveaux articles et notes produit, environ une fois par mois.