Blog
AI FinOps

Uber Ranked Its Teams by Token Spend. The Leaderboard Burned the Budget.

Uber put its engineering teams on internal leaderboards ranked by how much they used Claude Code and Cursor, then burned its entire 2026 AI budget in four months. Everyone points to the missing spend cap as the lesson. The real lesson is the metric at the center of the incentive. Here is the difference, and how to rank teams by outcome instead of activity.

Aug 19, 2026·8 min read
Uber Ranked Its Teams by Token Spend. The Leaderboard Burned the Budget.
Uber ranked its engineering teams by how much they used Claude Code and Cursor, then spent its entire 2026 AI budget in four months. The leaderboard rewarded exactly the behavior that broke the budget.

Uber put its engineering teams on internal leaderboards ranked by how much they used Claude Code and Cursor. Then it burned its entire 2026 AI budget in four months.

The leaderboard did exactly what it was built to do. It rewarded consumption, engineers consumed, and the number it was measuring went straight up. What went up with it was the bill, not a matching amount of shipped work. By the time leadership noticed, the annual budget was gone by April and the company was capping spend at $1,500 per engineer per tool to stop the bleeding.

The story got told this spring as a budget cautionary tale, and the lesson everyone drew was that Uber should have set a cap sooner. That is half the lesson. The cap stops the bleeding. It does not fix the thing that caused it, which was a leaderboard pointed at the wrong number.

Analytics leaderboard and cost charts displayed on a monitor in an office

What Actually Happened at Uber

The facts were widely reported. Uber rolled AI coding tools out to its engineering org and adoption climbed fast: 32% of engineers in February, 84% by March, roughly 95% using the tools monthly by spring. About 70% of committed code was coming through the tools. Per-engineer API spend landed somewhere between $500 and $2,000 a month, and one two-hour coding session reportedly cost $1,200 on its own.

To drive that adoption, Uber ranked teams and engineers on internal leaderboards by tool usage and told them to lean in. It worked. Usage exploded. And the 2026 budget set aside for these tools was fully spent four months into the year.

Uber's COO then said in public what a lot of leaders were thinking privately: the link between the AI spend and a matching increase in delivered output was not clearly there. In June the leaderboards were retired and the per-engineer, per-tool cap went in.

That is the story. It is a good story. But the popular reading of it is incomplete.

Why the Leaderboard Was the Problem, Not the Symptom

A leaderboard is an incentive machine. Whatever number sits at the top of it becomes the thing people optimize, because that is the entire point of putting it on a leaderboard. Rank teams by a number and they will move that number.

Uber ranked teams by token consumption. So teams consumed tokens. Not because anyone was acting in bad faith, but because the organization had drawn a line that said more usage is better and put it on a scoreboard. The engineers did precisely what they were rewarded to do. The budget burn was not a failure of the incentive. It was the incentive working perfectly on the wrong target.

This is the part that gets missed. The problem was never that engineers used the tools too much. The problem was that usage was the metric. Token consumption measures activity, and activity is not the same as output. A team can be at the very top of a token leaderboard while shipping less than a team near the bottom, because tokens spent tells you how hard the tools were worked, not how much finished work came out the other side.

Pro Tip: Before you put any AI usage number on an internal dashboard, ask one question: if a team optimized this number to the exclusion of everything else, would that be good for the business? If the answer is no, you are about to incentivize the wrong behavior.

Why the Spend Cap Is Not the Real Fix

The cap was the right emergency move and the wrong permanent lesson.

A cap at $1,500 per engineer per tool stops the runaway bill. It does nothing to tell you whether the money spent under the cap is producing anything. You can hit a cap and get enormous value, or hit the same cap and get very little, and the cap cannot tell the two apart. It is a limit on the input with no view of the output. It answers how much, never whether it was worth it.

So a company that only adds a cap has traded an uncontrolled bad metric for a controlled bad metric. Spend is bounded now, which is progress, but the organization still has no idea which teams are turning that spend into shipped work and which are lighting it on fire under the ceiling. The next budget conversation is just as blind as the last one, only cheaper.

The actual fix is upstream of the cap. It is changing the number the organization pays attention to, from tokens consumed to work delivered per dollar. Do that and the cap becomes a backstop rather than the whole strategy.

What a Leaderboard Ranked by Outcome Would Have Done

Here is the part the Uber story leaves on the table. Uber did not need to kill the leaderboard. It needed to change what the leaderboard ranked.

Imagine the same scoreboard, same competitive energy, same engineers who like being at the top of it, but ranked by cost per merged PR instead of tokens consumed. Now the incentive runs the other way. A team climbs by shipping more finished work per dollar, not by spending more. The engineer who finds a way to get the same output for half the token spend goes up the board instead of down it. The behavior the company rewards is the behavior it actually wants.

Token-Ranked Versus Outcome-Ranked: The Behavior Each One Rewards

  1. Token-ranked leaderboard · Rewards consumption. The way to win is to spend more. Budget burn is the built-in outcome, and the team that games it hardest looks like the best performer right up until finance runs the numbers.
  2. Outcome-ranked leaderboard · Rewards efficiency. The way to win is to ship more per dollar. A team that reduces waste climbs. The number at the top of the board is the number the CFO also cares about, so engineering and finance are finally optimizing the same thing.
  3. The cap alone · Rewards nothing. It only forbids. It caps the downside without creating any pull toward good behavior, which is why it is a backstop, not a strategy.

The difference between the first two is not the existence of the leaderboard. It is the metric. Same tool, opposite result.

How Do You Rank Teams by Outcome Instead of Activity?

You need the output number, not just the input number. Token spend you already have from the vendor. Delivered work you have to attribute, and cost per merged PR is the cleanest way to join the two.

A workable path:

  1. Get token spend attributed at the team and PR level, not as one blended monthly org total. A single number for the whole company cannot rank anyone.
  2. Join it to merged outcomes from your Git provider. Cost per merged PR is the score. Cost per resolved ticket works the same way if you track that.
  3. Rank on the ratio, never on the raw spend. The team that ships the most work per dollar is at the top. Raw token totals never appear on the board, because that is the number that started the fire.
  4. Show the trend, not just the snapshot. A team improving its cost per merged PR month over month is the real win, and a trend rewards progress from every starting point rather than just crowning whoever started cleanest.

Pro Tip: If you want the competitive energy of a leaderboard without the risk, rank teams by the change in their cost per merged PR quarter over quarter. It rewards every team for getting more efficient from wherever they are, instead of just rewarding the teams that happened to start in the best shape.

Two engineers reviewing delivery and cost metrics on a screen together

Key Takeaways

  • Uber's leaderboard was not a mistake in execution. It was a mistake in target. Ranking teams by token consumption rewarded spending, teams spent, and the annual budget was gone in four months. The incentive worked perfectly on the wrong number.
  • The spend cap stops the bleeding but fixes nothing. A cap bounds the input without telling you whether the spend under it produces anything. It answers how much, never whether it was worth it.
  • Activity is not output. Token spend measures how hard the tools were worked, not how much finished work came out. A team can top a usage leaderboard while shipping less than a team below it.
  • The fix is the metric, not the leaderboard. Rank teams by cost per merged PR and the same competitive scoreboard rewards efficiency instead of consumption. Engineering and finance end up optimizing the same number.
  • Rank on the ratio and the trend. Cost per merged PR, and its change over time, rewards shipping more per dollar from any starting point. Raw token totals never belong on the board.

Why Smart Companies Reward the Wrong Number

Uber is not a careless company. It ran the AI rollout the way a lot of well-run orgs would: drive adoption hard, measure the thing that is easy to measure, celebrate the teams that lean in. Every step of that is reasonable on its own. The trap is that token consumption is the number sitting right there in the vendor dashboard, and delivered work per dollar is a number you have to build. So the easy number becomes the scoreboard by default, and the scoreboard becomes the behavior.

The pattern I keep running into is exactly this shape. An organization wants people to adopt the tools, so it measures usage, because usage is what the tool reports. Then usage becomes the goal instead of the means, and a few quarters later someone in finance asks what all of it produced and the room goes quiet, because nobody was ever measuring that. The metric that was chosen for being easy to see quietly became the metric that ran the company.

The teams that avoid the Uber outcome are not the ones with the tightest cap. They are the ones who decided early that the number on the scoreboard would be the number they actually wanted people to move: work shipped per dollar, not dollars spent. That choice costs nothing extra to make at the start. Making it after the budget is gone, in the middle of a cost panic, costs a great deal more.

A leaderboard is not the villain here. Pointing it at consumption was. Point it at outcome and the same instinct that burned Uber's budget will fund the next one instead.

· Vukasin

How Yardstick Ranks Outcome, Not Activity

Yardstick joins token spend from your AI coding tools to merged outcomes from your Git provider and shows cost per merged PR by team, updated daily. The number you would put on an internal scoreboard is the ratio of work shipped to dollars spent, not the raw consumption total that started Uber's problem.

Because attribution is at the team and PR level, Yardstick can rank teams by cost per merged PR and by the change in that number over time, so a leaderboard rewards efficiency and improvement rather than spend. The team that ships more per dollar climbs.

For organizations that also need a backstop, Yardstick Treasury adds the cap Uber reached for, but as one layer under an outcome metric rather than as the whole strategy. The cap bounds the risk. The metric drives the behavior.

Yardstick - Measure the value AI creates

To see the platform or join the early access list, visit yardstick.fi/platform or review pricing for current plan options.

FAQ

What happened with Uber's AI coding budget?

As widely reported this spring, Uber rolled out AI coding tools including Claude Code and Cursor, drove adoption with internal leaderboards ranking teams by usage, and spent its entire 2026 budget for those tools in about four months. Adoption rose from 32% of engineers in February to roughly 95% using the tools monthly by spring, with around 70% of committed code coming through them. Per-engineer spend ran between $500 and $2,000 a month. The company later retired the leaderboards and capped spend at $1,500 per engineer per tool.

Why is ranking teams by token consumption a problem?

Because a leaderboard incentivizes whatever it measures. Ranking teams by token consumption rewards spending more, not shipping more, so teams optimize for usage and the budget burns while delivered output may not keep pace. Token spend measures activity, which is how hard the tools were worked, not output, which is how much finished work resulted. A team can top a usage leaderboard while shipping less than a team below it.

Is a spend cap enough to control AI coding costs?

A cap bounds the maximum spend but tells you nothing about whether the spend produces value. It answers how much was spent, not whether it was worth it, and it cannot distinguish a team getting enormous value under the cap from a team getting almost none. A cap is a useful backstop, but on its own it replaces an uncontrolled bad metric with a controlled bad metric. The durable fix is measuring output per dollar, with the cap as a secondary layer.

What should an AI coding leaderboard measure instead?

Cost per merged PR, or cost per resolved ticket, which express delivered work per dollar rather than raw consumption. Ranking on that ratio rewards efficiency: a team climbs by shipping more per dollar, and an engineer who gets the same output for less spend moves up rather than down. Ranking on the change in that ratio over time rewards improvement from any starting point, which keeps the competitive energy without punishing teams that began with messier codebases.

How do you attribute AI coding cost to specific teams?

You pull token spend at the team and PR level rather than as one blended monthly total, then join it to merged outcomes from your Git provider to produce cost per merged PR per team. That per-team, per-PR attribution is what makes an outcome-based ranking possible, because a single org-wide spend number cannot rank anyone or tell you where the money is turning into shipped work.

Recommended

Vukašin Kitanović · Co-founder

Leads commercial and go-to-market. Founder of Emberwood, an AI-driven cold-email agency for B2B, with a background in outbound, lead generation, and sales.

Partager

Continuer la lecture

Mesurez la valeur créée par votre IA.

Recevez les nouveaux articles et notes produit, environ une fois par mois.