블로그
Guide

Your Team Says AI Saves Them Two Hours a Day. That Is the Least Reliable Number You Have.

Ask your engineers how much time AI saves them and someone will say two hours a day. It feels like proof the spend is working. A randomized trial found developers felt 20% faster with AI while actually finishing 19% slower, a perception gap of almost 40 points. Here is why self-reported time savings cannot go in a budget, and what number to use instead.

Aug 26, 2026·7 min read
Your Team Says AI Saves Them Two Hours a Day. That Is the Least Reliable Number You Have.
The time savings your team reports feeling is the number that ends up justifying the AI budget. A randomized trial found the self-reported feeling can be off from measured reality by almost 40 points.

Ask your engineers how much time AI saves them. Someone will say two hours a day. It will sound certain, it will sound obvious, and everyone in the room will nod.

That number feels like proof the tools are working. It is the least reliable number you have, and it is usually the one that ends up in the deck that justifies the spend.

I am not saying your engineers are wrong to like the tools, or that AI does not help. I am saying the specific thing you are measuring when you ask how it feels is not the thing you think you are measuring, and you cannot put it in a budget.

Developer working at a desk with code on screen, looking focused

The Gap Between Feeling Faster and Being Faster

There is a study that should be taped to the wall of every engineering org running an AI pilot. In 2025, METR ran a randomized controlled trial with experienced open-source developers doing real tasks on codebases they knew well. Before the trial, the developers expected AI to make them about 24% faster. Afterward, they reported that it had made them roughly 20% faster.

They were actually 19% slower.

The measured result and the reported feeling were off by almost 40 points, and they pointed in opposite directions. The developers did not feel slower and hide it. They genuinely felt faster while genuinely taking longer. That is the part that should make you nervous, because it means the feeling is not a soft version of the truth. It is uncorrelated with the truth, and it is confident.

This is one study on one population, and I would not hang a company on it alone. But you do not need to believe the exact number to take the point. If self-reported speed can be that far off from measured speed among experienced engineers on familiar code, then the two hours a day your team feels is not evidence of anything you can bank on.

Why Does AI Feel Faster Even When It Is Not?

Because the parts of the work AI removes are the parts that feel like effort, and the parts it adds are the parts that do not feel like work at all.

AI takes away the blank page. It takes away the boilerplate, the syntax you half remember, the first draft you did not want to write. Those moments are where friction lives, so removing them feels like a large speedup even when it is a small one. What it adds is quieter: reading generated code to see if it is right, catching the subtle thing it got wrong, re-prompting when the first answer missed, and the review load downstream when more code shows up than before. None of that registers as effort in the moment. It registers as normal work. So the effort goes down while the clock stays the same or goes up, and your brain reports the effort, not the clock.

That is the whole illusion. You are remembering how it felt, and it felt easier. Easier and faster are not the same thing, and the budget only cares about one of them.

Pro Tip: Next time someone says AI saves them two hours a day, ask a gentle follow-up: two hours measured how? Not to catch them out, but because the answer is almost always a feeling, and naming that out loud is the start of measuring the real thing instead.

Why the Self-Reported Number Cannot Go in a Budget

A budget decision needs a number that means the same thing to the person spending it and the person paying for it. Self-reported time savings fails that test in three ways.

First, it is not comparable across people. One engineer's two hours is another's thirty minutes, and neither of them measured it. You cannot add up feelings from twelve engineers and get a team number that survives a finance conversation.

Second, it moves with mood, not output. A good week with the tools feels like a big saving. The same output in a stressful week feels like less. The number tracks how the work felt, which is exactly the thing a budget should not be built on.

Third, and this is the one that actually bites, it can be pointing the wrong way entirely. The METR result is not that the feeling is a bit optimistic. It is that the feeling was positive while the reality was negative. A team can report real, sincere time savings while their actual throughput per dollar is flat or down, and the self-reported number will never tell you, because it is measuring satisfaction, not delivery.

Two engineers discussing work in front of a monitor in an office

To Be Clear, This Is Not an Argument Against AI

I want to be careful here, because the easy version of this post is AI is secretly bad and everyone is fooling themselves, and that is not what I think.

The point is narrower and more useful. The perception gap runs in both directions. Your team feeling faster is not proof the spend is working, and it is also not proof the spend is wasted. It is not proof of anything, because it is not connected to the output. Plenty of teams are getting real gains from these tools. Plenty are not. The self-reported number cannot tell those two groups apart, which is the entire problem.

So the move is not to stop using AI, or to stop trusting your engineers. It is to stop asking the feeling to do a job it cannot do. Let people like the tools. Just do not let how much they like the tools be the number you take to the board.

What to Measure Instead of the Feeling

You replace a self-reported number with an observed one. The feeling asks how did it go. The measurement asks what came out. Those produce different data, and only one of them survives scrutiny.

From Reported Feeling to Measured Output

  1. Cost per merged PR, not hours saved · A merged PR is an unambiguous unit of shipped work. Total AI spend divided by merged PRs, tracked against a pre-AI baseline, tells you whether the money is turning into delivery. No one has to estimate anything.
  2. Lead time from commit to merge, not felt speed · If AI is genuinely making the team faster, work should move through the pipeline faster, not just feel faster to write. If lead time is flat while people report big savings, the gain is being eaten downstream in review.
  3. Correction rate, not confidence · How much AI-generated code gets materially rewritten before it merges. A high correction rate is the hidden work the good feeling is not counting.
  4. Throughput per dollar over time, not a one-time survey · The trend is the truth. A survey is a snapshot of a mood. Watching cost per merged PR move month over month tells you what a one-time feeling never can.

Pro Tip: You do not have to choose between the survey and the measurement. Run both and put them side by side. When the felt savings are high and the measured throughput is flat, that gap is not a contradiction to explain away. It is the most important finding you have, and it is pointing at where the work is leaking.

Self-Reported Signal Versus Measured Signal

What your team feelsWhat you can measure
UnitHours saved per dayCost per merged PR
Comparable across peopleNoYes
Survives a finance conversationNoYes
Can point the wrong directionYes, by ~40 points in one trialNo, it is the outcome
What it actually tracksHow the work feltWhat the work produced
Good forMorale, tool satisfactionBudget, ROI, renewal

Both columns are worth knowing. Only the right one belongs in the decision.

Key Takeaways

  • Self-reported time savings measures satisfaction, not output. Your team feeling faster is real as a feeling and unreliable as a metric, because it tracks how the work felt, not what it produced.
  • The perception gap can be enormous. In a randomized trial, developers felt 20% faster with AI while measuring 19% slower, a gap of almost 40 points pointing the opposite way. You do not need the exact number to stop trusting the feeling.
  • AI removes the parts that feel like effort and adds the parts that do not. Less blank page and boilerplate feels like a big speedup. The reading, correcting, and downstream review that replace it do not register as work, so the effort drops while the clock does not.
  • The gap runs both ways, so this is not anti-AI. The feeling cannot prove the spend works or that it is wasted. It is simply disconnected from output, which is why it cannot anchor a budget.
  • Replace the feeling with cost per merged PR, lead time, and correction rate. Measured against a baseline and tracked over time, those tell you what a survey never can: whether the money is turning into shipped work.

Why This Number Keeps Ending Up in the Deck Anyway

The self-reported number wins because it is free and it is available. You can get it by asking, today, with no instrumentation. The measured number takes work to build, so when a board date is coming and someone needs a slide about AI ROI, the feeling is what is within reach. It goes in the deck because it is there, not because it is right.

The pattern I keep running into is a leader who half knows this. They will say the two hours a day out loud and then, almost in the same breath, add something like but who really knows. They can feel that the number is soft. What they are missing is not the skepticism, they already have it. What they are missing is the other number, the measured one, that would let them stop guessing. So they present the feeling with a caveat and hope nobody pushes, and usually nobody does, until the quarter finance decides to push.

The teams that come out of that conversation well are not the ones whose engineers feel the most productive. They are the ones who can show that the felt productivity and the measured output agree, or, just as valuable, show exactly where they disagree and what they are doing about it. That second story is stronger than a confident two hours a day, because it is the one thing a survey can never be. It is checkable.

Let your team love the tools. Just measure what the tools ship.

· Vukasin

How Yardstick Measures Output Instead of Feeling

Yardstick joins token spend from your AI coding tools to merged outcomes from your Git provider and shows cost per merged PR, lead time, and correction rate on AI-touched code, updated daily. It is the measured number that sits next to the survey, so you can see whether felt productivity and shipped output actually agree.

Because the data is attributed at the PR level, Yardstick shows where the two diverge: high reported savings with flat throughput points straight at the review or correction load that the good feeling is not counting. That divergence is the finding most teams never get to see.

For the budget conversation, Yardstick replaces the caveated two hours a day with a number that survives finance: throughput per dollar, tracked against a pre-AI baseline, over time rather than in a one-off survey.

Yardstick - Measure the value AI creates

To see the platform or join the early access list, visit yardstick.fi/platform or review pricing for current plan options.

FAQ

Are self-reported AI productivity gains just wrong?

Not wrong, unreliable. Self-reported time savings accurately capture how the work felt, which is real information about tool satisfaction and morale. The problem is using that feeling as a measure of output. A randomized trial found developers felt about 20% faster with AI while measuring about 19% slower, so the feeling was both confident and pointing the wrong way. The feeling is valid as a feeling and unusable as a budget number.

What was the METR study and what did it find?

METR published a randomized controlled trial in 2025 in which experienced open-source developers completed real tasks on codebases they knew, with and without AI tools. The developers expected to be roughly 24% faster with AI and reported feeling about 20% faster afterward, but were measured to be about 19% slower. The headline is not the exact figure, which comes from one study on one population, but the size and direction of the gap between perceived and actual productivity.

Why does AI coding feel faster even when it is not?

Because AI removes the parts of the work that feel like effort, such as the blank page and boilerplate, and adds parts that do not register as effort, such as reading generated code, correcting subtle errors, re-prompting, and heavier downstream review. The brain remembers effort, not elapsed time, so the work feels easier even when the clock is flat or higher. Easier and faster are different things, and only the clock matters to a budget.

What should engineering teams measure instead of time saved?

Observed output rather than reported feeling: cost per merged PR against a pre-AI baseline, lead time from commit to merge, and the correction rate on AI-generated code. These are comparable across engineers, survive a finance conversation, and cannot silently point the wrong direction the way a self-reported figure can. Tracked over time rather than as a one-time survey, they show whether AI spend is turning into shipped work.

Does this mean AI coding tools are not worth it?

No. The perception gap runs in both directions, so a good feeling is not proof the spend works and not proof it is wasted. Many teams get real, measurable gains from these tools and many do not, and the self-reported number cannot tell the two groups apart. The conclusion is not to stop using AI, it is to stop using the feeling as the metric and measure the output instead.

Recommended

Vukašin Kitanović · Co-founder

Leads commercial and go-to-market. Founder of Emberwood, an AI-driven cold-email agency for B2B, with a background in outbound, lead generation, and sales.

공유

AI가 만드는 가치를 측정하세요.

새 글과 제품 소식을 한 달에 한 번 정도 보내드립니다.