Measure the value AI creates.

Impact, not adoption. Rank what AI delivered and govern the spend behind it. Coding agents, support bots, SDRs, and research workflows. One platform from the first metric to the budget that funds it.

app.yardstick.fi/
Overview
Acme · 6 agents · 4 repos measured
Last 30 days
Value created
$214k21%
ROI
7.2x14%
after review cost
AI spend / merged PR
$3.4018%
Merged PRs
1,2846%
Acceptance rate
92%
Trend
Merged PRs AI spend
AprMayJun
Recent activity streaming
claude-code
merged PR #318 · auth refactor
$3.40 / PR
cursor · web
resolved issue #2041 · checkout bug
94 score
Works with

Every agent. Every cloud. One ROI number.

Your agents don't live in one vendor's ecosystem, so your ROI number can't either. Yardstick measures across every coding agent, gateway, and cloud, not just what reaches a single vendor's dashboard.

The platform

Three suites. One spine.

Measure the value AI creates, optimize every agent, and govern the money behind it. 21 tools on one shared identity, ledger, and spine.

The metric that matters

Measure impact, not adoption.

Adoption dashboards tell you who used AI this week: seats, sessions, acceptance rates. None of it is value. The only number that survives a CFO review is what your agents shipped and what it cost. Yardstick measures exactly that: cost per merged PR, cost per resolved issue, change-failure rate, and the ROI behind every dollar of AI spend.

Adoption metrics
  • AI adoption %
  • Seats active this week
  • Tokens consumed
  • Lines of code
  • Suggestions accepted
Impact metrics
  • Cost per merged PR
  • Cost per resolved issue
  • Value created
  • Value per dollar
  • Change-failure rate
Time to value

90 days of history in minutes.

No waiting for data to accrue. Connect a repo and Yardstick backfills your recent merged PRs and reverts on the spot, then prices them.

1Minutes
Connect a repo

One-click GitHub, GitLab, or Bitbucket install. No SDK required to start.

2Instant
History backfilled

Up to 90 days of merged PRs and reverts pulled on connect, so the dashboard opens with your track record, not a blank state.

3Same day
First ROI report

Cost per merged PR, value created, and change-failure rate. Board-ready before your next standup.

Backfill covers your last 90 days of PRs. Cost capture starts streaming the moment an agent runs.

Observe · Measure

See what every agent is worth

Score AI in the units each vertical respects: cost per merged PR, resolution rate, citation accuracy. Activity metrics lie. Outcomes do not.

Explore Observe
app.yardstick.fi/agents/claude-code
Cost / merged PR
$42.10
▼ 9.4%
Value / dollar
2.3x
▲ 0.4x
91%
accepted
Acceptance rate
Outputs merged without rework, across 1,284 outcomes this month.
Optimize · Improve

Turn every score into a better agent

Observe gives you the score. Optimize tells you how to raise it: the prompt to tighten, the model to swap, the config to change, ranked by the lift it buys per dollar.

Explore Optimize
app.yardstick.fi/optimize/tuner
Predicted lift
+18%
outcome
At lower cost
24%
▼ spend
12
changes
Ranked changes
Highest-impact tweaks for this agent, ranked by outcome per dollar.
Treasury · Govern

Govern every dollar AI spends

Turn AI budgets into allowances, policies, and caps that respond to verified value. Less waste, a smaller compute footprint, a treasury that runs itself.

Explore Treasury
app.yardstick.fi/treasury
Quarter budget
$50k
on track
Projected payout
$21.6k
▲ 0.4k
68%
used
Budget burn
Spent against cap. Use-it-or-cash-it at quarter end.
How it works

One pipeline. Five stages.

The spine ships first so every new suite plugs in without rewriting infrastructure. Adding a vertical takes weeks, not years.

1
Measure

Universal LLM tracing plus suite collectors: repos, CRMs, support queues, conversations.

2
Rank

Versioned scoring engine. Per-suite score functions on one shared spine.

3
Report

Profiles, leaderboards, and CFO-ready reports on spend versus outcomes.

4
Budget

Allowances and policies per team. One ledger of cost and outcome.

5
Optimize

Cost-aware routing and waste controls. Less spend, same outcomes.

AI is finally doing real work. The proof of what it does, and the budget behind it, still run on screenshots and guesswork. We're replacing that with a yardstick.

Comparison

Three lanes. One product.

Adoption dashboards and engineering-intelligence tools count who uses AI and how developers perform. Gateways meter model calls. AI FinOps tracks the bill. Yardstick is where those lanes meet, scored on the work your AI coding agents actually shipped.

Yardstick
AWS CloudWatch
Revenium
Pensero
Olakai
Faros
Trace AI agent activity
Coding-agent wedge focus
Partial
·
Partial
Partial
Score agent outcomes, not just tokens
·
·
Partial
Partial
PR-level AI-vs-human attribution
·
·
Partial
Partial
Partial
Anonymous cross-org benchmarks
·
·
Partial
·
Partial
Budgets, allowances, kill switches
Partial
·
Partial
·
Cost-aware model routing
·
·
·
·
·
EU AI Act audit packs & SOC 2 evidence
·
·
·
·
Use-it-or-cash-it dev payouts
·
·
·
·
·
·
Open source SDK & attestation standards
Planned
·
·
·
·
·

Based on public docs as of mid 2026. Corrections welcome: hello@yardstick.fi

Security & trust
Privacy by default

We read only the tools you connect. Nothing leaves your stack without consent.

Encrypted throughout

All data is encrypted in transit and at rest.

SOC 2 in progress

Building to enterprise security standards from day one.

Your data is yours

Export anytime. We never sell it, and never train on it.

Why Yardstick

An instrument, not a dashboard

Outcomes, not activity

We score merged PRs and resolved tickets, never tokens or lines of code.

Your spend is your spend

We measure and govern budgets. We never mark up your provider costs.

Show the method

Every score is explainable. If we cannot show the math, we do not ship it.

Privacy by default

We read from the tools you connect. Nothing leaves your stack without consent.

Pricing

Launch pricing.

Free forever for three agents. Paid plans bundle Treasury so you only pay one thing per tier, not one per suite.

Save 20%
Free forever for 3 agents. Every paid plan starts with a 14-day trial. Verified scores and leaderboards stay permanent on every plan.
Free
$0
3 agents · alerts up to $5k governed
  • ·3 active agents
  • ·Unlimited viewers, always free
  • ·Universal LLM & tool tracing
  • ·Soft budget alerts
  • ·Public agent profile
  • ·7-day retention · community
Solo
$15/mo billed annually
Launch pricing
5 agents · any 1 suite
  • ·5 active agents
  • ·Vendor admin-API cost ingestion
  • ·Outcome value scoring
  • ·Public agent profile
  • ·90-day retention · email
  • ·then $5 / extra agent
Recommended
Pro
$63/mo billed annually
Launch pricing
25 agents · attribution & ROI
  • ·Everything in Solo
  • ·PR-level AI-vs-human attribution
  • ·SDK for per-PR token attribution
  • ·Reconciliation (SDK vs admin-API gap)
  • ·AI-tool ROI & recommendation engine
  • ·Treasury included up to $50k/mo governed
  • ·then $4 / extra agent
Scale
$239/mo billed annually
Launch pricing
150 agents · the CFO bundle
  • ·Everything in Pro
  • ·Per-team cost allocation
  • ·Higher Treasury cap ($250k/mo governed)
  • ·2-year retention · priority support
  • ·Shadow-AI detection
  • ·R&D capitalization reports
  • ·then $3 / extra agent
Enterprise
Unlimited · compliance & procurement

For orgs governing real AI spend, not just engineering. Unlimited agents, unlimited Treasury, with the compliance and procurement bundle on top.

  • ·Everything in Scale, unlimited
  • ·EU AI Act audit & evidence packs
  • ·Custom guardrails & org-wide AI gateway
  • ·SSO / SAML / SCIM
  • ·SOC 2 Type II (Coming soon)
  • ·On-prem option · dedicated CSM
FAQ

Fair questions.

Join the waitlist. Early signups get first access and design-partner slots across the suites while we're in private beta.

DX measures engineering broadly and ships an AI Measurement Framework. Yardstick measures the one thing it does not price: the cost and value of what your AI coding agents actually ship, with PR-level attribution and EU AI Act audit packs on top. Plenty of teams run both.

Claude Code, Codex, Gemini, and Cursor for dev tools, autonomous loop engineers like Devin, Factory, and OpenHands, plus the support, sales, and research agents your teams already run. The models underneath are metered too: Anthropic, OpenAI, Gemini, Kimi, DeepSeek, Mistral, Groq, Together, Fireworks, Perplexity, and OpenRouter all route through the gateway and get priced per token. A provider billing key or a lightweight connector, no rip-and-replace.

It is where we are taking Treasury, and it is not built yet. The idea: give each developer an annual AI allowance and let whatever they do not spend convert to compensation, gated on actually shipping outcomes, so nobody is taught to burn budget just to keep it. Today Treasury does the governance half, meaning budgets, burn tracking, projected unused and enforcing caps. The payout rail is a real design problem (it has to move money without us ever touching it) and we would rather say so than imply it ships.

No. Treasury produces statements, ledger entries, and payroll exports. Humans and existing rails move the money.

If AI does real work for you, measure it.

Early access opens with our first design-partner cohort. We're onboarding across every vertical now.

Just email is required. No spam. One email when we go live.

Yardstick · AI agent ROI platform: measure, optimize & govern AI spend