Measure the value AI creates.
Impact, not adoption. Rank what AI delivered and govern the spend behind it. Coding agents, support bots, SDRs, and research workflows. One platform from the first metric to the budget that funds it.
Every agent. Every cloud. One ROI number.
Your agents don't live in one vendor's ecosystem, so your ROI number can't either. Yardstick measures across every coding agent, gateway, and cloud, not just what reaches a single vendor's dashboard.
Three suites. One spine.
Measure the value AI creates, optimize every agent, and govern the money behind it. 21 tools on one shared identity, ledger, and spine.
Measure impact, not adoption.
Adoption dashboards tell you who used AI this week: seats, sessions, acceptance rates. None of it is value. The only number that survives a CFO review is what your agents shipped and what it cost. Yardstick measures exactly that: cost per merged PR, cost per resolved issue, change-failure rate, and the ROI behind every dollar of AI spend.
- AI adoption %
- Seats active this week
- Tokens consumed
- Lines of code
- Suggestions accepted
- Cost per merged PR
- Cost per resolved issue
- Value created
- Value per dollar
- Change-failure rate
90 days of history in minutes.
No waiting for data to accrue. Connect a repo and Yardstick backfills your recent merged PRs and reverts on the spot, then prices them.
One-click GitHub, GitLab, or Bitbucket install. No SDK required to start.
Up to 90 days of merged PRs and reverts pulled on connect, so the dashboard opens with your track record, not a blank state.
Cost per merged PR, value created, and change-failure rate. Board-ready before your next standup.
Backfill covers your last 90 days of PRs. Cost capture starts streaming the moment an agent runs.
See what every agent is worth
Score AI in the units each vertical respects: cost per merged PR, resolution rate, citation accuracy. Activity metrics lie. Outcomes do not.
Explore ObserveTurn every score into a better agent
Observe gives you the score. Optimize tells you how to raise it: the prompt to tighten, the model to swap, the config to change, ranked by the lift it buys per dollar.
Explore OptimizeGovern every dollar AI spends
Turn AI budgets into allowances, policies, and caps that respond to verified value. Less waste, a smaller compute footprint, a treasury that runs itself.
Explore TreasuryOne pipeline. Five stages.
The spine ships first so every new suite plugs in without rewriting infrastructure. Adding a vertical takes weeks, not years.
Universal LLM tracing plus suite collectors: repos, CRMs, support queues, conversations.
Versioned scoring engine. Per-suite score functions on one shared spine.
Profiles, leaderboards, and CFO-ready reports on spend versus outcomes.
Allowances and policies per team. One ledger of cost and outcome.
Cost-aware routing and waste controls. Less spend, same outcomes.
AI is finally doing real work. The proof of what it does, and the budget behind it, still run on screenshots and guesswork. We're replacing that with a yardstick.
Three lanes. One product.
Adoption dashboards and engineering-intelligence tools count who uses AI and how developers perform. Gateways meter model calls. AI FinOps tracks the bill. Yardstick is where those lanes meet, scored on the work your AI coding agents actually shipped.
Based on public docs as of mid 2026. Corrections welcome: hello@yardstick.fi
We read only the tools you connect. Nothing leaves your stack without consent.
All data is encrypted in transit and at rest.
Building to enterprise security standards from day one.
Export anytime. We never sell it, and never train on it.
An instrument, not a dashboard
We score merged PRs and resolved tickets, never tokens or lines of code.
We measure and govern budgets. We never mark up your provider costs.
Every score is explainable. If we cannot show the math, we do not ship it.
We read from the tools you connect. Nothing leaves your stack without consent.
Launch pricing.
Free forever for three agents. Paid plans bundle Treasury so you only pay one thing per tier, not one per suite.
- ·3 active agents
- ·Unlimited viewers, always free
- ·Universal LLM & tool tracing
- ·Soft budget alerts
- ·Public agent profile
- ·7-day retention · community
- ·5 active agents
- ·Vendor admin-API cost ingestion
- ·Outcome value scoring
- ·Public agent profile
- ·90-day retention · email
- ·then $5 / extra agent
- ·Everything in Solo
- ·PR-level AI-vs-human attribution
- ·SDK for per-PR token attribution
- ·Reconciliation (SDK vs admin-API gap)
- ·AI-tool ROI & recommendation engine
- ·Treasury included up to $50k/mo governed
- ·then $4 / extra agent
- ·Everything in Pro
- ·Per-team cost allocation
- ·Higher Treasury cap ($250k/mo governed)
- ·2-year retention · priority support
- ·Shadow-AI detection
- ·R&D capitalization reports
- ·then $3 / extra agent
For orgs governing real AI spend, not just engineering. Unlimited agents, unlimited Treasury, with the compliance and procurement bundle on top.
- ·Everything in Scale, unlimited
- ·EU AI Act audit & evidence packs
- ·Custom guardrails & org-wide AI gateway
- ·SSO / SAML / SCIM
- ·SOC 2 Type II (Coming soon)
- ·On-prem option · dedicated CSM
Fair questions.
Join the waitlist. Early signups get first access and design-partner slots across the suites while we're in private beta.
DX measures engineering broadly and ships an AI Measurement Framework. Yardstick measures the one thing it does not price: the cost and value of what your AI coding agents actually ship, with PR-level attribution and EU AI Act audit packs on top. Plenty of teams run both.
Claude Code, Codex, Gemini, and Cursor for dev tools, autonomous loop engineers like Devin, Factory, and OpenHands, plus the support, sales, and research agents your teams already run. The models underneath are metered too: Anthropic, OpenAI, Gemini, Kimi, DeepSeek, Mistral, Groq, Together, Fireworks, Perplexity, and OpenRouter all route through the gateway and get priced per token. A provider billing key or a lightweight connector, no rip-and-replace.
It is where we are taking Treasury, and it is not built yet. The idea: give each developer an annual AI allowance and let whatever they do not spend convert to compensation, gated on actually shipping outcomes, so nobody is taught to burn budget just to keep it. Today Treasury does the governance half, meaning budgets, burn tracking, projected unused and enforcing caps. The payout rail is a real design problem (it has to move money without us ever touching it) and we would rather say so than imply it ships.
No. Treasury produces statements, ledger entries, and payroll exports. Humans and existing rails move the money.
If AI does real work for you, measure it.
Early access opens with our first design-partner cohort. We're onboarding across every vertical now.
Just email is required. No spam. One email when we go live.