Live

Code.

Measure cost per outcome on AI dev tools. Outcomes, not LOC.

Code measures whether your AI dev tools are actually shipping work, not just burning tokens. It plugs into Claude Code, Codex, Gemini, and Cursor with a provider billing key and correlates token spend with the outcomes that matter to engineering leaders. It covers both assistive copilots and autonomous loop engineers like Devin and Factory, with runaway-loop guardrails so a spinning agent never quietly runs up the bill.

12 checks pass
PR #4821main ← feat/auth
-const ok = verifyToken(t)
+const ok = await verify(t, key)
if (!ok) return deny()
+return grant(session)
Merged Verified +$1,840
How it works

Connect, measure, govern.

01
Connect

Link Claude Code, Cursor, Codex, and GitHub with read-only OAuth.

02
Measure

Every merged PR and resolved issue is priced against the tokens it cost.

03
Govern

Cap spend per team, route to cheaper models, and pay out on value.

The measurements

What it measures.

01
Cost per merged PR

The headline efficiency metric. Token spend divided by PRs merged in the window.

02
Cost per resolved issue

Linked through Linear or Jira webhooks, attributed to the agent whose PR closed the issue. Outcomes that finance actually understands.

03
Acceptance and revert rates

How often the developer keeps suggested edits. How often AI-touched PRs get reverted later.

04
Outcome attribution

Which agent runs led to which PRs. Outcome metrics, NOT lines of code.

05
Per-team rollups

Cost allocation to teams, projects, and repos. Per-developer detail gated behind access policy.

06
Provider normalization

Claude Code, Codex, Gemini, Cursor metered differently, converted to one comparable value unit.

07
PR-level AI-vs-human attribution

Blame analysis at the line and PR level so you know what the agent actually wrote vs the human, and tie outcomes back to authorship.

08
Loop-engineer measurement

Autonomous coding agents (Devin, Factory, OpenHands, Cursor Agent) scored on autonomy rate, cost per run, and iterations-to-merge, not just tokens burned.

09
Runaway-loop guardrails

A per-run cost cap and a wasted-loop alert catch an agent spinning without merging, and page you before the bill lands.

10
Agent leaderboard

Rank every coding agent, assistive and autonomous, head-to-head on cost per merged PR and revert rate, with statistical significance on the gap.

11
DORA & SPACE composite score

DORA and SPACE rolled into one number per team. The score executive stakeholders already recognize, calibrated to AI-shipped work.

Who it's for

Who reaches for Code.

  • ·Engineering managers at series B+ tech companies
  • ·VPEs running Claude Code, Codex, Gemini, Cursor
  • ·CFOs evaluating AI dev tool ROI
Same suite

Pairs with the rest of Observe.

Want Code in your stack?

We're onboarding design partners now. Join the waitlist to be in the Code cohort.

Just email is required. One email when Code goes live. Nothing else.