Live

Replay.

Prove a change on your real traffic before you switch.

Replay answers the question a config change actually raises: would this hold up on our work? It mirrors live calls to the alternate, has a judge score both answers head to head, and reports the cost difference and a significance test. It does NOT replay stored prompts, because Yardstick keeps token counts and never the request itself, which is a privacy position worth more than the feature would have been.

18% cheaper
Replay · counterfactual$/PR
Current · gpt-4o$0.042
Alternate · haiku & verify$0.034
Same acceptance, 18% spend.
312 runs replayed
How it works

Record, replay, compare.

01
Record

Yardstick keeps your agent's full run history.

02
Replay

Re-run it against a new prompt, model, or configuration.

03
Compare

Ship the change only when the counterfactual wins.

The measurements

What it measures.

01
Mirrored live traffic

Every comparison is a real request your agent just made, sent to both sides. No synthetic prompts, no stored history.

02
Model, prompt or config arms

Change one thing per arm. Everything else about the request is byte-identical, which is what makes the pair comparable.

03
Judged head to head

An LLM judge scores both answers, in both presentation orders, so position bias cannot pick the winner.

04
Winners called on significance

A sign test on the paired verdicts, so 8-4 is reported as the coin flip it is rather than a result.

05
The experiment's own cost, shown

Mirroring spends your provider money. What the comparison cost is on the page, not buried.

Who it's for

Who reaches for Replay.

  • ·Teams iterating on prompts
  • ·Agent builders tuning configs
  • ·Teams shipping agent updates
Same suite

Pairs with the rest of Observe.

Want Replay in your stack?

We're onboarding design partners now. Join the waitlist to be in the Replay cohort.

Just email is required. One email when Replay goes live. Nothing else.