Replay.
Prove a change on your real traffic before you switch.
Replay answers the question a config change actually raises: would this hold up on our work? It mirrors live calls to the alternate, has a judge score both answers head to head, and reports the cost difference and a significance test. It does NOT replay stored prompts, because Yardstick keeps token counts and never the request itself, which is a privacy position worth more than the feature would have been.
Record, replay, compare.
Yardstick keeps your agent's full run history.
Re-run it against a new prompt, model, or configuration.
Ship the change only when the counterfactual wins.
What it measures.
Every comparison is a real request your agent just made, sent to both sides. No synthetic prompts, no stored history.
Change one thing per arm. Everything else about the request is byte-identical, which is what makes the pair comparable.
An LLM judge scores both answers, in both presentation orders, so position bias cannot pick the winner.
A sign test on the paired verdicts, so 8-4 is reported as the coin flip it is rather than a result.
Mirroring spends your provider money. What the comparison cost is on the page, not buried.
Who reaches for Replay.
- ·Teams iterating on prompts
- ·Agent builders tuning configs
- ·Teams shipping agent updates
Pairs with the rest of Observe.
Want Replay in your stack?
We're onboarding design partners now. Join the waitlist to be in the Replay cohort.
Just email is required. One email when Replay goes live. Nothing else.