Experiments.
A/B/n test variants against real outcomes.
Experiments turns tuning into evidence. Run prompt, model, and config variants head to head against the real outcome metric, on live traffic or replayed history, and let statistical significance pick the winner. Builds on Arena and Replay.
Vary, test, promote.
Set up prompt, model, or config variants for one agent.
Run them against the real outcome metric, live or on replayed history.
Significance picks the winner. Promote it, fully logged and reversible.
What it improves.
Test multiple prompts, models, or configs at once against the same outcome metric.
Two variants, head to head, scored by a rubric-driven LLM judge or human votes. The fastest way to settle 'is this better?'
Winners called on statistical confidence and sample size, not a one-day spike.
Mirrors a slice of your real production calls. Nothing you send is stored, so there is no recorded history to replay.
Promote the winning variant to production, fully logged and reversible.
Who reaches for Experiments.
- ·Teams shipping agent changes
- ·Operators who want proof before rollout
- ·Platform teams standardizing agents
Pairs with the rest of Optimize.
Want Experiments in your stack?
We're onboarding design partners now. Join the waitlist to be in the Experiments cohort.
Just email is required. One email when Experiments goes live. Nothing else.