Live

Experiments.

A/B/n test variants against real outcomes.

Experiments turns tuning into evidence. Run prompt, model, and config variants head to head against the real outcome metric, on live traffic or replayed history, and let statistical significance pick the winner. Builds on Arena and Replay.

p < 0.05
Experiment · 3 variantsacceptance
A control18%
B verify-first 24%
C cheaper model15%
Winner: B (+33%)
1,902 runs
How it works

Vary, test, promote.

01
Vary

Set up prompt, model, or config variants for one agent.

02
Test

Run them against the real outcome metric, live or on replayed history.

03
Promote

Significance picks the winner. Promote it, fully logged and reversible.

The levers

What it improves.

01
A/B/n variants

Test multiple prompts, models, or configs at once against the same outcome metric.

02
Pairwise eval with LLM-as-judge

Two variants, head to head, scored by a rubric-driven LLM judge or human votes. The fastest way to settle 'is this better?'

03
Significance built in

Winners called on statistical confidence and sample size, not a one-day spike.

04
Live traffic

Mirrors a slice of your real production calls. Nothing you send is stored, so there is no recorded history to replay.

05
One-click promoteUp next

Promote the winning variant to production, fully logged and reversible.

Who it's for

Who reaches for Experiments.

  • ·Teams shipping agent changes
  • ·Operators who want proof before rollout
  • ·Platform teams standardizing agents
Same suite

Pairs with the rest of Optimize.

Want Experiments in your stack?

We're onboarding design partners now. Join the waitlist to be in the Experiments cohort.

Just email is required. One email when Experiments goes live. Nothing else.