Live

Arena.

Prove an agent works before it touches production.

Arena is the proving ground. Run agents against curated scenario suites before they touch production, real customers, or real code. Replay tests against your own history; Arena tests against the situations you have not hit yet.

1,284 duels
Arena · head-to-headwin rate
Claude62%
GPT-4o38%
Judged on real merged outcomes.
winner: Claude
How it works

Curate, run, gate.

01
Curate

Build scenario suites per vertical, or record real incidents.

02
Run

Score every agent version against the suite.

03
Gate

Block the deploy until it clears, wired into CI.

The measurements

What it measures.

01
Your scenarios, your standard

The situations you have not hit yet: hostile tickets, malformed repos, adversarial prompts. Each one carries what a correct answer must do, in your words, and that is what gets scored.

02
Deploy gates

Block a release until the agent clears its suite. Wire into CI like any other test.

03
Custom scenarios

Record a real incident once, replay it against every future version forever.

04
Score history

Every run versioned. Watch capability rise and regressions appear across versions.

Who it's for

Who reaches for Arena.

  • ·Teams shipping their first agent
  • ·Operators gating deploys on eval scores
  • ·Platforms certifying internal agents
Same suite

Pairs with the rest of Observe.

Want Arena in your stack?

We're onboarding design partners now. Join the waitlist to be in the Arena cohort.

Just email is required. One email when Arena goes live. Nothing else.