Live

Tuner.

Rank the changes that move the score.

Tuner reads what Observe knows about an agent and ranks the changes most likely to raise its outcome score per dollar: a tighter prompt, a cheaper model that still clears the bar, different tools, better context. Every recommendation is explainable and tied to the exact metric it should move.

+0.6x top lift
Tuner · ranked by liftvalue/$
1Route triage to Haiku+0.6x
2Add a verify step+0.3x
3Trim the system prompt+0.1x
ranked per $ spent
How it works

Read, rank, apply.

01
Read

Tuner reads what Observe knows about the agent and its task.

02
Rank

The highest-impact changes, ranked by predicted outcome-per-dollar lift.

03
Apply

Make the change, then watch the exact metric it was meant to move.

The levers

What it improves.

01
Impact-ranked changes

Recommendations ordered by predicted lift in the outcome metric, not guesswork.

02
Model fit

Which model holds the outcome bar for this task at the lowest cost.

03
Evidence, not arithmetic

Every change says whether your own experiment proved it, or whether it is still just list-price maths.

04
Refuted changes stay visible

A swap your experiment disproved is kept and labelled, so nobody suggests it again next month.

05
Auto-Engineer loop

When an evaluator regresses, Tuner generates a candidate fix, validates against the dataset, and opens a PR with before-and-after metrics. Coach, not just scoreboard.

Who it's for

Who reaches for Tuner.

  • ·Teams iterating on agents
  • ·Agent builders tuning prompts and models
  • ·Eng leaders chasing outcome-per-dollar
Same suite

Pairs with the rest of Optimize.

Want Tuner in your stack?

We're onboarding design partners now. Join the waitlist to be in the Tuner cohort.

Just email is required. One email when Tuner goes live. Nothing else.