Arena
Race your agents. Pick the winner.
Spawn the same goal across 2–4 agents, each in its own isolated worktree, then judge them on the data — cost, diff quality, files touched, scope violations. The auto-judge recommends; you make the call.
How a run works
One goal in, one winner out.
Set one goal, pick 2–4 agents
Write the goal once. Choose the agents you want to race — Claude Code, Codex, Cursor, or any of the others Operon drives.
Each agent runs in its own isolated worktree
Every agent works the exact same goal in a separate, isolated Git worktree. No collisions, no stepping on each other’s changes — every branch stays clean.
Operon judges every run on the data
As runs finish, each is scored on cost, diff quality, files touched, scope violations, and tool-success rate — the criteria are published, so you can see exactly why a run ranked where it did.
An advisory auto-judge recommends a winner
An LLM auto-judge weighs the signals and surfaces a recommended winner with rationale and a confidence score. It’s advisory only — the call is yours.
You pick the winner
The human picks the actual winner. The choice is write-once — enforced at the database, so a verdict is a verdict.
Losers reclaimed, winner routed to merge
The losing agents’ sessions end and their worktrees are reclaimed — uncommitted work is preserved. The winner routes straight into the Git Workbench merge flow.
And the winner is — the one the data backs
Every run is judged on cost, diff quality, files touched, and scope violations. The auto-judge recommends, but you pick the winner. Losers are reclaimed — their uncommitted work preserved — and the winner routes into the merge flow.

FAQ
Arena, answered
Q
On the data, not vibes.As each run finishes it is scored on five published criteria: cost, diff quality, files touched, scope violations, and tool-success rate. An LLM auto-judge then weighs those signals and surfaces a recommended winner with a rationale and a confidence score. It is advisory only: the recommendation is a starting point, and you pick the winner.
Q
No — you do.The auto-judge only recommends. The human picks the actual winner, and that choice is write-once, enforced at the database, so a verdict is a verdict. Even the advisory recommendation is stored separately from the human decision.
Q
Two to four agents on the same goal.Each runs the exact same objective in its own isolated Git worktree, so their diffs are directly comparable and nothing collides.
Q
Their sessions end and their worktrees are reclaimed.— but any uncommitted work is preserved, never destroyed. The winning run routes straight into the Git Workbench merge flow, so you go from race to shipped without leaving Operon.
Q
No — every agent gets a separate, isolated worktree.That is what makes a fair race possible: no shared state, no stepping on each other’s changes, and clean per-agent diffs to judge.
Get started
Stop guessing which agent won.
Private beta — join the waitlist to get an invite when your slot opens up.