Skip to content

Arena

Race your agents. Pick the winner.

Spawn the same goal across 2–4 agents, each in its own isolated worktree, then judge them on the data — cost, diff quality, files touched, scope violations. The auto-judge recommends; you make the call.

How a run works

One goal in, one winner out.

One goal · 2–4 agents

Set one goal, pick 2–4 agents

Write the goal once. Choose the agents you want to race — Claude Code, Codex, Cursor, or any of the others Operon drives.

Isolated worktrees

Each agent runs in its own isolated worktree

Every agent works the exact same goal in a separate, isolated Git worktree. No collisions, no stepping on each other’s changes — every branch stays clean.

Judged on data

Operon judges every run on the data

As runs finish, each is scored on cost, diff quality, files touched, scope violations, and tool-success rate — the criteria are published, so you can see exactly why a run ranked where it did.

Advisory auto-judge

An advisory auto-judge recommends a winner

An LLM auto-judge weighs the signals and surfaces a recommended winner with rationale and a confidence score. It’s advisory only — the call is yours.

You pick the winner

You pick the winner

The human picks the actual winner. The choice is write-once — enforced at the database, so a verdict is a verdict.

Losers reclaimed

Losers reclaimed, winner routed to merge

The losing agents’ sessions end and their worktrees are reclaimed — uncommitted work is preserved. The winner routes straight into the Git Workbench merge flow.

Arena · judged

And the winner is — the one the data backs

Every run is judged on cost, diff quality, files touched, and scope violations. The auto-judge recommends, but you pick the winner. Losers are reclaimed — their uncommitted work preserved — and the winner routes into the merge flow.

Operon's Arena view after a judged run: four agents — Claude Code, Codex, Gemini and Cursor — each showing cost, files, traces and scope violations. The auto-judge recommends Claude Code at 100% confidence, and Claude Code is picked as the winner with a diff quality score of 100.

FAQ

Arena, answered

Q

On the data, not vibes.As each run finishes it is scored on five published criteria: cost, diff quality, files touched, scope violations, and tool-success rate. An LLM auto-judge then weighs those signals and surfaces a recommended winner with a rationale and a confidence score. It is advisory only: the recommendation is a starting point, and you pick the winner.

Q

No — you do.The auto-judge only recommends. The human picks the actual winner, and that choice is write-once, enforced at the database, so a verdict is a verdict. Even the advisory recommendation is stored separately from the human decision.

Q

Two to four agents on the same goal.Each runs the exact same objective in its own isolated Git worktree, so their diffs are directly comparable and nothing collides.

Q

Their sessions end and their worktrees are reclaimed.— but any uncommitted work is preserved, never destroyed. The winning run routes straight into the Git Workbench merge flow, so you go from race to shipped without leaving Operon.

Q

No — every agent gets a separate, isolated worktree.That is what makes a fair race possible: no shared state, no stepping on each other’s changes, and clean per-agent diffs to judge.

Get started

Stop guessing which agent won.

Private beta — join the waitlist to get an invite when your slot opens up.

Join the waitlistExplore orchestrationCommand the fleet