The Operon LibraryChapters 11 Published 11
Volume VII · Generation is cheap. Verification is the job.
Verification & Trust
Generation cost collapsed while verification didn't — the defining asymmetry of the era, and the second-biggest gap in the first edition. The volume covers evals as engineering practice, LLM-as-judge, adversarial verification, benchmark skepticism, test and merge gates, provenance, and agent security — a topic absent from every volume in v1 — closing by treating verification as a first-class architectural layer, not a phase.
Key sources: SWE-Bench Illusion paper + benchmark-audit wave · Anthropic LLM-as-judge findings · Faros review-time data · OWASP LLM Top 10 · OpenTelemetry GenAI conventions (for provenance plumbing)
Related
- Compare AI coding agents on the same taskRace two to four agents on one goal and judge them on cost and diff quality.
- Local-first security and data handlingWhat stays on your machine, what syncs, and how to control both.
- The AI-assisted software engineering glossaryEvery term this volume uses, with coinages marked and origins cited.
- The Library’s research sourcesThe dated evidence base behind every claim in the handbook.