Skip to content
The Operon Library

The evidence base

Sources behind the Operon Library

Every claim traces to a dated source. Cited skeptically, cited always — read the note before the number.

8 sources

Measurement

The research the library leans on for productivity, adoption, and ROI claims — read skeptically, always dated per the Evidence Policy.

The annual DORA report and its seven-capability AI Capabilities Model, framing AI as an amplifier of an organization's existing strengths and dysfunctions rather than a fix on its own.

METR 2025 RCTMETR (Model Evaluation & Threat Research)

The early-2025 randomized controlled trial that found experienced open-source developers were roughly 19% slower with AI assistance — the field's most-cited and most contested headline number.

METR 2026 updateMETR

The 2026 follow-up that revised the 2025 finding to an estimated ~18% speedup with newer tooling and selection-effect caveats — cited alongside the original RCT per the Evidence Policy's 'METR rule' that any headline number must carry its follow-ups.

GitClearGitClear

Static code-churn analytics tracking AI-era trends in copy/paste code, code churn, and refactor rates across large commit datasets.

Faros / DX / Jellyfish vendor telemetryFaros AI, DX, Jellyfish

Engineering-intelligence vendor telemetry on AI tool adoption and output, cited skeptically because vendor-reported metrics carry incentive bias.

Stack Overflow Developer SurveyStack Overflow

Annual developer survey data used for adoption and sentiment trends around AI coding tools.

SPACEForsgren et al., ACM Queue

A five-dimension developer-productivity framework (Satisfaction, Performance, Activity, Communication, Efficiency) predating the AI era but load-bearing for the measurement volume.

DX Core 4DX

A four-dimension developer-experience measurement framework (speed, effectiveness, quality, impact) used as a counterweight to raw activity metrics.

7 sources

Vendor canon

First-party documentation and engineering writing from the vendors building the harnesses the library analyzes.

Anthropic engineering blog — context engineeringAnthropic

Anthropic's essay on choosing what fills the context window, the primary formalization of context engineering as a discipline.

Anthropic engineering blog — effective harnesses for long-running agentsAnthropic

Anthropic's guidance on building agent harnesses that survive long sessions without losing coherence — informs the Harness volume's execution-environment material.

Anthropic's write-up of its own orchestrator/subagent research system, the primary source for the library's multi-agent and LLM-as-judge material.

Anthropic engineering blog — writing effective tools for agentsAnthropic

Anthropic's guidance on designing tool interfaces for agent reliability, cited in the Harness volume's tool-design material.

Claude Code best practicesAnthropic

Anthropic's own operating guidance for Claude Code, used as the baseline for harness and workflow comparisons across tools.

The vocabulary baseline the library's adopted-term definitions are checked against before any coinage is minted.

OpenAI / GitHub equivalentsOpenAI, GitHub

The competing vendor documentation and engineering blogs (Codex, Copilot) used to confirm which vocabulary and practices are vendor-specific versus industry-general.

3 sources

Methodology

The specifications-as-artifacts literature the Intent Architecture volume is built on.

The open-source toolkit that won the spec-driven-development naming war; the library maps its vocabulary (constitution → specify → plan → tasks → implement) rather than reinventing it.

Understanding Spec-Driven Development: Kiro, spec-kit, and TesslMartin Fowler — martinfowler.com

Fowler's comparative survey of the three leading spec-driven-development implementations, the primary secondary source for naming and comparing the field's competing shapes.

Addy Osmani essays — orchestration and self-improving agentsAddy Osmani

Osmani's essays on multi-agent orchestration and self-improving agent loops, cited for the Workflow Engineering and Multi-Agent volumes' methodology chapters.

4 sources

Critique

The skeptical counter-literature — benchmark audits and context-degradation research the library cites to keep its own claims honest.

The SWE-Bench IllusionarXiv preprint

The paper documenting memorization and construct-validity problems in the field's most-cited coding benchmark — required reading before citing any SWE-bench number.

Benchmark-audit literatureVarious academic and independent researchers

The broader wave of papers auditing agentic coding benchmarks for leakage, overfitting, and selection effects, cited alongside the SWE-Bench Illusion paper.

The empirical study showing model reliability degrades non-uniformly as input context grows — the primary source for the library's context rot term.

Geoffrey Huntley — Ralph postsghuntley.com

The blog posts that named and popularized the 'Ralph' while-loop agent pattern, the anchor for the library's loop-reductionist debate over what the loop framing gets right and what it misses.

3 sources

Books to position against

The existing shelf the library positions against without colliding with — narrative single-volume books and adjacent disciplines.

Vibe CodingGene Kim & Steve Yegge — IT Revolution, 2025

The current single-volume narrative account of AI-assisted coding; the library positions itself as the evidence-first, multi-volume reference this book doesn't attempt to be.

Beyond Vibe CodingAddy Osmani — O’Reilly, 2025

A practitioner-oriented follow-up covering orchestration and agent workflows; the library treats it as a peer text to cite, not a competitor to displace.

AI EngineeringChip Huyen — O’Reilly, 2025

A different discipline entirely — building LLM applications, not AI-assisted software engineering — cited only to keep the library's own terminology from colliding with Huyen's already-established 'AI Engineering' label.