The evidence base
Sources behind the Operon Library
Every claim traces to a dated source. Cited skeptically, cited always — read the note before the number.
8 sources
Measurement
The research the library leans on for productivity, adoption, and ROI claims — read skeptically, always dated per the Evidence Policy.
The annual DORA report and its seven-capability AI Capabilities Model, framing AI as an amplifier of an organization's existing strengths and dysfunctions rather than a fix on its own.
The early-2025 randomized controlled trial that found experienced open-source developers were roughly 19% slower with AI assistance — the field's most-cited and most contested headline number.
The 2026 follow-up that revised the 2025 finding to an estimated ~18% speedup with newer tooling and selection-effect caveats — cited alongside the original RCT per the Evidence Policy's 'METR rule' that any headline number must carry its follow-ups.
Static code-churn analytics tracking AI-era trends in copy/paste code, code churn, and refactor rates across large commit datasets.
Engineering-intelligence vendor telemetry on AI tool adoption and output, cited skeptically because vendor-reported metrics carry incentive bias.
Annual developer survey data used for adoption and sentiment trends around AI coding tools.
A five-dimension developer-productivity framework (Satisfaction, Performance, Activity, Communication, Efficiency) predating the AI era but load-bearing for the measurement volume.
A four-dimension developer-experience measurement framework (speed, effectiveness, quality, impact) used as a counterweight to raw activity metrics.
7 sources
Vendor canon
First-party documentation and engineering writing from the vendors building the harnesses the library analyzes.
Anthropic's essay on choosing what fills the context window, the primary formalization of context engineering as a discipline.
Anthropic's guidance on building agent harnesses that survive long sessions without losing coherence — informs the Harness volume's execution-environment material.
Anthropic's write-up of its own orchestrator/subagent research system, the primary source for the library's multi-agent and LLM-as-judge material.
Anthropic's guidance on designing tool interfaces for agent reliability, cited in the Harness volume's tool-design material.
Anthropic's own operating guidance for Claude Code, used as the baseline for harness and workflow comparisons across tools.
The vocabulary baseline the library's adopted-term definitions are checked against before any coinage is minted.
The competing vendor documentation and engineering blogs (Codex, Copilot) used to confirm which vocabulary and practices are vendor-specific versus industry-general.
3 sources
Methodology
The specifications-as-artifacts literature the Intent Architecture volume is built on.
The open-source toolkit that won the spec-driven-development naming war; the library maps its vocabulary (constitution → specify → plan → tasks → implement) rather than reinventing it.
Fowler's comparative survey of the three leading spec-driven-development implementations, the primary secondary source for naming and comparing the field's competing shapes.
Osmani's essays on multi-agent orchestration and self-improving agent loops, cited for the Workflow Engineering and Multi-Agent volumes' methodology chapters.
4 sources
Critique
The skeptical counter-literature — benchmark audits and context-degradation research the library cites to keep its own claims honest.
The paper documenting memorization and construct-validity problems in the field's most-cited coding benchmark — required reading before citing any SWE-bench number.
The broader wave of papers auditing agentic coding benchmarks for leakage, overfitting, and selection effects, cited alongside the SWE-Bench Illusion paper.
The empirical study showing model reliability degrades non-uniformly as input context grows — the primary source for the library's context rot term.
The blog posts that named and popularized the 'Ralph' while-loop agent pattern, the anchor for the library's loop-reductionist debate over what the loop framing gets right and what it misses.
3 sources
Books to position against
The existing shelf the library positions against without colliding with — narrative single-volume books and adjacent disciplines.
The current single-volume narrative account of AI-assisted coding; the library positions itself as the evidence-first, multi-volume reference this book doesn't attempt to be.
A practitioner-oriented follow-up covering orchestration and agent workflows; the library treats it as a peer text to cite, not a competitor to displace.
A different discipline entirely — building LLM applications, not AI-assisted software engineering — cited only to keep the library's own terminology from colliding with Huyen's already-established 'AI Engineering' label.