Skip to content
The Operon Library

Volume XI · Chapter 7

Engineering Culture & AI Leadership

Norms, incentives, and what leaders must model.2026-07-13 · 8 min read

A VP of engineering closes a quarterly all-hands with a slide the whole room can read from the back: forty thousand AI-assisted commits this quarter, up from six thousand the one before. The room applauds. By any conventional measure it is a rollout success — adoption went from a pilot team to the entire org in four months, and she has said, more than once, that this is the fastest technology shift she has ever led. What she has not said, because nobody has asked, is when she last opened an agent session herself and watched it fail.

This Library’s tenth volume has already named what tends to happen next inside a team under that kind of pressure: engineers running agents on throwaway work to keep a personal token count above the office average, quietly withholding their most effective prompts and workflows from colleagues to stay ahead in a competition nobody officially opened. DORA’s own term for it is tokenmaxxing, and the chapter that introduced it here — Team & Organizational Intelligence — treated it mainly as a measurement-aggregation problem: publish a number fine enough to walk back to a person, and the person games it. That diagnosis is correct and incomplete. Nobody games a number that was never celebrated in the first place. The forty-thousand-commits slide is not a report of what happened last quarter; it is an instruction about what to do next quarter, delivered from the front of the room by the one person whose approval everyone in it is already calibrated to read.

The metric a leader celebrates is a policy, not a fact

Healthy Metrics, Dangerous Metrics, two chapters later in the same volume, built a checklist for whether a given number survives being turned into a target — the cost of gaming it, whether it is reported alone or paired with a harder-to-fake signal. That checklist quietly assumes someone reasonably neutral is choosing which number to publish. In practice the choice is rarely neutral, because the person making it is the same person the team watches for cues about what actually matters here. DORA’s 2026 piece on tokenmaxxing puts the leadership question directly, not as a footnote: “Are we trying to build a culture of forced, raw adoption, or are we trying to build a culture of testing, learning, and innovation?” Some of the same companies running internal token leaderboards, the piece reports, treat the leaderboard itself as the intervention — the explicit goal is to force a mental-model shift toward heavier usage, adoption pursued as an end in itself. A leader who puts a raw usage number on a slide, unpaired with any outcome figure, has already answered DORA’s question, whether or not she meant to.

A leaderboard is not evidence of a culture. It is an instruction disguised as a report.

Culture follows what leaders do, not what they announce

Organizational-behavior research has a settled, decades-old answer for why a slide moves behavior faster than a values document does. Edgar Schein’s work on organizational culture — foundational enough that later researchers mostly refine it rather than replace it — identifies a small set of what he calls primary embedding mechanisms: the handful of things a leader does, repeatedly and visibly, that actually teach an organization what it values, as distinct from what a leader merely writes down. Chief among them: what a leader pays attention to and measures on a regular basis, how a leader reacts to a crisis, and — distinctly — deliberate role modeling, teaching, and coaching, meaning a leader doing the work itself, in view of the people watching. A memo about thoughtful AI adoption is what Schein calls a secondary mechanism: a plaque on the wall, reinforcing a culture rather than creating one. What a leader visibly measures and visibly does is primary. It is the instrument that actually writes the culture, whatever the plaque says.

Applied to AI-assisted engineering, Schein’s distinction has a blunt operational test: does the person setting the expectation actually run sessions — not a scripted demo, but real work with real failures in it — or has the mandate arrived from someone who has delegated even the trying? A 2026 SD Times analysis of technical leadership under AI adoption states the standard plainly: engineering leaders need to be “using the tools personally, in real workflows, on a regular basis. Not in a demo environment.” The distinction matters because a demo is curated and a real session is not. A staff engineer who runs an agent on a genuinely hard piece of the codebase, hits a wall, says so out loud, and narrates how she recovered has modeled something no training deck can substitute for: that the tool fails sometimes, that a failed session is not a personal verdict, and that the right response to being stuck is to say so rather than grind quietly toward a number that looks fine on someone else’s dashboard. A leader who has never done this has no comparable behavior to model, only instructions to issue — and a team reads the absence of the first as clearly as it would read the presence.

Leaders who avoid hands-on engagement are forced into rubber-stamping their team’s judgment — which isn’t leadership — or overriding it without sufficient basis.

David Feldman, SD Times, on managers who skip using the tools themselves

The tension leaders actually have to hold

None of this is an argument against pushing adoption. This Library’s first volume has already made the case, with real evidence rather than vendor enthusiasm, that AI-assisted engineering produces genuine value when it is measured against outcome rather than activity — the same discipline Session ROI names at the level of a single sitting, scaled up here to a whole rollout. A leader who never pushes adoption at all is failing a quieter, opposite test: leaving real productivity gains on the table because asking a team to change how it works is uncomfortable to ask. The honest tension is not adoption versus caution. It is between two different things a leader can choose to optimize a rollout for, and pretending there is no tradeoff between them is its own kind of dishonesty.

What adoption pressure quietly erodes

Three things this Library’s own volumes have already shown erode quietly under adoption pressure, unless a leader deliberately protects them. The first is skill formation — this volume’s own chapter on the subject, three chapters back, takes up directly what happens to a junior engineer’s judgment once an agent writes the boilerplate she would once have struggled through herself, and struggle turns out to have been load-bearing all along. The second is a genuine learning culture, as distinct from faster throughput: the Measuring Learning chapter, two volumes back, showed that velocity and pull-request counts can climb for a full year while an organization’s actual rate of learning — how quickly the same mistake stops recurring — sits flat or drifts the wrong way, because moving faster does not by itself require getting smarter. The third is psychological safety, in the specific sense this Library has used since its own treatment of blameless postmortems: whether an engineer can say, out loud and before a deadline forces the question, that a session is not working — rather than quietly grinding through a failing approach to protect a metric a leader is known to be watching. John Allspaw’s original case for blameless review at Etsy, and Google’s later formalization of the same discipline inside its own SRE practice, both rest on the mechanism this chapter has been circling from a different angle: people tell the truth about what is and is not working in rough proportion to how safe the person watching has made it to tell them.

What leadership behavior actually teaches

A leader can…What it visibly rewardsWhat it teaches, intended or not
Put a raw usage or output number on a company-wide slide, unpairedVolume, adoption for its own sakeTokenmaxxing is the safe career move here
Report cost or usage next to a durable outcome figure, every timeWork that survives review and shipsThe number that matters has a denominator
Mandate adoption without ever running a session personallyCompliance with the rolloutThis is something done to engineers, not with them
Run a real session visibly, including a failed one, and narrate the recoveryHonest trial and errorA stuck session is a normal event, not a confession
Reward whoever grinds a failing session to a green checkmarkProtecting the metricAdmitting a session failed costs more than hiding it
Ask “did this stay stuck, and did you say so” before “how much did you use it”Early, honest signalSpeaking up early is the metric that actually matters

None of the six rows above requires new tooling to check. Most of it is already visible in the same session-level telemetry a rollout is collecting for cost reasons — the open question is whether anyone is reading it as a leadership signal, or only as a spend report.

What AI leadership actually requires

Traditional engineering leadership development prepares a manager to run one-on-ones, calibrate a performance review, and defend a roadmap to a skeptical stakeholder. None of that curriculum currently answers the question this chapter keeps circling back to: what does it cost a team, structurally, for its leader to be visibly bad at something today, on purpose, in front of them? That is a genuinely new demand, not a rebranded old one. It requires treating one’s own agent sessions as a public artifact rather than a private convenience, choosing which number gets celebrated with the same care this Library asks of any dashboard, and holding, without flinching, a real tradeoff between honest adoption gains and the slower, quieter things — a junior’s judgment, an organization’s actual learning rate, an engineer’s willingness to say a session failed — that no velocity chart will show eroding until long after the erosion is done. Most leadership training was built for a world in which a leader’s own technical practice was optional information about her. It is not optional anymore, and the leaders who have noticed are the ones whose teams turn out to be worth watching.

For Discussion

  1. The last time your team’s AI usage was mentioned publicly by leadership — an all-hands, a Slack channel, a review — was it paired with an outcome figure, or reported alone? What did that pairing, or its absence, actually teach the room?
  2. When did the most senior engineer or leader on your team last run a real AI-assisted session on production code, in front of someone, and visibly hit a wall? If the honest answer is “I don’t know” or “never,” what is the mandate to adopt currently resting on?
  3. If an engineer on your team quietly ground out a failing session to a technically passing result rather than stopping and saying it wasn’t working, would anyone have known — and would it have cost them anything to say so instead?

References

  1. establishedPrimary embedding mechanisms of culture — what leaders pay attention to and measure, how they react to crises, and deliberate role modeling as the most direct, visible ways a leader teaches an organization what it values (distinct from secondary mechanisms like mission statements)Edgar H. Schein — Organizational Culture and Leadership (4th ed.) · 2010
  2. emergingEngineering leaders need to be "using the tools personally, in real workflows, on a regular basis. Not in a demo environment"; leaders who skip hands-on use are forced into rubber-stamping or overriding their team’s technical judgment without sufficient basisDavid Feldman — "Leadership in the Age of AI: Why Managers Need to Stay Technical," SD Times · 2026-04-06
  3. emerging"Are we trying to build a culture of forced, raw adoption, or are we trying to build a culture of testing, learning, and innovation?" — internal token leaderboards as a leadership-chosen intervention, not a neutral reportDORA — Evan Conaway, "Finding balance in the era of tokenmaxxing" · 2026-06-02
  4. establishedTokenmaxxing named and diagnosed as a measurement-aggregation problem: engineers gaming AI-usage metrics under pressure once a number is published fine enough to trace to a personThis Library, Volume X — "Team & Organizational Intelligence" · 2026-07
  5. establishedA checklist for whether an engineering metric survives being turned into a target — cost of gaming, isolation from a paired signal, and Goodhart’s law applied to AI-era dashboardsThis Library, Volume X — "Healthy Metrics, Dangerous Metrics" · 2026-07-13
  6. establishedVelocity and throughput can climb for a sustained period while an organization’s actual learning rate — how quickly the same root cause stops recurring — stays flat or worsens; learning and speed are not the same signalThis Library, Volume X — "Measuring Learning" · 2026-07
  7. establishedBlameless postmortems and a "just culture": engineers who fear individual blame withhold the detail an investigation needs; treating a failure as a systemic signal rather than a personal one is the precondition for an honest accountJohn Allspaw — "Blameless PostMortems and a Just Culture," Etsy Code as Craft · 2012-05-22
  8. establishedPostmortem review formalized around treating the system, not an individual engineer, as the unit of failure — the same discipline this chapter extends to a leader’s handling of a failed AI sessionGoogle — Site Reliability Engineering Workbook, "Postmortem Culture: SRE Practices" · 2018
  9. establishedVerified economics evidence that AI-assisted engineering produces real value when measured against outcome rather than activity — the grounding for Session ROI and this chapter’s case for genuine (not unconditional) adoption pressureThis Library, Volume I — "The Real Cost of AI Coding Isn’t Your Token Bill" · 2026-07-09