AI Writes the Code. What's Left for You?
For twenty-five years, writing the code was the expensive part. It just got cheap — so most of our process is up for audit. What dissolves, what transforms, and the one thing that's now the whole job.
For twenty-five years, one thing in software was reliably expensive: writing the code. Every method we built — Agile, Lean, XP, DevOps — was scaffolding around that single scarcity. Sprints rationed human hours. Estimation predicted them. Standups synchronized them. Code review protected the expensive artifact they produced. Strip away the vocabulary and every ritual was managing the same bottleneck: humans are slow and costly to turn intent into working code.
That bottleneck just moved.
Agents now write the majority of new code in a growing share of teams and open the pull requests themselves. When the scarce resource moves, everything built to ration it is suddenly up for audit. Not "up for disruption" — up for audit. Because most of what we do wasn't wrong; it was a rational response to a constraint that no longer holds.
So here is the question nobody is actually asking, the one that cuts cleanest:
Did this practice exist because humans have judgment — or because human throughput was the bottleneck?
Point that question at every practice you have, and four verdicts fall out.
The four verdicts
Dissolves. Story-point estimation existed to ration the scarce resource — human hours. The machine doesn't need rationing; it executes in minutes, near-free. The ritual outlived its constraint. (Careful: coordination still matters. What dissolves is the ceremony that used to deliver it, not the need itself.)
Transforms. The PRD stops being a persuasion document and becomes the primary artifact — the thing the machine runs against. Vague intent no longer slows a human down gently; it builds the wrong thing at machine speed. Test-driven development, which was always right and always economically awkward, reverses polarity: a failing test handed to an agent is machined intent, executable and ungameable by confident prose. Code review moves up a level — from reading lines to reading intent and blast radius. Editorship, not inspection.
Emerges. Writing the evals that grade the output. Orchestrating the agents. Naming the trust boundaries. None of this existed five years ago, and it has no clean pre-AI ancestor. It's new work the machine itself created.
Endures. Defining the problem worth solving. Holding the standard. Taste. The call. These were always rooted in human judgment, and AI amplifies them rather than replacing them.
Read the pattern down the list and it's unmistakable: the throughput scaffolding dissolves, the judgment work transforms and multiplies, and new judgment work emerges. Which raises the obvious question — what, exactly, is the shape of the thing that's left?
It's not a third loop
A lot of smart people are circling the same observation right now. Andrew Ng has named a "developer feedback loop." Martin Fowler's Thoughtworks retreat — held where the Agile Manifesto was written — concluded that a new kind of work is forming that "nobody has named yet." Everyone reaches for the same word: a middle loop.
I think naming it a third loop is a mistake, and an expensive one. A third loop is a third speed inserted in series — and three speeds in series is exactly the assembly line we spent twenty-five years trying to escape. Name it wrong and you rebuild the bottleneck you just got rid of.
It isn't a third loop. It's a mesh.
Picture a two-speed engine. An outer loop of human judgment — understand, frame, verify, learn — turning slowly, maybe once a day. Coupled to an inner loop of machine execution — generate, test, deploy, monitor — spinning many times per outer turn. They meet at exactly one point: the specification. Intent goes in; evidence comes out. One housing, two speeds, never two teams — not a fast team handing off to a slow one.
The middle work everyone senses is real. But it isn't a new tier of the assembly line. It's the gear mesh — the coupling where a slow, high-force turn of judgment drives a thousand fast turns of execution. Judgment is torque, not speed.
What's left for you
You're not the typist anymore. You're the answerable owner — the person who can say what changed, why it's safe, and what happens if it's wrong. And you're the one who originates: the problem worth solving, the bet, the taste. The machine executes; it does not invent the thing worth building.
That sounds like a promotion, and it is. But there's a catch the research is blunt about, and skipping it is how this goes wrong.
The catch: freeing judgment doesn't strengthen it
Map the work onto Bloom's taxonomy and the split is clean. The machine now does the lower cognitive modes — remember, understand, apply. What's left for you is the higher modes — analyze, evaluate, create. The comforting story is that offloading the lower frees you for the higher. The uncomfortable finding is that it doesn't happen automatically. Left unguarded, the higher modes atrophy.
- Lisanne Bainbridge described this in 1983 as the ironies of automation: automate the easy parts, and the human is left with the rare, hard judgment they're worst at maintaining — because the learn-by-doing loop that forged that judgment has been severed. The most-automated systems, she noted, need the most-trained operators.
- A peer-reviewed study of 666 people (Gerlich, 2025) found a significant negative correlation between frequent AI-tool use and critical-thinking scores, mediated by cognitive offloading — with younger users the most dependent and the lowest-scoring.
- A Microsoft and Carnegie Mellon survey of knowledge workers (2025) found a confidence paradox: the more you trust the AI, the less you engage critically — and effort dropped across every level of Bloom's taxonomy.
- A randomized trial by Anthropic (2026) put a finer point on it: engineers who delegated to the AI scored seventeen points lower on a comprehension quiz than those who hand-coded; engineers who used the same AI for conceptual inquiry — asking, exploring trade-offs — kept their edge. It isn't whether you use AI. It's how.
(One honest caveat: a widely-shared MIT "cognitive debt" EEG study points the same way, but it's a small, not-yet-peer-reviewed preprint — suggestive, not settled. I'd weight it accordingly.)
There's a second-order version of this too. Judgment used to be forged by doing the lower-mode work — juniors learned to reason about systems by writing the code. That apprenticeship is thinning from both ends: fewer juniors are hired (Stanford's payroll data shows employment for 22-to-25-year-old developers down roughly 20% since late 2022), and those who are hired learn less if they delegate rather than inquire. The pipeline that used to produce senior judgment for free is quietly closing.
So the disciplines that can feel like overhead — verification at the boundary, executable specs, an answerable owner, a real collision with reality — are not overhead. They are the antidote the research is asking for: structured, effortful engagement that forces the higher modes to fire instead of being silently offloaded.
How judgment actually gets forged
If judgment is the job, and judgment can atrophy, then you need a way to deliberately exercise it. That's the human half of the engine, up close — what I call the Judgement Loop. It's just the scientific method pointed at your own judgment on a piece of work, and its whole point is that each stage defeats one specific way you fool yourself.
- Sourcing — defeat omission. Prefer what you observed first-hand; reading and listening are second-hand testimony that can leave things out. The enemy is omission on both sides of the glass: what a source omitted, and what you didn't look at.
- Articulation — defeat self-deception. Write the belief as a claim reality could prove wrong: an observable signal, a threshold, a time bound, and what would count as refuted. If you can't write the version you could lose, you don't have a claim; you have a feeling.
- Calibration — defeat groupthink. Surface the single strongest objection and the cheapest way to test it. Seek the breaking point, not applause. Instant agreement is gravity, not proof.
- Reality Collision — the one judge that can't be bribed. Ship the smallest thing that could produce the refuting observation.
Judgment is the residual — the surprise that survives a hard calibration. And notice the shape: the loop is bookended by two direct contacts with reality (observe to gather, collide to test), with the symbolic, social stages in the cheap middle. Good judgment is keeping that cheap middle honest by anchoring both ends in the real world.
Proof, not theory
Theory is cheap, so I gave the framework a real app to build — a habit-planning tool called Sixty-Four: one goal in, sixty-four daily habits out, taken from a raw brief to a shipped product entirely through the framework's own tools.
The most instructive moment wasn't a success. The executable spec passed fifteen out of fifteen checks — while the interface was visibly wrong three separate times. The machine verified everything that was executable; I had to own the thing no test captured — legibility, taste, whether it actually felt right to a person. That gap between "passes the spec" and "is actually good" isn't a bug in the process. It is the mesh, lived. It's precisely where judgment stays mandatory.
There was a smaller moment I keep coming back to. Shaping the bet, the tooling flagged a dependency I was certain the model had hallucinated. Its rule was verify, don't assume. A two-minute check: it was real. The framework corrected its own author. Verify over authority — even mine.
What's left is enough
The build was never the hard part. It was expensive, so we organized everything around it and mistook that scaffolding for the work. Now the build is cheap, and what remains is the part that was always the actual job: deciding what's worth making, defining what "correct" means, and owning the answer when someone asks why it's safe.
The loops that decide whether the thing matters — those are yours. That's what's left for you.
And it's more than enough.
This is the written companion to a talk I gave at Pie & AI: Coimbatore (DeepLearning.AI). The full framework — the ledger of verdicts, the Judgement Loop, the activity map, and the worked example — lives at theproductguy.xyz/two-speed-engine.