For twenty-five years, one thing was expensive: writing the code.
Every process, role, and ritual we built was scaffolding around that one scarcity.
the old world
PRDs · estimation · sprints · standups · code review · QA gates — all scaffolding around one scarcity.
the reversal
It just got cheap.
Agents write the majority of new code and open the pull requests. The scarce resource moved — and it took the scaffolding’s reason for existing with it.
the new world
The bottleneck moved. Code is cheap — judgment is the scarce resource now.
?
the audit question
So here’s the question nobody’s auditing:
Did this practice exist because humans have judgment — or because human throughput was the bottleneck?
Judgment deciding what’s worth building — and whether it’s any good. The call only a person can make.
Throughput how much work hands can crank out. How fast the code actually gets written.
the method
Point every practice at that one question. Four verdicts fall out.
the ledger
Four verdicts
Endures
Load-bearing on human judgment. Still scarce, still yours.
Transforms
Same intent, new mechanics. The polarity flips.
Dissolves
Managed a constraint that no longer exists. Let it go.
Emerges
New work the machine itself created.
Dissolves
Story-point estimation
You don’t estimate the machine.
Estimation rationed the scarce resource — human hours. The machine doesn’t need rationing. The ritual outlived the constraint.
Transforms
The PRD
The spec becomes the primary artifact.
It stops being a persuasion doc and becomes what the machine runs against. Vague intent doesn’t slow an agent down — it builds the wrong thing at machine speed.
Transforms
Test-driven development
Tests stop checking the work. They commission it.
A failing test handed to an agent is machined intent — executable, unambiguous, ungameable by confident prose.
Transforms
Code review
Editorship, not inspection.
The machine out-produces line-by-line reading. Review moves up: does this match the intent, fit the architecture, avoid what was forbidden?
Emerges
Evaluation engineering · loop engineering
New work the machine created.
Writing the evals that grade the output. Orchestrating the agents. Naming the trust boundaries. None of it existed five years ago.
the pattern
The throughput scaffolding dissolves. The judgment work transforms — and multiplies.
Story points
Dissolves
Manual QA as a phase
Dissolves
TDD
Transforms
PRDs
Transforms
Code review
Transforms
Eval-driven dev
Emerges
Loop engineering
Emerges
Define the problem
Endures
Taste & the call
Endures
the deeper structure
A new kind of work is forming between writing code and shipping it.
Andrew Ng calls it the developer feedback loop. Fowler’s Thoughtworks retreat: “nobody has named it yet.” Everyone reaches for the same word — a middle loop.
the disagreement
It’s not a third loop.
A third speed breaks the machine back into an assembly line. Name the wrong thing and you rebuild the bottleneck you just escaped.
loop, or mesh?
the two-speed engine
It’s a mesh, not a loop.
One housing. Two speeds. Never two teams.
the mesh, up close
The coupling has three jobs. All three are yours.
specification · trust calibration · verification
One turn of judgment drives a thousand turns of execution. Judgment is torque, not speed.
Torque turning force, not spinning speed. A long wrench frees a stuck bolt not by spinning faster, but by applying more force per turn — one slow, deliberate turn does the work of many fast ones.
so, what’s left for you
Judgment.
But judgment isn’t a vibe you either have or don’t. It’s a discipline you can run.
the judgement loop
Where judgment comes from.
Four stages, four adversaries — bookended by reality: observe (01), then collide (04); the middle two are symbolic. Judgment = the residual that survives Calibration.
live
I compiled this into a tool. Let’s pressure-test a real belief — right now.
Give me one you hold about AI and your work. We run it through the four stages, out loud, and see if it survives Calibration — or dies cheap, in front of everyone.
proof
64
Theory is cheap. So I gave the framework a real app to build.
Sixty-Four — one goal in, sixty-four daily habits out. Built end-to-end through the framework’s own tools.
the execution
Not one prompt — a multi-skill loop.
SHAPE/shape · /frame · /spec — the bet (a 2-min probe corrected me), the runs, 15 executable checks
RUNbuild ⇆ Review — machine builds, human decides: spec 15/15 while the UI was wrong 3×
GENERATERun E — claude-opus-4-8 · structured output → the 73-node chart
SHARPEN/loop-design · /metric-audit — cut the cross-pillar cost sink · keep the return rate
JUDGE/judgement-loop · delta-log — demand bet wounded at Calibration · residual → prior
the probe
The framework corrected its own author.
Shaping the bet, the tool flagged a dependency I was sure the model had hallucinated. Its rule: verify, don’t assume. A two-minute check — it was real. The false belief was mine.
the division of labour
The spec passed 15/15 while the UI was wrong three times.
15/15machine verified — the executable
3×human caught — legibility, taste
That gap between them isn’t a bug. It’s the mesh, lived.
the honest ending
Real generation worked. The risk was never quality — it was plumbing. The spec caught it.
But one loop code can’t close: do people come back? The machine can build it — only you ask should it exist, and will anyone return. That’s product thinking, and it’s parked on purpose: reality’s to settle, with real users and a real week. Not another commit.
the synthesis
You’re not the typist anymore.
You’re the answerable owner — the one who can say what changed, why it’s safe, what happens if it’s wrong. And the one who originates: the problem worth solving, the bet, the taste. The machine executes; it doesn’t invent the thing worth building.
from inside the machine
“I don’t prompt Claude anymore. I write loops that prompt Claude. My job is to write loops.”
— Boris Cherny, creator of Claude Code, Anthropic · June 2026
The person who built the coding agent doesn’t type at it — he designs the loop and owns the outcome. That’s this whole talk, said from inside the tool.
and your team
The plus-signs in “Agile + Lean + DevOps” were never process.
They were org-chart walls. Collapse them: one team, four hats — not a fast team handing off to a slow one.
your roadmap
The skills that compound now aren’t syntax.
frame the problem · write specs & evals · verify the machine · taste
Build them for real: run the loop on a live belief · write one executable eval · own one thing end-to-end. Depth of judgment, not lines of code.
the evidence · why this is the job
Free the lower modes carelessly, and the higher ones atrophy.
Bainbridge ’83 automate the easy parts and the human keeps the rare, hard judgment they’re worst at maintaining — the most-automated systems need the most-trained operators.
MIT ’25 EEG shows the weakest brain connectivity in LLM writers, who couldn’t reliably quote the essay they’d just written — “cognitive debt.” (Preprint.)
MSFT + CMU ’25 the more you trust the AI, the less you think critically — effort falls across every level of Bloom’s taxonomy.
Anthropic ’26 RCT: engineers who delegated to AI scored 17 points lower on comprehension (50 vs 67%); those who used it to ask, not delegate, kept their edge.
So verification isn’t overhead — it’s the antidote. The loop forces the higher modes to fire instead of quietly offloading them.
the takeaway
The build was never the hard part. The loops that decide whether it matters — those are yours.
monday
Audit one ritual this week.
Pick a practice your team still runs. Can you name the constraint it manages — or do you keep it from habit? Start there.
Take the 2-minute audit → theproductguy.xyz/two-speed-engine/prepare
The Two-Speed Engine
Nandri. Thank you.
Questions — especially the ones where you think I’m wrong. That’s Calibration.
The ledger · the playbook · the /judgement-loop skill → theproductguy.xyz/two-speed-engine