The Judgement Loop’s four stages aren’t abstractions — each is a muscle you build with ordinary activities. But the rule underneath the map matters more than the map: it isn’t the activity that fixes the stage, it’s which adversary you’re using it to defeat. “Reading” isn’t a stage; “reading to find what a claim omits” is Sourcing. Tag the adversary, not the verb.
train your judgment · the activity map
What to actually do to develop judgment
Judgment isn’t a trait you have. It’s a residual you forge — by running the loop, on real activities.
01 · observe & read
Sourcing
Defeats: Omission — theirs & yours
Source from / do: Primary: watch real users, use the thing yourself, take the call, read the raw data. Secondary: reading, listening, dashboards — you inherit the observer's bias, so climb toward primary.
Counts when: Are you observing reality, or ingesting someone's testimony about it? Hunt what was left out — by them, or by you.
02 · write
Articulation
Defeats: Self-Deception
Source from / do: Journaling, blogging, a prediction journal, writing the bet as one falsifiable claim: signal · threshold · time · what would refute it.
Counts when: Can the thing you wrote turn out wrong? If it can't be refuted, it's a feeling, not a claim.
03 · talk & listen
Calibration
Defeats: Groupthink
Source from / do: Debate, seek critique, red-team, present to a skeptical room, ask “what would change my mind?”
Counts when: Are you seeking the breaking point or applause? Instant agreement is gravity, not proof.
04 · ship
Reality Collision
Defeats: Every illusion — reality is unbribeable
Source from / do: Ship the feature, run the experiment, publish, sell to a real customer, make a real bet with stakes.
Counts when: Does the outcome bite back no matter what you believe? Is it the smallest thing that could produce the refuting observation?
Observation is primary sourcing; reading and listening are secondary— you inherit someone else’s observation and their bias. The closer to raw reality, the less room for distortion. Drive load-bearing inputs up the ladder:
- Primary — observed.You saw it first-hand. No claimant, so Authority and the True Lie can’t touch it. Highest fidelity — and the tier people skip.
- Secondary — testified. Someone else observed and reported it (a paper, a metric, a summary). Ask: can I observe this directly instead?
- Tertiary — aggregated. A summary of summaries; the omissions have compounded. Distrust by default.
Notice how the loop is built. Stages 01 and 04 both touch reality directly — observe to gather, collide to test: the same act of unmediated contact, pointed in opposite directions. The two in between — Articulation and Calibration — are symbolic and social: cheap and fast to iterate, but exactly where self-deception and groupthink distort. So the whole discipline compresses to one line: use the cheap symbolic middle without letting it drift from the two reality anchors.
Most people over-invest in Stage 1 — consuming feels productive — and stall before 2 and 4. Articulation is the common bottleneck (reading endlessly, never writing the version you could lose). Reality Collision is the rarer one (debating forever, never shipping the test). The residual — which isyour judgment — only forms at Stages 3 and 4. You can’t read your way to it.
Write the claim you could lose → attack it yourself before anyone else does → let reality settle only the ones that survive → log the surprise. One real belief a week.
Here is the catch the research makes unavoidable: freeing the higher modes doesn’t automatically strengthen them. Offload the lower modes carelessly and the higher ones atrophy. Map the modes to Bloom’s taxonomy — the machine takes Remember, Understand, Apply; what’s left for you is Analyze, Evaluate, Create — and the evidence converges on the danger:
- The ironies of automation (Bainbridge, 1983): automate the easy parts and the human is left with the monitoring and rare-but-critical judgment they’re worst at maintaining — and the “learn-by-doing” loop that forged that judgment is severed. The most-automated systems need the most-trained operators.
- Cognitive debt (MIT Media Lab, 2025 · preprint): EEG showed the weakest neural connectivity in the group that wrote with an LLM, and they struggled to quote the essay they’d written minutes earlier. Not yet peer-reviewed (n=54) — weight it accordingly.
- The confidence paradox (Microsoft + CMU, CHI 2025): the more you trust the AI, the less you engage critically — and generative AI reduced self-reported effort across every level of Bloom’s taxonomy. Critical thinking relocates from doing to verifying, integrating, and stewarding.
- The correlation (Gerlich, 2025 · peer-reviewed, n=666): frequent AI-tool use correlated with significantly lower critical-thinking scores, mediated by cognitive offloading — and younger users showed the highest dependence and the lowest scores.
- Delegation vs. inquiry (Anthropic, 2026): in a randomized trial, engineers who delegated to the AI scored 17 points lower on a comprehension quiz (50% vs 67%); those who used it for conceptual inquirykept their edge. It’s not whether you use AI — it’s how.
So the discipline on this page isn’t overhead — it’s the antidote the research calls for. Verification-at-the-boundary, executable specs, the answerable owner, and Reality Collision all forcethe higher modes to fire instead of being silently offloaded. The Judgement Loop is, in cognitive terms, a deliberate higher-order-thinking regimen — it exists precisely because those modes don’t survive on their own once the lower ones are outsourced.
It also reframes the pipeline. Judgment used to be forged by doing the lower-mode work — juniors learned by writing the code. That apprenticeship is thinning (Stanford, 2025: employment for 22–25-year-old developers is down ~20% since late 2022), so judgment now has to be trained on purpose — which is what this page is for.