A year ago, ranking the 26 modes of thinking in this series felt like a solo exercise — one hierarchy, one build order, one voice. This update changes that. The same 26 modes were ranked by three different AI systems (Claude, ChatGPT, and Gemini), each with no visibility into the others' answers. Comparing the three rankings side by side turns out to be more useful than any single ranking on its own: where all three agree, that's a reason to treat a claim as well-supported. Where they diverge, the divergence exposes exactly which parts of the original hierarchy were confident guesses rather than settled conclusions.
One caveat before any of that: three AI outputs agreeing isn't the same as three independent experts agreeing. All three models were trained on overlapping bodies of human-written material about causality, probability, systems thinking, and the rest — so convergence here is best read as "these ideas are widely and consistently represented as foundational across the models," not as an objective proof that they are. That distinction matters for how much weight the rest of this article should carry, and it's worth holding onto throughout.
This is that comparison — plus a revised hierarchy that reflects what held up and what didn't.
Why cross-checking a subjective ranking works at all
There's no ground-truth answer key for "how important is inversion thinking relative to structural thinking." It's not like ranking countries by GDP. So the value here isn't that three AIs converging on an answer makes it correct — it's that independent convergence across systems with different training and different reasoning styles is a much stronger signal than one system's confident-sounding list. A single AI producing a clean 1–26 ranking can just as easily be false precision dressed up as insight. Three systems landing in the same neighborhood, unprompted, is harder to wave away.
One of the three systems put this well when asked to grade the others: it argued that a model willing to say "these shouldn't all be ranked on a single scale — here's a tier structure instead" deserves more credit than one that confidently outputs a tidy list from 1 to 26, because the tidy list is exactly the kind of false-precision error this whole framework is supposed to guard against. That critique applies to the original version of this article as much as to any AI's output.
Where all three independently agreed
Each AI was given the same 26 modes, cold, with no sight of the others' rankings. Here's how the top five landed:
| Type of thinking | Claude's rank | Gemini's rank | ChatGPT's rank |
|---|---|---|---|
| Causal thinking | 1 | 1 | 2 |
| Critical thinking | 3 | 2 | 1 |
| Probabilistic thinking | 4 | 3 | 3 |
| Systems thinking | 2 | 4 | 5 |
| Analytical thinking | 5 | 5 | 4 |
No two systems agreed on the exact internal order, but all three landed on the same five skills for their top five. That's a real result, but it needs two qualifiers rather than one triumphant sentence. First: this is the foundation among these 26 listed modes, not a foundation in some absolute sense — a different list of 26 (add memory, attention, language, or decision theory) could easily produce a different top five. Second: how tightly the three cluster on a given skill varies a lot, and that spread is itself informative. Causal thinking is a near-lock (1, 1, 2 — essentially no disagreement). Systems thinking has real spread (2, 4, 5). So "the original hierarchy's Tier 1 holds up" is true as a claim about relative ranking within this list, and it holds up more strongly for some of the five than others — it's a supported cluster, not a validated law.
Where they genuinely split — and what that means
The disagreements below aren't noise. They're places where the original ranking made a call that two out of three independent systems would contest.
Two of the three systems ranked metacognition in their top six, treating it as a multiplier skill on par with the core five rather than a mid-pack add-on. The reasoning: metacognition is the mechanism by which every other skill on this list gets monitored, corrected, and improved. A thinker with excellent causal reasoning but no ability to notice when that reasoning is going wrong doesn't actually benefit from having it. This is a legitimate case for promoting metacognition closer to Tier 1, not just Tier 2.
Two systems ranked it in the top eight; one ranked it closer to the middle. The case for promotion: first principles thinking is what lets a thinker escape bad inherited assumptions, and inherited assumptions are often the actual bottleneck in a stuck problem — not a lack of analytical horsepower. The case against: it's used rarely compared to the daily-use skills like causal and probabilistic thinking, so its average importance across all decisions is lower even if its peak impact is high.
One system promoted inversion thinking sharply, into the middle of its ranking. The other two kept it near the bottom, treating it as a sharp but narrow specialized technique — high-leverage when you remember to use it, but not something you reach for constantly. Since this promotion wasn't replicated, it looks like an individual judgment call rather than a real signal to update the hierarchy on.
Rankings ranged from the top five to the middle of the top ten. All three systems agreed it's closely related to systems thinking — seeing the architecture producing outcomes rather than the symptoms — they just disagreed on how much additional value it adds once systems thinking is already in place.
The self-grading trap
When each AI was asked to grade the other two on ranking quality, the results were exactly as unreliable as you'd expect: every system rated itself among the strongest performers. One gave itself a 96 out of 100 while grading a competitor at 88. Another gave itself the top score in a three-way comparison it was itself constructing. This is worth naming explicitly, because it's a clean example of a bias this series has already covered elsewhere: nothing in the model architecture prevents a system from unconsciously favoring its own reasoning style when asked to judge reasoning quality, and no amount of confident-sounding scoring language changes that. Those self-assessment scores were discarded entirely in building the revised hierarchy below — they measure how each AI sees itself, not how good any ranking actually is.
The one useful thing that came out of the self-grading round was a shared observation, made independently by more than one system: a ranking task like this one doesn't really have a defensible single-scale answer, and an AI willing to say so is giving you better epistemics than one that hands you a beautifully confident list.
One correction: divergence isn't always low confidence
The first draft of this revision labeled everything past roughly rank 10 a "low-confidence zone" — the assumption being that wide disagreement between the three rankings meant the models simply couldn't pin these skills down. That's true for some of them. But for a specific subset, the divergence has a cleaner explanation: the skill's importance genuinely depends on who's using it, not on any model's uncertainty.
Inversion thinking is the clearest case. It's a huge lever for an investor running pre-mortems or a risk manager stress-testing a plan, and a minor one for a copywriter or a customer-support rep. Quantitative reasoning splits the same way — indispensable for anyone working with data, secondary for someone whose work is mostly qualitative. When three independent rankings put a skill anywhere from the top ten to the bottom five, that spread isn't always a measurement problem to average away. Sometimes it's the correct answer for three different audiences, superimposed.
That reframes part of Tier 4 below: it isn't a single bucket of "we don't know," it's two different things wearing the same label — skills nobody has pinned down yet, and skills that are highly rated conditionally on your specific work.
One more distinction: what the models agreed on vs. what got synthesized
Tier 2 below groups five skills together, but they didn't earn that grouping the same way, and the article should say so rather than blur it. Bayesian reasoning landed at 10, 11, and 12 across the three rankings — tight clustering, real agreement. First principles thinking landed at 6, 7, and 12 — a much wider spread, closer to a genuine three-way disagreement than a consensus. Putting both in "Tier 2" is a reasonable editorial call, but it's a call, not a finding. Where a tier placement below reflects tight agreement across all three models, that's noted; where it reflects one model's synthesis of a real disagreement, that's noted too.
The revised hierarchy
Below is the updated ranking. Tier 1 is unchanged in composition, presented as a strongly-supported cluster among this list of 26 rather than an objectively validated law. Metacognition moves up. Tier 4 is now split in two: skills that are genuinely unresolved, and skills that are context-dependent rather than low-confidence — their rank should shift based on your own domain rather than sitting fixed on this list.
- Tier 1 — Foundation (order contestable, membership confirmed)
- Causal thinking, systems thinking, probabilistic thinking, critical thinking, analytical thinking.
- Tier 1.5 — Multiplier (promoted on cross-model evidence)
- Metacognition. Sits just below Tier 1 rather than in the middle of the pack — it's the skill that monitors and corrects the use of every other skill on this list.
- Tier 2 — Strategic intelligence
- Structural thinking, strategic thinking, second-order thinking (tight model agreement on all three); Bayesian reasoning (tight agreement: 10/11/12); first principles thinking (wide spread — 6/7/12 — placed here by editorial synthesis, not consensus).
- Tier 3 — Evidence and synthesis
- Abductive reasoning, synthetic thinking, integrative thinking, inductive reasoning, deductive reasoning, counterfactual thinking.
- Tier 4a — Context-dependent (rank depends on your domain, not on model uncertainty)
- Inversion thinking, quantitative reasoning, interdisciplinary thinking. These score as high as Tier 2 for the right kind of work — an investor or risk manager should rank inversion thinking much higher than this list does; a data-heavy role should do the same for quantitative reasoning.
- Tier 4b — Genuinely low-confidence (rankings scatter with no clear pattern)
- Structured thinking, meta-rational thinking, multimodal thinking, prefactual thinking, operational thinking, tactical thinking.
The honest note to end on: this hierarchy, like the original, isn't a settled fact — it's a working model, built to be updated when it meets better evidence. That's the whole point of the feedback-loop concept from the original article. This revision is that loop, run twice: once when Gemini flagged that "low confidence" was hiding a context-dependence distinction, and again when ChatGPT flagged that "three AIs agree" was being overclaimed as objective validation rather than reported as convergent, non-independent model judgment. Both corrections are folded into the version above — which is itself the point: an article about calibration should visibly get calibrated.
Mind Map 1 — "Cross-Checking a Subjective Ranking" Mind Map 2 — "The Revised Hierarchy"
No comments:
Post a Comment