Machine vs Human Minds
Snapshot · 27 Sep 2026 · public leaderboards

Machine vs Human Minds

How the top ~10 frontier models compare with people on IQ-style tests, on newer adaptability benchmarks, and on how fast each side improves. The short version: models now beat almost every human on static reasoning tests, still trail an ordinary person at learning from experience, and are improving on a timescale of months while human cognition is flat.

Top AI · Mensa Norway
151Claude Fable 5.1 and GPT-6 Astra (Vision). About 1 in 3,000 humans score this high.
Top AI · leak-proof offline test
135GPT-5.6 Terra Ultra. Roughly the top 1% of humans. Scores drop 15–21 pts off the public test.
ARC-AGI-3 · learning new games
62.7%GPT-6 Astra vs a human baseline of 100%. Was under 1% for every model in March 2026.
METR task horizon doubling
~131 daysLength of expert task AI finishes half the time. Human capacity: roughly constant.
1 · Where the models sit on the human curve

IQ distribution (mean 100, SD 15) with AI scores marked

The public Mensa Norway test is in training data, so treat it as a ceiling. The offline test was written by a Mensa member and never posted online, so it is the fairer comparison.

Human population Top AI, public test (151) Top AI, offline test (135)
2 · Top 10 models on IQ tests

Public vs offline IQ, TrackingAI

Where a model has both scores, the gap shows how much of the public number is memorisation. Claude Fable 5.1 drops 21 points; GPT-5.6 Sol Ultra (Vision) drops only 2.

Mensa Norway (public) Offline (never online) Human mean 100 / Mensa cut-off 130

Grok-4.20 Expert is the April 2026 reading; all others are the September 2026 snapshot. Blank bar = not tested on that variant.

3 · Depth vs adaptability

Same puzzle family, three generations: the gap reopens when the task demands learning

ARC-AGI-1 and 2 are static grid puzzles, and models have passed the average person. ARC-AGI-3 drops an agent into an unfamiliar game with no instructions and scores how efficiently it learns the rules against human action counts.

Best AI Average individual human Human panel / baseline

GPT-6 Astra's 62.7% on ARC-AGI-3 is the official standard-harness score. A 99.9% figure also circulates; it used OpenAI's own agent harness and is not comparable.

4 · Cognitive profile

Different shapes of intelligence, not one number

Nine dimensions scored 0–100, anchored to benchmarks where one exists. The frontier-AI profile is "jagged": superhuman on knowledge, speed and static reasoning, weak on continual learning, sample efficiency and embodiment.

Frontier AI (best of top 10) Average adult Top human expert (in their field)

Profile scores are an editorial index, not a published benchmark. Anchors are listed per dimension; where no benchmark exists (continual learning, embodiment) the score is a judgment call.

5 · Change the weights, change the winner

Weighted composite intelligence

Who is "smarter" depends on what you count. Pick a definition or drag the weights.

6 · Growth velocity

AI moves in months; human cognition moves in decades

Left: best public-test IQ over time. It rose about 50 points in 30 months. The Flynn effect, the fastest measured change in human IQ, is about 3 points per decade. Right: METR's time horizon, the length of expert software task an agent completes half the time (log scale).

Top AI Top AI, offline test Human reference

METR says its suite cannot measure reliably above about 16 hours, so Claude Mythos Preview's "≥16 h" is a floor. At 80% reliability the best horizon is about 3 hours: models attempt long tasks well before they finish them dependably.

Part II · The model behind the comparison

3D landscapes, population maps and the equations that link them

Every chart below is computed in your browser from the equations in section 10. Parameters are calibrated to the published numbers above; where a step is an assumption, the chart says so.

7 · Success landscape (3D)

Where a human master still wins: novel, long tasks

Height is the probability of finishing a task. One axis is how new the task is (0 = well-practised, 1 = never seen, like ARC-AGI-3). The other is how long it takes a human expert, from one minute to one year. Blue is the frontier AI; green is a human master in their own field. Drag to rotate.

Model: P = σ(k·log₂(H₅₀/h) − β(n − 0.2)). AI: H₅₀ = 16 h in May 2026 doubling every 131 days, k = 0.574 (fits METR's 16 h at 50% and 3 h at 80%), β = 6.0 (fits ~30% on 8-minute novel tasks, the median of the ARC-AGI-3 leaders). Master: H₅₀ = 2,000 h, k = 0.5, β = 1.5 (assumed). Moving the date only grows H₅₀; if novelty handling also improves, the blue surface rises faster on the right.

8 · Population map

20,000 simulated people vs the models: static IQ against adaptive learning

Each hexagon counts simulated people, drawn from a bivariate normal with correlation 0.6 between the two abilities, with marginal histograms on each axis. Models sit far right on static reasoning but low on learning new tasks, a corner of the map almost no human occupies.

Human population density AI, measured (filled) · estimated (hollow) Human master (top 0.1%, z = 3)

x = (offline IQ − 100) / 15. y = Φ⁻¹(½ · ARC-AGI-3 score), which assumes the 100% human baseline sits at the human median. Hollow markers estimate offline IQ as public Mensa score minus 18, the average public-to-offline gap measured for models with both.

9 · Capability space (3D)

Humans form a cloud; AI moves along one axis

The nine profile dimensions from section 4 collapsed into three: Knowledge & reasoning (dims 1–3), Adaptability (4, 6, 7) and Agency & grounding (5, 9). Each grey-orange point is a simulated person. The blue path is the frontier AI from 2023 to today.

The AI path before September 2026 is an editorial reconstruction, scaled to ARC-AGI-3 and METR progress; it is not a measured series.

10 · Equations

The maths, and a live adaptability-adjusted IQ

Five relations drive this page. The last two show why weighting matters: an arithmetic mean lets a superpower cover a weakness; a geometric mean treats the weakest ability as a bottleneck, which is closer to how general intelligence behaves in real work.

Task success
P(n,h)=σ(k·log2H50h−β(n−n0))

Logistic in log-time, as METR fits it, plus a penalty for novelty n.

Horizon growth
H50(t)=H0·2(t−t0)/Td

H₀ = 16 h (May 2026), Td ≈ 131 days (METR TH1.1, 2023 onward).

Placing AI on the human scale
zs=IQoff−10015,za=Φ−1(12SARC3)

Static and adaptive z-scores; 0 = average person, 3 = one in 740.

Arithmetic vs bottleneck composite
IA=∑wisi∑wi,IG=∏isiwi/∑w

Uses the weights you set in section 5.

Adaptability-adjusted IQ
IQ*=100+15[(1−λ)zs+λza]

λ is the share of weight you put on learning new tasks.

Composite with your section-5 weights

When a profile has one very weak dimension, IG falls well below IA. Frontier AI loses the most, because continual learning and embodiment score low.

11 · Live projection

AI task horizon, extrapolated to this second

The counters tick in real time from the growth equation. They are an extrapolation of the last measured point, not a live measurement. The chart shows the projection with a band for doubling times of 100–200 days, against human work lengths.

Extrapolated 50% horizon now
–From 16 h on 8 May 2026 (a floor).
Reaches 1 work-week (40 h)
–
Reaches 1 work-month (167 h)
–
Reaches a master's year-long project (2,000 h)
–

METR notes its task suite cannot measure reliably above about 16 hours, and trends like this bend. Treat dates as "if the 2023–26 trend holds", and remember the 80% horizon is about five times shorter.

12 · Leaderboard

Top frontier models across benchmarks

Dashes mean no published score on that test. Human rows give the reference point where one exists.

HLE = Humanity's Last Exam (Artificial Analysis run). AA Index = Artificial Analysis Intelligence Index v4.3 (ceiling 58 as of 22 Sep). Gemini 3.1 Pro's ARC figures are Feb and Mar 2026.

13 · Research foundations

The peer-reviewed and most-cited work behind this page

Every assumption above traces to one of these. They run from a century of human psychometrics to the 2025–26 benchmark papers that supply the AI numbers.

Established · 100–1k2017
The Measure of All Minds
Hernández-Orallo, J. · Cambridge University Press

Shows why human IQ tests transfer poorly to machines and argues for task-difficulty-based measurement.

Used in §1–2 contamination caveats
Highly cited · 1k–10k2017
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B. & Gershman, S. J. · Behavioral and Brain Sciences

People learn new concepts from a few examples by building causal, compositional models; machines of the time did not.

Used in §4 sample efficiency
Highly cited · 1k–10k2019
On the Measure of Intelligence
Chollet, F. · arXiv

Intelligence as skill-acquisition efficiency, not skill; introduced ARC.

Used in §3 and §5 adaptability weights
Recent primary source2025
A Definition of AGI
Hendrycks, D., Song, D., Szegedy, C., Lee, H., Gal, Y. et al. · arXiv

Scores AI on ten CHC domains at equal weight: GPT-4 27%, GPT-5 57%, with long-term memory storage the largest deficit.

Used in §4 jagged profile; §5 balanced preset
Landmark · 10k+2020
Language models are few-shot learners
Brown, T. B. et al. · NeurIPS

GPT-3: large models learn tasks from a few in-context examples without weight updates.

Used in §4 in-context learning vs memory
Established · 100–1k2023
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B. & Koyejo, S. · NeurIPS (outstanding paper)

Apparent jumps often come from the choice of metric; smooth metrics show smooth progress.

Used in §5 why weighting and metric choice matter
Recent primary source2025
Humanity's Last Exam
Phan, L. et al. · arXiv

2,500 expert-written questions built to resist saturation.

Used in §12 HLE column
Established · 100–1k2023
Navigating the jagged technological frontier
Dell'Acqua, F. et al. · Harvard Business School working paper

Field experiment with 758 BCG consultants: AI raised quality on tasks inside its frontier and lowered it on tasks outside.

Used in §4 jagged profile; §7 landscape

Citation tiers are approximate orders of magnitude from Google Scholar as of mid-2026, not exact counts; the citation databases could not be queried when this page was built. Recent papers are too new to rank by citations and are included because they are the primary source for the numbers on this page.

Reading this fairly

What the numbers do and don't say

IQ tests measure skill, not learning.Chollet's definition of intelligence is how efficiently a learner turns experience into new skills. A model trained on millions of puzzles scoring 151 shows skill; ARC-AGI-3 is closer to measuring the learning.
Contamination inflates public scores.The 15–21 point public-to-offline drop is the clearest evidence. Use offline or held-out scores for any human comparison.
Harness and effort matter.Max-effort runs, agent scaffolds and vendor self-reports move scores by large amounts. Compare like with like.
Humans keep what AI lacks.Learning on the job from a handful of examples, long-lived memory, physical skill and reliable judgment over weeks of work.
Sources