Found that scores on different mental tests correlate (the positive manifold) and proposed a general factor, g.
How the top ~10 frontier models compare with people on IQ-style tests, on newer adaptability benchmarks, and on how fast each side improves. The short version: models now beat almost every human on static reasoning tests, still trail an ordinary person at learning from experience, and are improving on a timescale of months while human cognition is flat.
The public Mensa Norway test is in training data, so treat it as a ceiling. The offline test was written by a Mensa member and never posted online, so it is the fairer comparison.
Where a model has both scores, the gap shows how much of the public number is memorisation. Claude Fable 5.1 drops 21 points; GPT-5.6 Sol Ultra (Vision) drops only 2.
Grok-4.20 Expert is the April 2026 reading; all others are the September 2026 snapshot. Blank bar = not tested on that variant.
ARC-AGI-1 and 2 are static grid puzzles, and models have passed the average person. ARC-AGI-3 drops an agent into an unfamiliar game with no instructions and scores how efficiently it learns the rules against human action counts.
GPT-6 Astra's 62.7% on ARC-AGI-3 is the official standard-harness score. A 99.9% figure also circulates; it used OpenAI's own agent harness and is not comparable.
Nine dimensions scored 0–100, anchored to benchmarks where one exists. The frontier-AI profile is "jagged": superhuman on knowledge, speed and static reasoning, weak on continual learning, sample efficiency and embodiment.
Profile scores are an editorial index, not a published benchmark. Anchors are listed per dimension; where no benchmark exists (continual learning, embodiment) the score is a judgment call.
Who is "smarter" depends on what you count. Pick a definition or drag the weights.
Left: best public-test IQ over time. It rose about 50 points in 30 months. The Flynn effect, the fastest measured change in human IQ, is about 3 points per decade. Right: METR's time horizon, the length of expert software task an agent completes half the time (log scale).
METR says its suite cannot measure reliably above about 16 hours, so Claude Mythos Preview's "≥16 h" is a floor. At 80% reliability the best horizon is about 3 hours: models attempt long tasks well before they finish them dependably.
Every chart below is computed in your browser from the equations in section 10. Parameters are calibrated to the published numbers above; where a step is an assumption, the chart says so.
Height is the probability of finishing a task. One axis is how new the task is (0 = well-practised, 1 = never seen, like ARC-AGI-3). The other is how long it takes a human expert, from one minute to one year. Blue is the frontier AI; green is a human master in their own field. Drag to rotate.
Model: P = σ(k·log₂(H₅₀/h) − β(n − 0.2)). AI: H₅₀ = 16 h in May 2026 doubling every 131 days, k = 0.574 (fits METR's 16 h at 50% and 3 h at 80%), β = 6.0 (fits ~30% on 8-minute novel tasks, the median of the ARC-AGI-3 leaders). Master: H₅₀ = 2,000 h, k = 0.5, β = 1.5 (assumed). Moving the date only grows H₅₀; if novelty handling also improves, the blue surface rises faster on the right.
Each hexagon counts simulated people, drawn from a bivariate normal with correlation 0.6 between the two abilities, with marginal histograms on each axis. Models sit far right on static reasoning but low on learning new tasks, a corner of the map almost no human occupies.
x = (offline IQ − 100) / 15. y = Φ⁻¹(½ · ARC-AGI-3 score), which assumes the 100% human baseline sits at the human median. Hollow markers estimate offline IQ as public Mensa score minus 18, the average public-to-offline gap measured for models with both.
The nine profile dimensions from section 4 collapsed into three: Knowledge & reasoning (dims 1–3), Adaptability (4, 6, 7) and Agency & grounding (5, 9). Each grey-orange point is a simulated person. The blue path is the frontier AI from 2023 to today.
The AI path before September 2026 is an editorial reconstruction, scaled to ARC-AGI-3 and METR progress; it is not a measured series.
Five relations drive this page. The last two show why weighting matters: an arithmetic mean lets a superpower cover a weakness; a geometric mean treats the weakest ability as a bottleneck, which is closer to how general intelligence behaves in real work.
Logistic in log-time, as METR fits it, plus a penalty for novelty n.
H₀ = 16 h (May 2026), Td ≈ 131 days (METR TH1.1, 2023 onward).
Static and adaptive z-scores; 0 = average person, 3 = one in 740.
Uses the weights you set in section 5.
λ is the share of weight you put on learning new tasks.
When a profile has one very weak dimension, IG falls well below IA. Frontier AI loses the most, because continual learning and embodiment score low.
The counters tick in real time from the growth equation. They are an extrapolation of the last measured point, not a live measurement. The chart shows the projection with a band for doubling times of 100–200 days, against human work lengths.
METR notes its task suite cannot measure reliably above about 16 hours, and trends like this bend. Treat dates as "if the 2023–26 trend holds", and remember the 80% horizon is about five times shorter.
Dashes mean no published score on that test. Human rows give the reference point where one exists.
HLE = Humanity's Last Exam (Artificial Analysis run). AA Index = Artificial Analysis Intelligence Index v4.3 (ceiling 58 as of 22 Sep). Gemini 3.1 Pro's ARC figures are Feb and Mar 2026.
Every assumption above traces to one of these. They run from a century of human psychometrics to the 2025–26 benchmark papers that supply the AI numbers.
Found that scores on different mental tests correlate (the positive manifold) and proposed a general factor, g.
Split intelligence into fluid (solving new problems) and crystallized (accumulated knowledge).
Reanalysed 460+ datasets into a three-stratum model, the basis of today's Cattell–Horn–Carroll (CHC) theory.
Documented population IQ gains of roughly 3 points per decade, the fastest measured change in human scores.
Expert performance comes from about a decade of deliberate practice; defines what a human master is.
85 years of data: general mental ability is the strongest single predictor of job performance, but far from the only one.
Defines intelligence as performance weighted across all environments, a formal version of a weighted composite.
Shows why human IQ tests transfer poorly to machines and argues for task-difficulty-based measurement.
People learn new concepts from a few examples by building causal, compositional models; machines of the time did not.
Networks overwrite old skills when trained on new ones; the core obstacle to continual learning.
Intelligence as skill-acquisition efficiency, not skill; introduced ARC.
Rates systems on two axes, performance and generality, against percentiles of skilled adults.
Scores AI on ten CHC domains at equal weight: GPT-4 27%, GPT-5 57%, with long-term memory storage the largest deficit.
Loss falls as a power law in model size, data and compute.
GPT-3: large models learn tasks from a few in-context examples without weight updates.
Chinchilla: scale parameters and data together for a fixed compute budget.
Some abilities appear abruptly past a scale threshold.
Apparent jumps often come from the choice of metric; smooth metrics show smooth progress.
The 57-subject knowledge benchmark that frontier models later saturated, prompting HLE.
2,500 expert-written questions built to resist saturation.
Harder static puzzles; average person 66%, human panel 100%.
Interactive games scored against human action efficiency; every model under 1% at launch.
Defines the 50% time horizon and its exponential growth.
GPT-3 matched or beat people on Raven-style analogy problems zero-shot.
Argued GPT-4 showed early general intelligence; widely debated on method.
Field experiment with 758 BCG consultants: AI raised quality on tasks inside its frontier and lowered it on tasks outside.
Citation tiers are approximate orders of magnitude from Google Scholar as of mid-2026, not exact counts; the citation databases could not be queried when this page was built. Recent papers are too new to rank by citations and are included because they are the primary source for the numbers on this page.