AnyLearn
All lessons
AIintermediate

What a Percentage Does and Does Not License

A model went from 27 percent to around 57 percent, so it is more than halfway to AGI and the rest arrives shortly. That inference is wrong in at least four ways, and working through why is more useful than the score itself. This lesson covers the linearity assumption, construct validity, contamination, and what the framework is good for once you stop reading it as a progress bar.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

The inference everyone makes

Two numbers, 27 and 57, invite a reading that the paper does not endorse and cannot support.

The reading is: the score roughly doubled between model generations, it is now past halfway, so at this rate AGI is a generation or two away.

Every step of that is shaky.

The scale may not be linear, so 57 percent may not mean more than halfway in any useful sense.

The rate is derived from two points, which is not a trend.

The remaining 43 percent is not the same kind of thing as the 57 already gained, because the domains that moved were the ones that respond to scale and the ones remaining largely are not.

And the score measures what the framework measures, which may or may not be what you care about.

None of this makes the framework useless. It makes it a diagnostic rather than a progress bar, and the difference matters because the second reading is the one that appears in headlines.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. The inference everyone makes

Two numbers, 27 and 57, invite a reading that the paper does not endorse and cannot support.

The reading is: the score roughly doubled between model generations, it is now past halfway, so at this rate AGI is a generation or two away.

Every step of that is shaky.

The scale may not be linear, so 57 percent may not mean more than halfway in any useful sense.

The rate is derived from two points, which is not a trend.

The remaining 43 percent is not the same kind of thing as the 57 already gained, because the domains that moved were the ones that respond to scale and the ones remaining largely are not.

And the score measures what the framework measures, which may or may not be what you care about.

None of this makes the framework useless. It makes it a diagnostic rather than a progress bar, and the difference matters because the second reading is the one that appears in headlines.

2. The remaining percent is harder than the gained percent

The most important objection to the progress-bar reading comes from the previous lesson's grouping.

The gains came from three sources. Scale raised knowledge, reading and writing, and mathematics. A new training objective raised on-the-spot reasoning. New input modalities raised visual and auditory processing.

All three were available. Scaling was already happening, verifiable-reward training was a method the field adopted, and multimodality was an engineering programme. The score moved because things already in motion arrived.

What remains is different in kind. Long-term memory storage sits at zero and requires weights that change after deployment, which is an unsolved research problem rather than a scheduled deliverable. Working memory at five and retrieval at four are partial and not obviously scale-limited. Speed at three reflects an architectural mismatch.

So the curve's shape is not set by how fast it has moved. It is set by whether continual learning gets solved, and nothing about the previous rate of progress bears on that.

A percentage invites extrapolation. This one should not be extrapolated, because the remaining domains are not on the same trajectory as the completed ones.

3. Construct validity, stated plainly

Behind the technical objections is one question that decides how much weight the score can carry: does the measurement measure the thing it names?

This is called construct validity, and it is the hardest problem in any measurement of something abstract. A tape measure has no construct validity problem because length is directly observable. Intelligence is not, so it must be measured through tasks assumed to depend on it, and the assumption is where the difficulty lives.

For human psychometrics there is a substantial answer. The factor structure replicates across populations and test batteries, scores predict outcomes over decades, and there is a century of accumulated evidence that the tests track something real about people.

None of that evidence transfers to machines. It was gathered from humans, and its validity is validity for humans. A machine scoring well on a test designed to measure a human ability has demonstrated it can produce the right outputs on that test, which is what the test measures for a person only because of assumptions about how a person would have to produce them.

That gap is real, and the paper's framework does not close it.

4. Why the factor structure may not transfer

There is a sharper version of that objection, aimed at the thing that made the framework attractive in the first place.

The broad abilities exist as a finding about correlations. Across human populations, someone strong on one verbal test tends to be strong on others, and moderately strong on non-verbal ones. Those correlation patterns are what factor analysis decomposes.

The correlations have causes: shared neural machinery, common developmental factors, general health, education, and a genetic contribution. The factors are downstream of facts about how humans are built and how they vary.

A machine shares none of that. Its abilities were determined by training data composition, objective, architecture and scale. There is no reason its abilities should covary the way human abilities do, and the jagged profile is direct evidence they do not: a human at full marks on mathematics and reading would not score zero on memory storage, because in humans those abilities are correlated.

So applying the structure to machines uses a decomposition whose empirical basis does not hold for the thing being measured. The domains remain a reasonable checklist of capabilities. Treating them as factors of a machine's general intelligence is the step that overreaches.

5. Four ways to misread the number

The failure modes are distinct, and separating them makes the score usable rather than either authoritative or worthless.

Reading it as linear treats 57 percent as more than halfway. The scale has no established linearity, and the remaining domains resist the methods that produced the gains.

Reading it as a forecast extrapolates two data points into a date. Two points establish a difference, not a rate, and the mechanism behind the difference does not apply to what is left.

Reading it as validated imports the credibility of human psychometrics into a machine setting. The evidence base for that credibility was gathered on humans and does not transfer.

Reading it as complete assumes the domains cover everything that matters. They cover what human cognitive testing varies on, which omits a great deal: physical action, long-horizon planning, social judgement, and knowing when you do not know.

What survives all four is the per-domain profile, which is a set of measurements of specific capabilities with a clear account of what was tested. That is genuinely informative. The total is the part that invites every misreading.

flowchart TD
A["A score of 57 percent"] --> B["Read as linear: more than halfway"]
A --> C["Read as a forecast: a date from two points"]
A --> D["Read as validated: borrowing psychometric credibility"]
A --> E["Read as complete: the domains cover everything"]
B --> F["Remaining domains resist what produced the gains"]
C --> G["Two points are a difference, not a rate"]
D --> H["That evidence was gathered on humans"]
E --> I["Omits action, long horizons, social judgement"]

6. Contamination, and why it bites here

One practical objection applies to any benchmark and applies with particular force to this one.

The framework is assembled from cognitive tests, and cognitive tests are published. Items from well-known instruments, discussions of them, and worked examples all exist in text that models are trained on. A model may have encountered the items, or close relatives of them, during training.

When that happens, the test measures recall rather than the ability it names. This is the contamination problem from the synthetic data path, and it is worse here because the whole framework's credibility rests on the tests measuring the constructs they were designed for.

Contamination is also unevenly distributed across domains, which distorts the profile rather than merely lowering its reliability. Widely published verbal and knowledge items are more likely to be in training data than novel reasoning items constructed for the evaluation, so the domains most likely to be inflated are the ones already scoring highest.

The defence is to construct fresh items, which the authors would be aware of and which is expensive and has to be repeated. It is a maintenance burden rather than a one-time fix, and it is the reason any static evaluation degrades over time.

7. What is genuinely good here

Having spent four steps on limitations, it is worth being clear that the paper is a substantial improvement on what preceded it, which was nothing.

It is falsifiable. A definition with a procedure can be argued with on method, which is a large step up from competing intuitions about what AGI means.

It is exhaustive by construction. Rather than measuring what was convenient, it requires every domain to be scored, so a capability gap shows up as a zero rather than as an absence nobody noticed. That property produced the paper's most valuable finding.

It makes the jaggedness legible. A single number hides the shape; a profile shows that progress has been uneven and identifies where.

And it names a bottleneck with a mechanism. Long-term memory storage at zero, attributed to weights that do not change after deployment, is a concrete research target rather than a vague call for more capability.

That last item is the paper's real contribution. The definition is contestable and the total is easily misread, and the finding that current systems cannot form new long-term memories, with a clear account of why, stands regardless.

8. How to use it

Four habits make the framework useful without overclaiming, and they generalise to any capability measurement.

Quote the profile, never the total. The per-domain scores carry the information; the aggregate carries the misreadings. If you cite this paper, cite which domains are strong and which are zero.

Ask what a score licenses. A high score on a domain means the system produced correct outputs on those tests. It does not establish it will do so on your task, which is a different distribution.

Treat zeros as more informative than highs. A high score is consistent with genuine capability and with contamination. A zero is hard to fake and points at something structural.

And separate the definition from the measurement. You can accept that versatility across cognitive domains is a reasonable definition of general intelligence while disputing whether these ten domains, these tests, and this weighting measure it. Most productive criticism of the paper does exactly that.

The next course is about what happens when a measurement is taken at face value and the experimental design turns out to have been measuring something else.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why should the jump from 27 to around 57 percent not be extrapolated to a date?
    • Because the tests were changed between the two evaluations
    • Because two points establish a difference rather than a rate, and the remaining domains resist the scale, objective and modality changes that produced the gains
    • Because the scale is capped below 100 percent in practice
    • Because the models were not evaluated under identical conditions
  2. What is the construct validity problem in applying this framework to machines?
    • The tests are too easy for current models
    • The domains overlap, so scores are double-counted
    • The evidence that these tests track something real was gathered on humans, and that validity does not transfer to machines
    • Machines cannot be tested under standardised conditions
  3. What does the jagged profile itself suggest about the factor structure?
    • That the domains do not covary in machines the way they do in humans, since a human strong on mathematics and reading would not score zero on memory storage
    • That the weighting should be adjusted to reflect difficulty
    • That more domains are needed to capture machine cognition
    • That the tests were administered inconsistently
  4. Why does contamination distort the profile rather than just lowering reliability?
    • Because contaminated items are scored twice
    • Because it affects only the total, not individual domains
    • Because contamination is randomly distributed across domains
    • Because widely published verbal and knowledge items are likelier to be in training data than freshly constructed reasoning items, inflating the domains already scoring highest
  5. Which part of the paper is most defensible against the criticisms?
    • The overall percentage, since it aggregates across many tests
    • The identification of long-term memory storage at zero, with a mechanism, since a zero is hard to fake and points at something structural
    • The equal ten percent weighting, since it is empirically derived
    • The claim that the ten domains cover everything intelligence involves

Related lessons

AI
intermediate

What This Teaches About Measuring Anything

The exchange is a case study with transferable rules. A conclusion resting on failures needs a failure taxonomy. Every instance must be verified solvable before anyone is scored against it. Output format is a confound whenever answers get long. And when two explanations fit the same data, the productive move is to find the prediction on which they differ, then test it.

8 steps·~12 min
AI
intermediate

Ten Domains, and a Profile That Is Not Flat

The framework scores ten cognitive domains at ten percent each. Running it produces something more useful than the headline totals of 27 percent for GPT-4 and around 57 for GPT-5: a jagged profile, where a model is at or near full marks on some domains and at zero on others. This lesson walks the domains, reads both profiles column by column, and shows what the jaggedness explains.

8 steps·~12 min
AI
intermediate

Why AGI Had No Definition, and What Borrowing One Costs

For a term that anchors company charters, safety policy and hundreds of billions in investment, AGI has been remarkably undefined. A 2025 paper with over thirty authors proposes fixing that by borrowing psychology's most validated framework for human intelligence. This lesson covers why the definition was missing, what the paper anchors to, and what the borrowing assumes.

8 steps·~12 min
AI
intermediate

The Rebuttal: Three Ways to Score Zero Without Failing

The response disputed none of the data and argued the experiment measured something other than reasoning. Models had to print move lists exceeding their output limits, and said so in the transcripts. Some instances had no solution and were scored as failures anyway. And asking for a program instead of a move list produced high accuracy on instances reported as total collapse.

8 steps·~12 min