AnyLearn
All lessons
AIintermediate

Why AGI Had No Definition, and What Borrowing One Costs

For a term that anchors company charters, safety policy and hundreds of billions in investment, AGI has been remarkably undefined. A 2025 paper with over thirty authors proposes fixing that by borrowing psychology's most validated framework for human intelligence. This lesson covers why the definition was missing, what the paper anchors to, and what the borrowing assumes.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

A load-bearing term with nothing under it

Artificial general intelligence appears in company mission statements, in contracts, in national policy documents, and in arguments about whether the current trajectory is dangerous. It is doing a great deal of work.

And until recently there was no operational definition anyone agreed on. Ask ten researchers and you get answers ranging from a system that can do any task a human can, to a system that can do most economically valuable work, to a system that can improve itself, to a system that would pass as human in unrestricted conversation.

Those are not variations on one idea. They pick out different systems, arrive at different times, and imply different policies.

The practical consequence is that claims about AGI cannot be checked. If someone says a system is close to AGI, or that AGI is decades away, there is no measurement that would settle it, so the disagreement is unfalsifiable in both directions.

This is the gap a paper published in October 2025, with Dan Hendrycks, Dawn Song, Christian Szegedy, Yarin Gal, Erik Brynjolfsson and more than thirty other authors, sets out to close.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. A load-bearing term with nothing under it

Artificial general intelligence appears in company mission statements, in contracts, in national policy documents, and in arguments about whether the current trajectory is dangerous. It is doing a great deal of work.

And until recently there was no operational definition anyone agreed on. Ask ten researchers and you get answers ranging from a system that can do any task a human can, to a system that can do most economically valuable work, to a system that can improve itself, to a system that would pass as human in unrestricted conversation.

Those are not variations on one idea. They pick out different systems, arrive at different times, and imply different policies.

The practical consequence is that claims about AGI cannot be checked. If someone says a system is close to AGI, or that AGI is decades away, there is no measurement that would settle it, so the disagreement is unfalsifiable in both directions.

This is the gap a paper published in October 2025, with Dan Hendrycks, Dawn Song, Christian Szegedy, Yarin Gal, Erik Brynjolfsson and more than thirty other authors, sets out to close.

2. Why benchmarks were not already the answer

The field has thousands of benchmarks, so it is worth being clear about why none of them answered this.

Benchmarks measure task performance. A model scores well on a set of exam questions, coding problems, or reading comprehension items. That tells you about the tasks in the set, and generalising from it requires an assumption that the set represents something broader.

Three problems follow. Benchmarks saturate: once models score near the ceiling, the benchmark stops distinguishing anything and a new one is needed, so the yardstick keeps changing. They are contaminable: a benchmark published on the internet may be in the training data, and a high score may measure recall. And they are unrepresentative by construction, since a benchmark contains what was easy to collect and grade rather than a sample of what intelligence involves.

Most importantly, benchmarks have no theory of coverage. A model that scores well on twenty benchmarks might still have a large hole, because nobody built a benchmark for the thing in the hole.

That last problem is the one the paper attacks. It wants a decomposition of intelligence that is complete in some defensible sense, and then measurement against each part.

3. The anchor: a well-educated adult

The paper's definition is a single sentence. Artificial general intelligence is a system matching the cognitive versatility and proficiency of a well-educated adult.

Both halves are doing work.

Versatility means breadth. Not excellence at one thing but competence across the range of cognitive tasks a person handles. This rules out a system that is superhuman at mathematics and unable to hold a conversation.

Proficiency means a level, and the level is deliberately modest. Not the best human in each domain, not a genius, but a well-educated adult. Something a great many people already are.

Choosing a modest target is a defensible move rather than a lowering of the bar. A definition pinned to the best human in every field would be satisfied by nothing for a long time and would tell you little along the way. A definition pinned to an ordinary competent person gives a threshold that is meaningful, reachable in principle, and measurable, because psychology has spent a century measuring exactly that population.

That last point is what makes the rest of the paper possible.

4. Borrowing a century of factor analysis

Having anchored to human cognition, the paper needs a decomposition of it, and rather than inventing one it takes the one psychology already converged on.

Cattell-Horn-Carroll theory is the product of roughly a century of a specific empirical exercise. Give large populations many different cognitive tests. Look at which scores correlate. Use factor analysis to find the smallest set of underlying abilities that explains the pattern of correlations. Repeat, refine, and merge competing models.

The result is a hierarchical structure of broad abilities that has been replicated across populations and test batteries, and the paper describes it as the most empirically validated model of human cognition available.

The attraction for this purpose is that the decomposition was not chosen by the authors. It emerged from data about how human abilities actually cluster, which is what gives the coverage claim its force. If those broad abilities exhaust what human cognitive test performance varies on, then measuring all of them is a defensible definition of measuring general intelligence.

That is the paper's central methodological bet, and it is also where the criticism lands.

5. What the borrowing assumes

The move from a human framework to a machine assessment involves an inference worth making explicit, because it is easy to miss.

The framework was derived from correlations among human test scores. It says that in human populations, performance on many tests can be explained by a smaller number of broad abilities, because those abilities vary together across people in a stable way.

Applying it to machines assumes that the same decomposition carves machine cognition at the same joints. That is not obviously true, and there is reason to doubt it. The factor structure exists because of how human brains vary, shaped by shared neural machinery, development and genetics. A transformer has none of those constraints, so there is no particular reason its abilities should cluster the same way.

The paper has a partial answer, and it is a good one: even if the factors do not carve machine cognition naturally, the domains still enumerate things humans can do, and a system claiming human-level general capability has to do them.

So the framework may be better read as a coverage checklist derived from human abilities than as a theory of machine cognition. That reading is weaker, and it survives the objection.

flowchart TD
A["A century of human cognitive test data"] --> B["Factor analysis finds broad abilities"]
B --> C["Structure holds because human brains vary in shared ways"]
C --> D["Applied to machines"]
D --> E["Strong reading: these carve machine cognition too"]
D --> F["Weak reading: a coverage checklist of things humans can do"]
E --> G["Contested: machines share none of those constraints"]
F --> H["Survives the objection"]

6. The move that makes it operational

Two design decisions turn the framework from a description into something that produces a number.

The first is exhaustiveness. Every broad ability in the framework must be measured. You cannot score highly by being excellent at the ones you happen to be good at, because the missing ones are counted as zero rather than omitted. This is what makes it a versatility measure rather than a capability measure.

The second is equal weighting. Each of the ten domains contributes ten percent, and the authors are explicit that this is to prioritise breadth. A system superb at four domains and absent on six scores forty percent, no matter how superb.

That second decision is a value judgment rather than an empirical finding, and the paper does not pretend otherwise. There is no fact about the world establishing that auditory processing matters exactly as much as mathematical ability for general intelligence. One could argue the weights should reflect economic value, or frequency of use, or difficulty.

Equal weighting is the choice that best serves the definition, since versatility is the thing being measured. It is a choice, and reading the score requires remembering that it was made.

7. What a definition is for

Before the results, it is worth asking what a definition like this is actually supposed to achieve, because judging it against the wrong purpose produces bad criticism.

It is not supposed to settle what intelligence is. That is a philosophical question the paper does not claim to close.

It is not supposed to predict when AGI arrives. It provides a measuring stick, and a measuring stick makes no forecast.

What it is for is making claims checkable. Once a definition and a procedure exist, someone saying a system is nearly AGI can be asked which domains it scores well on and which it does not. The conversation moves from assertion to evidence, and the number can be disputed on method rather than on intuition.

A secondary purpose is direction. A framework that scores every domain shows which are weak, and weak domains are research targets. The paper's headline empirical finding is exactly of this kind, and it is more interesting than the total.

That is the next lesson: the ten domains, the scores, and the one that comes out at essentially zero.

8. Reading a preprint carefully

One habit worth carrying into the next lesson, since this paper is a preprint and preprints move.

This one was submitted in October 2025 and revised in December. The headline figures shifted slightly between versions, and the per-domain breakdown in the first version sums to a total a percentage point away from the abstract's figure.

That is entirely normal for a preprint and it is worth noticing, because a number cited without a version is a number that may not be in the paper you find.

The general habit is to cite the version, and to treat exact figures from any preprint as provisional while treating the structure as the durable part. Here the structure is the definition, the ten domains, and the equal weighting. Those have not moved. The precise percentage for a given model is the part most likely to change, and also the part most likely to be quoted in a headline.

More than thirty authors on a paper is a signal that its framing has broad buy-in. It is not a signal that every number in it is final.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. What definition of AGI does the paper propose?
    • A system that can perform most economically valuable work
    • A system matching the cognitive versatility and proficiency of a well-educated adult
    • A system that can improve its own capabilities without human input
    • A system indistinguishable from a human in unrestricted conversation
  2. Why are existing benchmarks insufficient for this purpose?
    • They are too difficult for current models
    • They cannot be automatically scored
    • They saturate, can be contaminated by training data, and have no theory of coverage, so a hole can exist simply because nobody built a benchmark for it
    • They only measure knowledge rather than reasoning
  3. What makes Cattell-Horn-Carroll theory attractive as the decomposition?
    • It was designed specifically for evaluating artificial systems
    • It produces a single number rather than a profile
    • It is the simplest model of cognition available
    • It emerged from a century of factor analysis on human test data rather than being chosen by the authors, which is what gives the coverage claim force
  4. What is the main objection to applying a human factor structure to machines?
    • The factor structure exists because of how human brains vary, and a transformer shares none of those constraints, so its abilities need not cluster the same way
    • Machines cannot take psychometric tests reliably
    • The theory was developed before modern statistics
    • Human tests are culturally biased
  5. What is the status of the equal ten percent weighting across domains?
    • An empirical finding from the factor analysis
    • A value judgment the authors make explicitly, chosen to prioritise breadth over narrow excellence
    • A constraint imposed by the scoring software
    • A weighting derived from the economic value of each ability

Related lessons

AI
intermediate

Ten Domains, and a Profile That Is Not Flat

The framework scores ten cognitive domains at ten percent each. Running it produces something more useful than the headline totals of 27 percent for GPT-4 and around 57 for GPT-5: a jagged profile, where a model is at or near full marks on some domains and at zero on others. This lesson walks the domains, reads both profiles column by column, and shows what the jaggedness explains.

8 steps·~12 min
AI
intermediate

What This Teaches About Measuring Anything

The exchange is a case study with transferable rules. A conclusion resting on failures needs a failure taxonomy. Every instance must be verified solvable before anyone is scored against it. Output format is a confound whenever answers get long. And when two explanations fit the same data, the productive move is to find the prediction on which they differ, then test it.

8 steps·~12 min
AI
intermediate

The Rebuttal: Three Ways to Score Zero Without Failing

The response disputed none of the data and argued the experiment measured something other than reasoning. Models had to print move lists exceeding their output limits, and said so in the transcripts. Some instances had no solution and were scored as failures anyway. And asking for a program instead of a move list produced high accuracy on instances reported as total collapse.

8 steps·~12 min
AI
intermediate

The Experiment: Puzzles With a Difficulty Dial

Apple researchers built an evaluation designed to fix a real problem with benchmarks: puzzles where difficulty turns up smoothly while the logic stays identical, and every step can be checked. They found accuracy collapsing to zero past a threshold, and not improving when the solution algorithm was handed to the model. This lesson covers the design and why it was a good one.

8 steps·~12 min