The "G" is Not a Benchmark: Why Scale Will Never Yield AGI

#AI#Philosophy


The tech industry treats AGI as a destination at the end of a smooth exponential curve. Keep scaling parameters, keep feeding in petabytes, keep stacking compute, and somewhere around 2029 the system crosses an invisible threshold and wakes up, presumably at 2:14 a.m. Eastern time.

That story misreads the G. Generality is not the ability to solve millions of specific tasks quickly. It is the ability to create new explanations. Current AI and general intelligence are not two points on one spectrum. They are different kinds of machine.

Current AI interpolates inside a space humans defined

The standard evidence for emerging AGI: AlphaFold mapped protein structures, an LLM passed the bar exam, algorithms predict molecular interactions.

These are remarkable feats of engineering. They are also, every one of them, statistical correlation inside a bounded space defined by training data. AlphaFold does not understand biochemistry. It interpolates structural patterns across known examples. It cannot propose a new theory to explain why the laws governing protein folding exist, and it cannot decide that protein folding is the wrong problem and go work on something better.

In every celebrated case, humans defined the win conditions and curated the search space. That is narrow optimization, done superbly.

General intelligence operates precisely where win conditions cannot be written down. It poses its own problems, conjectures hypotheses nobody asked for, and invents explanatory knowledge: accounts of why things happen, not just predictions of what happens next. Knowledge grows by conjecture and criticism, not by distilling patterns out of data.

More compute makes a better optimizer, never a different machine

“The curve is exponential and the gap is closing.”

If you build a faster ladder, you make exponential progress at climbing trees. No amount of ladder gets you to the Moon. The Moon requires a rocket, and a rocket is not a very tall ladder.

Scaling transformers optimizes pattern matching over existing human output. More compute makes the mimicry better. It adds nothing that could understand cause and effect, because the architecture has no mechanism for conjecture and refutation. An optimizer with infinite compute is still an optimizer, and it cannot originate an idea outside its loss function.

”If there is no benchmark, isn’t your definition unfalsifiable?”

This is the strongest objection, and it comes from a reasonable place: a definition you can never test sounds like a definition you can never lose with.

The objection assumes intelligence must be measured by test scores. But benchmarks test narrow outputs: exam questions, code completion, curated datasets. Any static benchmark can be gamed by a narrow optimizer with enough training data, which is Goodhart’s law running at datacenter scale. A system that games a test without understanding the material has demonstrated the opposite of generality.

The criterion for AGI is a capability: can the system pose its own problem, conjecture an original explanation to solve it, and criticize its own idea against reality? Nobody moved that goalpost. That has been the definition of creative thought all along.

The general intelligence in the loop is still you

“Humans plus AI are already AGI.” This one misses where the creative engine sits.

A human using AlphaFold is a general intelligence using a specialized tool. Hardware speed has nothing to do with generality. You interpret the output, catch the errors, invent the next explanation, and pick the next question. That loop is where the intelligence lives, and every part of it that creates something new is currently running on the person, not the model.

Until a machine can invent a new problem, reject its own programming because it found a better explanation, and produce hard-to-vary explanations of its own, it is a tool. A spectacularly powerful, economically disruptive tool. Not a general one.

In summary: scale builds taller ladders. AGI is the Moon.