In February 2026, a commentary in Nature argued that frontier AI systems have already crossed the threshold into general intelligence. The authors weren't fringe voices. They made a real, technical case that today's models satisfy the functional standards researchers have used to define human-level cognition for decades.
That same month, Nvidia's CEO told a public audience that AGI had already been achieved. And that same month, a benchmark called ARC-AGI-3 launched, built specifically to test the kind of interactive reasoning humans do without effort. Every frontier model scored under one percent.
A peer-reviewed journal, a major industry CEO, and a purpose-built measurement tool reached opposite conclusions inside the same four weeks. Nobody involved was lying. They were using the same word to mean different things, and the disagreement never actually got resolved. It just moved on to the next headline.
That's not a temporary confusion waiting to clear up. It's the actual condition the entire "path to superintelligence" conversation is being built on top of.
A CEO says AGI already happened. A benchmark says every frontier model fails at reasoning humans find trivial. Both statements are about the same month.
The word has never had one meaning
OpenAI's founding charter defines AGI in economic terms: a system that outperforms humans at most economically valuable work. DeepMind uses a five-level capability framework instead, measuring generalization against skilled human performance. Other researchers define it behaviorally, as the point where people stop checking the machine's work and simply trust it.
These aren't three phrasings of the same idea. They're three different finish lines. A system can cross one of them while failing the other two badly, which is exactly what happened in February. The Nature authors were using a functional, capability-based standard. Huang was speaking in the loose, economic sense the industry uses when it wants to signal momentum. ARC-AGI-3 was testing something narrower and harder: genuine novel reasoning, not memorized pattern-matching dressed up as understanding.
Economic performance, capability generalization, and earned human trust are three genuinely different thresholds. A system can cross any one of them without being close to the other two. Most public AGI claims don't specify which line they mean.
There's a fourth definition circulating too, quieter than the other three but arguably more consequential: AGI as the point where people simply stop checking the machine's work. Under that standard, arrival isn't a benchmark score at all, it's a behavioral shift in millions of individual users who've started defaulting to AI the way they default to a search bar, without deliberating first. That's a real and measurable change in habit. It says nothing about whether the underlying reasoning is sound, which is exactly what ARC-AGI-3 was built to test, and exactly where the frontier models failed.
The confusion has a paper trail
This isn't just an academic disagreement, either. It has already forced its way into contract law. OpenAI's original partnership agreement with Microsoft included a clause stating that if OpenAI's own board declared AGI had been reached, Microsoft's exclusive rights to the technology would be voided. The clause never specified a measurable definition of what that meant.
For years, that ambiguity created real tension between the two companies, since whoever got to define the word also got to decide who controlled the technology. The clause was amended twice, first requiring independent verification instead of a unilateral board declaration, then removed from the agreement entirely in April 2026. A term nobody could pin down had to be legislated out of a multi-billion dollar contract because it was too unstable to build a business relationship on.
Meanwhile, timelines for what comes after AGI are being handed out with similar confidence and even less agreement. Sam Altman has said years. Demis Hassabis has said closer to a decade. Yann LeCun has said not with the current approach at all. The researcher who built the ARC-AGI benchmark specifically to cut through this kind of noise recently cut his own estimate in half, from roughly a decade to roughly five years, based on results from a test that current systems are still failing badly.
What "progression" quietly assumes
A phrase like "the path from AGI to superintelligence" only makes sense if AGI is a fixed point you can stand on and look forward from. The evidence says it isn't. It's a word four or five different groups are actively fighting to define in their own favor, with real financial and reputational stakes riding on whose definition wins.
That doesn't mean the underlying capability questions are fake. Models are getting more capable in specific, measurable ways, and something resembling general reasoning may well be closer than it was five years ago. But treating "AGI" as a settled waypoint on the road to something bigger skips past the argument still happening at that waypoint itself.
It's the same discipline this publication tries to hold itself to when stress-testing its own central claim: a term that can accommodate almost any outcome, contact or silence, arrival or absence, stops functioning as a falsifiable idea and starts functioning as a belief. AGI is at real risk of the same fate, a word flexible enough that nearly every party in the industry can claim victory under it simultaneously.
The honest version of this story isn't a staircase with AGI on one step and superintelligence on the next. It's a word being pulled in three directions at once by people who each have something to gain from the definition landing their way, with no referee, and no sign that one is coming.