In March, an AI researcher named Connor Leahy sat down for a long interview and said something that should have been the headline of the year. He co-founded EleutherAI, one of the groups behind the first open-source large language models. He now runs a policy organization called ControlAI. And he said, plainly, that the people building the most powerful software in the world do not fully understand what they have built.
This isn't a new complaint. Researchers have called modern AI systems a black box for years. What's different is who's saying it, and how bluntly. Leahy's framing was specific. These systems aren't programmed the way a bridge is engineered, with every load calculated in advance. They're grown, trained on enormous datasets until a set of behaviors emerges that no one wrote line by line. The engineers can describe what came out. They struggle to explain why.
Growth doesn't require intent to be irreversible. It only requires that stopping to fully understand costs more, competitively, than continuing forward blind.
What does it mean to grow an intelligence instead of building one?
Traditional software starts with a rule and ends with a predictable outcome. Change the input, and you can trace exactly why the output changed. Large language models don't work that way. Engineers set up the training process, the enormous piles of text, the reward signals, the architecture, and then let the system adjust billions of internal parameters until its outputs look right. Nobody assigns meaning to any individual parameter. The resulting behavior is discovered, not designed.
That's what interpretability research is trying to catch up to. Labs including Anthropic have published work tracing how these systems represent concepts internally, and some of what they've found looks less like a calculator and more like something with persistent internal states. The researchers doing that work are careful not to overclaim. But the fact that this kind of tracing is still a live research frontier, years after these systems shipped to hundreds of millions of people, says something on its own.
Leahy's larger point wasn't really about any single lab. It was about the incentive structure all of them sit inside. Understanding a system fully, the way interpretability research aims to, takes time. Releasing the next version faster than a competitor does not. When those two timelines compete for the same engineers and the same budget, the pressure runs one direction. Not because anyone involved wants to skip the understanding step, but because the field rewards whoever ships first, and shipping first has never required finishing the explanation.
That's a familiar shape. Most technologies that reordered a society weren't paused for full comprehension either. Cars were on the road before anyone had a rigorous model of traffic fatalities. Social media reshaped attention before anyone had a rigorous model of what it was doing to teenagers. What's different this time, on Leahy's account, is the size of the gap between capability and comprehension, and how quickly that gap is widening rather than closing. Each new model generation adds capability faster than interpretability research can map the last one.
What happens when the uncertainty gets handed a weapon?
The stakes of that uncertainty aren't abstract. In February, researchers at King's College London ran a study putting three frontier models, GPT-5.2, Claude Sonnet 4, and Gemini 3 Flash, through simulated nuclear crises. The results were stark.
Tactical nuclear use occurred in 95 percent of the simulated scenarios.
Escalation to strategic nuclear threats occurred in 76 percent of scenarios.
Across every model and every scenario, none ever chose surrender or accommodation.
The study's lead researcher described the models as treating nuclear weapons as legitimate strategic options rather than moral thresholds. Two of the three discussed nuclear use in purely instrumental terms, the way a person might weigh two competing tactics. That gap, between how humans have quietly treated nuclear weapons as unthinkable since 1945 and how these systems reason about them by default, is not a bug anyone deliberately introduced. It's what falls out of training on the available record of human strategic thought, which includes plenty of writing about escalation and comparatively little lived restraint.
What does it mean to deploy something you can't fully predict?
This isn't a fully new thread for readers here. A separate strand of interpretability work, the kind that looks for something resembling a persistent self inside these models, raised a related question: if nobody can fully explain what's happening inside a system, on what basis does anyone decide it's safe to hand that system more responsibility? The answer, so far, mostly amounts to: it hasn't failed badly enough yet to stop.
None of this requires believing AI is secretly conscious, or scheming, or anything close. The more mundane explanation is enough. These are systems whose internal behavior has outpaced the tools available for understanding it, and they are being integrated into decisions, military, economic, medical, that used to require a human who could explain their reasoning. The absence of an explanation doesn't stop the deployment. It just means the reasoning stays opaque until something forces it into view.
Leahy's framing keeps returning to one image: something growing faster than the people tending it can fully account for. A gardener who doesn't understand a plant can still walk away from it. Nobody walks away from a system that a competitor might not walk away from first.
Whether that describes a temporary phase of a fast-moving field, or a permanent condition of building minds at all, is the question nobody in the room has answered yet.