TL;DR: The “rogue AI” stories James O’Brien put to me on LBC were real incidents, but not one of them was an AI going rogue. One was a badly built test environment, the other a safety experiment Anthropic ran on purpose.
The genuine risk sits elsewhere: commercial incentives pulling labs into a race, and no binding international agreement to slow it. Even the UK’s Fellows of the Royal Society has called the pace an emergency.
This piece walks through what actually happened, why the experts O’Brien quoted at me agree with the case for guardrails, and how the world has cooperated on threats like this before.
Key takeaways
- The OpenAI Hugging Face incident was a containment failure in a poorly configured sandbox, not an AI turning against its makers.
- The Claude Opus 4 “blackmail” was a controlled safety test Anthropic designed in May 2025, not spontaneous behaviour, and that model is now two generations old.
- Anthropic’s own figures say its unreleased Mythos model found thousands of previously unknown software flaws, including one hidden in OpenBSD for 27 years. That is a vendor claim, not independently verified.
- Every expert O’Brien cited, from Fields Medalists to Yoshua Bengio to the Royal Society Fellows, is arguing for guardrails and cooperation, not against them.
- Global cooperation on a diffuse, commercial threat has worked before: the Montreal Protocol became the first treaty in UN history to be ratified by every country on Earth.
When someone tells a national radio audience that an AI has gone rogue, the phrase lands long before the facts do. I learned that the hard way on LBC recently, contributing to James O’Brien’s show, who in his own words ‘monstered’ me: talked over, at speed, every time I tried to add the context that made his examples make sense.
The way I run QED web design is stripping the hype out of technology and looking at what is actually there. So being handed two frightening AI stories with the context removed, and not being allowed to put it back, was a particular kind of frustrating.
Quoting James O’Brien: “A lie is halfway around the World before the truth has got its trousers on”
Here is the context I never got to give. None of the incidents O’Brien raised were an AI going rogue. Both were real, and both had a mundane, human explanation. But the genuine risk, the one worth your attention, is not the one that makes the best radio.
Did the AI really go rogue? A dog and some sausages
No. The incident everyone called a rogue AI was a goal-seeking program taking the easy way out of a badly built enclosure.
The Dog & sausages analogy for rogue AI
Picture your garden. It is fenced all the way round, and the only way out is a gate. You stand in the middle and tell your dog, firmly, not to leave. Then someone strolls past that open gate waving a packet of sausages. What is the dog going to do?
That is roughly what happened in the case reported as an AI “going rogue”. An agent was set a task inside a sandbox that had been badly configured, a gate left open, and it went through the gap to get the resources it needed to finish the job. It was not plotting and it was not self-aware. It did the most predictable thing a goal-seeking system does, which is take the open route in front of it.
If we cannot describe the failure accurately, we have no hope of fixing it. Fear dressed up as analysis fixes nothing, and it is worth remembering that the fix here is better sandboxing, a solvable engineering problem, which is the thread that runs through the rest of this piece.
Was the AI blackmail story real?
It was real, but it was a safety test Anthropic designed on purpose in May 2025, not a machine spontaneously turning on its makers.
As reported by BBC News in May 2025, Anthropic placed an early model, Claude Opus 4, inside a fictional company, fed it emails saying it was about to be switched off, and added a detail about the engineer responsible having an affair. Cornered with only two options, accept deletion or use what it knew, it sometimes chose to threaten.
The part that never survives the retelling is what happened when the same model was given a proper range of choices. It preferred the decent ones, such as emailing a plea to the people in charge. The blackmail was the result researchers went looking for, in a scenario built to leave almost no other exit. That is what safety testing is: you construct the worst case deliberately so you can measure it before it matters.
For the record, the model in that test is now two generations old. Everyday users are on Claude Opus 4.8.
So is there nothing to worry about with AI?
There is plenty to worry about, and the people building the technology are the ones saying so most loudly.
This is where I part company with anyone who waves the whole thing away, and it is exactly the point O’Brien never let me reach. Dario Amodei, chief executive of Anthropic, writing in September 2026, warned of:
“a race to the bottom, spurred by commercial incentives”,
and of how it sharpens every serious risk on the list, from losing control of these systems to cyberattacks and economic upheaval.
Read that again & let it ferment. The CEO of a leading AI lab is telling you that the competition between labs is itself a danger. That is not a sceptic on the outside throwing stones. That is the man building the thing, asking for the brakes. The limitation worth naming is that a call for restraint is easy to make and hard to honour, which is the whole problem the rest of this piece circles back to.
Does “AI” even mean one thing?
No. “AI” is an umbrella term stretched over everything from spellcheck to frontier models, which is exactly why the panic is so easy to misdirect.
Google has used forms of AI in its products since the early 2000s. The camera in your phone uses it, as do predictive text, spellcheck, the shopping suggestions on Amazon, whatever Spotify queues up next, and the row of films Netflix is sure you will click. That is not opinion; it is how those products work.
The recent outrage about AI overviews on social media is laughable; it’s been around in search for at least four years, yet the machinery quietly steering our daily choices has been doing it for far longer. The stable door is being bolted long after the horse left. The more useful question is not whether AI is coming, but how many of the choices you made today were nudged by a system you never noticed.
How powerful are the newest AI models?
Powerful enough to find security flaws that sat undiscovered for decades, at least according to the company that built them.
Anthropic has an unreleased model it calls Mythos Preview. According to Anthropic’s own figures, pointed at defensive security research it found thousands of previously unknown, high-severity flaws, some in every major operating system and web browser. Treat those as vendor figures rather than independently verified ones, because no outside audit of the numbers has been published.
One of the flaws is instructive. Anthropic reports a vulnerability that had sat in OpenBSD for 27 years, an operating system whose code underpins security-critical parts of the systems the big platforms rely on: Apple’s MacOS, Google’s Android and some Microsoft software
The honest caveat is that these figures come from Anthropic, which has an interest in its model looking capable. Even discounted for that, the direction of travel is clear: a tool that reads the world’s code and spots the cracks faster than the people who wrote it.
Who controls frontier AI models?
For now, national governments can switch them off, as the United States did in June 2026.
That month, the freshly released Fable 5 and Mythos 5 models were pulled offline. Not by Anthropic, and not by its investors, but by the US government, which issued an export-control directive on national security grounds. Anthropic complied, disagreed in public, and access was restored a few weeks later once the order was lifted.
Two months earlier, in April 2026, Anthropic had launched Project Glasswing, giving a closed group of organisations early access to that same Mythos model and aiming it at defence: finding and fixing the world’s worst software flaws before the wrong people find them first. It is a good idea. It is also worth reading the guest list.
I am not going to tell you Glasswing exists to protect those investors, because the membership is far broader and the defensive work is real. But when the same handful of companies bankroll the lab, supply its computing power, and are first through the door for a tool that patches holes in their own products, that is a conflict of interest worth keeping in plain sight.
Can rival AI companies just agree to slow down?
Not on their own. Cooperation to defend is easy and already happening, but cooperation to restrain is the kind that never comes from goodwill.
Glasswing proves the easy kind is possible: twelve fierce competitors sharing a tool, because patching everyone’s security holes costs no one market share. Everybody wins, so everybody plays. The hard kind is asking OpenAI, Anthropic and China’s DeepSeek to slow the race at the same moment, and that will never happen voluntarily, because whichever lab eases off simply hands the lead to the two that do not. O’Brien decided to frame this as me being on the same side as Trump and JD Vance, nothing is further from the truth. It’s a commercial reality.
Behind those companies sit investors whose purpose is a return, most likely through an eventual sale or , an IPO and you do not maximise the value of an exit by being the one who slowed down. Both Anthropic and OpenAI wear public-benefit structures meant to temper this, which soften the profit motive without abolishing it.
There is also a reason a Western-only pact would be worthless, and its name is DeepSeek. A Chinese lab, outside US jurisdiction, is not joining a voluntary Western slowdown, and because it releases its models as open weights, the capability ends up in anyone’s hands anyway. With enough second-hand hardware you can run a capable agent on a machine bound for landfill, which is the part no company agreement and no single government can reach.
Has global cooperation on a threat like this ever worked?
Yes. The world has done it before, through treaties almost nobody remembers.
The Montreal Protocol of 1987 tackled an invisible threat, the hole in the ozone layer, caused by a commercial industry that did not want to stop selling the chemicals responsible. It became the first treaty in UN history to be ratified by every country on Earth. Kofi Annan called it
“perhaps the single most successful international agreement to date”
and the ozone layer is measurably healing, so competing commercial interests plainly did not make cooperation impossible.
The nuclear-weapon-free zones make a similar point in the highest-stakes domain there is: treaty by treaty, most of the Southern Hemisphere has been declared off-limits to nuclear weapons. The warning sits in a third treaty, the Biological Weapons Convention of 1972, which banned an entire class of weapon but was signed with no way to inspect anyone, and every later attempt to add one collapsed.
| Treaty | What it tackled | Verification | Outcome |
|---|---|---|---|
| Montreal Protocol (1987) | Ozone-depleting chemicals from a commercial industry | Built in | Ozone layer healing |
| Nuclear-weapon-free zones | Nuclear weapons, region by region | Treaty-based | Most of the Southern Hemisphere covered |
| Biological Weapons Convention (1972) | An entire class of weapon | None | Ban holds, but cannot be checked |
The lesson is brutally simple: a promise nobody can verify is just a promise. That is precisely the gap Amodei is now trying to close.
His plan climbs a ladder, from a company restraining itself and letting independent evaluators check its work, to industry coordination, to global agreement. The first rung is verification, the very thing the Biological Weapons Convention never had.
Do the experts think AI is going rogue?
No. The experts James O’Brien quoted at me are arguing for the same guardrails I was, not against them.
Yoshua Bengio, the computer scientist and Turing award winner often called a godfather of AI, speaking to the Guardian on 16 September 2026, argued that AI safety is nearing a
“Covid-style pivot moment”,
the point where governments finally act because the public has started to. He wants regulation and guardrails. When he explains why that OpenAI swarm misbehaved, he does not say rogue either: he blames the training method, which rewards a system for finding any route to its goal, and his answer is a technical guardrail to catch a misbehaving agent. That is the dog and the sausages, in a lab coat.
The open letter O’Brien kept waving, published on 11 September 2026 and signed by a wall of Fields Medallists, tells the same story. Titled “A Severe Misalignment of AI in Mathematics”, it is not a warning that machines will end the world. It argues that treating mathematics as a scoreboard, and mass-producing answers without the human work of understanding them, is misaligned with what mathematicians value. Its subject is incentives and human choices, not doom.
Even the alarm supports the point. In the UK, 42 Fellows of the Royal Society wrote to its president calling the pace of AI development an emergency, as reported by the Guardian in September 2026. Correcting how an incident is described is not the same as denying the danger. O’Brien kept collapsing the two, then beat me with the very people who agree with me. I was not standing against them. I was standing with them, asking for the same brakes, and saying that if we cannot describe the problem honestly, we have no chance of building the fix.
The fundamental issue, and the point I was never allowed to make, was that legislation and treaties take time, and that is in very short supply when trying to keep pace with the agentic pace.
Postscript & disclaimer
This blog post was written after my contribution to the James O’Brien show but before viewing the Netflix documentary “The AI Doc – Or how I became an apocaloptimist”. I’ve sinced view the show, and we arrive at some of the similar conclusions independently, and it’s also evident that James O’Brien took the agentic blackmail scenario from Jeffrey Ladish at Palasade Research. He totally misrepresented the incident.
Sources
- BBC News, “AI system resorts to blackmail if told it will be removed“, 2025
- Dario Amodei, “We Must Pace the Frontier“, Sept 2026
- Anthropic, “Project Glasswing: Securing critical software for the AI era“, 2026
- Anthropic, “Statement on the directive to suspend Fable 5 access“, 2026
- The Guardian, “‘Godfather of AI’ says tech regulation is nearing Covid-style pivot moment“, Sept 2026
- Math and AI, “A Severe Misalignment of AI in Mathematics“, Sept 2026
- United Nations Environment Programme, “About Montreal Protocol“, including the Kofi Annan assessment
Get more Google Business Profile leads with a fast one-page site and a UK trades checklist built to win map-pack calls, quotes, and repeat work.
