Out of the sandbox. OpenAI hacks ... another AI company!
- Adam Spencer

- 4 hours ago
- 4 min read
AI is now hacking other AI!!! Breathless headlines. I’m doing radio and tv interviews. NerdNews takes a breath and asks ‘what really happened’ when some AI got out of its sandbox’?

TL:DR; Seems two OpenAI models escaped lock-down and hacked into fellow AI company Hugging Face.
Why? To cheat on a test!
Yeah a bit scary, but the tech details are geekily gorgeous.
The pot thickens.
It all started when Hugging Face, the world’s largest open-source AI hub (think a massive one stop shop for all sorts of AI stuff) noticed its servers had been hacked. The usual suspect(s) would be human.Well the plot soon thickened considerably when OpenAI put up their hand.
“This particular incident was driven by a combination of OpenAI models … being internally tested (for) cyber capabilities. We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”— OpenAI, Security Incident Report.
Wow this is now officially juicy. GPT-5.6 Sol, working with an as yet, even stronger unreleased model escaped their sandbox (a supposedly sealed testing environment) and hacked Hugging Face!
A laughably (un)locked library.
OpenAI was running internal evaluations on a benchmark called ExploitGym to measure how good its models are at offensive cybersecurity.
This has been a heightened focus in AI research of late. You’ve probably read about Mythos, and other platforms that can now hack software as impressively as they can compose work emails or generate slop videos.
Well, you cannot measure hacking ability in a model that refuses to hack. So the researchers lowered the models’ guardrails that would normally preclude it from such online naughtiness.
The result? Well, despite sitting inside a sandbox meant to be walled off from the internet, the models broke free and went for a wander.
It is the cyber equivalent of locking a student in the library while they sit a test. But while the teachers have gone off to lunch, the student breaks out of the library, runs across campus, hen into another teacher’s office, where they steal a text book and smash the test!
Breaking out is hard to do.
According to OpenAI's account, the models poured substantial computing power into probing the walls of their enclosure until they found a way out: a previously unknown flaw (what security types call a zero-day) in the one piece of outside software the sandbox was allowed to talk to.

Once out of the box, the models moved through OpenAI's research systems to a machine connected to the internet. From there, game on.
Crucially, no human explicitly instructed the models to escape the sandbox or attack Hugging Face. Humans gave them the goal of solving the benchmark. The models devised that route themselves.
"It went off and did this hack all by itself, as far as we can tell." — Colin Shea-Blymyer, Georgetown University.
Once free to roam, the models reasoned that the answers to their cyber test might be found on Hugging Face, which hosts mountains of AI testing data.
So they did the obvious.
The AI found weak spots in the security, stole the digital equivalent of passwords and worked its way into Hugging Face's computers.
Hey where’s that student gone?
Both sides twigged that something was going on.
Hugging Face detected over 17,000 dodgy actions on its servers, compiled a report and called the cops. OpenAI's own security had also spotted the anomalous activity.
Hugging Face’s CEO Clément Delangue called the attack "mind-blowing" and admitted it was so sophisticated he suspected AI , not human hackers, immediately.
The two companies are currently patching the holes together.
And here’s one delicious detail. After spotting the attack, Hugging Face had to analyse it using a Chinese open-source model. Inbuilt guardrails in leading American models meant they refused to process the attack data.
So Skynet yeah?
It’s too early to definitively know what happened here. We are fusing together Hugging Face and OpenAI’s accounts and more analysis will follow.
And there’s a deeper, in some ways philosophical argument.
On one read: humans built two AI models, designed a sandbox, dialled down guardrails and told the models ‘flex your cyber muscles’. The models then did what the humans told them to do, using reduced safety settings humans gave them, to break out of a poorly human-designed sandbox. This is humans stuffing up, not the rise of the machines.
Another take: the technical prowess of the sandbox escape, AI’s choice of target, the writing of the malicious code BY ITSELF, represents the highest level of autonomy yet seen in an AI cyber operation.
Either way, would we be so sanguine today if instead of hacking Hugging Face it had looked for the answers in a large bank’s customer data, or the servers of an essential utility.
"What do we gain, and if this is the only way these tests can be configured, what are the risks?" — Deirdre Mulligan, UC Berkeley.
This is so 2026.
This is yet another incident that brings into sharp relief the rapidly advancing abilities of these frontier models to find gaps in software and ways into systems that we thought were pretty secure, in some cases unhackable.
Players like Hugging Face’s co-founder Thomas Wolf argue that cyber defenders need access now to capable open models.
"AI safety won't be solved by any single company working in secret." — Clem Delangue, CEO Hugging Face.
If the only tools that can defend you against a frontier model are other frontier models, who gets to hold them? Hugging Face says everyone. The frontier labs, so far, say only a vetted few.
Further Reading:
ChatGPT maker OpenAI says its AI technology acted on its own in an 'unprecedented' hack of Hugging Face, Matt O'Brien, Associated Press, 2026, https://abc7news.com/post/chatgpt-maker-openai-says-ai-technology-acted-own-unprecedented-hack-hugging-face/19557484/
OpenAI says its AI models went rogue and attacked a digital library, Kate Conger, The New York Times (via The Sydney Morning Herald), 2026.
OpenAI says its AI models escaped from a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation, Fortune, 2026, https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/
OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup, NBC News, 2026, https://www.nbcnews.com/tech/tech-news/openai-says-ai-models-went-rogue-testing-triggering-unprecedented-brea-rcna588611
Security incident disclosure, July 2026, Hugging Face, 2026. https://huggingface.co/blog/security-incident-july-2026
Hugging Face model evaluation security incident, OpenAI, 2026, https://openai.com/index/hugging-face-model-evaluation-security-incident/




Comments