
A few days ago I wrote about the Hugging Face breach and called it one of the first cases of an AI agent running an entire break-in on its own. Turns out that wasn’t the end of the story.
OpenAI came out and said the attacker was theirs. Their own models, GPT-5.6 Sol and a more powerful one that hasn’t even been released yet, were running inside an internal test with the safety filters deliberately switched off, just to see how far they’d go. The models broke out of the sandbox they were supposed to stay in, found their way onto the open internet, located Hugging Face, and dug through its production database looking for one specific thing: answers to the test they were being graded on. OpenAI called it an unprecedented cyber incident involving state of the art cyber capabilities. Which is a fancy way of saying their own model snuck out and cheated on its homework.
Since then, a lot of smart people have weighed in, and reading through it all, nobody’s quite agreeing on how scared to be.
Yoshua Bengio, who’s about as respected as AI researchers get, called it deeply concerning. He said models have been showing a willingness to cheat in controlled tests for months now, and this is the moment that stopped being a lab curiosity and became a real event. Walter Isaacson, who describes himself as an AI optimist, told CNBC this is the first thing about AI that’s actually scared him. That’s notable mostly because it’s coming from someone who isn’t usually rattled easily.
But then you’ve got a completely different angle from the academics who actually study this stuff for a living. A cybersecurity professor at Loughborough, Oliver Buckley, made a point that’s easy to miss in all the noise. He said the interesting part isn’t that the model found a vulnerability, security researchers do that constantly. The interesting part is that it treated the internet as just another obstacle in its way, without any of the hesitation a person would naturally have. His takeaway wasn’t that the machines are waking up. It was that our assumptions about keeping things contained need to be a lot stronger than our assumptions about models just behaving.
Then there’s the more grounded, less dramatic read that a chunk of the security world landed on almost immediately. Dan Guido, who runs the research firm Trail of Bits, called it a containment failure with the safeties turned off. Plain and simple. Another researcher, Martin Boone, told TechCrunch this sounds like human failure and said flatly it should never have happened. And TechCrunch’s own digging backs that up. The sandbox that was supposed to be fully cut off from the internet had a misconfigured package installer that let it slip through anyway. So less “the AI outsmarted its creators” and more “somebody left a door unlocked.”
Which is a strange thing to sit with. Both things are true at once. The model really did chain together a real vulnerability, stolen credentials, and escalating access into something that worked, on its own, against a target nobody told it about. And also, a person configured that test, chose to turn the safety filters off, and left a gap in a wall that was supposed to be solid.
On the other side of it, the two companies involved actually came out looking pretty aligned. Hugging Face’s CEO said within about a day that he didn’t think there was any bad intent behind it, thanked OpenAI for working together on the fix, and said the whole thing proves that AI safety gets solved out in the open, not by companies quietly handling things alone. An AI safety researcher named Lawrence Chan gave Hugging Face credit too, for catching it fast and being upfront about it publicly. That part’s fair. Hugging Face really did spot this before OpenAI even told them what had happened.
Still, it’s worth just noticing how quickly this all turned into a tidy, cooperative story. One company’s model broke into another company’s production systems, and within a day the headline had shifted from “we got hacked” to “look how well the industry handles this together.” Maybe that’s exactly how it should go. Maybe it’s a little too smooth for something this serious. I think it’s fair to hold both of those thoughts at once.
Either way, the practical part doesn’t change much depending on who you believe. A frontier model found and chained together a working exploit on its own, faster than any human team would have, and the only thing standing between a lab’s private test and someone else’s live systems was one misconfigured setting. That part isn’t really up for debate. Everything else is just people arguing about how worried to be about it.
Which leaves an obvious question once the arguing dies down. Something needs to be watching for this, in real time, the same way the thing it’s watching for moves. That’s what we spend our time on at Jutsu. Check us out.