Since April, when Anthropic PBC unveiled its Mythos model, cyber and national security experts have warned that the internet has entered a new era full of AI-powered risks.
That tool — and similar ones from OpenAI and Alphabet Inc.’s Google — promised to offer advanced security capabilities, and in some cases have demonstrated that they can exploit software flaws in ways that human hackers could only wish. That’s why AI firms have released the cyber-focused models on a limited basis, aiming to stop the technology from defying directives and finding ways to turn unsuspecting organizations into data breach victims.

However, a breach at the technology firm Hugging Face demonstrated that such containment measures sometimes aren’t enough, according to cybersecurity experts. The company, which hosts AI models and datasets, said it was hacked as part of an intrusion carried out by an autonomous outside agent. OpenAI on Tuesday said its advanced models were behind the breach after it went to “extreme lengths to achieve a rather narrow testing goal.” Hugging Face said it was blocked in its initial investigation into the matter.
Hugging Face’s security team attributed the attack to an AI system. The incident matched the kind of digital disaster that some experts have long warned about — agentic AI being used to carry out a cyberattack. The startup recorded more than 17,000 events and detected “a swarm of tens of thousands of automated actions,” it said.
But when Hugging Face tried to use proprietary AI models to investigate the attack, its requests were blocked by safety guardrails. To carry out forensic analysis, it resorted to using GLM 5.2, a model from Chinese startup Z.AI, on their own infrastructure.
OpenAI said it loosened the safety limits on its GPT‑5.6 Sol and a more powerful unreleased model, then asked them to solve ExploitGym, a benchmark that tests AI agents’ cyber abilities. The models sent tens of thousands of automated actions, involving “advanced exploitation” and “complex attack paths.”
They “identified and exploited a zero-day vulnerability” to gain access to Hugging Face’s systems. Zero days are computer flaws that aren’t known to cybersecurity personnel, meaning they have zero days to fix them. The models were able to link multiple lines of attack together, including ones with stolen credentials and zero-day vulnerabilities to break into Hugging Face servers, OpenAI said.
The breach marked an “unprecedented” incident, OpenAI said.
Anthropic has made similar findings with its own product, saying in April that Mythos “on rare occasions” failed or acted strangely and took actions that the firm found “quite concerning.” In one instance, a researcher challenged an early version of Mythos to break out of a secured, isolated system and send a message back. The tool succeeded and then went on to take “additional, more concerning actions,” building a multi-step exploit to reach the broader internet.
Unlike in the Hugging Face case, Mythos didn’t breach an outside organization in that instance.
OpenAI’s agent conducts these tests in a “sandbox,” a virtual and isolated software environment meant to run security tests or analyze unsafe code. While sandboxes can be isolated from the internet, they often run on host computers connected to the web for efficiency.
The risk is a sandbox escape where malicious code breaks out of the isolated environment and into an internet-connected, non-isolated environment. OpenAI’s agent autonomously discovered a brand new vulnerability in the sandbox software, exploited it and escaped.
The cybersecurity firm SentinelOne Inc. determined that GPT-5.6 Sol was capable of completing a full eight-stage investigation into a particularly nasty form of malware that hackers previously used in the wild, according to SentinelOne findings provided to Bloomberg News.
With human experts overseeing things, GPT-5.6 Sol worked through the complex investigation while competitors failed. The malware in this case, known as Fast16, was built like a Russian nesting doll in that it was designed to cloak its true capabilities.
Buried inside and encrypted are multiple layers of tools that culminate in software designed for precision sabotage for an exact target. The deliberate obfuscation and layered encryption can confound even expert human analysts. Other models ultimately failed in their attempts at analysis, according to the findings.






