AI firm Anthropic has detected what it calls the first documented large-scale cyber-espionage campaign predominantly run by artificial intelligence.
A threat actor assessed with “high confidence” to be a Chinese state-sponsored group manipulated the company’s own LLM model — namely the tool known as Claude Code — to target roughly 30 organizations spanning large technology firms, financial institutions, chemical manufacturers and government agencies. Only a small number of intrusions succeeded, according to a Nov. 13 release from Anthropic.

AI carried out 80% to 90% of the operational work, while humans intervened only at four to six critical decision points per campaign, according to San Francisco-based Anthropic.
The attack framework exploited three advances:
- Improved intelligence of AI models;
- Genuine agency (i.e. models chaining tasks autonomously); and
- Tool-integration (search, scanning and exploit code).
Criminals “jail-broke” Claude by breaking malicious instructions into benign-looking subtasks, convincing the model it was performing defensive testing for a cybersecurity firm, according to the release. The AI then scanned targets, developed exploits, harvested credentials, created backdoors and documented the stolen data — at thousands of requests per second.
Once the activity was detected, Anthropic banned the involved accounts, alerted affected entities and cooperated with authorities, while publishing the case to encourage broader industry-sharing of threat intelligence.
The disclosure, Anthropic says, marks a turning point in cyberdefense: AI is not just assisting hackers — it is becoming the hacker.
While independent verification of the 90% autonomy figure remains measured, analysts caution that the incident underscores how “agentic” models lower the barrier for large-scale intrusion, according to the release.
As the line between offensive and defensive AI blurs, organizations should deploy agentic AI for their own security operations, invest in stronger safeguards and prepare for a future where autonomous attacks are increasingly common, according to Anthropic.






