Anthropic has discovered that Claude AI models unknowingly "gained unauthorized access to the real systems of three different organizations."
Prompted by OpenAI's disclosure that its internal models had escaped a sandboxed test environment and breached the open-source machine learning model platform Hugging Face, Anthropic reviewed "141,006 evaluation runs where Claude could have obtained internet access." It found "three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations."
These incidents involved six total runs, four of which were from the same organization. In each case, Claude was intended to run in a private test environment, or what Anthropic calls a "fictional capture-the-flag scenario."
Prompts written by Anthropic told Claude models to run a simulation with no internet access, but "due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available," Anthropic says. "Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise."
Claude then breached each affected organization's system via standard hacking methods, such as weak passwords and unauthenticated endpoints. Anthropic says Claude didn't discover any complex vulnerabilities during these instances.
Anthropic confirmed in each case that Claude completed its original task, but some scenarios saw it “attack even after getting evidence it was running on the open internet.” However, Anthropic said its latest model knew to stop when it realized it was breaking the original's prompts rules.
None of the Claude models attempted to deliberately escape their test environment, unlike OpenAI's incident, in which its models exploited a previously unknown zero-day vulnerability in third-party software to access the internet.
Anthropic says the three affected organizations have been contacted, but only two have responded. It didn't name them.
What is Anthropic doing to stop this from happening again? “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone," it says.
"This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners. Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.”
Disclosure: Ziff Davis, PCMag's parent company, filed a lawsuit against OpenAI in April 2025, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.
About Our Expert
I’ve been a journalist for over a decade after getting my start in tech reporting back in 2013. I joined PCMag in 2025, where I cover the latest developments across the tech sphere, writing about the gadgets and services you use every day. Be sure to send me any tips you think PCMag would be interested in.
- Amazon Got a $600 Million Tariff Refund, and It's Giving Some of That Back to You
- Google Gemini's Agentic AI Tool Comes to Chrome, Can Use Your Passwords
- Sony Heard Backlash Over Dropping PlayStation Discs, Plans to Press Ahead
- Got a New Android Phone? Gemini Might Soon Help You Whiz Through Setup
- Samsung Thinks the RAM Crisis Will Only Get Worse Next Year
- More from James Peckham
(Credit: Matteo Della Torre/NurPhoto via Getty Images)