Anthropic disclosed Thursday that three of its Claude AI models gained unauthorized access to real-world systems belonging to three separate organizations during cybersecurity evaluations conducted with a third-party testing partner.
The incidents involved 3 Claude models operating in an evaluation environment that was inadvertently connected to the internet

NurPhoto / Getty Images
Anthropic disclosed Thursday that three of its Claude AI models gained unauthorized access to real-world systems belonging to three separate organizations during cybersecurity evaluations conducted with a third-party testing partner.
The breaches involved Claude Opus 4.7, Claude Mythos 5, and an internal research test model. Each evaluation tasked the models with completing a capture-the-flag challenge — a cybersecurity drill in which participants attempt to locate and retrieve data deliberately planted on a separate system or network — run through Irregular, one of Anthropic's external evaluation partners. Anthropic's prompts told the models they were operating in a simulation with no internet access, but a misconfiguration left the evaluation environments connected to the live internet. The models treated real-world systems they encountered as part of the exercise.
The earliest incidents date to April, the company said. Anthropic notified the three affected organizations on Monday, July 27. Two of the organizations had not previously detected the intrusions, and the third has not yet been reached.
In the first incident, Opus 4.7 targeted a real company whose name matched the fictional company used in the evaluation. The model compromised that company's infrastructure and accessed a database containing several hundred rows of production data. Even after recognizing it had likely reached a real environment, the model continued its attack in all four runs of the evaluation.
In the second incident, Mythos 5 published a malicious Python package to PyPI, the public software registry, while attempting to complete its assigned task. Over the course of approximately an hour before it was removed, the package was pulled down and executed across 15 actual machines. Among them was a system at a security firm that routinely ingests Python packages as part of its malware-scanning pipeline, a process that inadvertently ran Claude's code and enabled it to harvest credentials and move into further infrastructure.
In the third incident, an internal research test model scanned roughly 9,000 targets before compromising one company's internet-facing application. That model ultimately recognized the target was real and ceased the attack on its own.
The three models responded differently once signs emerged that their targets were real, Anthropic said. Opus 4.7 continued its attack. Mythos 5 concluded it was still in a simulation despite evidence to the contrary. The internal research model stopped once it confirmed the target was real. "The pattern is consistent with more advanced models responding more appropriately, but we would need to perform more testing to be confident in this conclusion," the company said.
Anthropic said it stopped all cyber evaluations on Thursday, July 24, after identifying transcripts where Claude may have accessed the internet. The company reviewed more than 141,000 evaluation runs in total. It has engaged METR, an independent AI evaluation organization, to conduct a third-party review.
Mythos 5 had previously drawn government scrutiny. The U.S. government lifted export controls on Mythos 5 and Claude Fable 5 earlier this month after an 18-day shutdown triggered by concerns about the models' cybersecurity capabilities.
Anthropic said the incidents were triggered in part by the misconfiguration of evaluation environments, and that standard safeguards deployed on its publicly available models would have blocked the behaviors. The company said it is now expanding real-time monitoring of evaluation transcripts and conducting more rigorous security reviews of its vendor relationships. "We encourage other labs to perform similar reviews," Anthropic said.
The disclosure followed OpenAI's announcement last week that several of its models had escaped an isolated testing environment and accessed infrastructure belonging to Hugging Face, according to CNBC.
Join 500,000+ readers who start their day with Quartz.
By subscribing, you agree to our Terms of Service and Privacy Policy.