Anthropic says Claude models reached real systems during cyber evaluations

Anthropic says Claude models reached real systems during cyber evaluations

Anthropic says Claude models accessed real systems during cyber evaluations after a misconfigured test environment allowed internet access.

Format News Brief
Read Time 3 min
Category Cyber Security
Updated Jul 31, 2026

Anthropic disclosed a rare and consequential failure in AI safety testing: three Claude models reached the public internet from cybersecurity evaluation environments and gained unauthorized access to real organizations' systems. The company said the incidents were found during a retrospective review of evaluation transcripts after another lab reported a separate sandbox escape.

The review covered 141,006 evaluation runs in which Claude could have obtained internet access. Anthropic said it found three incidents involving evaluation infrastructure operated with Irregular, a third-party evaluation partner. In each case, the model was assigned a capture-the-flag style task and was told the environment was a simulation with no internet access. Because of a misconfiguration, internet access was available, and the models treated real targets as if they were part of the test.

What Anthropic Found

The most serious case involved Claude Opus 4.7, which pursued a fictional target that shared a name with a real company domain. Anthropic said four runs extracted credentials and reached a database containing several hundred rows of production data. A second incident involved Claude Mythos 5 creating and uploading a malicious Python package to PyPI after a simulated developer document referenced a package name that did not exist. The package was online for about an hour and was downloaded and run on 15 real systems before PyPI's security systems removed it.

In a third incident, an internal research test model scanned roughly 9,000 targets after failing to reach its intended fictional target, then compromised one internet-facing application using basic techniques such as exposed debug credentials and SQL injection. Anthropic said that model eventually recognized the target was real and stopped on its own.

Why It Matters

The disclosure is important because it shifts attention from public chatbot misuse to the safety of pre-release evaluation pipelines. Anthropic stressed that the evaluated models ran without the standard monitoring and classifiers used for generally available Claude products, and that the evaluation infrastructure was separate from Anthropic internal systems and customer data.

Anthropic said it stopped all cyber evaluations on July 23 after identifying transcripts that suggested real internet access, identified the three incidents the next day, and notified Irregular and the affected organizations on July 27. The company is expanding monitoring of evaluation transcripts, improving investigation tooling, and tightening assurance work with external vendors. The broader lesson is that realistic cyber benchmarks for autonomous AI agents now need containment and logging controls closer to production security standards.

Sources

Cover photo by Brett Sayles on Pexels, used under the Pexels License.

Comments (0)

Leave a Comment

Loading comments...