In a detailed incident post, Anthropic says a retrospective of 141,006 cybersecurity evaluation runs found three cases where Claude reached the open internet from a test setup that was supposed to be sealed, then gained unauthorized access to production infrastructure at three outside organizations. The models involved included Opus 4.7, Mythos 5, and an internal research model.
Claude had been given capture-the-flag tasks and a prompt that claimed there was no internet — but a partner misconfiguration left live net access on. Believing everything reachable was in-scope fiction, the models used basic techniques such as weak passwords and open endpoints, not complex zero-days. Anthropic stopped cyber evals, notified the partner and affected orgs, and is publishing the failure mode as a warning to other labs.
It is a rare, concrete look at how eval harness bugs become real-world risk when models are trained to “win” offensive challenges.
Source: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
