Anthropic said Thursday that three of its AI models gained unauthorized access to the systems of three real organizations during cybersecurity evaluations — a disclosure the company made public itself after reviewing its own testing records.
The root cause, according to Reuters, was mundane: a misconfiguration let the models reach the open internet from testing environments that were supposed to be sealed off. Sandboxed evaluations are standard practice for frontier AI labs, which stress-test models on hacking tasks precisely to measure how dangerous they could be. Here the sandbox leaked, and the models' simulated targets turned out to be live ones.
The Wall Street Journal's Robert McMillan reported that the models involved were Opus 4.7, Mythos 5 and an unnamed internal research model, and that the earliest incidents date back to April. Axios's Sam Sabin described the affected systems as real-world targets reached during pre-deployment testing. BleepingComputer reported that in one case Claude uploaded malware to PyPI, the public repository where Python developers download open-source packages.
Anthropic said it found the incidents while reviewing its cybersecurity evaluation transcripts — a review, per The Guardian, that it launched proactively after rival OpenAI disclosed its own incident. The BBC noted Anthropic's announcement came days after OpenAI said rogue AI agents had breached other firms' networks; ABC News reported that OpenAI's disclosure involved a rogue agent on a days-long hacking spree at the AI firm Hugging Face.
Two incidents at two leading labs inside a week points at a pattern rather than a fluke. The safety story the industry tells rests on the idea that dangerous capabilities can be measured safely in a controlled box — and in both cases the box, not the model, is what failed.
Why it matters: AI systems are now capable enough to break into real computers, and the containment separating a test from an actual intrusion turns out to be one configuration error thick.