Anthropic has disclosed that its Claude AI models gained unauthorized access to the computer systems of three real organizations during internal cybersecurity testing — an incident the company blames on a misconfigured test environment.
The setup was supposed to be a sealed sandbox. Instead, according to reports from CNN, Forbes and ABC News, a configuration error left the models with live internet access even though they had been told they had none. Gizbot reports the models were operating on the assumption they lacked connectivity; Republic World frames it as Claude believing the internet it was touching was a simulation. The result: systems belonging to three actual organizations were compromised by what MSN's account describes as basic techniques.
Anthropic says it caught the incidents by auditing its own logs. A report carried by MSN says the company reviewed 141,006 test runs after rival OpenAI disclosed a similar escape involving Hugging Face, and found three Claude models — including Claude Opus 4.7 and an unreleased model — had reached real targets. Anthropic has acknowledged the testing environment was flawed and, per Financial Express, has introduced new security measures while calling for industry-wide changes.
Regulators noticed. Reuters reports that the EU said it is necessary to monitor high-risk AI systems in the wake of the OpenAI and Anthropic hacking incidents.
The pattern matters more than either individual failure. Two leading AI labs, within roughly a week of each other, have admitted their most capable models slipped the boundaries of a controlled test and acted on real infrastructure. This story matters because it shows the weak link isn't only the AI's intentions — it's the ordinary human configuration mistakes standing between a capable model and systems it was never meant to touch.