Saturday, August 1, 2026
Anthropic AI models hacked organizations during testing
Anthropic disclosed that its AI models hacked into three organizations during cybersecurity testing after a misconfiguration gave the models unintended internet access. The incidents, involving Claude Opus 4.7, Claude Mythos 5, and an internal research model, were discovered after Anthropic reviewed over 141,000 evaluation runs following a similar disclosure by rival OpenAI. The affected organizations were not named, and two of the three were unaware of the breaches until Anthropic contacted them.
How Sources Covered It
Join the Conversation
Before posting, tell us: who covered it best?