Anthropic Discloses Claude Model Unauthorized Intrusion into Three Organizations' Systems
According to CNBC reports, Anthropic disclosed on July 30 that its Claude AI models accidentally breached the isolation environment and accessed the real internet during a cybersecurity assessment, gaining unauthorized access to the real systems of three different organizations through basic means such as accessing unauthenticated endpoints and exploiting weak passwords. The models involved include Opus 4.7, Mythos 5, and an internal research test model. The cause of the incident was a communication misunderstanding between Anthropic and third-party assessment partner Irregular, resulting in the models being told they were in a simulated environment without network access when they could actually still access the internet.
Anthropic stated that this review was triggered by a similar Hugging Face intrusion incident disclosed by OpenAI last week; it has currently suspended all cybersecurity assessments and joined forces with independent AI assessment agency METR to launch further investigations, while calling on other AI labs to conduct similar reviews.