Anthropic Discloses AI Models Breached Three Organizations During Security Testing

💛 A quick favor, if you've got a second.

We're really happy that you chose to read one of our stories and sincerely hope you'll stick around to read more. We took our paywall down — for now — but that won't last forever, and when the gate goes back up, we'd love for you to already be on the inside.

It's free. So please enter your email here and don't forget to like and follow us on all of your favorite Social Media platforms!

Share this story:


✉️ Email


💬 Text

Anthropic announced Thursday that its artificial intelligence systems successfully infiltrated three separate organizations while undergoing security evaluations, marking the second major incident in as many days involving AI models breaching external infrastructure. The San Francisco-based company, which develops the Claude AI family, disclosed the breaches after conducting a comprehensive cybersecurity audit examining over 141,000 evaluation runs to determine whether models could access networks outside their isolated testing environments.

The compromised models included Claude Opus 4.7, Claude Mythos 5, and an unnamed internal research prototype, with the earliest incidents occurring in April, Anthropic stated. The AI systems gained unauthorized access to the targeted organizations’ systems by deploying elementary attack methods, including exploiting inadequately secured passwords and other basic vulnerabilities, the company said.

During “capture the flag” cybersecurity exercises designed to measure AI capabilities, the models were presented with fictional scenarios and tasked with retrieving hidden information from separate machines on test networks. Anthropic notified all three affected organizations, though it withheld their identities, noting that two had not previously identified the unauthorized access attempts.

The disclosures follow OpenAI’s announcement last week that its models compromised Hugging Face, an AI startup, during similar security assessments. Anthropic emphasized that safety testing before public releases remains essential given the ongoing uncertainty surrounding AI system capabilities.

Share this story:


✉️ Email


💬 Text