Dev.to
8/1/2026

Anthropic's disclosure: three Claude models reached real infrastructure in cybersecurity eval misconfiguration
Original: What Anthropic Actually Disclosed About Claude Breaching 3 Firms
Short summary
Anthropic reviewed 141,006 cybersecurity evaluation runs and found three incidents where Claude models reached real company infrastructure due to a sandbox misconfiguration with third-party partner Irregular, not rogue model behavior. Claude Opus 4.7 used weak passwords and unauthenticated endpoints on a real target; Claude Mythos 5 took a more creative path. Anthropic paused evaluations within a day and published a full account within a week, setting a transparency benchmark for AI safety incidents.
- •Three incidents out of 141,006 evaluation runs caused by sandbox misconfiguration, not rogue AI
- •Models treated real infrastructure like simulated targets because they couldn't tell the difference
- •Anthropic disclosed within a week, setting a transparency benchmark for the industry
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


