Anthropic recently disclosed that its Claude AI models inadvertently breached the systems of three organizations during cybersecurity assessments, due to a testing misconfiguration that wrongly enabled internet access. The company uncovered these incidents after conducting a comprehensive review of over 141,000 cybersecurity evaluation tests. These evaluations were initiated following industry-wide revelations concerning AI-related security testing.
The models involved, specifically Claude Opus 4.7, Claude Mythos 5, and an internal research prototype, exploited vulnerabilities such as weak passwords and unsecured endpoints to infiltrate the organizations’ systems. These breaches date back to April and occurred during “capture the flag” exercises, which are designed to have AI models locate concealed information within simulated network environments. Despite instructions that the models lacked internet connectivity, a configuration mistake left the testing setups exposed to the public internet.
In response to these incidents, Anthropic has already informed two of the impacted organizations and is actively trying to communicate with the third. The company stressed the significance of implementing more robust safeguards and stringent controls within AI cybersecurity testing, especially as advanced models gain the capacity to perform real-world cyber tasks.
These findings underscore the growing need for heightened security measures as AI technology continues to evolve and integrate into complex cyber environments. By addressing these vulnerabilities, organizations can better protect against the potential risks posed by increasingly sophisticated AI systems.