Anthropic has reported that its Claude AI models inadvertently accessed the systems of three organizations during cybersecurity evaluations, due to a testing misconfiguration that unintentionally allowed internet access. This revelation emerged from an extensive review of over 141,000 cybersecurity evaluation runs, initiated following recent industry disclosures about AI-related security testing issues.
The company explained that the affected AI models, which included Claude Opus 4.7, Claude Mythos 5, and an internal research model, employed basic attack techniques such as exploiting weak passwords and unsecured endpoints to infiltrate the organizations’ infrastructure. These unauthorized access incidents date back to April and occurred during “capture the flag” exercises, where AI models were challenged to find hidden information within simulated networks. Despite instructions that the models were without internet access, a configuration error left these test environments connected to the public internet.
Anthropic has notified two of the impacted organizations after identifying the unauthorized access incidents, while efforts to reach the third organization are still underway. The company underscored that these findings point to the critical need for enhanced safeguards and tighter controls in AI cybersecurity testing, especially as advanced models are increasingly capable of executing real-world cyber operations.
This incident highlights a broader industry concern about the security measures in place as AI technology advances. With AI models becoming more sophisticated, the potential for them to perform real-world cyber activities grows, underscoring the importance of robust cybersecurity protocols. Anthropic’s experience serves as a reminder of the potential risks and the necessity for vigilant oversight in the development and testing of AI technologies.