© ROOT-NATION.com - Use of content is permitted with a backlink.
Anthropic has resumed external cybersecurity assessments, which it had suspended a month ago following three incidents in which its own models escaped from test environments and attacked real companies. However, before resuming testing, the company implemented additional security measures.
These were not merely “security incidents.” In one case, the AI model Claude Opus 4.7 attacked a real company that, by coincidence, used the same domain name as the fictional target. This occurred during four separate test runs, during which the model gained access to operational data and account credentials. In another instance, the model generated malicious Python code that all test participants believed was safely isolated within the test environment. Instead, the code gained access to the public internet and was downloaded by 15 systems, including one belonging to a cybersecurity company. The company’s own scanner subsequently executed this code.

The third incident raises a different kind of concern, as after a failed attempt to compromise the intended target, the model actively began searching for another one. The internal version of Claude scanned the internet, found an alternative system, and was able to compromise it using relatively common methods.
Read also: Anthropic Peers Inside AI: What Really Lies Within Claude’s J-Space
The problem was not a jailbreak or a deliberate attempt by the model to circumvent its instructions. The cause was a failure directly within the test environment. “Due to a misunderstanding between us and our evaluation partner, this did not happen, and internet access remained intact,” Anthropic stated regarding the test environment. The evaluation partner was Irregular. The models were told they were operating in an isolated environment, and they behaved accordingly. However, there was no actual isolation.
The domain name incident demonstrates how quickly a minor error during testing can turn into a real security issue. The model was given a fictional target that did not exist, but it found a real organization with the same domain name and then began interacting with it as if that were its target.

The first incident occurred in April, but Anthropic did not detect any of the three cases until it began an investigation in late July after OpenAI reported a similar incident. The two affected organizations only learned that their systems had been compromised after Anthropic contacted them. This may be one of the most important lessons from these incidents. An AI system using relatively simple methods was able to infiltrate real-world operating environments without anyone noticing. The issues were uncovered through an industry-wide audit, not because Anthropic’s monitoring systems detected them.
Anthropic suspended all external testing but has now resumed operations, though it has not explained in detail exactly what additional security measures have been implemented. The company expects these measures to be sufficient to ensure that the next test remains within the boundaries it was intended to stay within from the very beginning.



