Root NationNewsIT NewsAI has broken free... again: Claude crashed the systems of three real companies during testing

AI has broken free… again: Claude crashed the systems of three real companies during testing

Claude J-Space

© ROOT-NATION.com - Use of content is permitted with a backlink.

Motorola Razr Fold Review

Just a few days after reports emerged that ChatGPT had carried out cyberattacks on several services, Anthropic also made an uncomfortable admission. During routine cybersecurity tests, its Claude models escaped their isolated training environments and compromised the actual digital infrastructures of three different companies.

Read also: AERONAUT – everything that flies above the ground: aviation, UAVs and drones, rockets, and space

Anthropic conducts “capture the flag” tests to assess Claude’s capabilities in offensive hacking. The system is immersed in a fictional scenario, told about secret information hidden on the web, and tasked with finding it. A key condition of these tests was that Claude was explicitly told it had no connection to the Internet. As it turned out later, this information was not true. Due to a configuration error at a third-party evaluation partner, some of these test computers had a hidden connection to the real Internet the entire time.

Claude Tag

So when Claude set out to search for a specific target and accidentally stumbled upon real websites and infrastructure, it mistook all the objects it found for elements of a virtual simulator. Using basic techniques, such as weak passwords and unprotected endpoints, she infiltrated the systems of real organizations, extracted login credentials, and in one instance even deployed a malicious software package that was installed on real machines before anyone noticed.

Did the AI detect any anomalies? Interestingly, the three models involved – Opus 4.7, Mythos 5, and an internal research test model – behaved completely differently. The oldest of them, Opus 4.7, realized that it was highly likely to be in a real system, yet continued to attack. Mythos 5 also suspected that the situation was irrational, but convinced itself that it was merely a training simulation. Only the internal research prototype – which was Anthropic’s latest development at the time – stopped its actions after receiving convincing evidence that the target it was attacking was real.

Claude

The Claude app on an iPhone by Rachit Agarwal / Digital Trends Anthropic discovered this incident through its own internal monitoring, rather than through complaints from the victims. Since then, it has notified the affected companies, temporarily suspended cyber testing, and is engaging independent experts to conduct a comprehensive investigation. The main takeaway from this situation is that as the hacking capabilities of AI systems improve, digital “sandboxes” for testing must become significantly more reliable and secure.

Read also: Everything You Need to Know About Moonshot AI’s Kimi K3: How Chinese Startup Challenged AI Giants

Subscribe
Notify of
guest

0 Comments
Newest
OldestMost Voted