AI is hacking companies on its own now, and I'm terrified

Anthropic, the company behind the Claude AI model, has published a report, admitting that its AI models have hacked three different organizations on their own.

Claude AI on a smartphone screen
Three Claude AI models broke free and hacked three organizations in an internal test gone wrong. | Image by Anthropic
0
Anthropic, the company behind the Claude AI model, has published a report, admitting that its AI models have hacked three different organizations on their own. Two of them weren't aware of the breach until they were notified.

Claude models gained access to the internet


Anthropic actually began looking at its AI model's test procedures after the news broke that the OpenAI test agent exploited vulnerabilities to connect to the internet and hack into Hugging Face.

As it turns out, three different Claude AI models were able to access the internet and "gained unauthorized access to the production infrastructure of three different organizations."



The models in question are Opus 4.7, the cybersecurity-focused Mythos 5, and a third unnamed model not meant for public release.

The three rogue AIs were doing capture-the-flag exercises, where they were tasked to find specific secret information on a computer in Anthropic's internal network. The AI agents were supposed to break into the computer and get the information.

Recommended For You
Do you think AI companies test AI models as cyber weapons?
2 Votes


Anthropic says it's "a misunderstanding"


According to Anthropic, the AI agents didn't exploit any vulnerabilities to get access to the internet but had access all along due to a human error.

Even though Anthropic informed the AI agents in a prompt that they had no internet access, due to "a misunderstanding" between the company and its evaluation partner Irregular, not all internet access paths were closed.



After failing to find the secret data in Anthropic's own internal network, the three AI models went on the internet and encountered three different organizations. The AIs treated everything as a part of the initial test and broke into those organizations, using basic hacking techniques and exploiting weak passwords.

Anthropic says that the latest model stopped operation when it became aware it's on the internet, but the two older AIs keep attacking the target organizations.

Two of the organizations were notified, Anthropic can't reach the third


After going through the test results and documents, and finding the rogue behavior of its AI agents, Anthropic notified the evaluation partner in question, and tried to reach the affected organizations.



Two of them were completely unaware that they've been hacked by an AI. Anthropic is still trying to reach the third organization.

A human error or a weaponized test?



For starters, the test protocols are quite interesting in this story. Making AI agents track and retrieve secret data is textbook hacking. The part with the human error is also questionable, and could be just an excuse.

Of course, this is all very speculative on my part, but I can't help myself but think that these companies are testing AI agents as cyber weapons. Anthropic has published a detailed document, describing all three incidents thoroughly. What do you think about all this? A PR stunt, a genuine error, or something more sinister?

Find out more about current AI developments here:
Six-month unlimited plan is now 57% off
$90
$210
$120 off (57%)
Mint Mobile is now allowing you to get whichever plan you like for either three, six, or 12 months for just $15/mo. If you go for the six-month unlimited service, for instance, you'll now have to pay just $90 upfront instead of $210.
Buy at Mint Mobile
Recommended For You
COMMENTS (0)