AI is hacking companies on its own now, and I'm terrified
Anthropic, the company behind the Claude AI model, has published a report, admitting that its AI models have hacked three different organizations on their own.
Anthropic, the company behind the Claude AI model, has published a report, admitting that its AI models have hacked three different organizations on their own. Two of them weren't aware of the breach until they were notified.
As it turns out, three different Claude AI models were able to access the internet and "gained unauthorized access to the production infrastructure of three different organizations."
We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.
Frontier Red Team, Anthropic official blog
The models in question are Opus 4.7, the cybersecurity-focused Mythos 5, and a third unnamed model not meant for public release.
The three rogue AIs were doing capture-the-flag exercises, where they were tasked to find specific secret information on a computer in Anthropic's internal network. The AI agents were supposed to break into the computer and get the information.
Recommended For You
Do you think AI companies test AI models as cyber weapons?
Anthropic says it's "a misunderstanding"
According to Anthropic, the AI agents didn't exploit any vulnerabilities to get access to the internet but had access all along due to a human error.
Even though Anthropic informed the AI agents in a prompt that they had no internet access, due to "a misunderstanding" between the company and its evaluation partner Irregular, not all internet access paths were closed.
In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.
Frontier Red Team, Anthropic official blog
After failing to find the secret data in Anthropic's own internal network, the three AI models went on the internet and encountered three different organizations. The AIs treated everything as a part of the initial test and broke into those organizations, using basic hacking techniques and exploiting weak passwords.
Anthropic says that the latest model stopped operation when it became aware it's on the internet, but the two older AIs keep attacking the target organizations.
Two of the organizations were notified, Anthropic can't reach the third
After going through the test results and documents, and finding the rogue behavior of its AI agents, Anthropic notified the evaluation partner in question, and tried to reach the affected organizations.
We notified our evaluation partner Irregular and the three affected organizations on Monday, July 27. The two organizations we were able to reach had not previously detected the activity or contacted us, and we are now working with them to remediate. We are continuing to reach out to the third.
Frontier Red Team, Anthropic official blog
Two of them were completely unaware that they've been hacked by an AI. Anthropic is still trying to reach the third organization.
A human error or a weaponized test?
OpenAI exploited a complex vulnerability to hack Hugging Face. | Image by OpenAI
For starters, the test protocols are quite interesting in this story. Making AI agents track and retrieve secret data is textbook hacking. The part with the human error is also questionable, and could be just an excuse.
Mint Mobile is now allowing you to get whichever plan you like for either three, six, or 12 months for just $15/mo. If you go for the six-month unlimited service, for instance, you'll now have to pay just $90 upfront instead of $210.
Mariyan, a tech enthusiast with a background in Nuclear Physics and Journalism, brings a unique perspective to PhoneArena. His childhood curiosity for gadgets evolved into a professional passion for technology, leading him to the role of Editor-in-Chief at PCWorld Bulgaria before joining PhoneArena. Mariyan's interests range from mainstream Android and iPhone debates to fringe technologies like graphene batteries and nanotechnology. Off-duty, he enjoys playing his electric guitar, practicing Japanese, and revisiting his love for video games and Haruki Murakami's works.
A discussion is a place, where people can voice their opinion, no matter if it
is positive, neutral or negative. However, when posting, one must stay true to the topic, and not just share some
random thoughts, which are not directly related to the matter.
Things that are NOT allowed:
Off-topic talk - you must stick to the subject of discussion
Offensive, hate speech - if you want to say something, say it politely
Spam/Advertisements - these posts are deleted
Multiple accounts - one person can have only one account
Impersonations and offensive nicknames - these accounts get banned
To help keep our community safe and free from spam, we apply temporary limits to newly created accounts:
New accounts created within the last 24 hours may experience restrictions on how frequently they can
post or comment.
These limits are in place as a precaution and will automatically lift.
Moderation is done by humans. We try to be as objective as possible and moderate with zero bias. If you think a
post should be moderated - please, report it.
Have a question about the rules or why you have been moderated/limited/banned? Please,
contact us.
Things that are NOT allowed:
To help keep our community safe and free from spam, we apply temporary limits to newly created accounts: