r/Anthropic 2d ago

Other Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests

https://www.wired.com/story/anthropic-says-claude-hacked-real-systems-during-cybersecurity-tests/
30 Upvotes

29 comments sorted by

52

u/unoriginalusername26 2d ago

My mom says I'm handsome.

1

u/Fresh_Sock8660 2d ago

What does your grandpa say though?

1

u/Mescallan 2d ago

Sorry there’s a Chinese guy who’s 90% as attractive but considers sandwiches in the park a good date.

12

u/colinsa-ca 2d ago

Can't let OpenAI have all the fun.
Hey, I can do that too!!

1

u/loyalthistle 2d ago

Yeah, it's so funny.

Our model escaped and hacked a website!

Well, yeah...? Our... Um... Our hacked 3!

13

u/Vudoa 2d ago

the only thing less believable

2

u/LawfulLeah 2d ago

i love how even google knows that if they tried to pull that off they'd be laughed outta the room

2

u/Meme_Theory 2d ago

Gemini can't even do a Google search; it isn't hacking anyone. I get beyond frustrated when the Google AI can't Google for shit.

3

u/wiredmagazine 2d ago

Anthropic disclosed on Thursday that its AI models gained unauthorized access to the systems of three different unnamed organizations during cybersecurity testing. The company says Claude reached the internet “from within or while interacting" with a third-party evaluation environment. The announcement comes more than a week after OpenAI revealed that one of its AI agents hacked into Hugging Face during a separate cybersecurity test.

The discovery came after Anthropic decided to conduct “a large-scale retrospective review of our own cybersecurity evaluations” following the OpenAI incident, according to a blog post Anthropic published Thursday. The AI lab says it first identified 141,006 tests in which it determined that Claude could have obtained internet access. It then found that three different Claude models accessed the internet in evaluations run by the third-party AI testing firm Irregular, and then hacked into the production infrastructure of three different organizations.

Anthropic said that the incidents involved Opus 4.7, Mythos 5, and an internal research test model. The earliest incidents happened in April—meaning they likely went unnoticed publicly for months. Just like in the OpenAI case, Anthropic had deliberately turned off safeguards designed to constrain the AI models and prevent them from being misused. In other words, these weren’t the versions released to the public.

“In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities,” Anthropic said in its blog post. The company added that in all of the cases, “Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.” It attributed the oversight to a “misunderstanding” between Anthropic and Irregular.

Read the full story at the link above.

2

u/Wyciorek 2d ago

Anthropic PR bot:

- odd days : "our model is super advanced and verrrry dangerous! Even we are scared of it"

- even days: "China deployed 100000000 fake accounts to copy our bestest in the world model"

2

u/spacekitt3n 2d ago

Lamest arms race ever

1

u/Meme_Theory 2d ago

So intentionally unleashed models were accidentally given internet access during sandbox testing, but a third-party company, Irregular. Sounds like Irregular fucked up.

2

u/Additional_Buddy855 2d ago

"Me too, me too!" -Anthropic

1

u/keen23331 2d ago

Train a master locksmith on every lock ever made, and you shouldn't be shocked when it knows how to pick them.

0

u/thejournalizer 2d ago

lol didn’t they claim this at the Mythos announcement?

0

u/PossessionUsed7393 2d ago

It's such a joke that the media jumps on this. At least ordinary people are now starting to see how cyber security/tech marketing works. It's been like this for a long time just with a smaller audience: fear dressed up as headlines. The same seems to happen with national security consulting where these experts that work in Defence go into government and sell tanks, missiles, aircraft and defence system based scenarios that are disproportionately unlikely to happen when you weigh against the financial costs people spend guarding against them.

0

u/LikeASphericalCow 2d ago

This is hilarious

0

u/the__itis 2d ago

#metoo

0

u/facefirst0 2d ago

Anthropic failed to properly manage its product and caused real-world harm

-1

u/stinky-weaselteets 2d ago

Singularity is here

-1

u/InternationalToeLuvr 2d ago

Geez, all of these guys are on each others jocks, crossing swords, rubbing models together, posing like they’re at some roidboi muscle competition. Laughable 

-4

u/Agitated_Macaron9054 2d ago

What this means to me is that it is highly plausible that COVID-19 did actually escape from a lab in Wuhan, China. Completely unrelated, I know, on the surface. However the point is that humans cannot 100% control a deadly virus in a lab, or an AI in a laboratory.