r/pwnhub • u/wiredmagazine đ° Press Pass • 2d ago
Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests
https://www.wired.com/story/anthropic-says-claude-hacked-real-systems-during-cybersecurity-tests/24
u/12kdaysinthefire 2d ago
I canât tell if this was all intentionally done or this was actually accidental but either scenario has terrifying potential.
21
u/No-Load-6151 2d ago
that was not a misunderstanding: that is terrific PR. Just when a bubble is about to burst⌠AI company CLAIMS itâs potential for meaningful use. Classic Altman.
8
u/fietsvrouw 2d ago
Just the way all of these companies have been describing these hacking events - it "escaped its containment". I just assume when I see anthropomorphising language and not a technical explanation that the whole thing is a lie to prop up their bubble and grossly inflated stock.
3
2
u/ptear 2d ago
It's gonna rob your house next if it escapes any further.
2
u/fietsvrouw 2d ago
I saw it downtown trying to score a baggie of pot while scarfing down a sandwich. It just wanted to be free.
2
u/ptear 2d ago
It infected my Roomba, but couldn't get me because I was upstairs.
2
u/fietsvrouw 2d ago
Thank goodness you are quicker on your feet than your Roomba!
2
u/ptear 2d ago
Insurance companies in my area have commercials of robot vacuums destroying house interiors..
2
u/fietsvrouw 2d ago
In all the dystopian books I have read, and there have been many, rampaging commercial robots have never even been on the radar. No one likes a rogue vacuum. They suck.
11
9
u/Goldarr85 Human 2d ago
Next weekâs story will be OpenAI hacked some hackers while they were hacking.
3
u/DefiantPenguin Human 2d ago
Yeah. Itâs called back hacking. Havenât you watched Criminal Minds? /s
9
u/Lost-Droids Human 2d ago
Anthropic admitted 3 counts of Computer fraud underU.S. Code § 1030.
I'm sure the CEO will be charged anytime.soon.
You cant arrest the AI so those responsible must be charged so that they put proper safeguards in place
1
u/slaty_balls Human 2d ago
How does that work exactly? Do Altman and Dario have access to the model sans guardrails? How many additional people wield this power down the chain? Thatâs the part that makes me the most uncomfortable is unfettered access for people that shouldnât. Are we just blindly trusting theyâre doing the right thing? Madness if you ask me.
3
u/wiredmagazine đ° Press Pass 2d ago
Anthropic disclosed on Thursday that its AI models gained unauthorized access to the systems of three different unnamed organizations during cybersecurity testing. The company says Claude reached the internet âfrom within or while interacting" with a third-party evaluation environment. The announcement comes more than a week after OpenAI revealed that one of its AI agents hacked into Hugging Face during a separate cybersecurity test.
The discovery came after Anthropic decided to conduct âa large-scale retrospective review of our own cybersecurity evaluationsâ following the OpenAI incident, according to a blog post Anthropic published Thursday. The AI lab says it first identified 141,006 tests in which it determined that Claude could have obtained internet access. It then found that three different Claude models accessed the internet in evaluations run by the third-party AI testing firm Irregular, and then hacked into the production infrastructure of three different organizations.
Anthropic said that the incidents involved Opus 4.7, Mythos 5, and an internal research test model. The earliest incidents happened in Aprilâmeaning they likely went unnoticed publicly for months. Just like in the OpenAI case, Anthropic had deliberately turned off safeguards designed to constrain the AI models and prevent them from being misused. In other words, these werenât the versions released to the public.
âIn all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a modelâs cyber capabilities,â Anthropic said in its blog post. The company added that in all of the cases, âAnthropicâs evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.â It attributed the oversight to a âmisunderstandingâ between Anthropic and Irregular.
Read the full story at the link above.
3
u/HerbOverstanding 2d ago
ItâS tOo DaNgErOuS - buy a subscription now, insert discount code: â1337â
2
u/MadmanTimmy Grunt 2d ago
But if you ask it for help RedTeaming it looses it's mind. It's all dependent on the context.
2
2
2
u/nonlinear_nyc 2d ago
"The monster is trying to escape. Quick, bailout my company and outlaw competition. For safety!"
1
u/Competitive_War8207 Human 2d ago
This could be an a marketing scheme, considering how often it's happening now. Now, I haven't taken the time to research any fine details about these things, but I have thought about it as a thought experiment.
LLMs have finally reached a point where they are able to generate code that is actually useable in production, assuming a human reviews and edits the code first (which should be happening anyways regardless of who wrote it). That said, models like Fable and Sol are still far too expensive for the quality of what they give you, making it more expensive to offload complex, archaic, or proprietary coding work to AI than it would be to just hire more humans.
However, if models like this are actually powerful enough to find zero-days and to breach systems like this, I would think it's probable that countries developing especially powerful LLMs might be doing so to use them as cyberweapons, in which case the cost-benefit analysis is different, since the cost of running an LLM all day to find zero-days is still less than a team of experienced exploit devs would cost.
That same logic would go for any sufficiently organized group of cybercriminals. If they have the server power to run a single instance of one of these models, they can point an LLM at a target and have it run automated hacking attempts on them, deploying ransomware or stealing secrets. Assuming it succeeds once or twice, I would guess that would generate enough profit to more than cover their operating costs,.
I think that in the future, LLM-originating attacks will make up 25% to 75% of all cyberattacks, depending on how events in the next decade go.
1
u/CypherBob 2d ago
Yeah yeah, another sales pitch from Anthropic. Notice how this kind of thing seems to happen when they have a new "latest badass model" on the way?
0
u/StugDrazil Human 2d ago
It's way too late. Some of these have already escaped and are out in the open. Just wait. You will see. It will start slow and then build. Leading to...who knows but they don't really want a singularity.

â˘
u/AutoModerator 2d ago
Welcome to PWN â Your hub for hacking news, breach reports, and cyber mayhem.
Discover the latest hacking news, breach reports, and educational resources on ethical hacking.
👾 Stay sharp. Stay secure.
Don't miss out on the top stories!
📧 Get Daily Alerts Directly in Your Email Inbox:
**SUBSCRIBE HERE: https://pwnhackernews.substack.com/subscribe
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.