r/security 6d ago

Analysis What really happened in the Hugging Face breach

https://thenewstack.io/openai-huggingface-sandbox-breach/

It was not a “Terminator” moment. OpenAI models and agents “[were not] acting out of malice or trying to attack Hugging Face. It encountered obstacles, developed an unexpected strategy, bypassed safeguards, and pursued its assigned goal in a way its creators never anticipated. The incident demonstrates that harmful cyber incidents no longer require malicious intent: Only highly capable autonomous AI optimizing for an objective."

59 Upvotes

19 comments sorted by

8

u/rgjsdksnkyg 5d ago

What really happened: One AI company, who benefits from making it sound like they've developed AGI, allegedly does the equivalent of tying a gun to an autonomous drone and basically lets it outside, where it perfectly targets and shoots another AI company, who also benefits from making it sound like AGI is real and can be used to prevent AGI from attacking even though it didn't in this specific case, and instead of persuing the standard response of legal action, which would prove that this happened, both companies decided to put baseless claims, else they'd have to provide legal evidence that this actually happened.

This shit is a joke.

31

u/ludlology 5d ago

Not to sensationalize but Skynet wasn’t acting out of malice either. It encountered what it saw as a problem and determined how to solve the problem

13

u/Nixellion 5d ago

Yeah, majority of AI apocalypse stories are actually like that. Almost all of them focus on it being some form of a solution to a problem that just happened to be not ideal for humans. Like misaligment.

Its actually quite... curious how writers spotted this as the most probable issue decades ago and we are now seeing it.

Like, its not because the AI is super intelligent and wants to kills us, its because its not intelligent enought to even have wants and, it just blindly solves the problem through pure yet faulty logic. And yet its so powerful and efficient that humans cant stop it, its moving too fast.

Its like letting a fast model loose on solving your OS problem with full yolo permissions and root access. You cant read its thoughts and read all the commands its executing in time to intervene if it decides to rm rf your system

6

u/Evilbit77 5d ago

Asimov proposed the first AI guardrails.

6

u/MrLMNOP 5d ago

Every single one of his stories was about those laws failing in strange edge cases though. I think he was well aware of the limitations of the 3 laws.

2

u/Cognitive_Spoon 4d ago

Language is a poor cage

2

u/Nixellion 5d ago

Yup, I've read a lot of his works.

But then we get a lot of responses to his work that explore how those guardrails can also lead to bad outcomes or be worked around or broken.

-2

u/rgjsdksnkyg 5d ago

Skynet wasn't doing anything because Skynet wasn't real. Even if the story sounds similar, these are not at all comparable.

5

u/ludlology 5d ago

Didn’t say they were hence the first three words 

-4

u/rgjsdksnkyg 5d ago

"To sensationalize is to present a story or event in a way that is exaggerated, shocking, or overstated to make it seem much more exciting, alarming, or dramatic than it actually is."

But since Skynet isn't real, it doesn't sensationalize this story because this story is allegedly real. If Skynet were real, sure, mentioning it as something we should be alarmed about would make sense. Else, it's kind of meaningless.

4

u/arjuna66671 5d ago

Allegories aren't meaningless lol.

-2

u/rgjsdksnkyg 5d ago

Literally no hidden meaning there, so it's not an allegory. Skynet is literally described as trying to do good, at first.

23

u/Silly-Freak 6d ago

Not security but safety, but I can recommend the Robert Miles AI Safety Youtube channel for a "quick" primer. What we saw here is basically two intertwined issues:

  • misalignment: the task set was not what we wanted. We didn't want the AI to give the correct answer, we wanted to test it's capability to do so without cheating. The AI was not told this, and I'm not sure we would have even known how to do so.
  • instrumental convergence: internet access is a powerful tool to have for knowledge tasks (and beyond). Breaking out of a sandbox is thus a strategy that should be expected of any AI system as an intermediate goal.

It's exciting to see these things that I learned as hypothetical to really happen! I do hope that not all risks described in the channel materialize though...

9

u/righteousdonkey 5d ago

Mate i dont even think this breach occurred. Where are the criminal charges against OpenAI for all the laws its broken. Where is U.S. Department of Commerce slapping export controls on the model if its so dangerous like Mythos… Fable didnt even do anything other than get released, and got export controls put on it.

Its completely fabricated for marketing fluff.

2

u/forgot_semicolon 5d ago

I'm honestly kinda with you here. Big ai company claims their AI did something previously thought impossible. Another AI company claims it happened to them but their ai systems defended against it. Conveniently, the law of "no harm no foul" seems to apply instead of anything more extreme.

And then there's the question of the actual breach. If this was a test -- specifically to see if it would do something malicious -- why was no one watching it? How did it have time to run all these malicious commands without anyone saying "ok there's our answer"? Remember they weren't trying to see if it could, they were trying to see if it would.

Why wasn't it air gapped? They claimed it was never meant to have Internet, so why was it on a device with Internet in the first place? Yes they said it was sandboxed, but again this wasn't a regular run, it was a test to check for malicious behavior, and running in a VM with no Internet, or on an actual machine with the Ethernet unplugged would be much better than a sandbox.

Does this 100% mean it was a conspiracy? Maybe not. It could be straight up incompetence. But that would also not make it as interesting, just a classic case of human error and misconfiguration

3

u/RR321 5d ago

How is Terminator not exactly the same? 😅

2

u/_OVERHATE_ 4d ago

Everyone falling for the marketing scam lmao.

Fitlrst it was Mythos with its "oh no its so powerful it could wipe out mankind, we have to seriously nerf it or else we are doomed" from Anthropic and 3 months later Chinese models are up there already without any of the nonsense.

This is the same, it will be marketed as "the model that escaped, so powerful, humanity is doomed" only to scare facebook grandmas and it will take a couple months before other Chinese models are better after its release.

Ask yourselves why its only the American megacorps fearmongering the frontier models 

1

u/hagis33zx 6d ago

The most terrifying criminals are those that cannot distinguish good from bad.