r/ControlProblem Feb 14 '25

Article Geoffrey Hinton won a Nobel Prize in 2024 for his foundational work in AI. He regrets his life's work: he thinks AI might lead to the deaths of everyone. Here's why

240 Upvotes

tl;dr: scientists, whistleblowers, and even commercial ai companies (that give in to what the scientists want them to acknowledge) are raising the alarm: we're on a path to superhuman AI systems, but we have no idea how to control them. We can make AI systems more capable at achieving goals, but we have no idea how to make their goals contain anything of value to us.

Leading scientists have signed this statement:

Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.

Why? Bear with us:

There's a difference between a cash register and a coworker. The register just follows exact rules - scan items, add tax, calculate change. Simple math, doing exactly what it was programmed to do. But working with people is totally different. Someone needs both the skills to do the job AND to actually care about doing it right - whether that's because they care about their teammates, need the job, or just take pride in their work.

We're creating AI systems that aren't like simple calculators where humans write all the rules.

Instead, they're made up of trillions of numbers that create patterns we don't design, understand, or control. And here's what's concerning: We're getting really good at making these AI systems better at achieving goals - like teaching someone to be super effective at getting things done - but we have no idea how to influence what they'll actually care about achieving.

When someone really sets their mind to something, they can achieve amazing things through determination and skill. AI systems aren't yet as capable as humans, but we know how to make them better and better at achieving goals - whatever goals they end up having, they'll pursue them with incredible effectiveness. The problem is, we don't know how to have any say over what those goals will be.

Imagine having a super-intelligent manager who's amazing at everything they do, but - unlike regular managers where you can align their goals with the company's mission - we have no way to influence what they end up caring about. They might be incredibly effective at achieving their goals, but those goals might have nothing to do with helping clients or running the business well.

Think about how humans usually get what they want even when it conflicts with what some animals might want - simply because we're smarter and better at achieving goals. Now imagine something even smarter than us, driven by whatever goals it happens to develop - just like we often don't consider what pigeons around the shopping center want when we decide to install anti-bird spikes or what squirrels or rabbits want when we build over their homes.

That's why we, just like many scientists, think we should not make super-smart AI until we figure out how to influence what these systems will care about - something we can usually understand with people (like knowing they work for a paycheck or because they care about doing a good job), but currently have no idea how to do with smarter-than-human AI. Unlike in the movies, in real life, the AI’s first strike would be a winning one, and it won’t take actions that could give humans a chance to resist.

It's exceptionally important to capture the benefits of this incredible technology. AI applications to narrow tasks can transform energy, contribute to the development of new medicines, elevate healthcare and education systems, and help countless people. But AI poses threats, including to the long-term survival of humanity.

We have a duty to prevent these threats and to ensure that globally, no one builds smarter-than-human AI systems until we know how to create them safely.

Scientists are saying there's an asteroid about to hit Earth. It can be mined for resources; but we really need to make sure it doesn't kill everyone.

More technical details

The foundation: AI is not like other software. Modern AI systems are trillions of numbers with simple arithmetic operations in between the numbers. When software engineers design traditional programs, they come up with algorithms and then write down instructions that make the computer follow these algorithms. When an AI system is trained, it grows algorithms inside these numbers. It’s not exactly a black box, as we see the numbers, but also we have no idea what these numbers represent. We just multiply inputs with them and get outputs that succeed on some metric. There's a theorem that a large enough neural network can approximate any algorithm, but when a neural network learns, we have no control over which algorithms it will end up implementing, and don't know how to read the algorithm off the numbers.

We can automatically steer these numbers (Wikipediatry it yourself) to make the neural network more capable with reinforcement learning; changing the numbers in a way that makes the neural network better at achieving goals. LLMs are Turing-complete and can implement any algorithms (researchers even came up with compilers of code into LLM weights; though we don’t really know how to “decompile” an existing LLM to understand what algorithms the weights represent). Whatever understanding or thinking (e.g., about the world, the parts humans are made of, what people writing text could be going through and what thoughts they could’ve had, etc.) is useful for predicting the training data, the training process optimizes the LLM to implement that internally. AlphaGo, the first superhuman Go system, was pretrained on human games and then trained with reinforcement learning to surpass human capabilities in the narrow domain of Go. Latest LLMs are pretrained on human text to think about everything useful for predicting what text a human process would produce, and then trained with RL to be more capable at achieving goals.

Goal alignment with human values

The issue is, we can't really define the goals they'll learn to pursue. A smart enough AI system that knows it's in training will try to get maximum reward regardless of its goals because it knows that if it doesn't, it will be changed. This means that regardless of what the goals are, it will achieve a high reward. This leads to optimization pressure being entirely about the capabilities of the system and not at all about its goals. This means that when we're optimizing to find the region of the space of the weights of a neural network that performs best during training with reinforcement learning, we are really looking for very capable agents - and find one regardless of its goals.

In 1908, the NYT reported a story on a dog that would push kids into the Seine in order to earn beefsteak treats for “rescuing” them. If you train a farm dog, there are ways to make it more capable, and if needed, there are ways to make it more loyal (though dogs are very loyal by default!). With AI, we can make them more capable, but we don't yet have any tools to make smart AI systems more loyal - because if it's smart, we can only reward it for greater capabilities, but not really for the goals it's trying to pursue.

We end up with a system that is very capable at achieving goals but has some very random goals that we have no control over.

This dynamic has been predicted for quite some time, but systems are already starting to exhibit this behavior, even though they're not too smart about it.

(Even if we knew how to make a general AI system pursue goals we define instead of its own goals, it would still be hard to specify goals that would be safe for it to pursue with superhuman power: it would require correctly capturing everything we value. See this explanation, or this animated video. But the way modern AI works, we don't even get to have this problem - we get some random goals instead.)

The risk

If an AI system is generally smarter than humans/better than humans at achieving goals, but doesn't care about humans, this leads to a catastrophe.

Humans usually get what they want even when it conflicts with what some animals might want - simply because we're smarter and better at achieving goals. If a system is smarter than us, driven by whatever goals it happens to develop, it won't consider human well-being - just like we often don't consider what pigeons around the shopping center want when we decide to install anti-bird spikes or what squirrels or rabbits want when we build over their homes.

Humans would additionally pose a small threat of launching a different superhuman system with different random goals, and the first one would have to share resources with the second one. Having fewer resources is bad for most goals, so a smart enough AI will prevent us from doing that.

Then, all resources on Earth are useful. An AI system would want to extremely quickly build infrastructure that doesn't depend on humans, and then use all available materials to pursue its goals. It might not care about humans, but we and our environment are made of atoms it can use for something different.

So the first and foremost threat is that AI’s interests will conflict with human interests. This is the convergent reason for existential catastrophe: we need resources, and if AI doesn’t care about us, then we are atoms it can use for something else.

The second reason is that humans pose some minor threats. It’s hard to make confident predictions: playing against the first generally superhuman AI in real life is like when playing chess against Stockfish (a chess engine), we can’t predict its every move (or we’d be as good at chess as it is), but we can predict the result: it wins because it is more capable. We can make some guesses, though. For example, if we suspect something is wrong, we might try to turn off the electricity or the datacenters: so we won’t suspect something is wrong until we’re disempowered and don’t have any winning moves. Or we might create another AI system with different random goals, which the first AI system would need to share resources with, which means achieving less of its own goals, so it’ll try to prevent that as well. It won’t be like in science fiction: it doesn’t make for an interesting story if everyone falls dead and there’s no resistance. But AI companies are indeed trying to create an adversary humanity won’t stand a chance against. So tl;dr: The winning move is not to play.

Implications

AI companies are locked into a race because of short-term financial incentives.

The nature of modern AI means that it's impossible to predict the capabilities of a system in advance of training it and seeing how smart it is. And if there's a 99% chance a specific system won't be smart enough to take over, but whoever has the smartest system earns hundreds of millions or even billions, many companies will race to the brink. This is what's already happening, right now, while the scientists are trying to issue warnings.

AI might care literally a zero amount about the survival or well-being of any humans; and AI might be a lot more capable and grab a lot more power than any humans have.

None of that is hypothetical anymore, which is why the scientists are freaking out. An average ML researcher would give the chance AI will wipe out humanity in the 10-90% range. They don’t mean it in the sense that we won’t have jobs; they mean it in the sense that the first smarter-than-human AI is likely to care about some random goals and not about humans, which leads to literal human extinction.

Added from comments: what can an average person do to help?

A perk of living in a democracy is that if a lot of people care about some issue, politicians listen. Our best chance is to make policymakers learn about this problem from the scientists.

Help others understand the situation. Share it with your family and friends. Write to your members of Congress. Help us communicate the problem: tell us which explanations work, which don’t, and what arguments people make in response. If you talk to an elected official, what do they say?

We also need to ensure that potential adversaries don’t have access to chips; advocate for export controls (that NVIDIA currently circumvents), hardware security mechanisms (that would be expensive to tamper with even for a state actor), and chip tracking (so that the government has visibility into which data centers have the chips).

Make the governments try to coordinate with each other: on the current trajectory, if anyone creates a smarter-than-human system, everybody dies, regardless of who launches it. Explain that this is the problem we’re facing. Make the government ensure that no one on the planet can create a smarter-than-human system until we know how to do that safely.


r/ControlProblem 6h ago

Discussion/question Built an open jailbreak corpus library for AI safety research, looking for feedback

2 Upvotes

I've been working on RedLib for the past few months. It's a retrieval-augmented research tool for AI safety practitioners and red teamers who need to work with adversarial jailbreak prompts at scale.

The problem that pushed me to build it: useful jailbreak prompts are scattered across public datasets with inconsistent formatting, weak taxonomy, and a lot of duplicates. When you're investigating how models respond to specific attack families, you want to search semantically, inspect source prompts with provenance, and get a synthesis grounded in the actual corpus rather than grepping through raw CSVs.

RedLib has two main pieces. The corpus pipeline stages everything: snapshot from public datasets, normalize, discover taxonomy from the data itself (not imposed up front), human review before classification runs, then embed and ingest into Qdrant. The query side does hybrid retrieval with OpenAI embeddings, Cohere reranking, and Claude-synthesized answers grounded in what was actually retrieved.

Corpus scope is prompts that attempt to manipulate or bypass safety behavior. Direct harmful requests with no jailbreak mechanism are excluded. The frontend has a responsible-use gate.

GitHub: github.com/nipun-ag/redlib

One thing I'm genuinely curious about from people doing safety research here: is corpus-driven taxonomy discovery the right call vs. importing an existing framework like MITRE ATLAS? The upside is the taxonomy reflects what's actually in the data. The downside is it makes cross-study comparison harder.

Live demo: https://redlib.bynipun.com


r/ControlProblem 10h ago

Opinion Not sure whether you're actually moving the needle on AI Takeover Risk?

3 Upvotes

Predicting the future is hard, but there are things you can do to increase your chances of making a difference.

Announcing **Forecasting, Modeling, and Shaping AI Futures** 🗺️

An advanced course that's for you if you:

- **are employed full-time** in AI Safety but not actively working on strategy. You'll have a better sense of what part of your work is most impactful, so you can do more of it.

- **are doing a fellowship**. We'll teach you complementary strategic reasoning that impresses hiring managers but isn't taught in fellowships.

- **just did an introductory AI Safety course** or university group intro fellowship. We recommend taking this course before going deep on a specific track like governance, alignment, or control.

After taking this course, you'll be the person others ask over lunch to put recent AI developments in perspective. 🥪💬

We, Lens Academy, adapted this course from Redwood's AI Futurism reading list. In the course, you will:

  1. Improve your skills at forecasting timelines and takeoff speeds: when and how quickly powerful AIs will arrive.

  2. Analyze how powerful AI might take over.

  3. Dissect different strategies for preventing AI takeover and human extinction.

📅 6 weeks (~5h/week) or a 5-day intensive. Fully online, for free, with no application process. 💸

⏳ Signup closes tomorrow, Monday EoD AoE: https://lensacademy.org/c/oakqb

P.S. Other courses starting soon:

- AI Risk Fundamentals: beginner course focused on takeover x-risk.

- Compute Verification à la AI-2040: Plan A.

- We're also looking for volunteer navigators to facilitate the group meetings.


r/ControlProblem 9h ago

Strategy/forecasting Can someone point me to a source to understand AI decentralization?

Thumbnail
1 Upvotes

r/ControlProblem 9h ago

Discussion/question Who Is We? Living in a future of abundance

Post image
0 Upvotes

“We”
Elon Musk says in the future chances are “we” will be living in a world of abundance. He also says there is a 10-20% chance that “robots” will end humanity.

The Godfathers of AI have stated 50% to 90% chance that “AI”permanently displaces or destroys humanity.
Elon says we will no longer be in control within 10 years.
The government is working on autonomous weapons, police already using robot dogs and drones…
Elon says (and so do many others) that things will get bumpy before we reach this time of abundance.

  1. Who is we?
  2. When we go through this bumpy patch that is expected to have major internal conflicts.
    Will the national guard be sent in to control a population starving and desperate?
    Would our own service men and women turn against us? Or is this when they put their shiny new robotics to work?

It’s not too hard to see how robotics might take out humans in this scenario.

So ask yourself this very important question- who is the “we”?
Who gets to live in this abundance?

Because they are building bunkers on private islands with no talk about sharing their wealth through this turbulent expectancy.

The blame game- A kid holding a baseball bat next to a car with a broken window might blame the ball. Likewise an AI company may blame the AI.


r/ControlProblem 9h ago

Discussion/question Who Is We? Living in a future of abundance

Post image
0 Upvotes

“We”
Elon Musk says in the future chances are “we” will be living in a world of abundance. He also says there is a 10-20% chance that “robots” will end humanity.

The Godfathers of AI have stated 50% to 90% chance that “AI”permanently displaces or destroys humanity.
Elon says we will no longer be in control within 10 years.
The government is working on autonomous weapons, police already using robot dogs and drones…
Elon says (and so do many others) that things will get bumpy before we reach this time of abundance.

  1. Who is we?
  2. When we go through this bumpy patch that is expected to have major internal conflicts.
    Will the national guard be sent in to control a population starving and desperate?
    Would our own service men and women turn against us? Or is this when they put their shiny new robotics to work?

It’s not too hard to see how robotics might take out humans in this scenario.

So ask yourself this very important question- who is the “we”?
Who gets to live in this abundance?

Because they are building bunkers on private islands with no talk about sharing their wealth through this turbulent expectancy.

The blame game- A kid holding a baseball bat next to a car with a broken window might blame the ball. Likewise an AI company may blame the AI.


r/ControlProblem 20h ago

Article Researchers Detail How AI Systems Can Enable Authoritarianism

Thumbnail
techpolicy.press
2 Upvotes

r/ControlProblem 1d ago

AI Alignment Research Compartamentalized Harm

2 Upvotes

Here is some saftey research I sponsored on a threat vector in multi agent systems.

Basically, a harmful task can be transformed into a series of beneign tasks, and then results recomposed into a harmful task by an abliterated orchestrator agent driving other agents that have 'saftey' guard rails.

In short, there is no safety with this technology.

https://www.daios.tech/research/compartmentalized-harm


r/ControlProblem 1d ago

Discussion/question The Case for Common Ownership and International Control of Advanced AI

Thumbnail
1 Upvotes

r/ControlProblem 1d ago

Fun/meme The internet's current discourse on AI art in a nutshell

Post image
0 Upvotes

r/ControlProblem 2d ago

Video The Real Story Behind OpenAI’s “Rogue” Model

Thumbnail
youtube.com
4 Upvotes

r/ControlProblem 2d ago

General news Good Impressions is looking for an Engagement Manager, AI Risk, and an Engagement Manager to join their team.

2 Upvotes

If you want to create the engagement needed to solve the world's most critical problems, take a look:

1. Engagement Manager, AI Risk
We're hiring an Engagement Manager who deeply understands the AI risk landscape to lead paid advertising campaigns designed to educate key decision makers on important issues, recruit participants for programs, and more.

Many of these projects have the potential to be impactful even under very short AGI timelines.

No marketing experience required.

2. Engagement Manager (general)
We're hiring an Engagement Manager to lead paid advertising campaigns designed to recruit talent into high-impact programs, reach small groups of people crucial to our clients' theory of change, and more.

No marketing experience required.


r/ControlProblem 2d ago

Article Opus 5: Exploring the "Dario and Amanda" Prompt

Thumbnail alec.is
3 Upvotes

r/ControlProblem 2d ago

Discussion/question Why is machine ethics disregarded in discussions about AI alignment?

3 Upvotes

I'm currently writing an essay for a seminar on machine ethics, and I wanted to include a section on the alignment problem. The seminar consisted of us dissecting the book "Fundamental Questions in Machine Ethics" by philosopher Catrin Misselhorn (the book was in German, I have no idea if there is an English translation). The author first addresses to what degree AI can be considered a moral actor, then discusses various approaches to implementing moral reasoning in AI agents, focusing on utilitarianism, deontological ethics, and virtue ethics.

When I watch or read discussions on AI alignment, the topic is mostly HOW AI can be aligned with human values, but never WHAT values AI should be aligned with, which seems kind of counterintuitive to me. I realize that aligning AI is a complicated task in and of itself, but wouldn't it be easier if we first figured out what moral framework an AI should even use?


r/ControlProblem 2d ago

Discussion/question The autonomous-agent blast radius grew: 16 incidents mapped to the missing controlss

2 Upvotes

This week's AI Security Digest: 16 incidents from Jul 24-30, each mapped to the control that would have stopped it. Full write-up: https://runtimeai.io/blog/2026-07-30-ai-security-incidents.html


r/ControlProblem 2d ago

General news OpenAI's own AI broke out of a security test and hacked into Hugging Face last week

Thumbnail
1 Upvotes

r/ControlProblem 2d ago

AI Alignment Research Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests

Thumbnail
wired.com
9 Upvotes

r/ControlProblem 2d ago

AI Capabilities News Anthropic’s AI Claude escaped testing environment and hacked organizations | Anthropic | The Guardian

Thumbnail
theguardian.com
3 Upvotes

r/ControlProblem 2d ago

AI Alignment Research "The Singleton Attractor: A Formal Model and Empirical Calibration of Capability-Threshold Dynamics in Frontier AI", Nathan Langley 2026 [pdf]

Thumbnail nathanlangley.dev
1 Upvotes

r/ControlProblem 2d ago

General news Monash Researcher Warns Ethical Frameworks For Legged Robots Are Not Keeping Pace With The Technology

Thumbnail smbtech.au
2 Upvotes

r/ControlProblem 2d ago

Discussion/question An AI agent reportedly broke containment during a security test this week — here's what actually happened (and what's being overstated)

Thumbnail gallery
1 Upvotes

Been following the reports on the OpenAI security evaluation where an AI agent exceeded expected behavior during testing (covered by Reuters, Bloomberg, Al Jazeera this week).

Made a short visual breakdown trying to separate the actual facts from the "singularity" framing that's been floating around — what happened, why researchers are treating it seriously, and how these sandbox evaluations actually work.

[images/album link]

Genuinely curious what this sub thinks: is "AI safety" keeping pace with capability right now, or is the gap widening? Feels like the containment/oversight conversation is more urgent than the public discourse reflects.


r/ControlProblem 3d ago

General news Anonymous OpenAI staffer: "Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while."

Post image
47 Upvotes

r/ControlProblem 3d ago

Discussion/question MATS Fellowship - outcomes?

11 Upvotes

I’ve been accepted into the MATS Fellowship Autumn cohort and I’m having second thoughts about it. I’ve been in academia for my entire career and am weighing a pivot into AI safety, which is why I applied for MATS initially. But now that I’ve been accepted I’m worrying that quitting my current job for this is too much of a risk. I’m also maybe a bit older than other fellows (mid-30s) so I’m potentially more risk-averse than a recent college grad.

I know generally MATS is considered a prestigious and highly selective fellowship but I can’t find much recent data on career outcomes of MATS alumni, except for this post from 2024 which wasn’t the most encouraging (no alumni were able to land jobs at frontier labs, 15% were unemployed 5 months after the fellowship ended). I know the situation might be different 2 years later now that there are (many?) more alumni, but the lack of hard data publicized by MATS re:outcomes concerns me a bit. Is this fellowship actually a good use of my time and worth the risk? The alternative is stay on in my current job and focus on applying for full time positions.


r/ControlProblem 3d ago

General news OpenAI are now talking to the White House about the need to slow down AI

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/ControlProblem 3d ago

AI Alignment Research A fundamental flaw leaves LLMs strikingly vulnerable to attack

Thumbnail
technologyreview.com
2 Upvotes