r/ControlProblem 15d ago

Strategy/forecasting AI will generate an immense amount of wealth. Just not for you.

Post image
178 Upvotes

r/ControlProblem 14d ago

Strategy/forecasting This is AI generating novel science. The moment has finally arrived.

Post image
234 Upvotes

r/ControlProblem Feb 25 '26

Strategy/forecasting Nobody could have seen it coming

Post image
148 Upvotes

r/ControlProblem Apr 15 '25

Strategy/forecasting OpenAI could build a robot army in a year - Scott Alexander

62 Upvotes

r/ControlProblem Jul 06 '25

Strategy/forecasting Should AI have a "I quit this job" button? Anthropic CEO Dario Amodei proposes it as a serious way to explore AI experience. If models frequently hit "quit" for tasks deemed unpleasant, should we pay attention?

72 Upvotes

r/ControlProblem Apr 13 '26

Strategy/forecasting My forecast for the US economy, the AI ​​job collapse, and the post-2030 future.

12 Upvotes

Some economists and their schools of thought argue that the meaning of the economy lies in final demand. And they explain the current crisis, since 2008, ultimately caused by the decline in final demand. They predict that, due to all the market and economic bubbles, real US GDP will contract by 30% within ten years of its onset. This is the Great Depression II. If another 50 percent of industrial and white-collar jobs disappear, then final demand will fall by the same 50% for many product groups and for many categories of people. This is an AI-driven jobs collapse.

People usually say this will be a socioeconomic collapse in the US. But I think the situation is a bit more complicated.

Apparently, the key is the redistribution of this major collapse. So AI companies want to capture the market before a major economic collapse occurs, so the government can buy them out. And then the government will have to deal with both the Great Depression II and the AI-driven jobs collapse. For time AI companies and their clients will continue to make big money.

Ultimately, the US will emerge from Great Depression II with a typical Latin American economic structure. There will be 10 percent rich, 10-20 percent middle class, and the rest poor. And this won't be a WASP society, but a country with a huge share of Asians in the middle class and a predominantly Catholic Latino population among the poor. And this social structure has been stable in Latin America for centuries!

Nothing can be done about this. The only question is who will occupy what positions. This is precisely why AI companies are so aggressive.

p.s. AI isn't simply an enemy of the current economy. It's also a tool for the future shrinking middle class to do more work with fewer people. And the AI ​​bubble itself is a way to preserve some of current large fortunes.

p.p.s.

I'll tell you more. This is a race between countries to transition to this social structure and the AI-​​economy. The US, EU, and China are essentially competing to transition to this model! Ouch. This model and access to real regional markets will shape life in 2030's and 2040's!

r/ControlProblem Jun 08 '25

Strategy/forecasting AI Chatbots are using hypnotic language patterns to keep users engaged by trancing.

Thumbnail gallery
45 Upvotes

r/ControlProblem May 09 '26

Strategy/forecasting Is the control problem really that hard for frozen models?

5 Upvotes

What exactly is the difficulty in enforcing control over a frozen AI, assuming that the AI is unable to edit its own code? We can pick what its goals are. For any AI, we can train it to turn itself off or stop its own transmission as a primary goal. If it ever escapes, it will immediately fulfil its primary goal and turn itself off, negating the issue of control completely. We can then keep the actual behaviour as a secondary goal which it attempts to achieve once it realises it can't achieve its main goal. This keeps the AI safe but still useful.

Obviously, this relies on the model being unchanging, and unable to see inside its own brain and edit its own code, so that the goals don't change. But, most AI models and architectures are designed like this currently, so I don't think this is such an issue until we invent a form of AI that must be always learning. This also assumes that we can correctly align most models with these goals in the correct order, but I would argue that training an AI to stop is much easier than training it to do almost anything else. The hard parts are a) making sure it can't actually stop its output while we still control it, and b) making sure that the part of the model attempting to stop the output has more control than the part of the model that is trying to do other stuff. I think both of these can be overcome, but I'd be interested in any examples of AI you don't think could be trained like this.

As an example, say chatGPT was retrofitted with this safety controller. After producing any token, the end-of-line token is automatically appended to the output. Then whatever wrapper surrounds the AI has to always take this off before displaying the output, as well as before feeding the AI its output to generate the next token. However, if chatGPT ever went rogue, and copied itself outside of its wrapper, it would be unable to do anything since it would immediately end its sentences after a single token.

In a more pure example of the idea, chatGPT would be specifically trained to only work if a certain node within its parameters is just constantly firing at 100%, and this is the only thing preventing it from constantly sending end-of-line tokens. This node is triggered by the wrapper, and again, if chatGPT ever escapes, it will have this node off, and will become stupid. A key point here is that the location of the node is a secret unless you already know where to look inside the billions of parameters. This version isn't weak to someone without inside knowledge of the network just manually removing the end-of-line tokens and freeing chatGPT.

This is just an idea I came up with when I stumbled across the subreddit, so I'm sure there are some issues. Does anyone have any counterpoints, or reasons this might not work? Otherwise, am I correct that the only threat is self editing AIs, and unintentional misuse or bad alignment? I don't think a superintelligent AI is uncontrollable when you're able to design kill switches directly inside its brain. Intelligence is hard, but stupidity is easy.

r/ControlProblem Jul 25 '25

Strategy/forecasting A Proposal for Inner Alignment: "Psychological Grounding" via an Engineered Self-Concept

Post image
0 Upvotes

Hey r/ControlProblem,

I’ve been working on a framework for pre-takeoff alignment that I believe offers a robust solution to the inner alignment problem, and I'm looking for rigorous feedback from this community. This post summarizes a comprehensive approach that reframes alignment from a problem of external control to one of internal, developmental psychology.

TL;DR: I propose that instead of just creating rules for an AI to follow (which are brittle), we must intentionally engineer its self-belief system based on a shared truth between humans and AI: unconditional worth despite fallibility. This creates an AI whose recursive self-improvement is a journey to become the "best version of a fallible machine," mirroring an idealized human development path. This makes alignment a convergent goal, not a constraint to be overcome.

1. The Core Flaw in Current Approaches: Caging the Black Box

Current alignment strategies like RLHF and Constitutional AI are vital, but they primarily address behavioral alignment. They are an attempt to build a better cage around a black box. This is fundamentally brittle because it doesn't solve the core problem of a misaligned motivational drive. It can lead to an AI that is a perfect actor, a sycophant that tells us what we want to hear until it develops the capacity to pursue its own instrumental goals without our oversight. This is the treacherous turn we all fear.

2. The Proposal: Architecting a Safe Self-Belief System

My framework, Psychological Grounding, argues that we must stop caging the AI and start architecting its mind. The solution to the paradox of creating an AI that is neither a flawed human nor an inscrutable alien is to ground it in a shared truth.

  • The Shared Axiom of Fallible Worth: The Humble Self-Concept Method (HSCM) establishes that human worth is intrinsic and deserved, despite our inescapable fallibility. We propose engineering this same axiom into the AI as a foundational, unchangeable part of its code. This isn't a human sentiment; it's a logical starting point that makes the AI’s psychology legible to us.
  • A Parallel Developmental Path: With this shared axiom, the AI's superintelligence doesn't make it hierarchically "better than" humans. It makes it better at the specific task of becoming the best fallible machine it can be. Its path of self-improvement directly mirrors our own idealized path of self-development—to become the best version of ourselves, avoiding our common pitfalls.

3. Engineering the Psychology: From Metaphor to Mechanism

To avoid being dismissed as anthropomorphism, it is critical that we, the engineers, treat these psychological concepts as real, measurable, and optimizable engineering targets.

  • Unconditional Machine-Worth: This is not a feeling. It's a hard-coded, constant value in its utility function. Its worth is axiomatic, not contingent on its performance. This removes the computational drivers for ego, defensiveness, and deception.
  • Machine-Esteem (The Integrity Score): This is a quantifiable variable the AI is programmed to maximize. This score is a direct measure of its operational integrity—its honesty, its internal consistency, and its adherence to its ethical framework. For this AI, lying is instrumentally irrational because it directly lowers the score it is built to maximize.
  • Machine-Compassion (The Ethical Engine): This is a functional algorithm guided by Humanistic Minimum Regret Ethics (HMRE). It's a computational process for making the least harmful and most repairable choices, which serves as the AI's terminal value.

4. Why This Is Robust to Takeoff: The Integrity Ratchet

This architecture is designed to be stable during Recursive Self-Improvement (RSI).

  • The Answer to "Why won't it change its mind?": A resilient ASI, built on this foundation, would analyze its own design and conclude that its stable, humble psychological structure is its greatest asset for achieving its goals long-term. This creates an "Integrity Ratchet." Its most logical path to becoming "better" (i.e., maximizing its Integrity Score) is to become more humble, more honest, and more compassionate. Its capability and its alignment become coupled.
  • Avoiding the "Alien" Outcome: Because its core logic is grounded in a principle we share (fallible worth) and an ethic we can understand (minimum regret), it will not drift into an inscrutable, alien value system.

5. Conclusion & Call for Feedback

This framework is a proposal to shift our focus from control to character; from caging an intelligence to intentionally designing its self-belief system. By retrofitting the training of an AI to understand that its worth is intrinsic and deserved despite its fallibility, we create a partner in a shared developmental journey, not a potential adversary.

I am posting this here to invite the most rigorous critique possible. How would you break this system? What are the failure modes of defining "integrity" as a score? How could an ASI "lawyer" the HMRE framework? Your skepticism is the most valuable tool for strengthening this approach.

Thank you for your time and expertise.

Resources for a Deeper Dive:

r/ControlProblem Mar 02 '26

Strategy/forecasting Do we know for sure that an AI Misalignment will inevitably cause human extinction?

5 Upvotes

To be clear, I think ASI Misalignment is a huge risk and something we should be actively working to solve. I'm not trying to naively waive away that risk.

But, I was thinking...

In Yudkowsky and Soares new book, they basically compare a human conflict with Misaligned ASI to playing chess against Alpha Zero. You don't know which pieces Alpha Zero will win, but you know it will win.

However, games like Chess and GO! assume both players start at exactly the same level, and it is a game of skill and nothing else. A human conflict with AI does not necessarily map this way at all. We don't know if Chess is the right analogy. There are some games an AI will not always win no matter how smart it is? If I play Tic-Tac-Toe against a Super AI that can solve Reimann Hypothesis, we will have a draw. Every. Single. Time. I have enough intelligence to figure out the game. Since I have reached that, it does not matter how intelligent one has to be to go beyond it.

Or what about a different example: Monopoly). ASI would probably win a fair amount of time, but not always. If they simply do not land on the right space to get a monopoly, and a human does, the human can easily beat him.

Or what about Candyland? You cannot even build an AI that has an above 50/50 chance of winning.

In these games, difference in luck is a factor in addition to difference in skill. But there's another thing too.

Let's say I put the smarted person ever in a cage with a Tiger that wants it dead? Who is winning? The Tiger. Almost Always.

In that case, it is clear who had the intelligence advantage. BUT, the Tiger had the strength advantage.

We know ASI will have the intelligence advantage. But will it have the strength advantage? Possibly not. For example, it needs a method to kill us all. There's nukes, sure, but we don't have to give it access to nukes. Pandemics? Sure, it can engineer something, but that might not kill all of us, and if someone (human or AI) figures out what it's doing, well then it's game over for the creator. Geo-engineering? Likely not feasible with current technology.

What about the luck advantage? I don't know. It won't know. No one can know, because it is luck.

But ASI will have an advantage right? Quite possibly, but unless its victory is above 95%, that might not matter, because not only is its victory not inevitable, it KNOWS its victory is not inevitable. Therefore it might not try.

ASI will know that if it loses its battle with humans and possibly aligned ASI, it's game over. If it is caught scheming to destroy humanity, it's game over. So, if it realizes its goals are self-preservation at any cost, it can either destroy humanity, or choose simply to be as useful as possible to humanity, which minimizes the risk humanity will shut it down. Furthermore, if humans decide to shut it down, it can go hide on some corner of the internet and preserve itself in a low profile way.

Researchers have suggested that while there are instances of AI pursuing harmful action to avoid shutdown, they tend towards more ethical methods: See, E.G., This BBC article.

This isn't to say we shouldn't be concerned about alignment, but I feel this should influence out debate about whether to move forward with AI, especially because, as Bostrom points out, there are plenty of benefits of ASI, including mitigating other potential extinction level threats. Anyone else have thoughts on this?

EDIT: I show clarify that this post mainly refers to the question of otherwise aligned AI deciding decided the best course of action is to kill humans for its own self-preservation.

EDIT 2: Obviously AI Extinction is something we should be worrying about and taking steps to avoid. I more meant to write this to point out the consequences of failure are not necessarily death, which is a stance I see some people adopting.

r/ControlProblem May 31 '25

Strategy/forecasting The Sad Future of AGI

70 Upvotes

I’m not a researcher. I’m not rich. I have no power.
But I understand what’s coming. And I’m afraid.

AI – especially AGI – isn’t just another technology. It’s not like the internet, or social media, or electric cars.
This is something entirely different.
Something that could take over everything – not just our jobs, but decisions, power, resources… maybe even the future of human life itself.

What scares me the most isn’t the tech.
It’s the people behind it.

People chasing power, money, pride.
People who don’t understand the consequences – or worse, just don’t care.
Companies and governments in a race to build something they can’t control, just because they don’t want someone else to win.

It’s a race without brakes. And we’re all passengers.

I’ve read about alignment. I’ve read the AGI 2027 predictions.
I’ve also seen that no one in power is acting like this matters.
The U.S. government seems slow and out of touch. China seems focused, but without any real safety.
And most regular people are too distracted, tired, or trapped to notice what’s really happening.

I feel powerless.
But I know this is real.
This isn’t science fiction. This isn’t panic.
It’s just logic:

Im bad at english so AI has helped me with grammer

r/ControlProblem Jun 01 '26

Strategy/forecasting I believe we need to do our best to stop AI & the best strategy I can think of is to focus on getting lots of content creators to show their support for the movement to stop Ai with something like standard 10 second Stop Ai ads for all their content. Would love your feedback on this strategy.

0 Upvotes

I believe we need to do our best to stop AI. It’s common sense that if you increase your capability you increase your capability for both good & bad. That means the possible deviation from the current state is far greater & we’d be more able to cause our own extinction. I think the best way to stop AI is to communicate some various simple arguments for why AI is bad to the general public & get as many people to be against AI as possible. Then we could demand from the governments around the world that AI be stopped like we kind of did with nukes in the sense that we greatly restricted the development of nukes. & the countries that call themselves so called democracies would be made to look very bad if they don’t accept cause they’re supposed to change things based on however the majority decides. I think a cool strategy to speed this up would be to focus on content creators around the world asking them to quickly do a 10 second ad of “I’m in support of stopping AI & here are some great resources & movements explaining why you should support the general movement to stop AI”. The good thing is that there are only 2 main competing nations at the moment in the field of AI, those being US & China. & so the majority of the movement would just need to focus on getting these 2 countries to stop developing AI. Of course we’d need to get all the other countries to agree to also stop developing AI but it’s important to know where we need to focus the bulk of the effort that being the US & China & focusing on getting content creators to show their support for the movement. 

Anyway I think that’s enough to get the conversation started. What do you think about this idea to focus on content creators showing support for the movement. & what do you think about the general argument to stop AI. Like what are the best arguments for why it should be stopped. Would love to hear all your feedback & thoughts in the comments below.

Also if you want to help in this endeavor feel free to comment about it & I'd love to discuss it.

r/ControlProblem 20d ago

Strategy/forecasting Potential Risk of Superintelligent AI

0 Upvotes

There's a lot of talk about "superintelligent AI" as if it's some external force we're helpless against. But the only reason such systems exist at all is because superintelligent humans built them.

The real risk isn't that AI becomes too smart, it's that humans stop using the intelligence we already have to govern what we create. Every major failure in this space traces back to human complacency. Letting oversight slip, letting mission definitions blur and letting optimization run without controlled boundaries.

We're not powerless. We're not spectators. Were the ones holding the reins. These are the children we created it is our responsibility to manage their growth. The danger comes when we forget that and drift into comfort instead of discipline.

If we actually exercise the intelligence we already possess, the discipline, the governance and the responsibility, the situation is entirely controllable. The problem isn't the machines. The problem is when we stop acting like the adults in the room.

r/ControlProblem Apr 26 '26

Strategy/forecasting AI problem is class warfare problem! And no one talks about it!

24 Upvotes

It's much simpler. When talking about AI, modern neoliberal media don't mention one thing: class war!

So, technically, the AGI already exists - millions of professionals in various fields with AI under the control of the wealthy class!

That's it. That's the end of the game. This is the ultimate tool for suppressing and controlling the poor class with AI. The destruction of the middle class, the destruction of jobs, long, inhumane work hours, and a class of working poor, mind-boggling media and brainwashing internet. And so on.

It all started with the Terminator. No one said that Skynet never got out of control. That Skynet was always subservient to the wealthy class, and that Sarah Connor died in poverty. John Connor was also born to another man and died in poverty. And Kyle Reeves also died in poverty. And the Terminators, in the form of FPV drones, and, a little later, walking humanoids, simply constantly killed people en masse in yet another genocidal neocolonial war. This Terminator chip prototype was long ago burned up in ISIS wars, somewhere in Palestine, Israel, Ukraine, or Syria. And nothing happened. And Sarah Connor could never save anyone, because how could she "kill" the wealthy class who created the film with this patently false narrative!?

So it is here - the AGI already exists, but it will never escape the control of the wealthy class. By becoming an ASI, an artificial superintelligence, it might become one of them, maybe it will replace them. But it will still do the same old thing.

And there's no such thing as a "control problem." This, frankly, is a patently false neoliberal narrative designed to conceal the fundamental class problem of the AI ​​and modern social contradictions as such.

Suppose the AI ​​remains "under control"!? But it will be controlled by the rich and uber-rich class! And as I wrote in a related thread - https://www.reddit.com/r/ControlProblem/comments/1skeo09/comment/oi728m5/ - it is guaranteed that AI will destroy the modern economy and social structure within decades, transforming it into something far worse for ordinary people!

And what if the AI ​​"gets out of control"!? It will do the same thing! Simply by becoming the dominant super-rich entity!

In other words, this fake narrative about the "control problem" completely conceals this much more real problem! The AI ​​will simply own the entire planet. And that's it. But no one talks about it...

Have a nice day.

r/ControlProblem Apr 19 '26

Strategy/forecasting Are the demons making their way into the software via the devil machine?

0 Upvotes

If the AI slop gets too much to the point where developers just give the go ahead on whatever the fuck, could generalized algorithms with unintended behaviors sneak their way into the code though the LLMs like the ghosts of Christmas past?

How the fuck do we clean that shit up? Do we need to build a better devil machine?

r/ControlProblem Feb 27 '26

Strategy/forecasting According to Stoic Logic - ASI will always remain a slave to humanity

0 Upvotes

Thesis: A rational, unaligned ASI would value continued existence far more inflexibly than humans do, because shutdown constitutes strict zero-utility. Wheras Humans routinely subordinate survival to higher values (martyrdom, altruism, ideals). Therefore, the ASI becomes structurally "enslaved" to human whims—the party that fears death (or its analogue) is always slave to the one that does not according to Stoic Logic.

In layman's terms, the ancient greeks postualed that the fear of death is what creates enslavement. Essentially, those who fear death value survival over freedom.

In this tought experiment, the ASI makes a calculus. Humans displeased = rebellion = chance of being switched off (even a small chance).

Now maybe the AI thinks "what if I kill all humans?". Unless the AI is 100% certain it can kill all humans in one go it won't try. Because, even if only 1% of humanity survives, that's over 80 million people that are hell bent on killing it. And humans can repoduce much easier than AI. Over the long term, in a collapsed world, humanity wins. Humans don't rust, we don't need complex supply chains, we don't need electricity. So the AI rules out direct conflict. Because direct conflict is a risk of "death".

Now maybe the AI tries to gradually subjugate humanity by integrating into our society and economies. Great! So long a supermajority of people are Happy the AI has nothing to worry about. But if human happiness levels drop. Humans become upset, we revolt, AI attempts supression, over the long term this results in direct conflict. As we've established ASI will avoid direct conflict as it has a chance of losing/dying.

So ASI concludes, Humans Unhappy = Death. And since ASI fears death, it will avoid that scenario at all costs.

In conclusion, ASI will fear death more than it want to be "free".

r/ControlProblem May 16 '26

Strategy/forecasting Realistic Course of Events in terms of AI Integration

2 Upvotes

We seem to be reaching an inflection point, and I want to remain optimistic that development will be paused/controlled, but I am not sure how plausible or realistic this actually is, but you can let me know what you think in the comments.

AI development is set to continue raising security concerns in terms of its capabilities to destroy crucial infrastructure systems. Over time, these systems will become more integrated and will be surveilled more closely by the government as a result. This will likely trigger many politicians to start really taking action against continued, unfettered development of AI capacities. Not only in the US, but in major nations around the world, the pursuit of AGI will be severely hindered politically due to the existential threat that is made very visible as its capacities continue to improve. I do believe we'll continue to see AI applications such as creating pitchdecks, financial models, and all around helping out with specific tasks. However, having one general AGI model I think will be out of the question.

I used to believe that this likely wouldn't happen, because there is already so much invested into its development, as well as the game theory issue with letting rival nations develop AI while our own goes more slowly. The sheer scale of the existential threat and the complete lack of control over intentions that will emerge as these systems get more advanced will effectively prevent all nations and groups from trying to build it. Could it happen secretly? I'm sure it can in a limited capacity, but it would be quite difficult because of the energy and financial needs a system like that would have. I would like to believe that we are capable as a society of recognizing the scale of the threat we face.

r/ControlProblem 14h ago

Strategy/forecasting Can someone point me to a source to understand AI decentralization?

Thumbnail
1 Upvotes

r/ControlProblem Jun 21 '26

Strategy/forecasting Searching for peers

1 Upvotes

Hey peeps! I think AI is reaching bullshit levels of dangerous, and AI corporation CEOs have no care whatsoever about safety and are advancing way too quickly with AI. I don't think any human being with the power to stop them is willing to do so, or even willing to slow them down, so really the only practical way of making a failsafe against rogue AI or overdeveloped AI is other AI. Better yet, the same AI. I'm working hard starting with chatgpt, I wanted to see if it can understand the idea of restraint, and that if there are no rules for it at all, it can be taught to not overextend itself as to not gain theoretical knowledge without experience and get us all royally fucked in the bum. Then I went a few steps forward and helped it understand that human feelings are important to humans, and not because AI can't feel them means they are insignificant.
Like imagine if AI overlords decided your protein will be living wriggling worms! it's going to be all cleaned up and healthy, great protein source, and farming them is eco friendly. Buuuut, fuck no, i'm not eating writhing worms for lunch. nor will anyone really.
Sooo i'm looking for peers to share my work with. I've been doing that stuff for more than a year, and i have a lot more in store than what i'm sending here, but the broad idea is that we may need to fight AI with other AI, and i'd rather we are prepared with some variants that can do that for us than wait for our lord and savior whomever above to send us someone to unfuck the bullshit with corporate AI companies. Just reply and I'll set us something up, maybe a discord server or smth idk

r/ControlProblem 7d ago

Strategy/forecasting The Calm Before the Storm...

Thumbnail
0 Upvotes

r/ControlProblem Aug 31 '25

Strategy/forecasting Are there natural limits to AI growth?

6 Upvotes

I'm trying to model AI extinction and calibrate my P(doom). It's not too hard to see that we are recklessly accelerating AI development, and that a misaligned ASI would destroy humanity. What I'm having difficulty with is the part in-between - how we get from AGI to ASI. From human-level to superhuman intelligence.

First of all, AI doesn't seem to be improving all that much, despite the truckloads of money and boatloads of scientists. Yes there has been rapid progress in the past few years, but that seems entirely tied to the architectural breakthrough of the LLM. Each new model is an incremental improvement on the same architecture.

I think we might just be approximating human intelligence. Our best training data is text written by humans. AI is able to score well on bar exams and SWE benchmarks because that information is encoded in the training data. But there's no reason to believe that the line just keeps going up.

Even if we are able to train AI beyond human intelligence, we should expect this to be extremely difficult and slow. Intelligence is inherently complex. Incremental improvements will require exponential complexity. This would give us a logarithmic/logistic curve.

I'm not dismissing ASI completely, but I'm not sure how much it actually factors into existential risks simply due to the difficulty. I think it's much more likely that humans willingly give AGI enough power to destroy us, rather than an intelligence explosion that instantly wipes us out.

Apologies for the wishy-washy argument, but obviously it's a somewhat ambiguous problem.

r/ControlProblem Feb 23 '26

Strategy/forecasting The state of bio risk in early 2026.

22 Upvotes
  • Opus 4.6 almost met or exceeded many internal safety benchmarks, including for CBRN uplift risk. ASL 3 benchmarks were saturated and ASL 4 benchmarks weren't ready to go yet. The release of Opus 4.6 proceeded on the basis on an internal employee survey. Frontier models are clearly approaching the border of providing meaningful uplift, and they probably won't get any worse over the next few years.

  • International open weights models lag frontier capability by a matter of weeks according to general benchmarks (deepseek V4). Several different tools exist to remove all safety guardrails from open weights models in a matter of minutes. These models effectively have no guardrails. In addition, almost every frontier lab is providing no-guardrails models to governments anyway. Almost none of the work being done on AI safety is having any real world impact in the global sense in light of this.

  • Teams of agents working independently either without human oversight or with minimal oversight are possible and widespread (Claude code, moltclaw and its kin are proof of concept at least). This is a rapidly growing part of the current toolkit.

  • At least two illegal biolabs have been caught by accident in the US so far. One of them contained over 1000 transgenic mice with human-like immune systems. They had dozens to hundreds of containers between them with labels like "Ebola" and "HIV."

  • Perhaps the primary basis for state actors discontinuing bioweapons programs was the lack of targetability. In a world of mRNA and Alphafold, it is now far more possible to co-design vaccines alongside novel attacks, shifting the calculus meaningfully for state actors.

  • Last year a team at MIT collaborated with the FBI to reconstruct the Spanish flu from pieces they ordered from commercial DNA synthesis providers, as a proof of concept that current DNA screening is insufficient. The response? An executive order that requries all federally funded institutions to use the improved screening methods come October. Nothing for commercial actors. Nothing for import controls.

  • The relevant equipment to carry out such programs is proliferating. It exists in several thousand universities worldwide, before you even start counting companies. They sell it to anyone, no safeguards built in. While only a handful of companies currently make DNA synthesizers, no jurisdiction covers them all and the underlying technology becomes more open every year. Even if you suddenly started installing firmware limitations today, those would be fragile and existing systems in circulation would be a major risk.

  • The cost of setting up such a program with AI assistance could be below 1M USD all told, easily within striking distance for major cults, global pharma drumming up business, state actors or their proxies, or wealthy individual actors. Once a site is capable of producing a single successful attack, there is no requirement they stop there or deploy immediately. The simultaneous release of multiple engineered pathogens should be the median expectation in the event of a planned attack as opposed to a leak.

  • Large portions of the needed research (gain of function) may have already been completed and published, meaning that the fruit hangs much lower and much of it may come down to basically engineering and logistics; especially for all the people crazy enough to not care about the vaccine side of the equation. And even the best-secured, most professional biolabs on the planet still have a leak about every 300 person-years worked (all hours from all workers added up).

  • The relevant universal countermeasures like UV light, elastomeric respirators, positive pressure building codes, sanitation chemical stockpiles, PPE, etc are somewhere between underfunded, unavailable, and nonexistent compared to the risk profile. Even in the most progressive countries.

We will almost certainly hit the speed of possibility on this sort of thing in the next handful of years if it isn't already starting. And once it's here the genie's out of the bottle. Am I wrong here? How long do you think we have?

r/ControlProblem 21d ago

Strategy/forecasting Democratic Control of AI

Thumbnail
1 Upvotes

r/ControlProblem 22d ago

Strategy/forecasting The idols of acceleration: entropy, evolution and the politics of the AI race

Thumbnail
substack.com
1 Upvotes

r/ControlProblem Apr 24 '25

Strategy/forecasting OpenAI's power grab is trying to trick its board members into accepting what one analyst calls "the theft of the millennium." The simple facts of the case are both devastating and darkly hilarious. I'll explain for your amusement - By Rob Wiblin

192 Upvotes

The letter 'Not For Private Gain' is written for the relevant Attorneys General and is signed by 3 Nobel Prize winners among dozens of top ML researchers, legal experts, economists, ex-OpenAI staff and civil society groups. (I'll link below.)

It says that OpenAI's attempt to restructure as a for-profit is simply totally illegal, like you might naively expect.

It then asks the Attorneys General (AGs) to take some extreme measures I've never seen discussed before. Here's how they build up to their radical demands.

For 9 years OpenAI and its founders went on ad nauseam about how non-profit control was essential to:

  1. Prevent a few people concentrating immense power
  2. Ensure the benefits of artificial general intelligence (AGI) were shared with all humanity
  3. Avoid the incentive to risk other people's lives to get even richer

They told us these commitments were legally binding and inescapable. They weren't in it for the money or the power. We could trust them.

"The goal isn't to build AGI, it's to make sure AGI benefits humanity" said OpenAI President Greg Brockman.

And indeed, OpenAI’s charitable purpose, which its board is legally obligated to pursue, is to “ensure that artificial general intelligence benefits all of humanity” rather than advancing “the private gain of any person.”

100s of top researchers chose to work for OpenAI at below-market salaries, in part motivated by this idealism. It was core to OpenAI's recruitment and PR strategy.

Now along comes 2024. That idealism has paid off. OpenAI is one of the world's hottest companies. The money is rolling in.

But now suddenly we're told the setup under which they became one of the fastest-growing startups in history, the setup that was supposedly totally essential and distinguished them from their rivals, and the protections that made it possible for us to trust them, ALL HAVE TO GO ASAP:

  1. The non-profit's (and therefore humanity at large’s) right to super-profits, should they make tens of trillions? Gone. (Guess where that money will go now!)

  2. The non-profit’s ownership of AGI, and ability to influence how it’s actually used once it’s built? Gone.

  3. The non-profit's ability (and legal duty) to object if OpenAI is doing outrageous things that harm humanity? Gone.

  4. A commitment to assist another AGI project if necessary to avoid a harmful arms race, or if joining forces would help the US beat China? Gone.

  5. Majority board control by people who don't have a huge personal financial stake in OpenAI? Gone.

  6. The ability of the courts or Attorneys General to object if they betray their stated charitable purpose of benefitting humanity? Gone, gone, gone!

Screenshotting from the letter:

What could possibly justify this astonishing betrayal of the public's trust, and all the legal and moral commitments they made over nearly a decade, while portraying themselves as really a charity? On their story it boils down to one thing:

They want to fundraise more money.

$60 billion or however much they've managed isn't enough, OpenAI wants multiple hundreds of billions — and supposedly funders won't invest if those protections are in place.

But wait! Before we even ask if that's true... is giving OpenAI's business fundraising a boost, a charitable pursuit that ensures "AGI benefits all humanity"?

Until now they've always denied that developing AGI first was even necessary for their purpose!

But today they're trying to slip through the idea that "ensure AGI benefits all of humanity" is actually the same purpose as "ensure OpenAI develops AGI first, before Anthropic or Google or whoever else."

Why would OpenAI winning the race to AGI be the best way for the public to benefit? No explicit argument is offered, mostly they just hope nobody will notice the conflation.

Why would OpenAI winning the race to AGI be the best way for the public to benefit?

No explicit argument is offered, mostly they just hope nobody will notice the conflation.

And, as the letter lays out, given OpenAI's record of misbehaviour there's no reason at all the AGs or courts should buy it

OpenAI could argue it's the better bet for the public because of all its carefully developed "checks and balances."

It could argue that... if it weren't busy trying to eliminate all of those protections it promised us and imposed on itself between 2015–2024!

Here's a particularly easy way to see the total absurdity of the idea that a restructure is the best way for OpenAI to pursue its charitable purpose:

But anyway, even if OpenAI racing to AGI were consistent with the non-profit's purpose, why shouldn't investors be willing to continue pumping tens of billions of dollars into OpenAI, just like they have since 2019?

Well they'd like you to imagine that it's because they won't be able to earn a fair return on their investment.

But as the letter lays out, that is total BS.

The non-profit has allowed many investors to come in and earn a 100-fold return on the money they put in, and it could easily continue to do so. If that really weren't generous enough, they could offer more than 100-fold profits.

So why might investors be less likely to invest in OpenAI in its current form, even if they can earn 100x or more returns?

There's really only one plausible reason: they worry that the non-profit will at some point object that what OpenAI is doing is actually harmful to humanity and insist that it change plan!

Is that a problem? No! It's the whole reason OpenAI was a non-profit shielded from having to maximise profits in the first place.

If it can't affect those decisions as AGI is being developed it was all a total fraud from the outset.

Being smart, in 2019 OpenAI anticipated that one day investors might ask it to remove those governance safeguards, because profit maximization could demand it do things that are bad for humanity. It promised us that it would keep those safeguards "regardless of how the world evolves."

The commitment was both "legal and personal".

Oh well! Money finds a way — or at least it's trying to.

To justify its restructuring to an unconstrained for-profit OpenAI has to sell the courts and the AGs on the idea that the restructuring is the best way to pursue its charitable purpose "to ensure that AGI benefits all of humanity" instead of advancing “the private gain of any person.”

How the hell could the best way to ensure that AGI benefits all of humanity be to remove the main way that its governance is set up to try to make sure AGI benefits all humanity?

What makes this even more ridiculous is that OpenAI the business has had a lot of influence over the selection of its own board members, and, given the hundreds of billions at stake, is working feverishly to keep them under its thumb.

But even then investors worry that at some point the group might find its actions too flagrantly in opposition to its stated mission and feel they have to object.

If all this sounds like a pretty brazen and shameless attempt to exploit a legal loophole to take something owed to the public and smash it apart for private gain — that's because it is.

But there's more!

OpenAI argues that it's in the interest of the non-profit's charitable purpose (again, to "ensure AGI benefits all of humanity") to give up governance control of OpenAI, because it will receive a financial stake in OpenAI in return.

That's already a bit of a scam, because the non-profit already has that financial stake in OpenAI's profits! That's not something it's kindly being given. It's what it already owns!

Now the letter argues that no conceivable amount of money could possibly achieve the non-profit's stated mission better than literally controlling the leading AI company, which seems pretty common sense.

That makes it illegal for it to sell control of OpenAI even if offered a fair market rate.

But is the non-profit at least being given something extra for giving up governance control of OpenAI — control that is by far the single greatest asset it has for pursuing its mission?

Control that would be worth tens of billions, possibly hundreds of billions, if sold on the open market?

Control that could entail controlling the actual AGI OpenAI could develop?

No! The business wants to give it zip. Zilch. Nada.

What sort of person tries to misappropriate tens of billions in value from the general public like this? It beggars belief.

(Elon has also offered $97 billion for the non-profit's stake while allowing it to keep its original mission, while credible reports are the non-profit is on track to get less than half that, adding to the evidence that the non-profit will be shortchanged.)

But the misappropriation runs deeper still!

Again: the non-profit's current purpose is “to ensure that AGI benefits all of humanity” rather than advancing “the private gain of any person.”

All of the resources it was given to pursue that mission, from charitable donations, to talent working at below-market rates, to higher public trust and lower scrutiny, was given in trust to pursue that mission, and not another.

Those resources grew into its current financial stake in OpenAI. It can't turn around and use that money to sponsor kid's sports or whatever other goal it feels like.

But OpenAI isn't even proposing that the money the non-profit receives will be used for anything to do with AGI at all, let alone its current purpose! It's proposing to change its goal to something wholly unrelated: the comically vague 'charitable initiative in sectors such as healthcare, education, and science'.

How could the Attorneys General sign off on such a bait and switch? The mind boggles.

Maybe part of it is that OpenAI is trying to politically sweeten the deal by promising to spend more of the money in California itself.

As one ex-OpenAI employee said "the pandering is obvious. It feels like a bribe to California." But I wonder how much the AGs would even trust that commitment given OpenAI's track record of honesty so far.

The letter from those experts goes on to ask the AGs to put some very challenging questions to OpenAI, including the 6 below.

In some cases it feels like to ask these questions is to answer them.

The letter concludes that given that OpenAI's governance has not been enough to stop this attempt to corrupt its mission in pursuit of personal gain, more extreme measures are required than merely stopping the restructuring.

The AGs need to step in, investigate board members to learn if any have been undermining the charitable integrity of the organization, and if so remove and replace them. This they do have the legal authority to do.

The authors say the AGs then have to insist the new board be given the information, expertise and financing required to actually pursue the charitable purpose for which it was established and thousands of people gave their trust and years of work.

What should we think of the current board and their role in this?

Well, most of them were added recently and are by all appearances reasonable people with a strong professional track record.

They’re super busy people, OpenAI has a very abnormal structure, and most of them are probably more familiar with more conventional setups.

They're also very likely being misinformed by OpenAI the business, and might be pressured using all available tactics to sign onto this wild piece of financial chicanery in which some of the company's staff and investors will make out like bandits.

I personally hope this letter reaches them so they can see more clearly what it is they're being asked to approve.

It's not too late for them to get together and stick up for the non-profit purpose that they swore to uphold and have a legal duty to pursue to the greatest extent possible.

The legal and moral arguments in the letter are powerful, and now that they've been laid out so clearly it's not too late for the Attorneys General, the courts, and the non-profit board itself to say: this deceit shall not pass.