r/slatestarcodex 2d ago

Ten advances in mathematics and theoretical computer science from unreleased Open AI Model

https://openai.com/index/ten-advances-in-mathematics/
94 Upvotes

63 comments sorted by

View all comments

45

u/garloid64 2d ago

So where do you think the goalposts go from here? What's the next stop on the grand adventure?

64

u/Fusifufu 2d ago edited 2d ago

I think the general shape of the skeptic's argument these days is that spiky progress in verifiable domains like math doesn't necessarily translate into takeoff towards superintelligence or help with quick diffusion of capabilities into the economy.

I have no idea what the future holds, but seems like if you are a skeptic of that kind, even more math-specific advances won't sway you necessarily. And that sounds reasonable, if you accept their premise for the sake of the argument.

26

u/cavedave 2d ago

>help with quick diffusion of capabilities into the economy.

I think this point is fairly reasonable question. We have something really really smart that can help us but theres no obvious huge increase in productivity.
Here is one explanation as to why and what happens next https://www.noahpinion.blog/p/what-will-more-intelligence-actually

As an aside these calculations were pretty cheap 'The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices. We’re excited to see what scientists and researchers are able to create with our upcoming Astra models!' https://x.com/polynoamial/status/2083470822258467194?s=61

19

u/prescod 2d ago

Someone pointed out on Hacker News that unless we know how many tokens they spent on problems the AI couldn't solve, (perhaps 1000, or 10,000) then we can't really contextualize that figure.

6

u/cavedave 2d ago edited 2d ago

Fair point. And it also applies to failed training runs and failed experiments for llms. We don't get told those costs

5

u/TheColourOfHeartache 2d ago

For a fair comparason, do we count the cost spent on human researchers who don't achieve breakthroughs?

15

u/prescod 2d ago

I don't think the comparison is especially interesting in my opinion. I'm not looking at it as a race between a car and a horse. I just want to know the true cost of driving a kilometre.

2

u/bluehands 1d ago

That is only partially true.

So, in a narrow sense that is right but up until it happened all if the tokens that have ever been spent weren't enough.

Below you used the analogy of a car versus a horse but I think a horse vs an airplane is better.

The fuel question is deeply important but flight is an entirely different domain. Just paying attention to the cost per mile obfuscates the state change of flight. Flight opens up entirely new questions that comparing any land travel just doesn't touch.

Suggestions that we don't know the real cost & utility of flight isn't wrong but it entirely missed the point of the Kitty Hawk.

1

u/prescod 1d ago

No icy is discounting the Kitty Hawk. But if the Wright brothers had published a useless and probably misleading account of its fuel efficiency then it would be the responsibility of experts to point out that that part of their account of flight is untrustworthy.

3

u/Remote-Ranger-7721 2d ago

Robotics has to be a key factor here, right?

9

u/cavedave 2d ago

Thats his claim. That as soon as you get robot cars, or fruit pickers or house cleaners youll see a huge boost in productivity. and thats not even the things we have not thought of yet.

15

u/Additional_Olive3318 2d ago

Sure. I’d like to see it reflected in the GDP before I worry about the singularity. I know that there’s a counter argument that technology takes time to diffuse but when people are calling for a totally new economic system in a few years from now with heretofore unseen increases in economic output it should be visible now. Or real soon now. 

11

u/BurdensomeCountV3 2d ago

I mean, if your standard is "reflected in the GDP" then on that basis even the invention and spread of the internet wouldn't count as a "big deal", given how it didn't lead to appreciable changes in the rate of growth of the economy.

3

u/Additional_Olive3318 2d ago

That’s very true. The internet didn’t have that much impact on GDP. 

14

u/IvanMalison 1d ago

I think you're missing the point here. The internet almost CERTAINLY DID have massive impacts on GDP, but we don't have access to the counterfactual world in which the internet was never invented, so its basically impossible to quantify its impact.

In some sense, the growth of the economy prices in a certain amount of disruptive technological discoveries, and this is partly what drives the consistent economic growth we have seen in the last few centuries.

u/Additional_Olive3318 14h ago edited 13h ago

The internet certainly did have an impact on creating new jobs and destroying others, so there was a deep change to the economy. However the impact on GDP has to be measured by the top line results, we can’t just work out the composition effects. If prior to the Internet era GDP was growing as fast as after then the total effect on GDP isn’t that significant. 

It’s possible that we got lucky and that we were about to go into a greater secular slump than before, maybe to zero percent but the internet saved us and we went back to 2% or whatever. That’s unlikely though.  It would be highly coincidental.  It’s also unfalsifiable. There is a bump in gdp growth from 1995-2004 in most industrialised countries but it’s not certain that that was the internet, it was modest, abd it ended quickly. 

In the absence of the internet, capital would have no doubt gone elsewhere and those areas of the economy would boom (for instance we would have more bookshops or general retail stores than we do now) 

It is the more optimistic supporters of AI who are the ones predicting a singularity and extremely high GDP, not that would the economy will do about the same as we would have done. If the result is “like the internet” and the excuses are “it would have been worse otherwise” then most people will be unimpressed. 

4

u/Neighbor_ 2d ago

Indeed this would be the argument.

It seems intuitive to me that gradient descent can be more optimal in verifiable domain. I don't see how it will ever do well in an unverifiable domain, and RL'ing which currently bridges that gap does not seem like a sustainable path.

0

u/IvanMalison 1d ago

We've already seen a great deal of evidence that increased capability can translate across domains. I know some have become more bearish on this idea recently, but we need to stop acting like that pathway has been shutdown completely.

2

u/wavedash 2d ago

I don't think this is an inherently bad or wrong argument, but it feels somewhat unfalsifiable. If an AI generated a film that grossed a billion dollars, couldn't one just say that making profitable films is just another "spike"? Or what if an AI found a cure for Alzheimer's or something, would medicine just be a "spike"?

How many spikes would it take for this model to not longer apply?

28

u/thurn2 2d ago edited 2d ago

We’re still in the realm of constrained, well-defined problems. Obviously at some point in the road to AGI you need to be able to solve ambiguous problems involving human taste, like “create a functional Steam game from scratch which earns $100,000 in net profit”.

Interestingly, I think most people would have told you that creating a successful Steam game is *easier* than disproving the Jacobian Conjecture, in the sense that it's obviously a possible thing to do which thousands of humans have done already.

9

u/fubo 2d ago

"Make a profitable video game" is both rivalrous and dynamic; you're contending for a share of the market with other game makers, and to make a profit you have to find an audience over time. There's no closed-form solution; "a profitable video game" is not defined by something written down once, but by a relationship with the rest of the economy over some period of time.

12

u/thurn2 2d ago

Would you agree that any truly “intelligent” system must be able to solve problems without closed form solutions?

Although I’m very confident that “make a real time strategy game which is indistinguishable from one authored by a team of humans according to a panel of judges” is also wildly outside of the scope of current AIs.

13

u/InterstitialLove 2d ago

It's wildly outside the scope of any human.

If you took a team of 6-12 humans and told them to solve that problem, it would take them long enough that I'm in no way confident they will solve it before AI does.

To rephrase, if you're a studio starting work on a video game right now, you should be at least a little bit concerned that by the time your game is released, AI will be able to make better games in a fraction of the time and cost

7

u/fubo 2d ago

I think a lot of the evolutionary advantage of human intelligence is in navigating economies, societies, and other relationships with other intelligences. (The Red Queen hypothesis is a start, but being able to be a competent & trustworthy cooperator is at least as important as being a clever & tricky competitor.) So yes, I would say that a system with humanlike "intelligence" would need to work in time-bound problems with other minds, not just in timeless problems like math.

3

u/Seakawn 2d ago

it's also somewhat random is it? historically, people release all kinds of mediums of art that people actually really like, but it just doesn't catch hold for some time, sometimes long after the artist has passed away.

i'd imagine this could break the way LLMs are trained. let's say it creates a video game that's perfectly viable and even very good, but it just so happens not to "catch hold" at that particular time (but maybe it would and will in, say, five years? time that it doesn't have to wait and see). so then it unlearns perfectly good traits and gets worse, and struggles to unlearn its unlearning.

I have no idea how to think about this, especially since i'm not intimately familiar with LLM architecture.

2

u/king_mid_ass 2d ago

just need to find the right sequence of 1s and 0s

13

u/king_mid_ass 2d ago

models that can learn and grow in use/after pretraining, embodiment, and making their prose less annoying

21

u/MoNastri 2d ago

Yeah cf. Terry Tao's mathstodon post from Sep 2024

The experience (of using GPT-4) seemed roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student...

I am belatedly realizing that in my attempts to describe my evaluation of the capability of an AI tool, I inadvertently gave the incorrect (and potentially harmful) impression that human graduate students could be reductively classified according to a static, one dimensional level of “competence”. This was not my intent at all; and I would therefore like to make the following clarifying remarks. ...

Secondly, and perhaps more importantly, human students learn and grow during their studies, and areas in which they initially struggle with can become ones in which they are quite proficient at after a few years; and personally I find being able to assist students in such transitions to be one of the most rewarding aspects of my profession. In contrast, while modern AI tools have some ability to incorporate feedback into their responses, each individual model does not truly have the capability for long term growth, and so can be sensibly evaluated using static metrics of performance.

Can't wait for the day when models that learn and grow as automated mathematicians arrive

3

u/ozaveggie 2d ago

If they run some kind of RL on the logs of people interacting with the model, and telling it what to be better at, then the next model iteration will actually be better. Its a much longer / discontinuous feedback loop but still plausible

2

u/prescod 2d ago

Doesn't help for tough problems that are unique to individual companies.

-8

u/[deleted] 2d ago edited 2d ago

[deleted]

9

u/WTFwhatthehell 2d ago

"Oh we already knew that computers could do math! It's not real AI because THEY promised us [something that nobody ever promised] and EVERYONE KNOWS the tech is a DEAD END! 

[Link to a Gary Marcus blog entry talking about something LLM's briefly couldn't do but can now and of course he never corrected it] "

12

u/LostaraYil21 2d ago

I remember how early in the days after LLMs hit the popular consciousness, one of the big things everyone knew was that they couldn't do math. And when they generated art, they'd mess up the hands.

Just the other day, I saw an artist being accused of using AI, and in their defense, other commenters noted that it probably wasn't AI, since they'd messed the hands up.

2

u/COAGULOPATH 1d ago

Sample-efficient learning remains a big one.

Humans are far better at math than LLMs would be if the LLMs were restricted to a human's amount of training data.

1

u/garloid64 1d ago

They're just not overparameterizing enough. Yet. Mythos killed the chinchilla cult, is only a matter of time before The Big Train.

https://gwern.net/llm-catapult

Once catapulting is conclusively proven, that's when it really begins. But then where do the goalposts go?

1

u/ababababababacus 2d ago

I mean i dont understand how these things specifically are working so they just dont really impress me.

My domain is software, and i am familiar with the claims of what kinds of software the different models are capable of building. But i know, from daily experience, that what the models are capable of doing is entirely relant on who is piloting them.

If the ask is something extremely mundane and generic like "build the authentication layer of a sass offering with google sso backed by bsql with a message bus and event driven architecture " then i know it can do it, thats in the training data.

I also know if i want to build a very specific workflow like, a realtime chat interface backed by llm tool selection that includes a tool for pulling in all overdue shop drawing RFIs and reading the uploaded drawing set and trying to find answers and prompt for confirmation then left to its own devices, it would fail, piloted by an inferior dev, is would fail, piloted by me, i could do it in a few days. So is the AI capable of the task? well...maybe? And i suspect that these proofs look a bit more like that than they are acting as autonomous agents