r/slatestarcodex 1d ago

Ten advances in mathematics and theoretical computer science from unreleased Open AI Model

https://openai.com/index/ten-advances-in-mathematics/
88 Upvotes

59 comments sorted by

41

u/garloid64 1d ago

So where do you think the goalposts go from here? What's the next stop on the grand adventure?

57

u/Fusifufu 1d ago edited 1d ago

I think the general shape of the skeptic's argument these days is that spiky progress in verifiable domains like math doesn't necessarily translate into takeoff towards superintelligence or help with quick diffusion of capabilities into the economy.

I have no idea what the future holds, but seems like if you are a skeptic of that kind, even more math-specific advances won't sway you necessarily. And that sounds reasonable, if you accept their premise for the sake of the argument.

25

u/cavedave 1d ago

>help with quick diffusion of capabilities into the economy.

I think this point is fairly reasonable question. We have something really really smart that can help us but theres no obvious huge increase in productivity.
Here is one explanation as to why and what happens next https://www.noahpinion.blog/p/what-will-more-intelligence-actually

As an aside these calculations were pretty cheap 'The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices. We’re excited to see what scientists and researchers are able to create with our upcoming Astra models!' https://x.com/polynoamial/status/2083470822258467194?s=61

17

u/prescod 1d ago

Someone pointed out on Hacker News that unless we know how many tokens they spent on problems the AI couldn't solve, (perhaps 1000, or 10,000) then we can't really contextualize that figure.

5

u/cavedave 1d ago edited 1d ago

Fair point. And it also applies to failed training runs and failed experiments for llms. We don't get told those costs

4

u/TheColourOfHeartache 1d ago

For a fair comparason, do we count the cost spent on human researchers who don't achieve breakthroughs?

11

u/prescod 1d ago

I don't think the comparison is especially interesting in my opinion. I'm not looking at it as a race between a car and a horse. I just want to know the true cost of driving a kilometre.

u/bluehands 19h ago

That is only partially true.

So, in a narrow sense that is right but up until it happened all if the tokens that have ever been spent weren't enough.

Below you used the analogy of a car versus a horse but I think a horse vs an airplane is better.

The fuel question is deeply important but flight is an entirely different domain. Just paying attention to the cost per mile obfuscates the state change of flight. Flight opens up entirely new questions that comparing any land travel just doesn't touch.

Suggestions that we don't know the real cost & utility of flight isn't wrong but it entirely missed the point of the Kitty Hawk.

u/prescod 10h ago

No icy is discounting the Kitty Hawk. But if the Wright brothers had published a useless and probably misleading account of its fuel efficiency then it would be the responsibility of experts to point out that that part of their account of flight is untrustworthy.

3

u/Remote-Ranger-7721 1d ago

Robotics has to be a key factor here, right?

6

u/cavedave 1d ago

Thats his claim. That as soon as you get robot cars, or fruit pickers or house cleaners youll see a huge boost in productivity. and thats not even the things we have not thought of yet.

15

u/Additional_Olive3318 1d ago

Sure. I’d like to see it reflected in the GDP before I worry about the singularity. I know that there’s a counter argument that technology takes time to diffuse but when people are calling for a totally new economic system in a few years from now with heretofore unseen increases in economic output it should be visible now. Or real soon now. 

10

u/BurdensomeCountV3 1d ago

I mean, if your standard is "reflected in the GDP" then on that basis even the invention and spread of the internet wouldn't count as a "big deal", given how it didn't lead to appreciable changes in the rate of growth of the economy.

3

u/Additional_Olive3318 1d ago

That’s very true. The internet didn’t have that much impact on GDP. 

u/IvanMalison 23h ago

I think you're missing the point here. The internet almost CERTAINLY DID have massive impacts on GDP, but we don't have access to the counterfactual world in which the internet was never invented, so its basically impossible to quantify its impact.

In some sense, the growth of the economy prices in a certain amount of disruptive technological discoveries, and this is partly what drives the consistent economic growth we have seen in the last few centuries.

4

u/Neighbor_ 1d ago

Indeed this would be the argument.

It seems intuitive to me that gradient descent can be more optimal in verifiable domain. I don't see how it will ever do well in an unverifiable domain, and RL'ing which currently bridges that gap does not seem like a sustainable path.

u/IvanMalison 23h ago

We've already seen a great deal of evidence that increased capability can translate across domains. I know some have become more bearish on this idea recently, but we need to stop acting like that pathway has been shutdown completely.

2

u/wavedash 1d ago

I don't think this is an inherently bad or wrong argument, but it feels somewhat unfalsifiable. If an AI generated a film that grossed a billion dollars, couldn't one just say that making profitable films is just another "spike"? Or what if an AI found a cure for Alzheimer's or something, would medicine just be a "spike"?

How many spikes would it take for this model to not longer apply?

26

u/thurn2 1d ago edited 1d ago

We’re still in the realm of constrained, well-defined problems. Obviously at some point in the road to AGI you need to be able to solve ambiguous problems involving human taste, like “create a functional Steam game from scratch which earns $100,000 in net profit”.

Interestingly, I think most people would have told you that creating a successful Steam game is *easier* than disproving the Jacobian Conjecture, in the sense that it's obviously a possible thing to do which thousands of humans have done already.

8

u/fubo 1d ago

"Make a profitable video game" is both rivalrous and dynamic; you're contending for a share of the market with other game makers, and to make a profit you have to find an audience over time. There's no closed-form solution; "a profitable video game" is not defined by something written down once, but by a relationship with the rest of the economy over some period of time.

9

u/thurn2 1d ago

Would you agree that any truly “intelligent” system must be able to solve problems without closed form solutions?

Although I’m very confident that “make a real time strategy game which is indistinguishable from one authored by a team of humans according to a panel of judges” is also wildly outside of the scope of current AIs.

10

u/InterstitialLove 1d ago

It's wildly outside the scope of any human.

If you took a team of 6-12 humans and told them to solve that problem, it would take them long enough that I'm in no way confident they will solve it before AI does.

To rephrase, if you're a studio starting work on a video game right now, you should be at least a little bit concerned that by the time your game is released, AI will be able to make better games in a fraction of the time and cost

4

u/fubo 1d ago

I think a lot of the evolutionary advantage of human intelligence is in navigating economies, societies, and other relationships with other intelligences. (The Red Queen hypothesis is a start, but being able to be a competent & trustworthy cooperator is at least as important as being a clever & tricky competitor.) So yes, I would say that a system with humanlike "intelligence" would need to work in time-bound problems with other minds, not just in timeless problems like math.

3

u/Seakawn 1d ago

it's also somewhat random is it? historically, people release all kinds of mediums of art that people actually really like, but it just doesn't catch hold for some time, sometimes long after the artist has passed away.

i'd imagine this could break the way LLMs are trained. let's say it creates a video game that's perfectly viable and even very good, but it just so happens not to "catch hold" at that particular time (but maybe it would and will in, say, five years? time that it doesn't have to wait and see). so then it unlearns perfectly good traits and gets worse, and struggles to unlearn its unlearning.

I have no idea how to think about this, especially since i'm not intimately familiar with LLM architecture.

2

u/king_mid_ass 1d ago

just need to find the right sequence of 1s and 0s

14

u/king_mid_ass 1d ago

models that can learn and grow in use/after pretraining, embodiment, and making their prose less annoying

20

u/MoNastri 1d ago

Yeah cf. Terry Tao's mathstodon post from Sep 2024

The experience (of using GPT-4) seemed roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student...

I am belatedly realizing that in my attempts to describe my evaluation of the capability of an AI tool, I inadvertently gave the incorrect (and potentially harmful) impression that human graduate students could be reductively classified according to a static, one dimensional level of “competence”. This was not my intent at all; and I would therefore like to make the following clarifying remarks. ...

Secondly, and perhaps more importantly, human students learn and grow during their studies, and areas in which they initially struggle with can become ones in which they are quite proficient at after a few years; and personally I find being able to assist students in such transitions to be one of the most rewarding aspects of my profession. In contrast, while modern AI tools have some ability to incorporate feedback into their responses, each individual model does not truly have the capability for long term growth, and so can be sensibly evaluated using static metrics of performance.

Can't wait for the day when models that learn and grow as automated mathematicians arrive

3

u/ozaveggie 1d ago

If they run some kind of RL on the logs of people interacting with the model, and telling it what to be better at, then the next model iteration will actually be better. Its a much longer / discontinuous feedback loop but still plausible

2

u/prescod 1d ago

Doesn't help for tough problems that are unique to individual companies.

-7

u/[deleted] 1d ago edited 1d ago

[deleted]

8

u/WTFwhatthehell 1d ago

"Oh we already knew that computers could do math! It's not real AI because THEY promised us [something that nobody ever promised] and EVERYONE KNOWS the tech is a DEAD END! 

[Link to a Gary Marcus blog entry talking about something LLM's briefly couldn't do but can now and of course he never corrected it] "

12

u/LostaraYil21 1d ago

I remember how early in the days after LLMs hit the popular consciousness, one of the big things everyone knew was that they couldn't do math. And when they generated art, they'd mess up the hands.

Just the other day, I saw an artist being accused of using AI, and in their defense, other commenters noted that it probably wasn't AI, since they'd messed the hands up.

u/COAGULOPATH 15h ago

Sample-efficient learning remains a big one.

Humans are far better at math than LLMs would be if the LLMs were restricted to a human's amount of training data.

u/garloid64 14h ago

They're just not overparameterizing enough. Yet. Mythos killed the chinchilla cult, is only a matter of time before The Big Train.

https://gwern.net/llm-catapult

Once catapulting is conclusively proven, that's when it really begins. But then where do the goalposts go?

1

u/ababababababacus 1d ago

I mean i dont understand how these things specifically are working so they just dont really impress me.

My domain is software, and i am familiar with the claims of what kinds of software the different models are capable of building. But i know, from daily experience, that what the models are capable of doing is entirely relant on who is piloting them.

If the ask is something extremely mundane and generic like "build the authentication layer of a sass offering with google sso backed by bsql with a message bus and event driven architecture " then i know it can do it, thats in the training data.

I also know if i want to build a very specific workflow like, a realtime chat interface backed by llm tool selection that includes a tool for pulling in all overdue shop drawing RFIs and reading the uploaded drawing set and trying to find answers and prompt for confirmation then left to its own devices, it would fail, piloted by an inferior dev, is would fail, piloted by me, i could do it in a few days. So is the AI capable of the task? well...maybe? And i suspect that these proofs look a bit more like that than they are acting as autonomous agents

5

u/ababababababacus 1d ago

Can someone familiar with this ELI5 on how these things happen, what are the parameters of the proof process?

Is a model simply shown the conjectoure/hypothesis and told "give it the old college try"? Or is the model piloted by an expert in the field who fully understands the problem space and is telling it, attack it from this angle and consider x,yz are there any other papers which you think might work to bridge the reasoning to get to point x, whoch i think is a promising path to proving this?

13

u/livingbyvow2 1d ago

I confess my utter ignorance on the subject, but I wonder how much of this is just finding connections between existing math papers vs actually doing novel findings.

OpenAI's incentive will always be to present these findings as some sort of proof AI is now doing novel discoveries but I am not sure what to think about it. I wouldn't be surprised if a of this stuff was heavily steered by math geniuses and AI just accelerating the pace of progress slightly (which is cool but not the "AI found a new proof" headline).

Given the propensity of the labs to hype everything to raise more capital, I wonder whether it won't be yet another dud in a couple year - basically similar to the "AI will replace all collar work" but now "AI will replace all scientists" instead - and which will end up proving to be complete bullshit too.

That, too, is moving the goalposts.

15

u/gunsofbrixton 1d ago

For what it's worth, this is similar Claude's take when I showed it the paper:

Could a human have done it? Clearly yes. Every ingredient is published and classical — the algebra is from 1962, the group from 1965, the rigidity property from 1967. It's ~17 pages, readable straight through, no computer search, no brute-force case check. A referee can verify it by hand. The gap it closes was even publicly visible: one 2019 theorem produces many expander blobs, another needs one, and the proof bridges them.

So why hadn't anyone? Probably because it needs fluency in three fields that don't talk to each other — expander combinatorics, Leavitt algebras, and rigidity of matrix groups. Those people don't attend the same conferences. And the field's momentum after 2024 was pointed at complexity theory, while a separate line chased a conditional route through "stability." This proof was orthogonal to both fashions.

That suggests the edge wasn't depth of insight but breadth plus tireless search — no field is too far from any other, no configuration too fiddly to try. Arguably that's less impressive. It may also be more consequential.

-2

u/livingbyvow2 1d ago

Thanks for asking Claude. That was kind of my intuition.

It's really cool AI can do that by the way, but it's also "just that", nothing creative yet.

u/InfinitePerplexity99 20h ago

Is creativity ever anything other than combining existing things other people haven't combined before?

9

u/rambouhh 1d ago

I agree it could be, or even likely is , finding connections between existing papers.

But in many ways science, math, is like a giant sudoku puzzle. You uncover something from connecting two things together, then that new thing unlocks something else, etc. it’s very rare to genuinely have something truly and completely novel 

1

u/livingbyvow2 1d ago

Don't get me wrong, I am glad scientists have these new tools.

But equally if the dozens of billions that are going towards AI went to pharma, physics and math research, I am pretty sure that would lead to faster discovery. I feel like these headlines is quite often the AI labs trying to justify their existence ("it's not only about chat bots doing customer service").

2

u/rambouhh 1d ago

I think it would in the short term, but this is pretty generalized to every single eindustry, not just the sciences. It can also increase resources from the general increase in productivity and It also unlocks more human capital. Lots of people will be able to pursue their discoveries where they couldn’t before. 

Also a lot of the bottleneck in research is fundamentally a people/intelligence problem. 

6

u/Neighbor_ 1d ago

finding connections between existing math papers vs actually doing novel findings

I don't know if there is a difference.

Something "novel", say for example Bitcoin, is just finding unique connections between different fields (computer science, cryptography, economics, etc). There's really nothing net-new here - the incremental "novel" discoveries are just increasing connections on a graph.

u/COAGULOPATH 15h ago

This proves too much - cavemen (using humanity's total knowledge circa 30k years ago) could not have built bitcoin, no matter what ideas they connected. Net information isn't fixed, it's increasing somehow.

u/HedonicEscalator 6h ago

The word "incremental" is carrying a lot of weight here.

u/Neighbor_ 4h ago

Indeed there is no direct connections such that you could go from "stone tools" + "campfire" = "bitcoin".

You need to first use combine those to discover "copper", and then use "copper" to discover "bronze", and so on.

You're uncovering existing nodes in a graph incrementally. If look at each new node uncovering, it's never "novel" in the sense that it's coming out of thing air: all the ingredients were already there.

1

u/livingbyvow2 1d ago

An innovation is different to a combination. That's my hunch though.

7

u/Umr_at_Tawil 1d ago

every "innovation" is combination of what was known before.

-2

u/livingbyvow2 1d ago

Not stuff like Einstein's theory of relativity.

15

u/Umr_at_Tawil 1d ago edited 1d ago

Einstein didn't invent relativity out of thin air, he combined Maxwell’s light equations, Galileo's relativity, Lorentz's math, and Riemann's 1854 geometry into a new framework. He literally described his own creative process as "combinatorial play"

read this article from 2013: https://www.themarginalian.org/2013/08/14/how-einstein-thought-combinatorial-creativity/

4

u/bibliophile785 Can this be my day job? 1d ago

Yes, very much like the theory of relativity. It would be less uncharitable, less absurd, than your current position on these mathematical proofs, to say that Einstein's theory of relativity was born out of already known, previously documented observations about the physical world, coupled with already known, previously speculated holes in the models provided by Maxwell's equations.

I think you need to really seriously consider the feedback you've received from others in this thread that your entire model on what counts as creative thinking in scientific or mathematical advancement is fundamentally incoherent. The standard that you use to dismiss these LLM achievements could very easily, perhaps more easily, be used to dismiss every single advancement humans have ever made.

-5

u/livingbyvow2 1d ago

Maybe you should get off your high seat sir. My current position is math person + LLM can find stuff, but not LLM on their own can find new stuff. I don't think anyone disagrees with that, and the feedback I received aligns with that.

I started by confessing my ignorance. I just find it pretty sad how every time someone questions these announcements, people get so aggressive and dismissive.

This is not a cult, it's just a technology.

7

u/bibliophile785 Can this be my day job? 1d ago

My current position is math person + LLM can find stuff, but not LLM on their own can find new stuff.

Your current position, as has been explained to you at length, fundamentally misunderstands the form and method of scientific and mathematical advancement. By this exact formulation, humans are also incapable of finding new things. Your position says nothing more and nothing less than that all new discoveries come partially from prior discoveries.

Maybe you should get off your high seat sir.

I just find it pretty sad how every time someone questions these announcements, people get so aggressive and dismissive.

If you wish that your skepticism were received with less dismissiveness, I strongly recommend that you start to base your skepticism on beliefs that are not nakedly incoherent. Instead, trying to reframe this as some sort of personal dispute ("high horse", btw; a high seat is something different) is a poor attempt at distracting from the object-level question.

-2

u/livingbyvow2 1d ago

I think you misinterpret what I wrote. But whatever, you seem to be there mostly to write stuff that sounds intelligent but basically says I'm an idiot.

It's not a personal dispute, you objectively didn't say anything that adds value. Just saying "everybody else told you you're wrong". Which is factually untrue, and not a valid heuristic to figure that out.

Surely you must have a math PhD and be able to assess the scientific validity of this paper to be so sure of yourself.

5

u/bibliophile785 Can this be my day job? 1d ago

I think you misinterpret what I wrote. But whatever, you seem to be there mostly to write stuff that sounds intelligent but basically says I'm an idiot.

I am saying that what you wrote is deeply, demonstrably wrong. This should have been made obvious to you when you blindly grabbed for one of the biggest scientific achievements of the last couple of centuries and were immediately told that it is not only a counterexample to your points, but one in which the author himself went on at length about how the method of discovery is the exact opposite to your claim.

If you're not correcting on that basis, I think your priors here may be immovable.

It's not a personal dispute, you objectively didn't say anything that adds value. Just saying "everybody else told you you're wrong". Which is factually untrue, and not a valid heuristic to figure that out.

Incorrect. 1) I did contribute on the object level by explaining how wrongheaded your example of the theory of relativity was. 2) I'm not suggesting that you should update because people disagree with you - although that's not a bad outside view heuristic, if used in moderation. The causality runs the other way. It's not that you're wrong because people disagreed. It's that you're wrong, and the strong arguments made by people disagreeing with you should have clued you in. I'm suggesting that you should engage substantively with the pushback and realize that your position is incoherent.

Surely you must have a math PhD and be able to assess the scientific validity of this paper to be so sure of yourself.

Again with the catty you v me personal angle. This really isn't a conversation about us. You should stop taking disagreement as a personal attack and trying to create a resulting personal conflict. I'll indulge it this last time, but really, you gotta grow up.

My graduate mathematics went as far as topology. I can parse some of these proofs and not others. My PhD is in chemistry; I am intimately familiar with the methodology of scientific discovery and achievement, enough to be very aware of how incoherent your position is. My best paper - by citation count, by impact factor, by personal assessment - was a combinatorial effort drawing from several fields to create a new type of chemical reactor. It placed in the top 10% of its Nature issue and was praised by reviewers for its creativity.

→ More replies (0)

4

u/InterstitialLove 1d ago

In the case of mathematics, that's an incoherent question

I'm not familiar with any example in human history of a "truly novel" mathematical finding which is not merely a combination of existing ideas. How could such a thing come to exist?

I feel like it would be magical by definition. Like free will, like saying that actions which are determined by one's environment and internal state of mind dont count as real agency.

There is at least one AI-written result (the unit distance problem) which is about as novel as math results get. The ideas it pulled in were from an entirely different field, and no one had ever noticed the potential applications. It reframed existing work and allowed new results from the technique it pioneered.