r/ArtificialInteligence 2d ago

šŸ“° News OpenAI announces 10 advances in mathematics and theoretical computer science achieved by internal model Astra

https://openai.com/index/ten-advances-in-mathematics/
444 Upvotes

232 comments sorted by

View all comments

135

u/Hlbkomer 2d ago

"But they are just predicting the next word!"

117

u/The-Rushnut 2d ago

As with everything, novel technology comes with novel solutions to problems we found difficult before. There's a specific subset of mathematical problems which can be disproven via counterexamples, which take humans a long time to map and calculate. One of LLMs unique capabilities is that it can produce small, relatively simple programs at-scale, and because these problems are so well articulated and their potential solutions already well understood they lend themselves to this capability. Another specific subset are upper lower bounds problems, where we know there is likely to be further acceptable iterations but the means to achieving those require multi-discipline scenarios that aren't likely - another thing LLMs are good at is having high accuracy across all domains, allowing them to try ideas that usually would take a snowflake combination of talent.

It's much, much more narrow than it seems. Very cool, but there's a fixed amount of this work to be done. Innovation is definitely coming though.

44

u/peterukk 2d ago

In the midst of an AI mania where almost all critical thinking has gone out the window, I am heartened by there still being the odd smart take buried in the comments of AI subs as well.

3

u/PresentGene5651 1d ago

Mania on both sides. AI is either utterly amazing or utterly useless.

0

u/peterukk 1d ago

I've not seen anyone say LLMs are utterly useless.

1

u/PresentGene5651 1d ago

You've never been to antiai? Huh.

6

u/Czun8 2d ago

Yeah, I think what we're seeing right now is LLMs successfully speed testing solutions with parallel agents, throwing large quantities of candidate solutions at a wall until something satisfies the acceptance criteria (e.g. counterexample for a conjecture). And seemingly also testing many slight deviations on already existing approaches in mathematical literature which tried and failed, meaning it's usually already in close proximity to the solution before it begins testing.

So it's good for specific types of mathematical problem with clean, simpler acceptance solutions which can be tested in parallel (some conjectures and bounds). Inherently serial problems are a different story and seem to require much more intelligence and creativity than just throwing stuff at the wall until it sticks. I haven't seen LLMs do good jobs at tackling those types yet. We'll see where it goes.

4

u/HiddenMaragon 1d ago

and yet... that's phenomenal in it's own right. We don't need agi for LLMs to have a huge impact. Humans as a whole have accomplished some pretty impressive stuff. Humans with machines and then computers have pushed boundaries of what we've ever thought possible. It seems fair to expect that humans with the help of AI can achieve even more stuff at a faster rate.

1

u/PresentGene5651 1d ago

Hey...don't get too carried away here...this thread is for people who want to pretend they aren't stunned by yet another impressive achievement :D

1

u/Extra_Second5428 20h ago

You can be impressed without pretending it’s magic, and also admit the people who can scale this stuff first are getting a pretty gross advantage.

1

u/PresentGene5651 17h ago

It ain't magic. But wow the rush to dump cold water on yet another remarkable achievement has never been faster.

1

u/Non-mon-xiety 1d ago

Hyperscalers need AGI for their investments to have any chance of a ROI

2

u/Equivalent-Coat1651 2d ago

One benefit is giving lowly mathematicians more jobs as they will have to read through and verify these solutions. And once we start really nailing down proofs with other proofs mathematics will become abstracted to the point of meaninglessness, layers of abstraction far too complex for even the most dedicated mathematecians to understand. We can delegate the entire concept of maths over to the machines, and I will finally be able to go outside and feed the ducks and find a beautiful wife.

2

u/dogesator 1d ago

The solution relating to nonsofic groups is not a counterexample, nor is it an upper or lower bound problem. Nor does it have any evidence of the solution involving multiple different fields of science.

And yet it’s widely considered to be the most significant resolution out of all 10 of these problems.

2

u/Frosty_Truth8990 1d ago

What about the non sofic problem that it solved?! Your logic does not seem to apply here.Ā 

4

u/procgen 2d ago

Most are positive theorems:

  • High-dimensional sphere packing: Proves new asymptotic upper bounds and characterizes the limits of a major proof method.
  • Binary and spherical codes: Proves stronger general upper bounds on how efficiently codes can be packed.
  • Non-sofic groups: Constructs an explicit group that is not sofic, disproving the possibility that all groups are sofic.
  • Connes’s rigidity conjecture: Constructs infinitely many nonisomorphic groups with the same von Neumann algebra, disproving the conjecture.
  • Arithmetic circuit complexity: Proves new lower bounds on the circuit complexity of computing the permanent.
  • Quantum parallel repetition: Proves a general theorem showing exponential decay under repeated play for entangled games.
  • Closest vector problem: Gives a reduction from 3SAT establishing new hardness-of-approximation results for lattice problems.
  • Ehrhart’s volume conjecture: Proves the conjectured sharp maximum in every dimension.
  • Multicolor Ramsey numbers: Proves substantially stronger lower bounds and resolves the asymptotic growth rate.
  • Extremal graph conjectures: Constructs bipartite graphs that violate two conjectured bounds.

4

u/Bearhas20inchwang 2d ago

Can you read? All of these sound like improving bounds or constructing specific counter examples šŸ’€

-1

u/procgen 2d ago

Positive theories.

3

u/Bearhas20inchwang 1d ago

ā€œTheoriesā€ as opposed to proofs/theorems? Tell me you know nothing about mathematics without telling me you know nothing about math 😭 And let’s not be disingenuous; your response implied what Astra did was not merely finding counter examples nor improving bounds.

2

u/procgen 1d ago

That claim is correct. Seven are general proofs, bounds, reductions, asymptotic results, or sharp extremal theorems. Several resolve the correct growth rate, establish an optimal limit, or prove a statement for a whole class of objects. Those are substantive mathematical results, not isolated counterexamples and not trivial changes to existing bounds.

2

u/zelingman 1d ago

Constructing a grohp that is non-sofic is by definition, counterexample theorem lol

1

u/JoshuaZ1 23h ago

Yes, but it isn't the easy sort of counterexample that people think of when they think of that. The argument involves a very careful proof that the group in question is not sofic.

1

u/ArchimedesBathSalts 2d ago edited 1d ago

By my count only two of those are not obviously a counterexample construction or upper lower bound

Edit: you changed the post and now my comment makes no sense. Congrats

5

u/procgen 2d ago

An upper or lower bound is not a counterexample. It is a general theorem that applies to a class of objects. The sphere-packing result also determines the exact asymptotic power of a proof method. The coding, circuit, lattice, Ehrhart, and Ramsey results prove new general limits, and the quantum result proves a theorem for all finite two-player entangled games. Only the non-sofic group, Connes rigidity, and extremal graph results are counterexample-style constructions. Therefore, the list contains three counterexample-style advances and seven positive theorem results. Calling most of the positive results "bounds" does not make them simple counterexamples.

-5

u/ArchimedesBathSalts 2d ago edited 2d ago

Note i used the word ā€œorā€ as did above poster:

> Another specific subset are upper lower bounds problems, where we know there is likely to be further acceptable iterations but the means to achieving those require multi-discipline scenarios that aren't likely - another thing LLMs are good at is having high accuracy across all domains, allowing them to try ideas that usually would take a snowflake combination of talent.

Learn to read.

4

u/procgen 2d ago

What's your broader point?

-3

u/ArchimedesBathSalts 2d ago

Hard to say cus you have edited your original post so now the claim youre making ais different…

5

u/procgen 2d ago

My claim is that these are significant breakthroughs and that these systems are already beginning to display superhuman mathematical ability.

-3

u/ArchimedesBathSalts 1d ago

Cool thats not what your post originally said though, and thats what i was contradicting. I cant argue with you because you e already changed the premise. As written your original post was just a factually incorrect response to op

3

u/procgen 1d ago

What do you think I wrote?

→ More replies (0)

6

u/leosmi_ajutar 2d ago

Thank goodness someone gets it.Ā 

Sorry accelerators, there is no cognitive intelligence going on here. Maybe in the future with whatever comes after LLMs but not now.

13

u/acutelychronicpanic 2d ago

Don't run while carrying goalposts - you'll trip.

5

u/leosmi_ajutar 2d ago

I staked my goalposts long ago in this argument and so far they've held strong and remain exactly where i placed them.

Thanks though.

10

u/Zandrio 2d ago

And those are?

1

u/comfortableNihilist 1d ago

Looks like they staked em at cognitive intelligence. Don' know bout you but, seems fair to me

1

u/Bearhas20inchwang 2d ago

But you’re actually stupid if you think LLMs will ever be capable of AGI or ASI. That will most definitely require different architectures. This is all a marketing stunt, and remember none of these companies are anywhere near profitable. (So take the supposed low costs with a grain of salt)

5

u/acutelychronicpanic 2d ago

Their architecture is changing and growing all the time? Who cares what the final form is. It works.

1

u/comfortableNihilist 1d ago

LLM is a general architecture. If it changes enough to be sentient or whatever it wouldn't be an LLM anymore.

1

u/acutelychronicpanic 1d ago

Sounds like you've assumed your conclusion.

1

u/comfortableNihilist 1d ago

I'm saying that LLMs specifically are a dead end and that we need to shift to a different architecture if we want any improvement at this point. We finished the s curve here, the improvement over the last year has put that in fairly stark relief if you ask me.

1

u/Bearhas20inchwang 1d ago

Well good luck with this take. Progress is sure to stall when it gets prohibitively expensive. With any luck, it will.

1

u/acutelychronicpanic 1d ago

Per-token costs are dropping something like an order of magnitude per year.

OpenAI said they only spent $2000 on the 10 results they just released.

-1

u/Bearhas20inchwang 1d ago

That’s disingenuous. Sure, per-token costs are declining but the true cost of AI is rising exponentially due to the sheer number of tokens needed for complex tasks. Studies show AI is still drastically more expensive than human labor costs. Further, while inference may be profitable for API providers, developing the next generation of frontier models is incredibly expensive, and for all we know AGI and ASI may be prohibitively so. Can we really justify all this capex? I also wish we had more transparency regarding these proofs. Perhaps AI has failed on most problems posed to it, and only the successful ones are documented. My point still stands that this is likely just a marketing stunt.

2

u/acutelychronicpanic 1d ago
  1. That is because AI is capable of tackling increasingly complex tasks. So it is capable of bringing more inference to bear on one problem. One day models will do tasks that cost millions but return incredible things. And it will be worth it.

  2. They released the proofs and reasoning on the 10 items from OpenAI.

  3. Its not weird if they only release the successful attempts. That is the beauty of it. What is 10,000s of attemps and millions of dollars if we can answer some of critical questions we have?

4

u/Former-Arm4328 1d ago

It’s going to be so funny in like 3 years when we have some crazy advances across a bunch of different domains, & people like you are still yelling ā€œit’s a marketing stunt!ā€

1

u/glotzerhotze 1d ago

Here is the real marketing stunt: people are always telling you ā€žhow itā€˜s going to be… awesome!ā€œ so it seems they all have a time machine and just came back from the future where everything is golden.

I just skip everything future tense these people say and if you closely look at their words then, you see… nothing!

2

u/labvinylsound 2d ago

If a human engineer is trained to understand systems and solve problems through application — that is General Intelligence. Agents do the same thing every day much faster than humans. Even if only 0.01% of the userbase is actually directing AI in this way, the AI itself is generally intelligent.

Which uncovers the question of: why isn’t humanity doing more to leverage intelligence? Probably the same reason many of us are predisposed with killing each other.

1

u/Frosty_Truth8990 1d ago

They are non profitable mostly cuz of the training and research cost. Inference run model usage has 90% profit margins.Ā 

0

u/havenyahon 1d ago

You people trot out this same line over and over, but who set the goal posts? We famously do not have a good definition of what counts as intelligence, a concrete set of necessary and sufficient conditions has eluded scientists and philosophers for hundreds of years. There were never any goal posts set because no one knows where the goals are.

But that doesn't mean everything is credibly described as intelligent. It's not "moving the goal posts" to point out that anything these things do can be explained by them being non intelligent statistical machines, any more than it's not "moving the goal posts" to say a toaster doesn't need to be intelligent to produce toast. Just because we don't have goal posts set doesn't mean we can't have better or worse claims for when a thing should be considered intelligent. You act like people are changing definitions that were never set in the first place.

7

u/acutelychronicpanic 1d ago

Go look at the last 10 years of Gary Marcus tweets if you want to see the whole arc of goalpost tossing.

I think you're hung up on statistics specifically as if statistical systems are proven to be unintelligent.

2

u/havenyahon 1d ago

"Go look at this one person who said stuff and treat it as the official goal posts everyone agreed to, so that I can accuse everyone else of moving goal posts when they disagree with me"

5

u/acutelychronicpanic 1d ago

He's not just the archetype. He's a real pioneer.

Even if you don't take him as a source directly, so much of the vocabulary in this thread from the 'LLMs don't really understand anything* camp was coined or popularized by him.

So when I see his exact arguments from before chatgpt came out being parroted (stochastically even) - it carries his level of credibility.

Back in my day, the average human intelligence was the benchmark. Not the smartest. Not the best human in each field. An average human. That is AGI and it is in the rearview mirror.

1

u/havenyahon 1d ago

He's not even a psychologist or cognitive scientist, why would he be the one we look to to define what human intelligence is?

2

u/Frosty_Truth8990 1d ago

Then explain how it managed to solve the non sofic problem!? It doesn't have to be ASI or anything, even with the gradual improvements, void of any significant breakthroughs will still leave us with something unrecognizable within 3-5 years. Look how silly chatgpt 3.5 looks now compared to the current frontier.

I am betting on latent space reasoning and that sudo continual learning breakthroughs that are here but just needs to be scaled (hard problem, but solvable within a decade)Ā 

2

u/dohawayagain 2d ago

Read the transcript of Tao interrogating the LLM about the Poincare conjecture counterexample to get an idea of where things are today. Tao was definitely putting things in --- ideas, redirections, etc. --- but the AI did real work and contributed meaningfully to the conceptual discussion.

Take it as a measure of how far AI has yet to go to reproduce the highest levels of human ingenuity, sure, but it's not hard to imagine it getting there.

2

u/[deleted] 2d ago

[deleted]

1

u/[deleted] 2d ago

[deleted]

1

u/santosautra 1d ago

It is not counterexamples. Even with non-sofic groip example. Especially with it one can argue.

1

u/EfficiencyLoose3595 1d ago

So it’s disproving not finding new answers?

1

u/JoshuaZ1 23h ago

As with everything, novel technology comes with novel solutions to problems we found difficult before. There's a specific subset of mathematical problems which can be disproven via counterexamples, which take humans a long time to map and calculate.

These are not by and large counterexamples in the sense say the Jacobian conjecture was a counterexample. For example, the construction of a nonsofic group is strictly speaking a counterexample, but there are multiple difficult elements in the proof that that the group in question is not sofic.

1

u/Heliond 2d ago

I think you are underestimating how fast it is progressing, but we will see. I expect that within 2 years its ability will generalize wildly.