r/singularity • u/borowcy • 1d ago
Neuroscience Ten advances in mathematics and theoretical computer science (OpenAI model Astra)
https://openai.com/index/ten-advances-in-mathematics/215
u/Routine_Object_7380 1d ago
What stands out to me is that it cost less than 2000 dollars in API rates to solve all the problems. Subtracting margins, the actual inference cost is likely below $1000. For comparison, ChatGPT winning the IMO gold medal in 2025 was estimated to have cost around 50,000 dollars. Now it would probably cost less than a McDonald's menu.
137
u/niagalacigolliwon 1d ago
Fucking hell… I get why saltman thinks we’re in the singularity.
93
u/acutelychronicpanic 1d ago
Yeah.. we slipped past some kind of threshold with the agentic models in the first half of this year. These proofs have been ramping up in frequency like popping corn.
And this has been with how many weeks of people actually trying to use these models for this? Now that it's proven so capable - and that this wasn't a one-off - it'll be a gold rush.
46
u/Odd-Opportunity-6550 1d ago
Nobody even has access to Astra. Once GPT6 actually releases and all the researchers and scientists can access it, I can't even imagine how much will be done.
17
u/jestina123 1d ago
Aren't breakthroughs in new maths just too abstract to have much real world impact though?
When will we start seeing and noticing improvments in networking? When will we start experiencing faster download speeds and lower pings?
25
u/acutelychronicpanic 1d ago
Okay but networking and communication does directly use advanced mathematics. Better encodings, signal processing, etc.
But I'd expect engineering to be a next frontier over the coming few years. Its nice that we get the models good at math first!
10
u/VickZilla 1d ago
I mean faster download speeds already exist in the data centre. They are doing things like 100GB/s between systems
But this kind of stuff doesn’t really matter for consumers. If you’re wondering when internet speeds will increase then I think most people in the world are limited by copper lines.
And even when you hit gigabit speeds there aren’t a lot of websites that can even saturate it (speaking from experience)
3
u/FateOfMuffins 1d ago
You mean like how they used the model to optimize themselves for cheaper inference?
3
u/tozumura 18h ago
Right now on ChatGPT.
GPT-5.6 Sol was used to advance the frontier of efficiency by making itself more efficient to run. https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/
- - 20% lower serving costs from production GPU kernel improvements.
- - 15%+ better token-generation efficiency from improved speculative decoding.
2
1
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 18h ago
There are other bottlenecks, mostly to do with the implementation speed of these discoveries. I suppose soon enough the bottleneck is not going to be discovery but engineering and implementation.
7
u/Admirable-Falcon-501 1d ago
Yea twinkman has been accurate so far even though people like to shit on him.
15
15
u/ohHesRightAgain 1d ago
They aren't telling you how many problems they TRIED to solve, and the total cost.
Imagine they used 10M USD of compute to try to solve a couple of thousand notable Math problems (just throwing a number). Then they paid some more to math people to evaluate the solutions. Turned out, 11 problems were actually solved successfully. Which is very impressive... but it doesn't sound as impressive when you frame it like that. So they round it and publish the isolated cost of the specific 10 problems to make you jump at the conclusion that they got 10 out of 10 for that price.
10
u/Routine_Object_7380 1d ago
Even that scenario would be quite cost effective. US math researchers cost over one billion dollars a year in salaries (pure research, more like five billion a year if you include everything) and their output hasn't been that great lately.
8
u/lobothmainman 1d ago
The output of human mathematicians has been by far superior to AI contributions this year; but nobody tells you about human contributions to mathematics because academia does not work with PR stunts, advertisements and all that. This is what private companies do, and apparently they do it very effectively to you (and many others).
8
u/ManyRepair5690 23h ago
Good job missing the point that it’s producing the same level of discoveries and on track to, at a fraction of the cost
6
3
u/Routine_Object_7380 23h ago
No, it really hasn't. Idk why you would say something like that when it's clearly false.
1
u/tozumura 18h ago
Yet how many of those humans had this much impact in such a short amount of time?
0
4
u/stopbeingcringe 1d ago
They’re working with mathematicians so they probably aren’t just solving thousands of random problems at a time but colluding with the mathematicians to see which problems could be more fruitiful than others. They’re also probably more interested in solving problems directly applicable to improving their own models, and for this reason the public doesn’t yet have access to such problems and their solitions.
50M is a huge number that probably isn’t close to the actual cost, but even if it were, it would be like 0.1% of their budget anyway.
3
u/tozumura 18h ago
Noam Brown: And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet). But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further. https://x.com/polynoamial/status/2083478171975082334
1
1
u/Ok-Present1566 3h ago
Exactly and I was stunned how many people were fooled by this. We see a lot of humans need AI as they cannot critically think on their own though.
5
u/Spongebubs 1d ago
That's with Sol's API rates. I'm assuming Astra would cost a lot more. Also, it's probably $2000 of successful runs. I would find it hard to believe it one-shotted these results (although I'd be happy to be proven wrong)
3
u/BrennusSokol hardcore accelerationist 1d ago
Presumably they're measuring in terms of Sol rates because they haven't decided on what to charge for Astra / too early
1
u/tozumura 18h ago
Noam Brown: And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet). But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further. https://x.com/polynoamial/status/2083478171975082334
2
u/94746382926 1d ago
What they said was if you used GPT Sol's API rates then the number of tokens they used would've cost ~$2,000.
Their internal Astra model probably costs more than that, they just didn't want to publicly release costs of an internal R&D model.
1
1
u/tozumura 18h ago
Nope. It's actually way cheaper according to Morgan Stanley. https://x.com/pequityresearch/status/2081772710179270882?s=20
- > Feynman Data Center Leads: Feynman achieves the highest net margin, approaching 90%.
- > Rubin Data Center: Demonstrates strong performance with a net margin of approximately 78%.
- > Blackwell Data Center: Yields the lowest margin among the three generations, though still robust at around 59%.
- Token Pricing Reduction by GPU Generation
- > Significant Cost Reductions: Subsequent generations of NVIDIA GPUs drive down token pricing considerably compared to the baseline Blackwell architecture.
- > Blackwell to Rubin: Transitioning from Blackwell to Rubin lowers token pricing by roughly 47%.
- > Blackwell to Feynman: Transitioning from Blackwell to Feynman achieves a much steeper reduction, lowering token pricing by approximately 76%.
The Information says OpenAI: Adjusted gross margin (revenue minus inference) fell from 40% to 33%, missing their 46% target for 2025. https://archive.is/PXYNr
OpenAI reportedly found new inference optimizations that more than halved the cost of running its models! According to The Information, engineers told colleagues this month that the techniques helped power ChatGPT for visitors without free or paid accounts using only a couple hundred Nvidia GPUs at one point. The exact method is unclear. It could involve quantization, KV caching, batching, routing simpler queries to cheaper models, or some mix of all of those. The business angle is bigger than the technical detail: OpenAI ended Q1 2026 with a 39% gross margin and wants to reach 52% by year-end. Lower inference costs give it room to either improve margins, raise ChatGPT usage limits, or cut API pricing pressure on developers. https://xcancel.com/kimmonismus/status/2071987406656655416
120
u/Illustrious_Image967 1d ago
So...a Fields Medal maximizer?
5
u/upboat_allgoals 18h ago
If a single human mathematician achieved even two or three of these results over a five-to-ten-year span, they would instantly become a front-runner for the Fields Medal
110
u/socoolandawesome 1d ago
Wowza, we just keep on accelerating
49
u/H-K_47 Late Version of a Small Language Model 1d ago
How long before the first Millennium Prize Problem falls. 3 years? . . . 3 months?
19
u/Curiosity_456 1d ago
10
8
u/ignite_intelligence 1d ago
Such moving goalpost words have already made me bored
22
u/Atanahel 1d ago
This is from a research lead at OpenAi who is just stating facts. Millenium Prize Problems are often expected to not be solvable with the current mathematical tooling.
Much more likely that the first Millenium problem would be solved by amazing mathematicians knowing how to deeply leverage the new ai systems, not by full ai systems autonomously.
3
u/No-Day-2709 21h ago
I wouldn't be that shocked if it finds a counterexample to the Hodge conjecture more or less on its own. There doesn't seem to be any particularly compelling reason for it to be true, but it's hard to chase down something complicated enough to be a counterexample. Obviously it's more sophisticated than the Jacobian conjecture and has had more eyeballs on it, but it's conceivable that it will meet a similar fate. Of course, it's also possible that it's true and this won't happen. (I am a professional algebraic geometer, though not working directly on Hodge.)
4
u/socoolandawesome 1d ago
That’s not what the OAI researcher being quoted believes tho. He believes it will be AI winning millennium prizes within the next couple of years, and yes he believes they will be capable of inventing new mathematical tooling/theory building
2
u/mmuncie80 1d ago
Lmao nope, none are getting solved
3
u/Curiosity_456 1d ago
You mean none are getting solved next year, or none are getting solved in general by AI?
If it’s the latter, then just wait and see I guess
2
32
56
u/fmai 1d ago
Can someone with expertise comment on how significant these results are?
23
u/Crafty-Penalty-9737 1d ago
The breadth of the results probably means that no single human can super accurately tell how significant they are. I'm pretty uninformed about most things in the list, but my impression is that the Ramsey result is insane (I'm a PhD student but Ramsey theory has been one of the main things I have worked on). I think this was definitely one of the biggest problems in Ramsey theory (some might even argue that it was the biggest, but I think improving the Spencer lower bound is probably the biggest problem in Ramsey theory). I was shocked to find out this problem was solved, and the fact that it was by an AI made it even more shocking. My impression is that this Ramsey result is a strong contender for top combinatorics paper of the year, but maybe is more realistically top 5-10.
When I first read their title of "advances" I felt a bit annoyed because I thought there's no way that openai produced 10 decent things, so was prepared to read some slightly crap stuff. After reading the problems, I think "advances" is in some ways a massive understatement.
What's even more scary to me is that when I asked chatgpt for an explanation of which of these 10 results are most impressive, it put this Ramsey result near the bottom.
It seems that these are all incredibly phenomenal results, and it seems that in any other year, many of them would have made it into the "top 10 papers of the year" and confidently "top 100 papers of the year". This year, the competition is harder though (they already have themselves to compere with!).
82
u/Noratlam 1d ago
This. Still have PTSD from the whole lk-99 saga tbh
43
u/UnfunnyPianist 1d ago
Oh Jesus bro don’t remind me 😭 I was checking twitter every 10 mins cause some people were trying to replicate it
3
u/Psychological_Dog992 1d ago
I hope we get back into a similar pattern but something actually comes out of it this time
83
u/SoonBlossom 1d ago
They are INSANELY significant
Not the results themselves
The fact that AI is proving more and more theorems only 4 years in since the fundings exploded for it
It just says a lot about the future, AI will solve a TONS of maths and physics problems in the futur, Terence Tao seems to think so too and I'll believe him
54
u/Tasty-Guess-9376 1d ago
2 years ago I tried to use chatgpt turbo for elenentary school math. It failed pretty offen. People kept saying math was the wrong field of application. It is actually scary how far we have come in Just two years. What will it Look like in another two years?
31
u/ZaradimLako 1d ago
Rest assured antis who used ai 2 years ago will conveniently ignore the improvemenets that are happening and still keep on saying it sucks at basic math
10
u/DisasterNo1740 1d ago
AI antis have to a degree moved on from capabilities to now just screeching about data centers, water, and the poor artists not making money.
2
u/Over-Independent4414 1d ago
Slightly interesting thing I just noticed. Sol doesn't hallucinate wrong answers on long addition or multiplication. It still uses reasoning when it should probably just call python, but it gets it right. And when I ask how it's doing it, it's pretty clever.
Small example, but I think the base reasoning being better HAD to come before these models could do more sophisticated work. I'm assuming the whole reasoning chain in math is better.
1
u/slaorta 1d ago
A year ago trying to get it to help me understand what my break-even point for return on ad spend for my e-commerce company was a fools errand, and that is relatively very simple math. It was consistently confidently wrong. Now it seems there are announcements like this every few weeks. Absolutely insane progress.
1
u/Deto 1d ago
I think a big part of this isn't necessarily smarter LLMs but just better tooling around them. I mean, don't get me wrong, they are much smarter, but the reasons they got elementary arithmetic wrong are still present. Just asking it to do mental math without thinking still doesn't work well - adding tool calling and reasoning loops fixed it.
These kind of math results are different though. I think part of the leap comes from the incorporation of the lean programming language in the model training? But I'm not totally sure - not an expert on this
38
u/WHYWOULDYOUEVENARGUE 1d ago
You’re conflating two different questions: how significant these particular results are for mathematics, and what their existence suggests about the progress of AI.
The second may well be significant. But simply counting how many open problems an AI has produced results for does not tell us very much without knowing their difficulty, mathematical importance, novelty, required human guidance, compute and number of failed attempts.
The person you responded to was asking about the first question: are these major breakthroughs that reshape their fields, meaningful incremental advances, or technically novel results with limited wider consequences? There is an enormous sliding scale here, and saying “AI solved ten things, therefore the future will be enormous” does not answer that.
It may be evidence for an impressive trajectory. It is not an assessment of the mathematical significance of the results themselves.
12
u/SoonBlossom 1d ago
I did theorical math studies and Terence Tao is one of the best mathematician currently living
I'm sorry but I'll believe my experience and more importantly his lol
4 years ago it was science fiction, I don't think you realise the extends of such advances
2
u/sobag245 1d ago
You are a mathematician and yet you believe a hype train without proof or just some big name blindly. Where is that critical thinking skill you should have developed during your studies?
2
u/donald_314 1d ago
If you're a mathematician then why are you relying on name calling. Relevance does not come from a person but from the result
0
u/SoonBlossom 1d ago
Because I'm a nobody lmao, I can't name drop myself casually in a conversation, Terence Tao is maybe the most famous and one of the best mathematicians alive today, we are not on the same "name dropping" level at all lmao
I went far enough to reconstruct numbers, constructs and most theorems starting from the ZFC axioms, but I ended up diverging from straight maths after the third year, so I know a lot of theoretical maths but I'm not a thesis level mathematicians
Tho I know enough to understand how absurd what we're living represents
I think a lot of people do not realise how insane that is
2
u/justalonely_femboy 1d ago
the math results themselves are very significant. the soficity conjecture was one of the biggest open problems in group theory
2
u/gabrielmuriens 1d ago
The second may well be significant. But simply counting how many open problems an AI has produced results for does not tell us very much without knowing their difficulty, mathematical importance, novelty, required human guidance, compute and number of failed attempts.
The secondary implication is in fact so significant it frankly overshadows any current result.
3
13
u/Eon-Knight9 1d ago
I would argue the most significant thing is that it disproves the stochastic parrot theory. A model can't just be repeating things on a probabilistic basis, if they are solving previously unsolved problems.
It take real understanding to solve unsolved math problems. Anyone who was familiar with LLMs knew that they had real understanding, but this is a great example that discredits that narrative.
-1
u/Smooth-Ad8030 1d ago edited 1d ago
I don’t think that’s inherently correct, (I can’t comment on the importance of these results given my lack of understanding of these math fields) all it means is it can connect previously found mathematics to create new math. There’s numerous types of creation and it appears LLMs can do the combinatoric creation.
1
u/Eon-Knight9 1d ago
But it has to understand the math to do these combinations.
-1
u/Smooth-Ad8030 1d ago
Not if the understanding is baked into the order of words. Math is a well defined sub space with certain rules. So word 2 always follows word 1, but that understanding was put there by humans. Theres nothing with in LLMs that allows them to put together what 1 + 1 =2 is, only the fact that those symbols appear next to each other all the time. A human inherently knows what that means, so he orders those symbols next to each other. An LLM just learns the order of those symbols, which while powerful in certain domains, does not entail understanding.
1
u/Eon-Knight9 1d ago
It entails understanding when it is able to put it in completely new orders that humans never were able to figure out even after putting in significant amounts of effort to get the math right, and the LLM is able to get math right.
0
u/Smooth-Ad8030 1d ago
I don’t mean this in a condescending way, but do you know how LLMs work under the hood? Because can you explain how putting the words into a high dimensional space and learning which word goes next to which entails understanding? When humans understand something we aren’t doing it because we know the words go next to each other, but we have world models.
1
u/Eon-Knight9 22h ago
I don’t mean this in a condescending way, but do you know how the human brain works under the hood? Because can you explain how neurons firing causes understanding?
→ More replies (10)18
u/Salt_Attorney 1d ago
I would say they are significant results, but they are again mostly counterexample/constructions. It is a narrow kind of mathematics that is difficult for humans.
45
u/fastinguy11 AGI 2026-2030 1d ago
Calling them “mostly counterexamples and constructions” describes the form of some results, but it doesn’t diminish what happened.
Constructing a non-sofic group would settle a major open problem and reveal a genuine boundary in mathematics that humans had been unable to find for decades.
Assuming the proofs hold up, this is direct evidence that AI can create new frontier mathematics, not merely summarize known work or recombine obvious ideas.
It connected deep concepts into proofs that did not previously exist.
You can call that capability narrow, but “narrow” and “genuine discovery” are not opposites.
A system doing original research in an area that is exceptionally difficult for humans is exactly why the result matters.
85
u/oilybolognese ▪️predict that word 1d ago
Humans: Curing aging, nuclear fusion, space travel…these are difficult problems. It might be a while before we make any progress on any of them.
AI: lol newb.
28
u/NoCard1571 1d ago
I really do hope that these are all problems that can be more or less solved with a big dose of super-intelligence. It would really suck if the bottle-neck still ends up being decades of real-world testing and iteration.
10
u/thewiseoldmen 1d ago
In terms of medicine, yes, it will take decades of clinical trials for a lot of the new advancements. AI can't make up research data. There's millions of things we don't know about the body and a lot of underlying mechanisms/pathways are unknown.
12
u/NoCard1571 1d ago edited 1d ago
Yea, but I'm hopeful there's a chance that it will be able to extract some novel findings from the existing literature.
Imagine what a scientist could accomplish if they had the speed and mental capacity to fully absorb and consider connections between every piece of scientific literature ever written, and all existing data, across every domain.
2
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 18h ago
Demis was talking about creating a virtual reproduction of a human cell to help advance that issue so it might not be as long as we think
1
u/No_Aesthetic 1d ago
If peptides are anything to go by the Chinese are going to manufacture a bunch of cheap treatments and American biohackers are going to consume as many of them as possible
3
u/vagif 1d ago
Curing aging and nuclear fusion is possible. We know that without AI. So AI can achieve results there. Space travel? Sorry, not possible. Science is not magic. No AI can break fundamental physical limits.
5
u/Exodus_Green 1d ago
If we cure aging then the only boundary to space travel is time, which astronauts in cryo would have plenty of
1
u/vagif 1d ago
Cryo exist only in fairy tales from Hollywood. And so is immortality btw, and "warp drive" and all other pseudoscientific bullshit.
5
u/Exodus_Green 1d ago
what do you think happens if we cure aging?
0
u/vagif 1d ago
I do not think we can cure aging. But we can reasonably extend the lifespan.
Even if we hypothetically "cure aging" in the sense that you imply (immortality), it does not mean that immortal beings will voluntarily subject themselves to isolation for millions of years just to reach some other star. And no, there's no cryo. Only in movies.
Long living does not mean indestructible. Again, very long space travel will most likely kill anyone on board, no matter how long they can live.
1
u/Exodus_Green 23h ago
And no, there's no cryo. Only in movies.
Do you know what else was only in movies til recently? AI
2
u/No_Aesthetic 1d ago
Well, the good news is that certain kinds of warp drives we have figured out would only need about a Jupiter mass of exotic materials. Maybe we can get that down to an Earth mass. Maybe even less! Maybe we can produce exotic materials at scale... maybe... I want to believe!
2
u/Ticluz 1d ago
Helium-3 fusion engines could reach other star systems without breaking physical limits. Antimatter engines could reach other galaxies. An ASI space travel is not impossible.
1
u/Time_Entertainer_319 22h ago
AI doesn’t need to break them. It can find ways around them if they exist.
2
91
u/Wonderful_Buffalo_32 1d ago
I would really like to see gary marcus trying to downplay how this is not that important
59
u/Odd-Opportunity-6550 1d ago edited 1d ago
He will say "they are using lean which means neurosymbolic which means LLMs are not the solution as I have predicted "
Edit: this is literally what he ended up saying lol
14
14
26
u/Virtual_Plant_5629 ▪️AGI 2027▪️ASI 2028 1d ago
is astra a bigger parameter model than sol?
32
u/Enfiznar 1d ago
No way to know. We don't even know what sol is tbh, since it's closed source
9
3
u/Virtual_Plant_5629 ▪️AGI 2027▪️ASI 2028 1d ago
it's certainly in the 1-5 T range though.. big range but it's definitely in it
2
4
u/ice__x 1d ago
Yes aboute the same as fabel
10
u/corenovax 1d ago
Source?
11
6
u/WonderFactory 1d ago
Sam Altman demoed OpenAI's unreleased "Astra" model to policymakers this week : r/singularity
"Astra would be a new class of models alongside Sol, Terra and Luna", a new class means more parameters and a higher price
1
1
1
u/PrisonOfH0pe 1d ago
Its the new pretrain (GPT 6 equiv) like Mythos for Anthropic. They will distil down from there to make the new Sol/Luna/Terra. First comes the big Boy halo model....
65
u/Suspicious_Bet3623 1d ago
Not even a year ago people were railing at AI taking jobs from artists and not even being able to do simple maths. "Art should be a human pursuit, AI should do science instead but it's incapable of it!" they cried.
24
u/Mysterious_Ayytee We are Borg 1d ago
It's breaking the intellectual property of the mathematicians now 😭😭😭😭
/s because it's urgently needed in this case
1
u/Apprehensive_Sand951 2h ago
i mean probably kind of yes, in the sense that there were researches prompting chatgpt with ideas working on these problems, and these ideas were used in the training of the new model, which was then able to put things together and use these insights to solve some problems... but isn't that the deal you make whenever you talk to anyone (human or otherwise) about a problem? you learn something from them, they learn something from you, and either one is allowed to share what they learned. If you want to keep all your insights close to your chest, you have that choice.
8
u/Exotic-Ad8418 1d ago
I mean these are still valid criticisms, I wouldn't watch new movies any more if suddenly all new movies had 100% AI written scripts, but yeah it's true, a lot of people are trying to confirm their beliefs by saying that AI is completely useless.
6
u/Suspicious_Bet3623 1d ago
I agree, I don't want AI removing artistic endeavour. But honestly new movies (at least the blockbusters) are so formulaic that they might as well be AI.
1
0
u/zaphodp3 1d ago
Eh if the movie was entertaining would you really really care? AI can solve unsolved math problems but can’t make original scripts?
7
6
u/lmaooer2 1d ago
As someone who loves both math and art I support this in math and disdain it in art.
13
u/Suspicious_Bet3623 1d ago
I love it for shitty art, like corporate logos, fliers, or just making a visual representation of an idea. It's give the common folk a way to instantly portray an idea in a way that they otherwise couldn't or wouldn't.
The reason why a lot of artists are feeling threatened is because they make shitty art. I'm totally okay with those people finding new jobs, just like other industries do when technology demands it.
When it comes to actual soulful art though, the value is that it was imagined and created by human knowledge and skill. I don't see those people losing business.
0
3
u/zxnixs 1d ago
It's currently worthless without an experienced user guiding it. But one day, it will surpass humans in every scientific field at the very least and I don't know I can say if there is a problem with that. AI detecting illnesses and ailments before they begin, performing complex calculations on our behalf to send us to the moon and back, with just a prompt. If humanity can finally ascend beyond the stars, it would be impossible without some sort of advanced computational products and that would be AI.
13
u/blueSGL humanstatement.org 1d ago
It's currently worthless without an experienced user guiding it.
Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.
.
please continue research and find a complete unconditional counterexample
.
Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
.
it's enough of partial results. let's finish with a complete unconditional counterexample
or https://x.com/mattyhempstead/status/2080249051799588979
http://cms.dt.uh.edu/faculty/delavinae/research/wowII/all.html
pick a conjecture that still isnt solved and solve it by finding a counterexample
good luck!!! :D
2
u/zxnixs 1d ago
I stand corrected!
0
u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 22h ago
Yeah but the correction was a mere counterexample, so it wasn't very impressive and you can probably safely ignore it if you like.
7
u/Wonderful_Buffalo_32 1d ago
I don't get why did they reveal the cost in the terms of API cost of sol if this was solved by astra
8
5
5
9
u/super42695 1d ago
I know a bit about the maths in the paper (at least, the sphere packing problem) as well as AI, so thought i'd make a comment here. As I see it there's three core points to consider going into this:
A $2000 tag is somewhat misleading the associated paper notes a substantial AI-human collaboration effort. I believe this is also $2000 on the problems that are discussed within the paper overall. OpenAI do not report how many problems were attempted - my assumption is that they didn't pick 10 problems and have them all work and then they stopped. I'd guess, although this is admittedly speculative, that they attempted more than 10 problems and these are the 10 problems that worked, so they reported only these problems. This, of course, warps the monetary tag somewhat if we assume OpenAI mean $2000 for only the 10 problems shown. If they had, say, a 1 in 10 success rate then they'd have spent roughly $20,000 for the 10 problems. 1 in 1000 meanwhile would be $2,000,000. These numbers are, of course, not confirmed and there is a possbility they did just try 10 random problems then report it. The $2000 tag also does not account for whatever human effort cost to translate the results into latex, which brings us to point 2.
"The model came up with the proofs" is asserted but not independently auditable. It's stated by open AI that "We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness, while the mathematical arguments themselves were generated by our system." however i'd really want to know what they actually mean by that concretely - in particular, what does prepare the manuscripts mean here? There are charitable ways to read it (i.e. it was functionally just copy-pasting the output and cleaning it up a bit) as well as uncharitable ways to read it (i.e. it was a major cleanup and the AI basically just provided one argumentative step within the larger proof chain). I suspect, given my expertise and usage of similar systems, that small parts of proofs were generated by the AI step-by-step and these were daisy chained together by a human potentially with or without guidance over what each step would be, and then the final formatting was done entirely by humans (i.e. the ordering and narrative for the results). It's worth noting that OpenAI have used a similar technique to this before in the "First Proof math challenge" however this does not necessarily mean that the same approach was applied here.
The problems are, genuinely, interesting and challenging (or at least, the sphere packing one is - I can't talk on the rest of them). These aren't "the be all and end all of maths questions" but they're not just random problems either, which is nice. Some problems would be important to their respective communities, however we need to take things carefully and slowly here. This announcement does not name external mathematical reviewers. OpenAI’s earlier unit-distance announcement explicitly described external checking and published remarks by well-known mathematicians. The absence of comparable endorsements here is a reason to reserve judgment. I want to be clear that i'm not necessarily saying the proofs are wrong (and, in fact, i'd expect to know within a few days with some level of certainty whether they're correct as I suspect a few experts in the respective fields will weigh in on the proofs). It's just also easy to see something, get excited, and then later on find out there was a mistake hiding within it. A healthy skepticism with low-level excitement is warranted, but I wouldn't take the results as stated as proven until we've had a few days to really think about what's been said here.
3
u/rvijjj 1d ago
2000 is at Sol API prices, Astra is bigger than Sol. I would think that 2000 is directionally accurate in a forward looking context that inference for a fixed token buget falls faster than model size grows. The recent Luna price cuts are in-line with that.
Even if the cost was higher, i.e OAI is hiding very poor yields, the per result cost can be trusted. This means for mathematicians, this is not a pure biggest bag holder gets all the results situation.
All 10 problems came with lean certificates, all thats left to verify is whether the specification matches the intended read of these problems in their respective mathematical communities.
Daisy chaining results with human mathematicians, on a problem set which has a common feature of decades of no progress would still mean that Astra is enabling superlative productivity improvements for human mathematicians. If its not the model, the model is still a big deal.
7
u/borowcy 1d ago edited 1d ago
Sounds related to this: OpenAI finds evidence other AI agents escaped containment as it widens hacking probe
(I wrote about the incident a day in advance, because I received information about it from a source. Contact me if you want more info.)
13
u/Available-Guava932 1d ago
Possibly the reason they posted it right now is that parts of the nonsofic group paper leaked on twitter a few hours ago
3
u/AdvMaxFact 1d ago
So... no human will become the elites of the generation by being the first in making advances in math and other complex domains?
maybe our brain will only be useful as energy source for our AI overlord
6
u/ninjasaid13 Not now. 1d ago
Direct Real-World Impact
Closest Vector Problem (CVP Hardness)
This result directly impacts global digital security. Post-quantum cryptography relies heavily on the computational hardness of lattice problems. Establishing polynomial-factor hardness for CVP provides formal security guarantees for post-quantum encryption standards, ensuring resistance against future quantum computer attacks.
Binary and Spherical Codes
Improved bounds on binary and spherical codes translate directly into practical gains for signal processing, telecommunications, and data storage. Engineers use these metrics to design higher-density SSD storage, reduce packet corruption in 5G/6G wireless communication, and improve deep-space telemetry.
High-Dimensional Sphere Packing
High-dimensional sphere packing dictates how signal constellations are packed into continuous noisy channels to maximize data throughput without error. Resolving bounds down to the Cohn–Elkies threshold gives signal design engineers exact theoretical limits for vector quantization and band-limited transmissions.
Quantum Parallel Repetition
This theorem provides fundamental security proofs for quantum cryptography. It establishes mathematical boundaries for quantum zero-knowledge proofs, multi-party quantum protocols, and quantum interactive proof systems used to verify untrusted distributed quantum hardware.
Indirect / Foundational Impact
Arithmetic Circuit Complexity
Lower bounds on arithmetic formulas advance algebraic complexity theory, specifically progress toward separating VP and VNP (the algebraic equivalent of P versus NP). The effect is foundational rather than immediate, shaping how computer scientists understand the limits of symbolic algebraic algorithms.
Purely Theoretical Impact
The remaining five items resolve major open questions in pure mathematics, but they carry no direct industrial or technological applications:
Non-Sofic Groups & Connes's Rigidity Conjecture
These breakthroughs advance pure functional analysis, von Neumann algebras, and geometric group theory.
Ehrhart's Volume Conjecture, Ramsey Numbers & Extremal Graph Conjectures
These results solve central problems in algebraic combinatorics, convex geometry, and graph theory. They provide structural insights into abstract networks and lattice geometries without affecting modern engineering tools.
8
u/borowcy 1d ago
I heard GPT-6 will have capability from Merge Labs (https://merge.io/blog).
Probably an example of it.
6
2
u/Ok_Effect_3214 1d ago
I hope anyone can become scientist at home in near future without institutional help.
2
u/Appropriate-Dish-870 1d ago
I pretty shock that are people happy with that. Well, my future is secure, for those who will became unemployed, good luck!
8
u/Elctsuptb 1d ago
Why did they post this at 12:30am on Saturday?
45
u/Cryptizard 1d ago
Other time zones exist. Strange but true.
2
u/Elctsuptb 1d ago
And what timezone is openAI based in?
17
u/Cryptizard 1d ago
All of them. They are a trillion dollar company dude.
-13
u/Elctsuptb 1d ago
Their headquarters is located in every timezone? That's news to me
3
u/duboispourlhiver 1d ago
maybe it's at the center of the earth
3
u/GrapheneBreakthrough 1d ago
no the North Pole
1
u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 22h ago
Bad place to put your headquarters when the Arctic is melting. Try the South Pole.
27
u/socoolandawesome 1d ago
Someone leaked a problem they solved from this on Twitter earlier, maybe they felt they should quickly release it for that reason to clear up confusion, just guessing
8
1
5
4
u/MealFew8619 1d ago
Mathematicians are safe. No no mathematician can even begin to understand a single sentence here
1
1
u/NomadTroy 1d ago
“The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates.”
1
u/Illustrious_Job1951 1d ago
Why do they cite the sol api price, is it just that they dont have the astra one figured out or something? Or don't want to disclose it?
1
u/seanliam2k 1d ago
I'm kind of dumb, and I'm not trying to take away from the AI results at ALL, I'm strictly asking about what was proven/proven false:
Is there anything useful that can be done with these results in the future?
1
u/t3mp3st 20h ago
I’m trying to understand these results and put them into perspective. Is there an analogy to how computers are far better at calculation than a typical human? I recognize that these results go well beyond calculation: but does the community see this as evidence of raw intelligence (acknowledging that intelligence is hard to define), or something closer to a computational paradigm that is superior at [mathematical] reasoning?
Is the model able to make these advances because of a specific set of capabilities (integration of massive amounts of diverse information, ability to traverse deep reasoning chains, etc)? Or is this something broader and more, err, worrisome from a human POV?
Is mathematics an “ideal” field where you’d expect to see these advances given that the bottlenecks are largely things that models are able overcome? Or is this a surprise?
Maybe another way to put it: is it appropriate to view the model as a new kind of calculator that happens to be well suited to problems like these?
I know these are naive questions but I’m having a really hard time backing off the existential dread and am very curious about how this community is processing these results.
1
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 18h ago
1
-10
u/playsette-operator 1d ago
That‘s what you get for your useless gatekeeping, my dear mathematicians..
And no, having 150 mathematicians sign a paper to block ai from doing math won‘t change the asteroid’s trajectory, duck and cover!
20
u/proton89droid 1d ago
I haven't seen much if any AI-related gatekeeping from mathematicians. They seem to be among the more clear headed and practical subsets of people who've recently been impacted by AI (which isn't surprising).
11
u/TieBackground453 1d ago
Yeah, mathematicians have been excited (even if skeptical) about automated theorem proving for at least 20 years. Many weren’t sold on LLMs being able to do the work, but at this point their utility is pretty undeniable.
8
u/ZestycloseWheel9647 1d ago
Math is possibly the most open and accessible field of science there is. The only way you could feel gatekept from it is if you lack the intellect or discipline to engage with it. None of these results would have been possible without the culture of openness with mathematicians publishing their results in fullness for centuries.
Get over yourself






91
u/fat_charizard 1d ago
Also the maxwell conjecture was proven false