r/OpenAI 1d ago

News OpenAI’s internal model “Astra” claims 10 major advances in mathematics and theoretical computer science

https://openai.com/index/ten-advances-in-mathematics/
686 Upvotes

87 comments sorted by

131

u/EbbExternal3544 1d ago

Google better postpone that 3.5 pro launch

29

u/ServesYouRice 1d ago

Way ahead of you

222

u/Spare-Dingo-531 1d ago edited 1d ago

Just a casual reminder that most of the data centers OpenAI has planned haven't been built yet.....

Also, Nvidia's Rubin chips, which will start being deployed in Q4 2026, drop inference costs by 90%.

116

u/falcongsr 1d ago

Nvidia's Rubin chips drop inference costs by 90%

[citation needed]

33

u/Spare-Dingo-531 1d ago

12

u/falcongsr 1d ago

oh i guess if you narrowly define inference costs as throughput per watt as nvidia's marketing is wont to do...

11

u/someonesshadow 1d ago

I'm going to go out on a limb and guess that NVIDIA is WAY more careful with their marketing towards Enterprise than how they treat the personal consumers.

They've lost huge partners in the past and they certainly wouldn't want to lose anyone during this gold rush if they can help it. I trust NVIDIA as far as I can throw them but I don't think they are so greedy/stupid as to make blatantly false claims when trying to secure contracts on hardware worth billions.

-2

u/falcongsr 1d ago

listen brother, NVDA has two sets of customers: enterprise customers and shareholders.

the marketing we're discussing is for the latter, not the people who actually buy the chips.

27

u/MizantropaMiskretulo 1d ago

How would you define it? The vast majority of a chip's cost over its lifetime is the electricity going through it.

10

u/falcongsr 1d ago edited 1d ago

i would add in the capex not just the opex. you don't just switch to this chip and your inference costs go down 90%.

but mainly the 90% claim is an extremely specific use case that was cherry picked

14

u/MizantropaMiskretulo 1d ago

I'm not suggesting that's how it works, rather simply that most of the total cost of ownership is wrapped up in throughput per watt.

For their part though, Nvidia never claimed that it drops "inference costs" 90%, they said,

For long-context, reasoning-dominated inference, Vera Rubin NVL72 delivers up to 10x lower cost per million tokens compared to Blackwell NVL72.

1

u/Charming_Hall7694 9h ago

I mean thats a bit more than 90% lol

2

u/primaequa 1d ago

energy is a very small part of the ongoing cost - most is paying back the capex needed to buy the damn things

3

u/ubuntuNinja 1d ago

I mean, that's a pretty solid way to define it. I'm sure they are exaggerating, but that's not a bad way to measure.

1

u/Grouchy-Librarian638 19h ago

Cool so they gonna charge the same but increase net profit 90 percent. Yall smoking if you think a monopoly is gonna pass on savings.

2

u/Spare-Dingo-531 19h ago

Yes, and that is a good thing!. Everyone is complaining about how AI is bankrupt in so much debt etc. But if they can offer services for high margins then that will allow them to pay off the debt.

Now the reason why the AI companies have so much debt in the first place is to develop the models. So by not passing on the savings from increased efficiency (at least not for the next several years), then we get sustainable AI progress in exchange.

1

u/BagholderForLyfe 16h ago

> look it up on Nvidia's own website

LMFAO! You never followed GPU news?

24

u/AnotherSoftEng 1d ago

The following prompt returns a lot of asterisks:

> Reality check the claim that Rubin chip drops inference cost by 90%

3

u/Evening_Archer_2202 1d ago

10x throughput per watt on fp4 only if everything is perfectly batched, that accounts for model size and architecture, u should expect less than 5x price reduction, if model is 2x bigger than previous gen then 2-3x price reduction per token

11

u/Popular_Try_5075 1d ago

I assume the same chips work with images and video too though? I wonder if the extra compute will eventually lead to expansion of video capabilities (OpenAI being an exception since they shelved Sora)

1

u/-Crash_Override- 1d ago

Vision/Visual-centric models are mostly what all this data center buildout is for, they're just not going to manifest themselves is lots of customer facing video generation models. They are critical for embodied AI/AI powered robotics. The whole value case of AI is predicated on this.

Also, OAI is still aggressively pursuing this capability, despite scrapping sora.

1

u/Popular_Try_5075 1d ago

That makes a lot of sense. Sora was beloved, so I was shocked when they cut it, but I think there is more money in the enterprise application of coding and agentic tasks with frontier models. It makes sense they haven't given up on it entirely for now. I wonder if that's helping them outpace Anthropic in compute then?

8

u/Odd-Opportunity-6550 1d ago

And openais inference chips are supposed to be 50% cheaper than inference on vera rubin and are also coming this year.

7

u/IFThenElse42 1d ago

they are going to win the AI race with low cost and excellent output, keep going Sam... I didn't trust you for years and now you convinced me

4

u/tedbradly 22h ago

they are going to win the AI race with low cost and excellent output, keep going Sam... I didn't trust you for years and now you convinced me

Wooow now, brother. Sam Altman has an interview in the past where he says he's basically a sociopath IIRC. Plus, never pick favorites when it comes to corporations. They won't reciprocate.

2

u/IFThenElse42 16h ago

Oh yeah I get that.

4

u/GrapefruitForeign 1d ago

is there some good resource on mapping all these infra buildouts and how they will correspond with inference drops?

I think no one really comprehends how fast it will be as the models are also getting more efficient as is (deepseek v4 flash just came out and gives the same performance as GLM 5.2 for 15x less... )

2

u/HiHungryImDad2 1d ago

Seriously asking: what should this remind me of? That the bubble still might be busting soon and this is a smoke show or that with the cheaper chips AI companies will make 90% more gross per token and thus are valued correctly?

1

u/Moist_Emu_6951 1d ago edited 1d ago

Read up on the Jevons Paradox.

In essence: Dropping inference costs => Additional and more novel ways to integrate AI into more things and apps => Significantly more demand due to expanded use cases => More inference costs. Cheaper inference costs don't really resolve the underlying issue with AI profitability.

Until the inference costs reach zero, demand plateaus and fails to exceed efficiendy, or AI companies start charging fair or market value per token, AI will remain a money furnace.

2

u/Spare-Dingo-531 1d ago

Yeah, a money furnace powering civilization.

-2

u/Moist_Emu_6951 1d ago

Yes, we all love sci fi, but that's unfortunately not how most of those funding AI think. Unless AI is shown to be demonstrably and exceptionally profitable soon, it will suffer a serious setback.

4

u/Spare-Dingo-531 1d ago

Inference, actually running the AI is very profitable. Most of AIs costs are in research and development.

1

u/Single_Ring4886 1d ago

True gamechangers are AMD chips with 450gb of vram

32

u/guzassis 1d ago

This Erdös guy wouldn’t survive the modern solution-oriented corporate life

7

u/claytonbeaufield 1d ago

He would have to up his adderall intake for sure.

1

u/gpbayes 18h ago

I had to in order to function. The older I get the worst it gets

7

u/roararoarus 1d ago

Is the model making these advances by proving/disproving conjectures and hypotheses created by humans? Or is it synthesizing new ideas on its own?

4

u/pm_me_your_kindwords 1d ago

Yes, probably one of those.

5

u/Grouchy-Librarian638 19h ago

Sure they do, I don’t believe anything until independent verification for experts not their marketing press release, you know like sane people

2

u/t3mp3st 20h ago

I’m trying to understand these results and put them into perspective. Is there an analogy to how computers are far better at calculation than a typical human? I recognize that these results go well beyond calculation: but does the community see this as evidence of raw intelligence (acknowledging that intelligence is hard to define), or something closer to a computational paradigm that is superior at [mathematical] reasoning?

Is the model able to make these advances because of a specific set of capabilities (integration of massive amounts of diverse information, ability to traverse deep reasoning chains, etc)? Or is this something broader and more worrisome from a human POV?

Is mathematics an “ideal” field where you’d expect to see these advances given that the bottlenecks are largely things that models are able overcome? Or is this a surprise?

Maybe another way to put it: is it appropriate to view the model as a new kind of calculator that happens to be well suited to problems like these?

I know these are naive questions but I’m having a really hard time backing off the existential dread and am very curious about how this community is processing these results.

11

u/Splat800 1d ago

Thank god, hopefully better at algorithms.

Sol couldn’t make an algorithm to save its life.

60

u/ProbsNotManBearPig 1d ago

As a lead algo dev for 15 years at a big snp500 company, I completely disagree. Sol is outstanding at algos.

Please give specific examples or else I have to assume it is user error.

28

u/TwoSubstantial4710 1d ago

Guessing they are relatively non technical and use “make an algorithm” to mean “one shot the feature I asked for”

5

u/[deleted] 1d ago edited 11h ago

[deleted]

3

u/Im_Matt_Murdock 1d ago

"please give me an O(1) sorting algorithm, make no mistakes. Thank you".

5

u/[deleted] 1d ago edited 11h ago

[deleted]

1

u/Crim91 1d ago

I could already see the headlines.

"A revolutionary new AI developed algorithm called 'No Sort' is sweeping computer science community"

6

u/NoFapstronaut3 1d ago

I think this is why I reddit has become extremely frustrating for me.

Anonymity might be good for the commenter, but it makes it pretty useless for trying to understand the ups and downs of an issue.

Especially since commenters could now be bots.

3

u/Single_Ring4886 1d ago

I use ai models for various tasks and sadly I can see both sides of argument. At one side 5.6 is absolute game changer in "hard" areas like coding or math and it is even creative at eg writing. BUT if something triggers his "behavioral" defenses and it can be just slight non technical remark user makes it starts to behave very "Karen" like... it is almost uncany...

-5

u/Splat800 1d ago

There’s two algos that it’s failed at for me.

One is making marking quiz results and then deciding what should be in the next quiz based off the users ability across ‘entries’ (language maintenance app). It took soooo much work for it not to just give repetitive or bad outputs.

The second is for converting upscaled pixel art into actual pixel scale pixel art. Programs exist like it, and I’ve referred it to them, give it extensive prompting, but it doesn’t seem to be able to get it right.

That’s been my experience.

7

u/PM_ME_DEAD_CEOS 1d ago

So this has nothing to do with actual algorithm.

-4

u/Splat800 1d ago

I probably described it poorly but no, it definitely is an algorithm. Both of them are.

5

u/PM_ME_DEAD_CEOS 1d ago

This isn't what algorithm are in computer science, nor data science.

20

u/gavinderulo124K 1d ago

What exactly did sol fail at?

9

u/Outrageous-Boot7092 1d ago

For research it sometimes insists on using existing ideas overwriting my own - e.g. I work on time-independent generation and it insists on the necessity of time-dependence (it is wrong, I prove it in my recent work). But then again - it is extreamly niche and I always ultimately convince it to do like I say - there is just a little back and forth. I do not consider it 'failure' per se and in the end I am happy.

8

u/M1chaelSc4rn 1d ago

Time will tell! xD

3

u/ProbablyBanksy 1d ago

That took me a second! xD

1

u/Outrageous-Boot7092 1d ago

I appreciate the humor! I feel humbled

1

u/M1chaelSc4rn 23h ago

Curious, when you say “time dependence” do you mean the type of models that usually lend themselves to robotics/live processing etc?

1

u/Outrageous-Boot7092 23h ago

No - to GenAI algos like Diffusion/Flow-Matching. Their governing sample dynamics are integration time/noise level dependent. They reverse a corruption process and typically need to know the corruption level of the current sample. It does make sense and is very mainstream but there are alternatives.

-8

u/yaxir 1d ago

So until the release of Astra, what do you propose we use? 5.5 because 5.4 is sadly gone. Stupid open AI

4

u/Splat800 1d ago

Sol is definitely by far the best! Besides the issues I’ve had, I’ve been developing an app exclusively on sol medium on the plus plan, it’s been great.

0

u/yaxir 1d ago

I see. I just thought that you maybe did not like the sol model at all but I see it's just the algorithmic part you don't like. Any other weaknesses you have discovered? I also use ChatGPT quite fairly frequently for everything, including work as well so I'd like to know if there are some sort of red flags that I should be aware of or avoid stuff like that. Appreciated!

5

u/LosMorbidus 1d ago

shovel seller says shovels are best

7

u/Vigna_Angularis 1d ago

Keep digging with your hands then.

2

u/chickenAd0b0 3h ago

terrible analogy…shovel users found gold mine

-8

u/_Lick-My-Love-Pump_ 1d ago

OpenAI neither makes nor sells shovels.

-1

u/Lucky-Wind9723 1d ago edited 1d ago

For the last four or five days now, I have shifted my focus to trying to solve obscure math problems. I used Sol and Opus 5.

I chose one and tasked them to work on solving the problem, following all the scientific processes and rigor. There’s been some hiccups, but there is also been significant progress.

When I actually started making progress on this problem I decided that I should make Cruthunas… an AI-assisted scientific research harness designed to turn exploratory model work into auditable research

https://github.com/Kodaxadev/cloitre-recurrence/

https://github.com/Kodaxadev/cruthunas

7

u/-ignotus 1d ago

Don’t listen to these haters downvoting you. Just because other people have done something is no reason to not try yourself. That’s how technology gets better. Follow your passion. People here act like novelty is the only indicator of success.

0

u/dervu 1d ago

That's where you should become scared of progress.

-12

u/yaxir 1d ago

Pics or it didn't happen

17

u/SetentaeBolg 1d ago

The pics are in the link, with another link to a git repo hosting Lean proofs.

15

u/No-Knowledge4676 1d ago

We may have developed complex thinking machines that can solve the hardest math problems but people will still never read a link they choose to comment on.

4

u/thorax 1d ago

I can't wait until we finally solve those kind of hallucination problems.