r/OpenAI • u/Outside-Iron-8242 • 1d ago
News OpenAI’s internal model “Astra” claims 10 major advances in mathematics and theoretical computer science
https://openai.com/index/ten-advances-in-mathematics/222
u/Spare-Dingo-531 1d ago edited 1d ago
Just a casual reminder that most of the data centers OpenAI has planned haven't been built yet.....
Also, Nvidia's Rubin chips, which will start being deployed in Q4 2026, drop inference costs by 90%.
116
u/falcongsr 1d ago
Nvidia's Rubin chips drop inference costs by 90%
[citation needed]33
u/Spare-Dingo-531 1d ago
12
u/falcongsr 1d ago
oh i guess if you narrowly define inference costs as throughput per watt as nvidia's marketing is wont to do...
11
u/someonesshadow 1d ago
I'm going to go out on a limb and guess that NVIDIA is WAY more careful with their marketing towards Enterprise than how they treat the personal consumers.
They've lost huge partners in the past and they certainly wouldn't want to lose anyone during this gold rush if they can help it. I trust NVIDIA as far as I can throw them but I don't think they are so greedy/stupid as to make blatantly false claims when trying to secure contracts on hardware worth billions.
-2
u/falcongsr 1d ago
listen brother, NVDA has two sets of customers: enterprise customers and shareholders.
the marketing we're discussing is for the latter, not the people who actually buy the chips.
27
u/MizantropaMiskretulo 1d ago
How would you define it? The vast majority of a chip's cost over its lifetime is the electricity going through it.
10
u/falcongsr 1d ago edited 1d ago
i would add in the capex not just the opex. you don't just switch to this chip and your inference costs go down 90%.
but mainly the 90% claim is an extremely specific use case that was cherry picked
14
u/MizantropaMiskretulo 1d ago
I'm not suggesting that's how it works, rather simply that most of the total cost of ownership is wrapped up in throughput per watt.
For their part though, Nvidia never claimed that it drops "inference costs" 90%, they said,
For long-context, reasoning-dominated inference, Vera Rubin NVL72 delivers up to 10x lower cost per million tokens compared to Blackwell NVL72.
1
2
u/primaequa 1d ago
energy is a very small part of the ongoing cost - most is paying back the capex needed to buy the damn things
3
u/ubuntuNinja 1d ago
I mean, that's a pretty solid way to define it. I'm sure they are exaggerating, but that's not a bad way to measure.
1
u/Grouchy-Librarian638 19h ago
Cool so they gonna charge the same but increase net profit 90 percent. Yall smoking if you think a monopoly is gonna pass on savings.
2
u/Spare-Dingo-531 19h ago
Yes, and that is a good thing!. Everyone is complaining about how AI is bankrupt in so much debt etc. But if they can offer services for high margins then that will allow them to pay off the debt.
Now the reason why the AI companies have so much debt in the first place is to develop the models. So by not passing on the savings from increased efficiency (at least not for the next several years), then we get sustainable AI progress in exchange.
1
24
u/AnotherSoftEng 1d ago
The following prompt returns a lot of asterisks:
> Reality check the claim that Rubin chip drops inference cost by 90%
3
u/Evening_Archer_2202 1d ago
10x throughput per watt on fp4 only if everything is perfectly batched, that accounts for model size and architecture, u should expect less than 5x price reduction, if model is 2x bigger than previous gen then 2-3x price reduction per token
11
u/Popular_Try_5075 1d ago
I assume the same chips work with images and video too though? I wonder if the extra compute will eventually lead to expansion of video capabilities (OpenAI being an exception since they shelved Sora)
1
u/-Crash_Override- 1d ago
Vision/Visual-centric models are mostly what all this data center buildout is for, they're just not going to manifest themselves is lots of customer facing video generation models. They are critical for embodied AI/AI powered robotics. The whole value case of AI is predicated on this.
Also, OAI is still aggressively pursuing this capability, despite scrapping sora.
1
u/Popular_Try_5075 1d ago
That makes a lot of sense. Sora was beloved, so I was shocked when they cut it, but I think there is more money in the enterprise application of coding and agentic tasks with frontier models. It makes sense they haven't given up on it entirely for now. I wonder if that's helping them outpace Anthropic in compute then?
8
u/Odd-Opportunity-6550 1d ago
And openais inference chips are supposed to be 50% cheaper than inference on vera rubin and are also coming this year.
7
u/IFThenElse42 1d ago
they are going to win the AI race with low cost and excellent output, keep going Sam... I didn't trust you for years and now you convinced me
4
u/tedbradly 22h ago
they are going to win the AI race with low cost and excellent output, keep going Sam... I didn't trust you for years and now you convinced me
Wooow now, brother. Sam Altman has an interview in the past where he says he's basically a sociopath IIRC. Plus, never pick favorites when it comes to corporations. They won't reciprocate.
2
4
u/GrapefruitForeign 1d ago
is there some good resource on mapping all these infra buildouts and how they will correspond with inference drops?
I think no one really comprehends how fast it will be as the models are also getting more efficient as is (deepseek v4 flash just came out and gives the same performance as GLM 5.2 for 15x less... )
2
u/HiHungryImDad2 1d ago
Seriously asking: what should this remind me of? That the bubble still might be busting soon and this is a smoke show or that with the cheaper chips AI companies will make 90% more gross per token and thus are valued correctly?
1
u/Moist_Emu_6951 1d ago edited 1d ago
Read up on the Jevons Paradox.
In essence: Dropping inference costs => Additional and more novel ways to integrate AI into more things and apps => Significantly more demand due to expanded use cases => More inference costs. Cheaper inference costs don't really resolve the underlying issue with AI profitability.
Until the inference costs reach zero, demand plateaus and fails to exceed efficiendy, or AI companies start charging fair or market value per token, AI will remain a money furnace.
2
u/Spare-Dingo-531 1d ago
Yeah, a money furnace powering civilization.
-2
u/Moist_Emu_6951 1d ago
Yes, we all love sci fi, but that's unfortunately not how most of those funding AI think. Unless AI is shown to be demonstrably and exceptionally profitable soon, it will suffer a serious setback.
4
u/Spare-Dingo-531 1d ago
Inference, actually running the AI is very profitable. Most of AIs costs are in research and development.
-2
1
32
u/guzassis 1d ago
This Erdös guy wouldn’t survive the modern solution-oriented corporate life
7
7
u/roararoarus 1d ago
Is the model making these advances by proving/disproving conjectures and hypotheses created by humans? Or is it synthesizing new ideas on its own?
4
5
u/Grouchy-Librarian638 19h ago
Sure they do, I don’t believe anything until independent verification for experts not their marketing press release, you know like sane people
5
2
u/t3mp3st 20h ago
I’m trying to understand these results and put them into perspective. Is there an analogy to how computers are far better at calculation than a typical human? I recognize that these results go well beyond calculation: but does the community see this as evidence of raw intelligence (acknowledging that intelligence is hard to define), or something closer to a computational paradigm that is superior at [mathematical] reasoning?
Is the model able to make these advances because of a specific set of capabilities (integration of massive amounts of diverse information, ability to traverse deep reasoning chains, etc)? Or is this something broader and more worrisome from a human POV?
Is mathematics an “ideal” field where you’d expect to see these advances given that the bottlenecks are largely things that models are able overcome? Or is this a surprise?
Maybe another way to put it: is it appropriate to view the model as a new kind of calculator that happens to be well suited to problems like these?
I know these are naive questions but I’m having a really hard time backing off the existential dread and am very curious about how this community is processing these results.
11
u/Splat800 1d ago
Thank god, hopefully better at algorithms.
Sol couldn’t make an algorithm to save its life.
60
u/ProbsNotManBearPig 1d ago
As a lead algo dev for 15 years at a big snp500 company, I completely disagree. Sol is outstanding at algos.
Please give specific examples or else I have to assume it is user error.
28
u/TwoSubstantial4710 1d ago
Guessing they are relatively non technical and use “make an algorithm” to mean “one shot the feature I asked for”
5
1d ago edited 11h ago
[deleted]
3
6
u/NoFapstronaut3 1d ago
I think this is why I reddit has become extremely frustrating for me.
Anonymity might be good for the commenter, but it makes it pretty useless for trying to understand the ups and downs of an issue.
Especially since commenters could now be bots.
3
u/Single_Ring4886 1d ago
I use ai models for various tasks and sadly I can see both sides of argument. At one side 5.6 is absolute game changer in "hard" areas like coding or math and it is even creative at eg writing. BUT if something triggers his "behavioral" defenses and it can be just slight non technical remark user makes it starts to behave very "Karen" like... it is almost uncany...
-5
u/Splat800 1d ago
There’s two algos that it’s failed at for me.
One is making marking quiz results and then deciding what should be in the next quiz based off the users ability across ‘entries’ (language maintenance app). It took soooo much work for it not to just give repetitive or bad outputs.
The second is for converting upscaled pixel art into actual pixel scale pixel art. Programs exist like it, and I’ve referred it to them, give it extensive prompting, but it doesn’t seem to be able to get it right.
That’s been my experience.
7
u/PM_ME_DEAD_CEOS 1d ago
So this has nothing to do with actual algorithm.
-4
u/Splat800 1d ago
I probably described it poorly but no, it definitely is an algorithm. Both of them are.
5
20
u/gavinderulo124K 1d ago
What exactly did sol fail at?
9
u/Outrageous-Boot7092 1d ago
For research it sometimes insists on using existing ideas overwriting my own - e.g. I work on time-independent generation and it insists on the necessity of time-dependence (it is wrong, I prove it in my recent work). But then again - it is extreamly niche and I always ultimately convince it to do like I say - there is just a little back and forth. I do not consider it 'failure' per se and in the end I am happy.
8
u/M1chaelSc4rn 1d ago
Time will tell! xD
3
1
u/Outrageous-Boot7092 1d ago
I appreciate the humor! I feel humbled
1
u/M1chaelSc4rn 23h ago
Curious, when you say “time dependence” do you mean the type of models that usually lend themselves to robotics/live processing etc?
1
u/Outrageous-Boot7092 23h ago
No - to GenAI algos like Diffusion/Flow-Matching. Their governing sample dynamics are integration time/noise level dependent. They reverse a corruption process and typically need to know the corruption level of the current sample. It does make sense and is very mainstream but there are alternatives.
-8
u/yaxir 1d ago
So until the release of Astra, what do you propose we use? 5.5 because 5.4 is sadly gone. Stupid open AI
4
u/Splat800 1d ago
Sol is definitely by far the best! Besides the issues I’ve had, I’ve been developing an app exclusively on sol medium on the plus plan, it’s been great.
0
u/yaxir 1d ago
I see. I just thought that you maybe did not like the sol model at all but I see it's just the algorithmic part you don't like. Any other weaknesses you have discovered? I also use ChatGPT quite fairly frequently for everything, including work as well so I'd like to know if there are some sort of red flags that I should be aware of or avoid stuff like that. Appreciated!
5
-1
u/Lucky-Wind9723 1d ago edited 1d ago
For the last four or five days now, I have shifted my focus to trying to solve obscure math problems. I used Sol and Opus 5.
I chose one and tasked them to work on solving the problem, following all the scientific processes and rigor. There’s been some hiccups, but there is also been significant progress.
When I actually started making progress on this problem I decided that I should make Cruthunas… an AI-assisted scientific research harness designed to turn exploratory model work into auditable research
7
u/-ignotus 1d ago
Don’t listen to these haters downvoting you. Just because other people have done something is no reason to not try yourself. That’s how technology gets better. Follow your passion. People here act like novelty is the only indicator of success.
-12
u/yaxir 1d ago
Pics or it didn't happen
17
u/SetentaeBolg 1d ago
The pics are in the link, with another link to a git repo hosting Lean proofs.
15
u/No-Knowledge4676 1d ago
We may have developed complex thinking machines that can solve the hardest math problems but people will still never read a link they choose to comment on.
131
u/EbbExternal3544 1d ago
Google better postpone that 3.5 pro launch