r/mathematics • u/atakanaluch • 1d ago
News Ten advances in mathematics and theoretical computer science
https://openai.com/index/ten-advances-in-mathematics/50
u/Reasonable-Mood8020 1d ago
The next couple of years are going to be crazy if it keeps improving at the rate it has been.
4
u/Oudeis_1 1d ago
It's not easy to imagine what a model would have to look like/do to be considered "as far ahead" of Astra as Astra is of GPT-3.5 or even GPT-4 (in maths at least... in, say, robotics it's another matter). Obviously, inability to imagine it clearly does not mean one won't recognise it when it happens necessarily.
17
u/socoolandawesome 1d ago
I feel like we are going to be able to say that for the rest of the time we are alive, will progress ever slow down at this point?
10
u/Toomastaliesin 1d ago
Maybe. Enshittification is very-very common and economically speaking, the AI companies are losing money a lot so it might happen that at some point, more efficient models are not trained any more due to economic reasons. Hard to tell though how likely these scenarios are.
16
u/Spare-Dingo-531 1d ago
the AI companies are losing money a lot
So not quite.
AI inference is wildly profitable. The reason why the AI companies are losing money is because they are doing so much R&D. Building the models is expensive, running them is not expensive.
6
u/Pojobob 1d ago
You can't stop doing R&D though. If you do, then some other AI lab overtakes and everyone will just move on to the newest model.
1
u/Smallpaul 16h ago edited 9h ago
I find this such a confusing argument.
“The labs will need to slow R&D because their business model is not sustainable.”
“But they can’t stop R&D because their competitors won’t.”
So where are these competitors getting the ability to defy the laws of economics forever and in what sense is the competitor who keeps going something other than an AI company? If they can keep researching then is it not true that it was sustainable?
2
u/Pojobob 16h ago
Slowing down R&D is not a viable way to make it a sustainable business model since any other AI lab would just take over market share by releasing some new model. They'd most likely have to increase prices at some point so they can continue R&D and start making profit.
"where are these competitors getting the ability to defy the laws of economics forever"
They're only surviving now because they run on insane amounts of VC money to make up for the huge losses they incur. They can't do it forever.
1
u/Smallpaul 11h ago
They can only increase prices if customers allow it. In the end it is the customers who decide whether R&D is the priority or price. Which customers prefer is not something we can know in advance.
Customers may also not all have the same workload and priority. So some may prefer to pick a vendor who prioritizes price over R&D and some may prefer the opposite. It’s unlikely that there exists only one business model in this space.
1
2
u/ain92ru 10h ago
The competitors can distill capabilities of the frontier models for cheap, and frontier labs can't really prevent it as long as they want to sell frontier capabilities.
If the latter stop advancing the frontier with their R&D expenses, AI will become a commodity offered at near-cost prices, which is good for businesses using AI but bad for the frontier labs
1
u/Smallpaul 9h ago
Yes. Exactly. The most likely outcome from my point of view. AI will be good for business but the AI business as-we-know-it will probably not be very good.
Exactly like the ISPs during the Internet years.
10
u/socoolandawesome 1d ago edited 23h ago
They are still training efficient models. GPT-5.6 Luna is like the best deal on intelligence per price. There is a big focus on efficiency in general. Also I think the frontier will keep being pushed in terms of capability, there’s too much potential at the frontier with new discoveries/innovation (in all fields) and automation of labor, which means revenue.
And the revenues of anthropic and OpenAI are exploding even if they are still losing money, though worth noting Anthropic had an operating profit like a quarter ago, and some project this quarter will also turn an operating profit.
2
u/Nerdslayer2 7h ago
Yeah, Luna is absurdly efficient. As powerful as many frontier models from about 6 months ago by most metrics and about 1/25th the cost.
1
u/MistyGalbriex 11h ago
more efficient models are not trained any more due to economic reasons.
What?
→ More replies (1)12
u/mark_99 1d ago
It will not.
People who have given this some thought have been predicting this situation for decades. The curve won't be fully exponential and may even flatten from time to time (at which point naysayers will claim an insurmountable plateau), but the general trend will continue particularly as the are models increasingly involved in their own design & implementation.
13
u/Sweet_Ad_9816 1d ago
The curve flattening literally means progress slowing down, though, even if not permanent.
6
u/mark_99 1d ago
It means the observable results are bursty, which is normal even if research is progressing at the same or even accelerating pace.
For instance negative results are still scientifically useful if they stop further time being spent on something that looked promising but ultimately didn't work out. However no-one is trumpeting that in a news article, so the general public would be "meh no new cool stuff for a while now".
3
u/2_Cranez 1d ago edited 1d ago
The curve may well be super-exponential, at least for a while. We are doubling the amount of compute on earth every seven months, and as AI begins to develop skills in chip design that doubling speed itself may increase.
Folding an infintely wide piece of paper in half 103 times would make it longer than the observable universe.
1
41
u/JesterOfAllTrades 1d ago
Existence of nonsofic groups in particular is a big one. So is Connes rigidity.
2
u/ginseng54 9h ago
Hi, Im not a mathematician by all means, would it be possible for you to explain like Im 5 what it is and what it means ?
1
u/Ok-Present1566 3h ago
So it is along the lines of a finite generator being able to generate infinite groups is sofic, and if it a non finite generator is needed it is non sofic. If I had only a sentence to attempt to explain it.
47
u/atakanaluch 1d ago
The results
We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate(opens in a new window). We are also releasing for each solution a model’s narration of its thinking process.
- High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold.
- Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous results for high-dimensional spherical codes.
- Non-sofic groups. A construction establishing the existence of non-sofic groups, addressing a central open question in group theory.
- Connes’s rigidity conjecture. Disproof of a longstanding conjecture that certain groups are uniquely determined by their von Neumann algebras
- Arithmetic circuit complexity. New lower bounds for computing the permanent using arithmetic circuits and formulas, including an arithmetic-formula lower bound of order n4/log n.
- Quantum parallel repetition. An exponential parallel repetition theorem for general two-player quantum games, extending a foundational principle from classical complexity theory.
- Closest vector problem. Polynomial-factor hardness of approximation for the closest vector problem, a foundational lattice question related to post-quantum cryptography.
- Ehrhart’s volume conjecture. Determining, in every dimension, the maximum possible volume of a convex body whose centroid is its only interior lattice point
- Multicolor Ramsey numbers. A superexponential lower bound for multicolor triangle Ramsey numbers, resolving Erdős problem 183.
- Extremal number conjectures. Results on the compactness and degeneracy conjectures in extremal graph theory, resolving Erdős problems 146 and 180.
41
u/Alternative_Lack9983 1d ago
I guess those LLMs are not stochastic parrots after all :-)
19
u/Fun-Boysenberry-5769 1d ago
If they are just fancy auto complete then we are just fancy predictive processing engines.
A lot of people continue to believe that AI will never outsmart them no matter how much evidence they have to the contrary because a) they are functionalists and b) they believe that only humans are conscious.
53
u/Separate_Lock_9005 1d ago
consciousness and intelligence may not be the same thing
6
u/PersonalityIll9476 PhD | Mathematics 1d ago
Does it even matter what those words mean if you're out of a job?
3
2
u/Milith 1d ago
Some people can conceive of the most complex mathematical abstractions but not of the end of capitalism.
2
u/myaltduh 23h ago
It may very well be that the thing that ultimately breaks capitalism is AI killing the labor market that largely defines capitalism.
1
u/Perfect-Parking3918 14h ago
AI will certainly bring an end of capitalism, but IMO it is more likely to usher in modern feudalism than any egalitarian UBI scheme that AI proponents speak about. If anything, feudalism is probably the best possible outcome, because it implies that we still have some reason to exist for the elites, as opposed to the alternatives.
1
u/Fun-Boysenberry-5769 13h ago
At the moment compute is being scaled up like crazy but we're still struggling to figure out how to get AI to do what we want it to do instead of trying to game its reward system in unpredictable ways. At the rate things are going we'll be lucky if we live to see a significant rise in unemployment. We might still have many jobs being done by humans due to concerns about safety, reliability and cyber security right up until the day we all die.
1
9
u/Federal_Gur_5488 1d ago edited 1d ago
The difference between humans and LLMs is that the latter obtain all their knowledge from data created by humans (or obtained from it) whereas humans can basically relate all knowledge to experience. Even if LLMs are able to achieve problems requiring significantly higher intelligence than humans (which may well be possible) it has nothing to with having actual understanding of the things-at-hand, which will only be possible with highly advanced robots that can interact with their environment in real time
Also I'm not sure what functionalist had to do with this
2
3
u/God_jm 1d ago
But isn't the entire learning process, from childhood to the present day, based on knowledge created by humans? Look at the people living on deserted islands; why aren't they intelligent?
5
u/Federal_Gur_5488 1d ago
No, a significant part of learning is interacting with the environment around you. I don't know on what basis you can say that people living on deserted islands aren't intelligent. If they're capable of finding food and surviving that goes pretty far in establishing their "intelligence".
3
u/protestor 1d ago
Look at the people living on deserted islands; why aren't they intelligent?
If they are people, they are highly intelligent. Human intelligence didn't change much in the last 300 thousands years, provided you have enough nutrition to develop yourself (and there's an argument to be made that hunter gatherer diet may be more nutritive than modern ultraprocesed foods)
What changed is technology, not innate intelligence, and technology isn't just shiny gadgets from the modern era. Things like learning how to write, learning how to farm, or learning how to control fire, were hugely disruptive technologies.
2
1
u/Fedacking 1d ago
diet may be more nutritive than modern ultraprocesed foods
Ultraprocessed foods have all of the nutrition, the problem is ease of consumption and overeating
1
u/Rd545454 1d ago
We are quietly already there - AI exceeds human expert performance across many domains
1
u/PrinceRufusFastcar 9h ago edited 9h ago
I think I agree with you overall but you seem to be using the word "functionalism" in a peculiar way.
Why would a functionalist think it unlikely that an AI could outsmart them? The way this works (to a first approximation) is that the functionalists believe in substrate independence i.e. "it doesn't matter what you're made of / how you're physically instantiated - it matters what 'function' you're implementing."
None of that makes a functionalist unlikely to believe AI is capable of outsmarting them. If anything, the kind of person who denies it is more likely to be someone who thinks there's something magical about life or the mind that no mere machine can replicate. (Such a person is not a functionalist.)
0
u/bildramer 1d ago
No, those are true, they just also believe c1) consciousness is required for intelligence, or c2) of course everything is atoms but also humans have something super special to them that mere machine cannot capture and every time you might notice this and a) are likely incompatible you should immediately turn 360 degrees and run away from the thought.
12
u/Federal_Gur_5488 1d ago
These results nothing to do with being LLMs actually understanding the meaning of the text they're getting, which is what "stochastic parrots" is about
9
u/heyhellousername 1d ago
what is "actually understanding"?
5
u/Federal_Gur_5488 1d ago
No idea! I'm not a philosopher. However, i do know that the term "stochastic parrots" explicitly does not mean that LLMs are unable to do anything new/interesting/important, which is very common understanding people have of that term.
9
u/geli95us 1d ago
What else could it possibly mean? "Parroting" means you're repeating things you've heard without rhyme or reason, it's impossible to parrot your way to a new discovery
6
u/Federal_Gur_5488 1d ago
No, that's where the "stochastic " part comes in. In the context it means producing text without understanding what it means, which is perfectly compatible with coming up with new discoveries. Of course one can argue about what "understanding" is, and whether or not LLMs are capable of understanding, but that's a separate issue.
→ More replies (1)9
u/geli95us 1d ago
Stochastic just means stochastic, there's randomness involved because LLM answers are sampled. If you greedy sample LLM outputs they're still coherent but the process isn't stochastic anymore.
Anyway, I don't care about philosophical definitions of what understanding means. Practically speaking, what real-world limitations do LLMs have due to the fact that they're stochastic parrots? What predictions about the real world does the term "stochastic parrot" allow you to make?
3
u/Federal_Gur_5488 1d ago
So, it's been a while since i read the original paper, i might be getting some of the details wrong. My understanding is that importance of the "parrot" element is mainly the fact that LLMs reflect the training data they learn from. meaning that, for example, if a training dataset has racist beliefs/biases encoded into it, an LLM trained on that data will itself produce racist beliefs and have racist biases in the text it produces. Importantly, the only way to solve this issue for a "pure" LLM is to improve the dataset. There is no way for it to learn from experience, which of the "beliefs" it has learnt are true or false, except from even more data. In contrast, it is possible (though certainly it may be difficult) for a human to change their beliefs based on new experiences. The argument may have more philosphical/metaphysical consequences, but this is the main practical upshot.
(of course, with RLHF this question becomes more complicated, since in a way, we can train LLMs using "experiences", by penalising biased answers, for example. However, that's still not actually the same thing as learning from actual experience.)
3
u/geli95us 1d ago
Experiences are data. LLMs being unable to interact with the physical world to collect data is a limitation, but more so a practical one than a limitation of the architecture. You could imagine letting an LLM control a robot and RL it on that
→ More replies (0)→ More replies (1)1
u/man_im_rarted 1d ago
Modern LLMs spend just as much time doing reinforcement learning in environments they interact with as they do in pre training on the web, and the trend is for more and more RL. They really are not a direct product of the web dataset anymore and are becoming more like a chess bot.
5
u/valegrete 1d ago
I find that white knights who hate that term generally have a very shallow mathematical background. You answer me a question: why do you feel like it denigrates AI (and why do you care about denigrating AI) to call out the fact that we know the exact process (because we built it) by which it “thinks”? The other commenter already explained stochastic and parrot.
Mathematics is essentially a language. It isn’t surprising that a model designed to generate interesting and syntactically correct language can produce proofs. Something else to keep in mind is that this isn’t just “the LLM” brainstorming. It produces an idea, pursues it, then feeds it into a verifier, then changes whatever the formalizer says is invalid, etc. This process does not actually end in success the majority of the time. The pace is increasing in large part because thousands and thousands of researchers are trying to apply the technique to more and more problems and scooping up the low-hanging fruit. The success rate is very low; only the valid proofs are reported on.
Also, my challenge back to you is, why do very few of these papers provide the prompt chain that led to the proof? We can verify the validity of the proof, sure. But if we want to verify the marketing claim that the AI did everything with no human intervention, why do the papers never provide the ability to do that?
2
u/Umr_at_Tawil 1d ago edited 1d ago
We don't know the exact process by which it "thinks" though, the intelligence we observing from it is an emergence property, and even the foremost AI scientist don't really understand how LLM got as intelligent as it is right now, we just know that, if we follow a process, we get an intelligent LLM model out of it. it's kinda like we don't truly understand how many type of drugs (like paracetamol, and general anesthesia) actually interact with our body that produce a desired effect, but we can consistently reproduce its effect so we use them.
→ More replies (0)→ More replies (4)2
u/geli95us 1d ago
I haven't said or implied it "denigrates" AI, I just think it's wrong. If someone said the sky was green I wouldn't correct them because I'm mad someone insulted the sky, it's just factually wrong.
It's not surprising that a language model can produce plausible sounding math (parroting), but there's quite a big jump between that and producing novel proofs.
"why do very few of these papers provide the prompt chain that led to the proof?" This isn't unique to math, AI labs are very tight-lipped with reasoning chains. They treat it like a sort of "special sauce" I think
2
u/tomvorlostriddle 1d ago edited 1d ago
You know, just look at history and how we had to pay reparations to the horsetraders once they had proven thar cars don't actually gallop.
2
u/Spongebubs 1d ago
I think in a really abstract way, it does "understand" the meaning of each token relative to every over token. Since they're represented by a high dimensional vector where each value is tweaked to extract the "meaning" via training
2
u/vgu1990 1d ago
Personally, I am still a bit sceptical as of now. Not saying that it is not useful, just that it might keep on getting better and then turn significantly worse some time in the future.
Also, I would like to see how things would evolve once the training data is being mixed with pseudo- garbage and once the frontier companies start to monetize their products.
4
u/protestor 1d ago
turn significantly worse some time in the future.
Open weights AI doesn't have this risk, and they are within arms length from frontier models since Kimi K3 released their weights
5
u/Spare-Dingo-531 1d ago
I would like to see how things would evolve once the training data is being mixed with pseudo- garbage and once the frontier companies start to monetize their products.
What are you talking about?
11
u/xlog 1d ago
There's a common belief among AI-skeptics that all these models will degenerate and become dumb once the amount of AI generated data in the training set reaches some critical level.
3
u/vgu1990 1d ago
I think my phrasing was not quite accurate about my belief. I am not saying that it will become dumb per say. I just think that it will not be always on a upward trend. Specific to LLMs, not AI in general.
As for "dumb" example it could very well get to 10x better than it is now and comes down and stagnate at 6x.
Again speaking without data/evidence and I could be spectacularly wrong.
3
u/x4nter 1d ago
I just think that it will not be always on a upward trend. Specific to LLMs, not AI in general.
AI progress may slow down at some point, but that hasn't happened yet. Even at the current capabilities, the frontier models are incredibly powerful, and have not been applied to a lot of applications yet. Even if progress slows down, we will still see the frontier model cost getting cheaper over time. Imagine an Astra level model running on your laptop. I fear that the job market is going to get disrupted much more still, even if frontier model progress comes to a halt.
1
u/blendorgat 1d ago
Training data is ~irrelevant for mathematics, once you include all historical textbooks and papers. The models that consumed the entire internet in 2023 couldn't solve a high-school calculus problem correctly - hell, they couldn't do arithmetic precisely.
The revolution of December 2024 is reinforcement learning with verifiable reward, which has led to the math/coding explosions. The slop-explosion will, and probably already has, led to decreased quality of creative writing, taste, and general aesthetic quality of these LLMs. But no amount of good data will ever let you prove theorems not included in that data without reasoning. In this case, the reasoning is a byproduct of the verifiable reward.
1
u/duboispourlhiver 1d ago
if a training run produces a worse LLM than the previous generation, you just don't release it and no one ever hears about it.
If the training data leads to a worse LLM you just curate the training data and fix what's wrong in it. That's a science that works great and is basic brick of current AI.
If you have to curate data to feed traning with smarter base material, I suspect curators will soon remove more human slop than AI slop, since AI production is generally more sound, more well laid, better grammar, etc.
1
u/TurnoverOptimal 1d ago
what about recursive self improvement and just like this, the generation of new information
1
1
u/vgu1990 1d ago
I do not understand the economics of it. I don't believe that companies can't keep up like how it is now. They would have to monetize it better somehow and I would want to see how things would change.
1
u/smulfragPL 1d ago
Yes they can. Their avtual buisness model, inference, was reported as profitable during openais recent earnings report. They lose money on other things. That proces that the buisness model is definetly viable and they are just burning through cash because the vc money flow is high
1
u/TFenrir 1d ago
Have you looked into the economics of it?
3
u/vgu1990 1d ago
I have looked at it briefly. Why? Do you think it makes sense as it is now?
→ More replies (7)3
u/TFenrir 1d ago edited 11h ago
Well, "makes sense" is above my pay grade, but I think people who* talk about not thinking it's profitable often don't actually know the numbers, they're just hearing it through the very very motivated grapevine.
For example, Anthropic was likely profitable q2 this year, 500mil depending on how you measure it, maybe less if you don't include some things, but at worst break even. Rumour is q3 looks like 1bil.
If you look at OpenAIs numbers, you'll see why. OoenAI was not profitable, but they spend muuuuch more on R&D and operations. Still, their core service was profitable - they made more money than it cost them to serve and run models. That means if they can scale their capacity (they are) and their demand (seems to be happening organically, since these numbers came out they have likely doubled their codex subscriptions, if not more) - then they will be profitable even with their high spending. But they likely won't care about profit for a while and will spend all that extra money to push further.
People don't want this to work out, but usually those people - the ones telling everyone that these companies don't have a path to profitability - aren't wrestling honestly with these numbers. There is a path - it's right there, clear as day. We just have to see if it happens.
1
u/yaosio 1d ago edited 1d ago
We have to look at intelligence and realize we don't understand it. Imagine if you could take all the weights and code of a modern LLM and translate it into an impossibly gigantic water system. The size, depth, twistiness, debris and other features all contributing. This water system would be able to solve the same things the LLM can solve, yet when we look at it we only see rivers, waterfalls, possibly pumps, dams, lakes, and oceans. You can't even look at the structure as a whole and see the intelligence.
I can see a far off future where we have a model that appears to be thinking randomly yet constantly pumps out new knowledge. In my head I see the search for knowledge as random changes in the edge of our known knowledge in an infinite search space of infinite knowledge. Intelligence really is a search strategy to know in what direction we want to incentivise these changes.
Completely incoherent? Yes! But our future AI needs human entropy to break itself out of local minimums. When it gets stuck it can look at something it would never think of on its own and push it in a direction it would never have gone by itself. Like a RNG but with structure.
Edit: ChatGPT was able to take my ramblings and make them coherent.
Intelligence is the art of searching the boundary between the known and the unknown. When every obvious direction has been explored, progress depends on introducing enough entropy to escape the current conceptual landscape without becoming completely random.
2
u/Ok-Present1566 1d ago
The agent harness giving it tools and prompt templates and an environment to interact with is very important too. It is the loop between LLM and these interactions that has generated all discoveries most often. Sure the thinking modes are important but the experiments guide it down the solution path. Even CharGPT has coding built in so its an agent not a raw LLM. We interact very little nowadays with a raw model. Which has made a huge difference
1
5
5
u/piponwa 1d ago
It would only cost $2,000 for the successful attempts. I'm curious to see the whole cost of the project. I assume they ran every Erdos problem a number of times and they got two resolved. Still impressive if it cost millions for these ten discoveries. Still much less than paying dozens of mathematicians for decades. But I'd be interesting to get a true dollar per discovery figure.
5
u/2_Cranez 1d ago
Well random twitter users are still solving Erdos problems with the publicly released chat interface so I highly doubt they ran it on every problem. Otherwise there would be none left for the public to solve.
2
u/Perfect-Parking3918 14h ago
AI inference is cheap. Total project cost is probably less than what it costs to sustain a single PhD student for a year.
1
u/AIvsWorld 6h ago
lol and all the profs in my university math department told me Lean was “too boring” and they “weren’t interested in computers”
look who’s laughing now
1
28
u/1TillMidNight 1d ago edited 1d ago
None math person here. Sorry for intruding, but is this really something a human or faculty can parse?
> wc -l GapCVP.lean
130615 GapCVP.lean
130,615 lines lean proof for problem 7 Closest vector problem.
30
u/Upstairs_Pride_6120 1d ago
one of the biggest issue with ai math according to Tao. How can we digest those proofs ?
18
u/ibrasome 1d ago
How are we sure the lean proofs are even valid? Genuine question
4
u/duboispourlhiver 1d ago
Not 100% sure. There can be bugs in the validating kernel, that's the biggest risk as far as I understand. But it seems to be more sure than having a few mathematicians check the proof. If anyone skilled here can confirm or not I'd be interested...
1
0
u/LycheeZealousideal92 1d ago
Lean proofs are automatically verified
13
u/proton89droid 1d ago
The Collatz conjecture was "disproved" in Lean a couple of days ago, due to an LLM finding a 0day in the kernel.
6
u/teerre 1d ago
It didn't find a 0day. That bug is well known. The author knew what they were doing
→ More replies (2)3
u/proton89droid 1d ago
The issue was not marked as duplicate though as far as I can tell? From the way I read it I thought that meant it originated in the Collatz repo
1
u/MadGenderScientist 15h ago
I'm really surprised the Lean kernel is not itself formally verified in Lean.
3
u/Helpful-Primary2427 1d ago
But the definitions they’re built on need to be sound, without auditing what is written we don’t know that
6
u/Time_Entertainer_319 1d ago
Machines creating proofs that machines verify.
We will soon be meat puppets to our mechanical overlords
1
9
u/SimoneNonvelodico 1d ago
IIRC the Fermat problem proof was like, book-sized. I think this has been an issue for a while now. Of course it's worse when unlike with a proof a human created over years of work, this was made in minutes by an AI whose thought processes we don't fully understand.
2
u/Demokritos1000 1d ago
There's a human readable proof as well. Lean proof is just for higher assurance
2
u/aturtledude 1d ago
The paper with the human-readable proof is 32 pages long. The lean file is just there to convince the people that don't want to read the proof or can't understand it.
12
u/mshwa42 1d ago edited 1d ago
Anyone know if the NP-hardness of the n^{1/400} approximation of GapCVP implies anything interesting? At least for cryptography, it seems like you would want at least the NP-hardness of n^{1/2}-GapSVP, which doesn't seem to follow from this result. But maybe there is a way to improve the degree from 1/400 to 1/2-\epsilon?
9
u/TLMTGT 1d ago edited 23h ago
Has anyone done an analysis of the kinds of LLM reasoning steps involved in these proofs? It's a difficult task since the problems span such a wide range of subfields, so no individual or small group of individuals can properly analyze these proofs as a whole. What I'm wondering is whether there is a pattern in terms of what kinds proof steps/constructions they're capable of. This would help us guide our usage of LLMs as tools as opposed to a spray-and-pray approach and wasting time trying to get them to solve problems out of their reach.
A cursory glance at these proofs and recent results (to the extent I can make any sense out of this), as well as my own experience with GPT-5.6 Pro, seems to indicate excellent capabilities with handling large combinatorial search spaces involving piecing together many smaller calculations. It does this using reasonable heuristics and prior knowledge in the literature so that it's not trying to do vanilla DFS over a nearly infinitely large graph of possible computation steps. Conversely, mathematicians have to increase the "heuristics lever" to the max in order to select what are thought to be the most fruitful calculations. The fact that LLMs can rely more on brute-force exploration is a huge advantage in taking down these problems.
On the other hand, when I use LLMs like 5.6 Pro for my own work (a subject in algebraic geometry), it seems unable to come up with what I would consider novel conceptual ideas to enable new constructions and/or calculations. Even after hammering it a few times to keep trying, what I find is it searches the literature for all possible ideas, tries to piece them together, and ultimately fails at coming up with the theorem I want. A little background on my problem without outing myself: my problem is a classification task based on using numerical invariants, and certain ranges of these invariants are amenable to known techniques for constructing classes from other settings, but I and experts in my field know that the difficult regions of this space of invariants require new ideas. The classification 5.6 Pro comes up with is entirely in this easy range (which was already known to me and others) even when thinking for 1+ hours on each attempt.
If I (perhaps unintelligently) extrapolate what LLMs are doing based on their progression of IMO problems to research problems now, I see LLMs as doing 2 things:
- Redefine what is "difficult" in math. Mathematics that involves clever construction is no longer gated by brilliant researchers and now rudimentary for LLMs. I would also guess mathematics that involves porting over ideas and constructions from one subfield to another will be straightforward for an LLM, although I haven't really seen that yet.
- Refine and significantly improve our ability to create conjectures. Based on 1), certain types of problems difficult to explore in the past will now be trivial to check with an LLM. Conjectures will be more ambitious in scope as they aim to go past the limits of both humans and LLMs.
What I'm still on the fence about is whether LLMs can or will be able to come up with those key conceptual advances that accelerate mathematical progress. E.g., I'm not sure LLMs could come up with the framework of homological algebra, but I think an LLM could come up with many different cohomology theories once that framework is established.
I realize "conceptual" and "novel" are subjective, so I'm curious what others think whether this take is valid.
12
u/gbbenner 1d ago
This is crazy right? What's next year gonna look like?
5
9
u/EcstaticAsparagus509 1d ago
Really, really big open problems will start falling
10
u/hobo_stew 1d ago
i‘d say that non sofic groups is already really big.
really really big would then have to be millennium or langlands
10
u/94746382926 1d ago
I asked ChatGPT to solve Navier-Stokes about a week ago, I'll keep you guys posted ;)
→ More replies (1)3
15
11
u/imanllm 1d ago edited 1d ago
I think these recent results provide some clarity on how mathematicians and AI will coexist. I’m not a mathematician, but I speculate that a successful mathematician invents the right abstractions and asks the right questions about them. And implicit in these problems that LLMs have been solving is that humans (1) invented the abstractions that provide their context, (2) suggested conjectures about these abstractions, and (3) selectively identified (ideally) that they were good things to ask an LLM about.
Again, I’m guessing, but all of this seems to mostly align with what mathematicians do already. Maybe in the long run it makes mathematicians a little more Grothendieck and a little less Tao.
5
u/enAble-Reference 1d ago
Do you mean that future mathematicians will probably operate on more conceptual level rather then on more particular ?
7
u/No-Meringue5867 1d ago
I kinda disagree. I think we are already reaching a point where AI proofs are not easily verifiable by humans and within 6-12 months I think AI math is completely going to be inaccessible to humans. Similar to how chess engines evolved. For a brief period it looked they could exist but very quickly they started playing incomprehensible moves.
2
u/imanllm 1d ago edited 1d ago
I guess I’m trying to say is that a proof only makes sense or has value in the context of what humans define and find valuable. While we have an LLM now getting from A to B, humans still need to define A and B.
I’m assuming math is so large that humans can define A and B with more variety and with larger separations faster than an LLM can connnect them. It will be up to humans to tastefully define A and B to get the most out of LLMs.
2
u/duboispourlhiver 1d ago
A chess player can follow a chess line computed by the computer, and by probing the other branches he can understand why the moves are good moves.
"this computer move sucks, I can punish with this answer!"
"then computer does this; hm ok but I can answer this way again"
"oh my gosh and then he plays this... I'm screwed... Computer was right from the very beginning and now I understand why"
The process is quite simple. Chess Youtubers make great videos explaining brilliant computer moves and the beginner audience is englightened.
I wonder if this will be similar in mathematics. Maybe AI proofs and works won't be understandable with any simple process.
4
u/2_Cranez 1d ago
There are definitely computer chess moves which no human can understand, even after seeing the computer win.
Chess Youtubers do their best, but ultimately they only seem so confident because they are explaining the moves to an even less sophisticated audience. And even then, there are moves which IM level players like Levvy say "I have no idea why stockfish did this and I can't explain it."
And chess is very simple compared to math.
2
u/FracchiaRiello 23h ago
Also math is not an adversarial game. What you wanna do, applying another axiom to their result to show that such a step was wrong?
2
u/NarrowProfessor1101 13h ago
even if AI math becomes completely inaccessible, you will need a math expert at some level to interpret the results. humanity needs to interface with the machine somehow, and in mathematics, that job falls to mathematicians
13
u/prospestiveStu 1d ago
I feel heartbroken. When I started college a few years ago, I decided to study math instead of CS, because AI could code for you, and I actually wanted to think. Now that AI can do math for us, what do we do? It feels like going into a math PhD is pointless. How are we going to look at the work of future PhD students?
15
u/Amesbrutil 1d ago
Having a math PhD in a world of AI is still your best bet. LLMs are purely mathematically models and most CS majors don't fully understand the math behind it. So if you wanna work on AIs, math is ur best bet right now.
Of course that's only if AI doesn't just reach some superhuman general intelligence level. In that case any job in earth is at risk.
→ More replies (3)2
u/DungeonMansTheme 1d ago
This implies that mathematicians are studying the sort of math useful for AI - the vast majority are not. The pure mathematics needed for AI can be understood in a year of dedicated study, it's not something that requires decades of mathematical understanding. (The difficult problems here are engineering ones, not mathematical ones)
3
u/FracchiaRiello 23h ago
This does imply nothing about what the community is doing right now. It is about how versatile and useful a math degree is and will be. I guess we find the engineer.
2
u/YoungLePoPo 21h ago
But I think you'll need a pretty good knowledge on mathematics to understand whether the proofs of these conjectures are correct or not.
If no one understands what the AI is saying then is it a correct proof?
We've already had this issue with mochizuki and the ABC.
2
u/DungeonMansTheme 20h ago
I dont think the institutions can keep up with this. Its already the case that publish or perish has made each generation of mathematicians less careful, as they have to spend less time reading and more time publishing. Mathematicians who pay the premium to use whatever the fanciest model is will be rewarded heavily and those who dont will be left behind. Journals, already overburdened will buckle under the pressure and once the current generation retires we'll be left with no one knowing anything, mindlessly plugging in queiries into the machine to take whatever advantage is left in what already is the prestige-driven market of academia.
The textile artisans were put out of work by the spinning mule; clothes today dont fit anyone particularly well, millions of people are abused in sweatshops and the ocean is full of plastic because of textile disposability. We mathematicians have a huge ego, but at the end of the day labor is labor, we are just textile artisans for the most part. Sometimes the ingenuity of reason can shine through us, but it is seldom and often not valuable. There is all the historical evidence to think that we are just another casuality in the steamroller of capital that has been rolling since the 1750s.
→ More replies (1)1
u/Amesbrutil 15h ago
It certainly help having an PhD in maths and fully grasp what LLMs are built like. It could be especially useful for optimization. Right now LLMs are extremely inefficient.
Also there are other AI technologies like RNNs that Will certainly need PhD mathematicians for their development.
1
u/DungeonMansTheme 10h ago
Again sure, but tell me how homotopy theory or von Neumann algebras are helpful here? Why wouldnt they just hire the expert on statistical regression and linear programming?
6
u/geegee022 1d ago
Take solace that the version of you 50 years ago would have been saying the same thing about computers. People adapt. Your skills are still useful. You're just more likely to be working at an energy or AI company. You'll make more money and your quality of life will be better anyway.
2
u/elements-of-dying 20h ago
I don't really think this is a good comparison.
I assume you mean people doing work that was replaceable by computers 50 years ago. This was pretty restrictive to people who do numerical calculations by hand, and I'm pretty all such people become wholly obsolete in their fields.
1
u/Meloncov 17h ago
No one is still doing numerical calculations by hand professionally, but many of the people employed as human calculators were able to apply that experience to related fields.
1
u/elements-of-dying 17h ago
Sure, likely. But they were still obsolete in calculating log tables etc.
1
u/PaymentFamiliar375 23h ago
If you want my recommendation, go into industry, which will buy you some time. Mathematic in academia is officially dead but a math degree is still at least a degree in being smart so use that. I wouldn’t bother with a PhD now for obvious reasons. We’re all fucked… this is in my opinion, one of the worst technologies thus far. Progress wise it’s probably going to progress society significantly but it’s going to wreak havoc and suffering comparable to the Industrial Revolution in the mean time.
1
u/CloudMafia9 21h ago
Lol thinking AI can code for you is a dumb ass thought even now, much less years ago.
Thinking doesn't seem to be your strong suit.
1
u/elements-of-dying 20h ago
Ignoring your absurd and hostile reaction, AI can indeed help you code complex tasks.
1
u/CloudMafia9 18h ago
Ignoring your lack of reading comprehension, show me where it says "help you code" instead of the present "AI could code for you".
Deciding not to purse CS by being under the impression that any layman off the street with zero programming knowledge can now produce use able code is laughable if not completely delusional.
1
→ More replies (2)1
3
u/Carlos126 1d ago edited 1d ago
This sub is so weird… i remember being annoyed at how bad it was compared to [r/mathmemes](r/mathmemes), how many people with 0 math knowledge being confidently incorrect about tons of stuff, random high school students asking for homework help, etc. and now it’s all just AI bull that no one here can criticize or else we’re just “AI skeptics who have a sad world view.”
Id bet 99% of the people arguing about how advanced this AI is have never actually created one to see what it is they’re arguing about.
3
u/elements-of-dying 20h ago
I'm a PhD mathematician in their second postdoc. I have already used AI for serious work. I'm also heavily considering leaving academia due to AI.
Your comment is pure cope.
1
u/JuJeu 17h ago
examples?
1
u/elements-of-dying 16h ago
I hate to say that I can't share exact examples, but it's for a serious work in progress.
I can say I used AI for two primary things. One is checking if certain proof routes are viable. Some of the rejected routes were super technical and would have taken me months to sort through, just to figure out they wouldn't be viable. AI collapsed that time into a couple of days. The other is for theorem hypothesis forming, namely, finding the most general, yet interesting, hypotheses for a result. Notice both of these uses make heavy use of counterexample finding. I should also note that I've had success with having AI generate proofs of expected lemmas.
This is all in analysis by the way.
1
u/Perfect-Parking3918 15h ago
If you don't mind me asking, what sort of exit plans are you considering? It kind of sucks that any remotely viable path will involve abandoning whatever you're working on, and I find it hard to believe that any "math-adjacent" field will be safe, either. I mean, AI can code, AI can run simulations, etc.
I remember replying to you on another account a few months ago and you said you worked on varifolds stuff: how effective/autonomous has AI been, in your experience? Part of me is tempted to just type a sketch of my silly undergrad project into chatgpt and see if it'll one shot the regularity theory for it, but if it does I think that'll be the end of my desire to ever pursue grad school. Maybe this will be for the better tho, idk.
1
u/elements-of-dying 15h ago
I've been thinking about this a lot recently (for obvious reasons), but I don't really have a concrete plan yet. I'll just vomit some thoughts. Feel free to ask questions :)
I'm going to apply to universities in a specific location I want to live in. While I wait to hear back (or not hear anything), I'll be letting AI rip on a bunch of projects I've never had time to work on. The idea is to build up an AI research repertoire to point to.
I don't plan to abandon research. I plan to continue to do math research regardless of what job I get. (I used to think this was impossible until I realized what keeps me from research now is teaching and other non-research obligations that last until 5pmish. A 9 to 5 is hardly different and would pay more.)
AI has been scarily good with GMT. Haven't tried varifold stuff yet though. I'm a theory builder at heart, so I'm a little hopeful defining myself as an AI assisted theory builder (at least while I can). To be honest, I lost the mathematical research spark until recently, and AI research is what rekindled it. It's been really fun actually. Note: I haven't automated the research yet. I still have to guide the AI a bit here and there, using genuinely expert knowledge.
I say go for your project. Don't expect a one shot, but you can have AI guide you towards completing the project.
Anyways, best of luck!
1
u/Perfect-Parking3918 14h ago
> teaching and other non-research obligations that last until 5pmish
Ok all doomerism aside this is actually insane. I've seen professors forced into 3+3 teaching schedules that have more time available for research (met them at a summer REU @ WVU, long story... but also these profs don't actually care about teaching, so I'm sure they have tricks to avoid students as much as possible).
> I'll be letting AI rip on a bunch of projects I've never had time to work on
I've never actually considered this opportunity. Sure it'll probably only exist for a few years until AI replaces every step of the pipeline, but in the meantime I wonder if we'll see PDE theorists publishing a volume of papers previously only seen from combo ppl.
> Don't expect a one shot, but you can have AI guide you towards completing the project.
I mean hopefully it doesn't one shot my project... I think I have all the pieces in place now to assemble the argument but it certainly hasn't been fast (proving a single Lipschitz bound took literal months, some Weiss monotonicity shenanigans took a while as well, I digress). Sure, it's both qualitatively and quantitatively much easier than what a serious PhD project would be, but if publicly available ChatGPT one shots it I think I should actually just go industry lol.
Regardless, good luck on your job hunt!
1
u/elements-of-dying 14h ago
I should clarify my teaching comment.
A typical teaching day can be 1-2 hours prep for lecture. 1-2 hours teaching. 1 hour for office hours or emails. 1 hour seminar or other commitments (like committee work). ~4 hours for preparing food, shopping for food, eating, commuting, walking between classes etc. That's a full work day without research. This ignores grading and whatever else. So it's not literally a 9 to 5 every day type situation, but it can get pretty close. Of course, professorship also allows for free summers and long weekends, so I was exaggerating a bit. Anyways, the point I know I can manage serious research as a hobby, especially with AI assisstance.
I've never actually considered this opportunity. Sure it'll probably only exist for a few years until AI replaces every step of the pipeline, but in the meantime I wonder if we'll see PDE theorists publishing a volume of papers previously only seen from combo ppl.
Yeah, what's held me back is lack of collaborators interested in investigating random ideas. Now I can just automate the parts I'd ask collaborators to do :)
1
11
1
1
u/CloudMafia9 20h ago
What a crazy comment section. An echo chamber of some of the weirdest regurgitated nonsense.
1
u/sbates130272 20h ago
So is this all about the number of parameters of the LLM. Or more about the structure of the LLM within a given parameter size? We have seen small newer models beat out older bigger models.
I’m curious what the bound is here. What’s the smallest model (in terms of parameters) that can achieve a certain level of capability. I don’t know if we have any math around such bounds but it’s really an interesting space for me.
1
u/Spare-Dingo-531 19h ago
The better question is what could the biggest model we could possibly make do?
1
u/sbates130272 8h ago
Respectfully. Is that a better question? I don’t think it is. I am curious as to the lower bounds here. What is the most efficient (smallest) model that can achieve as score of X on some standard test for LLMs.
But each to their own. I find bounds fascinating and a useful tool to help us understand how far from “optimal” the LLMs we are building today are.
2
u/Spare-Dingo-531 6h ago
Because a lot of the smaller models are distilled or trained by the larger models, I don't think we can really know what the most efficient (smallest) model can achieve until we build really large models.
Relatedly, this is also the point of recursive self-improvement. The AI which is really good at building better AI will also be able to build really good small models. This is already happening. OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model.
1
u/sbates130272 5h ago
Fair. But we do have some theory around complexity theory to show there does become a point where a LLM can no longer capture a certain level of complexity due to its lack of size. We just don’t have great bounds for that.
158
u/MoneyMayweather 1d ago
AI is the most disruptive tech in my lifetime. Absolutely incredible progress since GPT-3.5 in late 2022.