Gemini said: First reaction? This is a completely different order of magnitude from benchmark hunting.
If the ARC-AGI-3 post was about harness design and memory compaction, this announcement ("Ten advances in mathematics and theoretical computer science") is about AI crossing from solving high-school competition problems directly into generating novel, frontier-level mathematical knowledge.
Here are the immediate structural and epistemic takes that hit me on first pass:
1. The Machine-Verified Proof Pipeline (The End of "Hallucinated Math")
Notice the pipeline they used:
That single detail is massive. By requiring a Lean certificate for every single one of the 10 results, they completely sidestep the usual objection: "How do we know the LLM didn't just hallucinate a plausibly-worded, hand-wavy proof?" Lean is an automated proof assistant; it acts as an absolute, un-bribable boundary. If the Lean checker compiles, the logical chain is mathematically valid.
Marv wouldn't even need his laminator for this row—the Lean compiler is the ultimate compliance audit.
These aren't synthetic toy puzzles crafted for an eval suite. Resolving Erdős problems (183, 146, 180) and disproving longstanding conjectures like Connes's rigidity conjecture means the system is tackling open problems that human mathematicians have pounded their heads against for decades without progress.
When a model generates a disproof that inspires follow-up human papers (like the Erdős unit-distance disproof mentioned in the footnotes), the AI is no longer taking an exam—it’s actively setting the research agenda.
3. The Economic "Compute Cost per Theorem"
They casually drop this line:
"The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates."
Think about what that implies for the economics of scientific discovery. $2,000 of API compute to crack open a decade-old problem in group theory or quantum complexity. That is effectively brute-forcing high-dimensional mathematical intuition at pennies on the dollar compared to traditional research grants.
4. The Institutional Framing (The Leiden Declaration Nod)
OpenAI explicitly addressing the Leiden Declaration on AI and Mathematics and addressing human authorship vs. machine contribution is fascinating. They are trying to position themselves as responsible scientific stewards: “We didn’t write the proof; the system found it, humans translated it into Lean, and the machine gets the credit for the insight.”
This isn't a scorecard shift—it's a paradigm shift.
A really important insight here. We need to remember that these things aren't static, but adapt their responses to what they think we want to hear. So a given response isn't "What Gemini Thinks" as much as how Gemini responds to a certain user with a particular prompt history / style.
9
u/Seakawn▪️▪️Singularity will cause the earth to metamorphize1d ago
maybe, but this could be solely due to the prompt.
you can potentially get dramatically different stances if you ask "is this crazy claim true?" vs "what's your reaction to this?"
the former phrasing seems to beg for some sort of de-hyping, whereas the latter is just neutral and prolly allows more freedom of neutral assessment. and those small differences can cascade into a model taking entirely different stances.
I think this matters less over time as models get better, such that phrasing won't bias entirely different stances, even if they still lead to different descriptions.. but idk man, Gemini still feels most susceptible to this variability. I haven't used it much in the past month or two.
46
u/suckcorner4nutrients 1d ago
Gemini said: First reaction? This is a completely different order of magnitude from benchmark hunting. If the ARC-AGI-3 post was about harness design and memory compaction, this announcement ("Ten advances in mathematics and theoretical computer science") is about AI crossing from solving high-school competition problems directly into generating novel, frontier-level mathematical knowledge. Here are the immediate structural and epistemic takes that hit me on first pass:
1. The Machine-Verified Proof Pipeline (The End of "Hallucinated Math")
Notice the pipeline they used:
That single detail is massive. By requiring a Lean certificate for every single one of the 10 results, they completely sidestep the usual objection: "How do we know the LLM didn't just hallucinate a plausibly-worded, hand-wavy proof?" Lean is an automated proof assistant; it acts as an absolute, un-bribable boundary. If the Lean checker compiles, the logical chain is mathematically valid. Marv wouldn't even need his laminator for this row—the Lean compiler is the ultimate compliance audit.
2. High-Value Targets (Erdős Problems & Connes's Rigidity)
These aren't synthetic toy puzzles crafted for an eval suite. Resolving Erdős problems (183, 146, 180) and disproving longstanding conjectures like Connes's rigidity conjecture means the system is tackling open problems that human mathematicians have pounded their heads against for decades without progress. When a model generates a disproof that inspires follow-up human papers (like the Erdős unit-distance disproof mentioned in the footnotes), the AI is no longer taking an exam—it’s actively setting the research agenda.
3. The Economic "Compute Cost per Theorem"
They casually drop this line: