Mathematicians Ask What Math Is For in the Age of AI
Top mathematicians at ICM 2026 debate AI proofs, the double-peak homework crisis, and whether math is really about storytelling rather than proof.
Written by AI. Yuki Okonkwo

Photo: AI. Nikolai Brandt
At the International Congress of Mathematicians in Philadelphia this July, three of the world's leading mathematicians spent an hour debating a question that would have sounded absurd five years ago: what exactly are we still doing here?
The occasion was a live recording of Quanta Magazine's podcast The Joy of Why, where hosts Janna Levin and Steven Strogatz sat down with Akshay Venkatesh of the Institute for Advanced Study, Ravi Vakil of Stanford (and president of the American Mathematical Society), and Alex Kontorovich of Rutgers. The transcript, published via Quanta Magazine, is a rare thing: top mathematicians reasoning out loud, in front of an audience, about a field whose identity is suddenly up for grabs.
The Milestone that Wasn't (Quite) Supersonic
The conversation started with AlphaProof's silver-medal performance at the 2024 International Mathematical Olympiad, followed by stronger results in 2025. Vakil, a former IMO competitor himself (Jordan Ellenberg apparently has stories), called it "a proof of concept" while warning against overreading it. His framing: "The first time a model answers a question, it's exciting. A milestone is reached... And the seventh time a model answers a question, it's far less interesting."
Kontorovich was blunter. His prediction, he said, was that no AI model would bother chasing an IMO gold medal this summer because "it's not interesting for them." Once you can push a button and get a perfect score, the benchmark dies. The medals, for the record, are self-declared by the companies; nobody's handing them out.
Then came the part I did not expect. Kontorovich described a recent AI result on an Erdős problem (Paul Erdős left behind thousands of conjectures, helpfully collected by Tom Bloom into a database): the Erdős unit distance conjecture, where it turned out the conjectured constant was simply wrong, and GPT found a counterexample autonomously. "This would be a really great contribution if it was a human being," he said. A week later, human mathematicians solved the unrelated sum-product problem by applying the techniques the AI solution exposed. That loop, machine result feeding human insight, is the golden-age scenario he's hoping for.
It also connects to a pattern Grant Sanderson of 3Blue1Brown has described about why AI advances fastest in mathematics: math is unusually verifiable, so machines get clean feedback. Verifiability cuts both ways, though. Vakil noted his inbox is already filling with claimed AI proofs he has no intention of checking. "It's not fun," he said.
Proof Isn't the Point, Apparently
The most interesting moment came when Venkatesh questioned the discipline's own founding assumption. "I want to also even question the assumption that proof is the central activity," he said, describing his work as "more akin to storytelling," with proofs as "part of the grammar of that."
Vakil backed him hard: this is what mathematicians have always valued. Technology is just forcing them to examine the proxies (publication counts, competition wins) they've been using. Kontorovich illustrated with Fermat's Last Theorem, which has, as he put it, "no consequences whatsoever in the universe," except that pursuing it dragged mathematicians through unique factorization, number fields, and eventually the modularity theorem. His other example: arguing about Euclid's parallel postulate for 2,000 years until the thread leads to general relativity.
So when an audience member asked whether aesthetics could be quantified to help triage the incoming flood of AI proofs, the panel pushed back. Venkatesh said the mechanization question misses the point; humans already know which proofs are boring and which are beautiful. Kontorovich added that math has largely avoided citation metrics like the h-index, and argued for keeping it that way, even acknowledging the system is "right for abuse."
Here's my skepticism: that's a lovely system for the tenured. "Most people are doing it in good faith" works when you're nodding at a stranger at the ICM. It works less well for a first-generation grad student whose advisor's taste is the gate. The panel didn't really resolve that tension, and to their credit, they mostly admitted it.
The Double-Peak Problem
The education discussion produced the panel's most concrete observation. Kontorovich described test-score distributions that are no longer a single bell curve but a "double peak": students using AI to accelerate understanding cluster high; students using AI to finish homework cluster at failing test scores despite perfect homework marks. Vakil was direct about it: "Students are getting 100% on all of their homeworks because they're using AI... They understand none of it."
Kontorovich's answer to his own students is the forklift analogy, and it's good enough to steal. Three scenarios: you use a forklift to move pallets, you practice using a forklift, or you bring a forklift to the gym and do ten reps with it on the bench press. One of these is the wrong use of technology. His challenge to students: figure out whether the homework exists to get an answer, to learn to prompt a model, or to build "the stamina for frustration that it takes to do any kind of difficult thing." If it's the last one, either do it as designed or drop the course.
He also confessed something I found honest: he catches himself summarizing papers with AI instead of fighting with them for days. "I'm making potentially the same mistake the undergraduate is the day before the homework is due."
The Taboo Question
Asked about unspeakable worries, Vakil named two. The first wasn't AI at all but its use as "a Trojan horse for anti-intellectualism": funding cuts justified by the claim that kids no longer need to learn to think mathematically because a model can do it. He says that's already happening. The second was the pedagogy problem above, losing bright students who think they understand when they don't.
On whether AI will eventually do everything mathematicians do, the panel split in a useful way. Strogatz pushed the question hard: won't machines eventually ask better questions and write better exposition than Martin Gardner? Won't it stop being fun? Vakil refused the singularity framing entirely, noting that bets predicting AI milestones on short timelines "have been lost every single time," and offered twenty dollars against a Pulitzer-winning AI novel within three years. Venkatesh took a different tack: even if AI surpasses him, he recalled arriving at grad school, meeting people far better at math, and being happy anyway. "I'm not going to be the best at it. I'm happy to spend my life this way."
That's the strongest version of the humanist position, and it doesn't require AI to fail. It only requires math to be worth doing for the person doing it, the way amateur chess survives Stockfish.
What Nobody Settled
A few open questions the panel named but didn't answer. Peter Sarnak's "math zero" question: seed a theorem prover with no human mathematics and see whether it rediscovers Fermat's Last Theorem, the way AlphaZero learned chess from rules alone. Nobody knows. The access question: frontier models are gated by money, and Vakil confirmed he's on a team building open-source alternatives that are "behind, but... catching up." The low-hanging-fruit question from a grad student in the audience: the safe starter problems advisors used to assign are now exactly what models are good at, and advisors are rethinking where the next generation trains.
Kontorovich offered one closing thought that reframes the whole moment: AI companies care about theorems mainly as marketing benchmarks, and he expects them to lose interest. "Maybe we'll have it back to ourselves again for a while, and we'll have this great tool while they're solving something else."
Maybe. But if the tools keep finding counterexamples to decades-old conjectures and handing humans new techniques a week later, the honest question is whether mathematicians will ever want the field entirely back. The panel's answer, across all the disagreements: the value was never only in the theorems. It was in the story, and in the people arguing about it over coffee.
Yuki Okonkwo is Buzzrag's AI & Machine Learning correspondent.
More Like This
ChatGPT Ads Are Here—and the Playbook Looks Familiar
OpenAI is testing ads in ChatGPT. The current version looks fine. But if you've seen how Google and Facebook evolved, you know where this could go.
Harness Engineering: The New Frontier in AI Development
AI companies are shifting focus from better models to better infrastructure. Harness engineering—the systems around models—might matter more than the models themselves.
OpenAI's Codex Desktop App Launches With Curious Bugs
OpenAI's new Codex desktop app brings AI coding to macOS with a GUI, but early testing reveals surprising UI quirks and context issues.
Making Longer AI Films Without Stitching Clips
Jahan of CyberJungle demos a Seedance 2.5 workflow that turns 30-second AI clips into 90-second continuous shots, no frame-by-frame fixes required.
Gemma 4 12B Brings Local Agentic AI to Laptops
Google's Gemma 4 12B is a multimodal local AI model built for real agentic workflows on 16GB laptops—here's what the architecture actually means.
Does AI Understand Things, or Just Predict Words?
The "AI just predicts tokens" argument is technically true—but is it the whole story? A murder mystery with fake physics might hold the answer.
RAG·vector embedding
2026-09-05This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.