OpenAI this week published a massive trove of 722 mathematical papers it says were produced largely by an unreleased artificial intelligence model, covering 372 open problems across algebra, geometry and theoretical computer science. The release, variously labeled a "drop," a "dump" and a "mathocalypse" by researchers, has sent shockwaves through the mathematical community, according to The Conversation.
Several papers claim significant progress on high-profile Millennium Prize Problems, each carrying a $1 million bounty. These include the Riemann hypothesis and the Birch–Swinnerton-Dyer conjecture. However, at least three papers have already been retracted due to elementary errors, and multiple others have been amended after mistakes invalidated their results.
Proofs described as unintelligible by experts
The sheer volume is not the only concern. A senior mathematician who found one of his favorite problems among those claimed solved told The Conversation the accompanying paper was so unintelligible that, had he received it as a journal editor, "it would have gone straight into the bin." Even OpenAI's own large language model, identified as GPT-5.6 Sol, described at least one high-profile result as "a serious hallucination … which should never be cited, submitted, or circulated as a proof without a complete expert audit."
This raises a fundamental issue: despite the potentially groundbreaking nature of some results, the way the papers are written is not comprehensible even to specialists in the relevant fields. The traditional mechanism for establishing mathematical truth — expert peer review — can take years for long, technical papers. OpenAI's release bypasses this process entirely.
Previous controversy and AGI ambitions
The release follows a controversial episode in September when OpenAI published a claimed solution to a case of the Navier-Stokes problem, another Millennium Prize challenge. Mathematicians Tristan Buckmaster and Levent Alpöge, who had been working on the problem themselves, alleged OpenAI had accessed their data and used their ideas, then devoted approximately $15 million of computing power to find a solution. OpenAI has denied these allegations.
Observers suggest a broader motivation may be the pursuit of artificial general intelligence, or AGI — a hypothetical AI that could surpass humans in virtually all cognitive tasks. Mathematics, especially pure abstract mathematics, is widely seen as a domain requiring extremely complex reasoning. Demonstrating research-level mathematical capability could bolster claims that OpenAI's models are approaching AGI, potentially benefiting the company ahead of a planned stock market launch targeting a valuation of up to $1.4 trillion.
Formalization efforts and verification challenges
To address verification concerns, OpenAI claims to have formalized 300 of the main results using software systems such as Lean, which build mathematics from basic axioms and can check each logical step. This "autoformalization" — automatically translating results into machine-checkable proofs — has gained traction over the past decade. An early milestone was the 2008 formalization of the four color theorem in graph theory.
However, not all formalizations are implemented correctly. Past evidence, including issues with the Navier-Stokes formalization, suggests AI-generated formalizations can contain errors. It remains unclear whether the mathematical community will accept OpenAI's formalizations as sufficient verification.
Community reaction splits between excitement and alarm
Reactions have been sharply divided. Some researchers call it "obviously the most significant moment in mathematical history," while others describe it as "a giant pile of turd dumped at our doorstep." There is genuine excitement about potential progress on important problems across multiple fields, as experts begin combing through the results.
Yet this excitement is mixed with uncertainty. Mathematicians are grappling with what the volume and speed of AI-generated discoveries mean for the future of their discipline — and who should bear the labor of making sense of arguments that initially read, in the words of one researcher, as "AI slop." The episode has forced a confrontation with the changing nature of mathematical practice itself.
Why this matters for mathematics and AI
The OpenAI release represents a stress test for the centuries-old system of mathematical consensus. If AI can generate plausible-looking but flawed proofs at scale, the bottleneck shifts from discovery to verification — a task that still requires human expertise. For the global research community, including mathematicians in Bangladesh and across South Asia who contribute to international collaborations, the episode underscores both the promise and the peril of AI-assisted research. The coming months will reveal whether this "mathocalypse" yields lasting advances or becomes a cautionary tale about the limits of automated reasoning.