Anthropic AI Makes Significant Progress on 150-Year-Old Riemann Hypothesis
Anthropic has announced that an unreleased AI model made significant progress on the Riemann hypothesis, one of mathematics’ most enduring unsolved problems. The achievement, which saw the model autonomously test 650 different solution strategies over a day and a half, demonstrates how large language models are beginning to function as genuine research collaborators rather than simple pattern-matchers.
The Riemann hypothesis has puzzled mathematicians since 1859, concerning the distribution of prime numbers and carrying a $1 million prize for a correct general proof. While Anthropic’s model hasn’t claimed the bounty, its work significantly advanced the known lower bound of solutions for which the hypothesis holds true. More importantly, the method of discovery raises compelling questions about how AI might reshape mathematical research.
The team plans to publish formal verification of their findings using the open-source proof assistant Lean. The company’s in-house mathematicians have already confirmed the results, which will likely spark renewed debate about AI’s role in fundamental scientific discovery. Notably, the process required minimal human mathematical expertise to initiate the breakthrough.
How the AI Made Its Discovery
An Anthropic staff member without significant mathematical training prompted the model to “take a real stab” at proving the hypothesis. The model then coordinated across 60 subagents, spending 31 million output tokens over roughly 36 hours. This wasn’t simply generating variations on existing approaches but coordinating a structured research effort among specialized AI agents.
The breakdown of subagent responsibilities reveals the sophistication of this approach. Two subagents developed the key mathematical ideas that led to progress. Thirteen contributed supporting ideas to those leads, while 30 attempted but couldn’t generate new approaches. A dedicated group of 13 served as validators, checking argument correctness before two agents wrote the initial paper. This division of labor mirrors how human research teams operate, with some members generating hypotheses, others refining them, and designated reviewers verifying their validity.
The validation group is particularly significant. Mathematics demands rigorous proof, and models have historically struggled with consistent reasoning. By building in verification agents, Anthropic addressed a core weakness in AI-generated mathematics without requiring human oversight at every step. The company’s in-house mathematicians ultimately confirmed the work, providing the stamp of approval the mathematical community requires.
A Growing Trend in AI Mathematics
This breakthrough isn’t an isolated achievement. Large language models have been increasingly productive in mathematics throughout 2026, solving several Erdos problems and producing noteworthy results. OpenAI recently released a set of ten major results proved by its internal “Astra” model, while a separate Anthropic effort disproved the longstanding Jacobian conjecture.
What distinguishes Anthropic’s Riemann hypothesis work is the minimal human intervention. Previous AI mathematical achievements typically required significant human guidance, with researchers crafting prompts, selecting approaches, and validating results. Here, a non-expert initiated the process, and the model coordinated its own research effort. This suggests the technology is approaching a threshold where researchers can set problems and let AI systems work toward solutions with limited supervision.
The computational cost, however, remains substantial. 31 million output tokens represents significant resources, raising questions about whether this approach scales practically for broader research. The discovery required considerable time and processing power, suggesting AI-assisted mathematics won’t replace traditional approaches entirely but will serve as a complementary tool.
The Mathematical Community’s Divided Response
These developments have generated significant tension within the mathematical community. In June, a group of prominent mathematicians signed a declaration expressing concern that AI could undermine fundamental mathematical values, particularly the principle that “true mathematical proofs should be attributable to specific authors who take credit for their discovery and assume responsibility for their correctness.”
The declaration reflects anxiety about shifting professional norms. If AI becomes essential for producing significant results, what happens to the human dimension of mathematical practice? The collaborative dynamic, the personal satisfaction of discovery, and the attribution of credit could all face disruption. For early-career mathematicians, the rise of AI capable of generating substantial results raises legitimate questions about training, professional recognition, and the future of the field.
However, Fields Medal winner Timothy Gowers offered a compelling counterperspective in response to the declaration. He questioned whether the influence of AI might change mathematics in a more complex and positive way, suggesting that “if we arrive at a world where mathematical theorems are no longer associated with mathematicians, maybe that won’t be any more problematic than the fact that stars aren’t named after astronomers and most aren’t named at all.”
Gowers’ analogy highlights a fundamental question: Does the value of mathematics lie in the final result or the human activity of discovery? His measured response suggests the shift may be less threatening than some fear.
What the Riemann Hypothesis Progress Really Means
The broader implication here extends beyond mathematics to how we evaluate AI capabilities. Progress on the Riemann hypothesis matters less as a standalone achievement and more as evidence that language models can autonomously coordinate research efforts. The model generated original ideas, validated its own work, and produced results at a level that required minimal human mathematical expertise to initiate.
Anthropic’s deliberate approach matters. The company published methodology transparently, including the breakdown of subagent contributions and the validation process. This clarity enables the research community to replicate and build upon the work, setting a constructive precedent for similar efforts.
The progress on the Riemann hypothesis also demonstrates that AI can contribute to problems requiring substantial computational effort while not replacing human creativity. The model tested 650 different approaches, far more than a human researcher could reasonably examine. But the key insights came from two subagents, suggesting that even within the model’s own operations, quality matters alongside quantity.
Practical Implications for Researchers
For mathematicians and researchers, the practical consequence of this development is a potential rethinking of research workflows. The success of an AI model guided by a non-expert suggests that the primary obstacles to AI-driven research may be less about specialized knowledge and more about effectively setting problems and interpreting results.
Junior researchers in computationally intensive fields may gain new tools for exploring broad problem spaces. Instead of spending years developing expertise in a narrow subfield, they could leverage AI to rapidly test hypotheses and identify promising directions. However, the discipline’s emphasis on rigorous proof and attribution suggests adaptation will likely be gradual.
The model’s self-verification mechanism offers a model for trustworthy AI research. Rather than simply generating potential solutions, the system cross-checked its own work, reducing reliance on human oversight. This suggests current AI limitations can be addressed through architecture design rather than external intervention.
What Could Happen Next
The immediate next step is publication and formal verification of the findings. Anthropic’s mathematicians have confirmed the work internally, but the broader mathematical community will scrutinize the proof carefully. This period of review will determine whether the progress stands as a genuine advance.
The company’s choice to release details about an unreleased model aligns with a growing trend of AI developers sharing research progress while retaining proprietary advantages. Anthropic has revealed enough to generate excitement and credibility while reserving details of the model itself.
Future applications beyond pure mathematics are likely. The structured approach to problem-solving, validation, and coordination used here could be relevant to fields such as physics, chemistry, and engineering, where complex problems require extensive computational exploration. If a model can make progress on a 150-year-old mathematical mystery without expert prompting, similar approaches may be applicable to many other longstanding scientific challenges.
We may also see the emergence of specialized mathematical AI models that combine language capabilities with formal proof assistants like Lean. The integration of these tools could accelerate research significantly, but the human role in designing problems, interpreting results, and advancing the conceptual frontier remains essential.
The conversation about authorship and attribution in AI-generated mathematics will likely intensify. Institutions such as universities and academic journals will face decisions about how to treat AI-generated contributions. The Gowers perspective suggests the field may adapt in unexpected ways, but the transition will require careful navigation.
Anthropic’s achievement represents another step toward a research environment where mathematicians and AI systems collaborate as partners. The work on the Riemann hypothesis suggests the potential of this relationship while making clear that the technology, though powerful, operates as a tool that enhances human capabilities rather than superseding them.
For further reading on AI capabilities, see our coverage of AI advancements and the latest developments in Anthropic’s research.

