OpenAI’s Astra Solves 10 Math Problems That Stumped Experts

OpenAI announced on Aug. 1, 2026 that an internal version of its next major model, Astra, has resolved or made substantial progress on 10 long-standing open problems across eight fields of mathematics and theoretical computer science — with machine-checkable proofs published on GitHub.

OpenAI announced Aug. 1, 2026 that an internal, unreleased version of its next major model — called Astra — has produced new results on 10 open problems that had stumped mathematicians for decades. The problems span eight distinct fields: high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics. For each result, OpenAI published a 249-page manuscript alongside machine-checkable Lean 4 certificates on GitHub, allowing independent verification rather than requiring the math community to simply take the company’s word for it.

The scope is broader than anything OpenAI has previously claimed for an AI system. In May 2026, an unreleased model produced an  AI-generated disproof of the Erdős unit-distance conjecture — a single result in one subfield that has since inspired at least five follow-on human-authored papers. The 10-result release covers entirely unrelated mathematical domains, which observers say makes it harder to dismiss as a one-off. Among the highlights: a construction establishing the existence of so-called non-sofic groups, a disproof of Connes’s rigidity conjecture in operator algebras, an exponential parallel repetition theorem for quantum games, and resolutions of two longstanding Erdős combinatorics problems (problems 183 and 180/146).

Human researchers prepared the AI-generated arguments into manuscripts; Astra then formalized each proof as a Lean certificate. OpenAI is also releasing a narration of the model’s reasoning process for each solution. The company was explicit about attribution in its announcement.

“We believe attribution should honestly reflect how a result was produced: claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work,” wrote OpenAI.

One figure from the announcement has circulated widely: OpenAI said the token compute needed to find these solutions would cost roughly $2,000 at Sol API rates. That number is worth contextualizing. It does not represent the full cost of the research — it excludes training Astra, running the underlying infrastructure, paying the researchers who organized the manuscripts, or the work of formalizing the arguments. It is a measure of inference cost only.

How This Fits the Competitive Landscape

OpenAI’s results land in an already active AI-for-math arena. Google DeepMind has been the most prominent competitor: its AlphaProof system achieved a silver-medal standard at the 2024 International Mathematical Olympiad, and more recently Gemini equipped with Deep Think reached gold-medal performance, solving five of six IMO problems perfectly. DeepMind’s AlphaProof Nexus autonomously solved 9 Erdős problems out of 353 attempted — including two open for 56 years — at an inference cost of a few hundred dollars per problem.

The meaningful distinction is breadth and domain novelty. DeepMind’s systems have excelled at competition-style math and have made inroads into combinatorics, but largely within stylistically similar problem types. Astra’s 10 results span eight structurally different subfields in a single release. If outside review confirms these results, it shifts the conversation: the question is no longer whether AI can occasionally assist with a proof, but whether a model has developed a repeatable research capability that generalizes across mathematics.

What This Means for Students and Researchers

A few days before the Astra announcement, OpenAI launched ChatGPT for Academic Researchers — a program offering free access to frontier models for 100,000 scientists, mathematicians and engineers. The rollout starts with 10,000 participants this summer and expands through 2027. Participants receive access across ChatGPT, ChatGPT Work and Codex, including the GPT-5.6 model family, with expanded deep research tools, higher usage limits and larger context windows. For grad students and postdocs who would otherwise pay significant API costs, that is a concrete cost reduction for research workflows ranging from coding to grant writing.

The Lean 4 angle matters separately. Every Astra proof was verified via a Lean certificate — a machine-checkable formal proof that anyone can run against the published GitHub repository. That means verification is no longer gated behind peer review timelines. For students considering careers in mathematical research or theoretical computer science, fluency with formal proof systems like Lean 4 is becoming a differentiating skill. Understanding how to work alongside AI on formally verifiable proofs — rather than relying on narrative LLM outputs — may soon be a baseline expectation in research environments.

The results are not yet externally peer-reviewed in the traditional sense, and the mathematical community will need time to engage with all 10 problems in depth. But the Lean certificates lower the barrier to that engagement considerably. Whether Astra’s results hold up to scrutiny, the combination of a broad multi-domain release and publicly verifiable proofs represents a new benchmark for what AI-assisted mathematical research looks like in practice.

Source: OpenAI

Additional research sources