← Back to Insights

OpenAI Says Astra Solved Ten Open Math Problems for About $2,000

Nils Liu
AI OpenAI Astra Mathematics News

TL;DR

OpenAI published Astra solutions to ten open problems in mathematics and theoretical computer science with Lean proofs, while external review and product timing remain unresolved.

OpenAI Says Astra Solved Ten Open Math Problems for About $2,000

OpenAI’s claim has a clear condition under which it should be revised. Independent mathematicians need to confirm the original questions, assumptions, and novelty of all ten results, while rebuilding the proofs in the same Lean 4.32.0 environment. If the files compile but a problem was already solved, the formal statement is narrower than the recognized open question, or a decisive assumption differs from the one used by the field, the claim that Astra solved ten open problems would not survive in its present form.

On August 1, 2026, OpenAI published “Ten advances in mathematics and theoretical computer science.” The company says an internal version of Astra solved ten problems across high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics. According to OpenAI, none had seen progress for at least a decade, and most had remained open much longer. The Decoder reviewed the announcement and identified Astra as OpenAI’s next major model family. It also reported that there is no public release date and that OpenAI has not decided whether the eventual product will carry the GPT-6 name or remain within the GPT-5 line.

The ten outcomes are separate research claims, not a benchmark average. One constructs a non-sofic group, addressing whether every group admits finite permutation approximations. Another gives an arithmetic-formula lower bound of $n^4 / \log n$ for computing the permanent. The collection also includes results tied to Erdős problems 183, 146, and 180. Because the problems span different fields and were selected by the company, ten successes do not establish a success rate for arbitrary research questions. OpenAI did not disclose the full candidate pool, a complete list of failed attempts, or the selection procedure.

Lean moves the audit to statements and assumptions

OpenAI released Lean formalizations for all ten results, alongside a paper and reasoning walkthroughs. Lean can check whether every formal step follows from encoded definitions, axioms, and earlier steps. That makes it harder for an algebraic omission or an invalid citation to hide inside a long argument. Formal verification still answers a narrower question: whether a formal statement follows from a particular set of assumptions. Domain experts must determine whether that statement faithfully represents the original open problem, whether the result is genuinely new, and whether its assumptions are accepted.

The company says human researchers helped prepare the papers and formalize the proofs, while Astra generated the mathematical arguments; OpenAI accepts responsibility for their accuracy. The Decoder quotes University of Manchester mathematician Thomas Bloom describing the results, including the non-sofic-group construction, as big news. The same report records an important limitation from OpenAI researcher Noam Brown: the team attempted other major problems and failed, solved no Millennium Prize Problems, and did not spend very much test-time compute on each problem. This makes the release a selected body of research rather than a blinded evaluation of a publicly available model.

$2,000 covers tokens, not the research program

OpenAI estimates that the tokens used to generate all ten solutions would have cost about $2,000 at Sol API rates. That figure does not include problem selection, researcher review, Lean formalization, unsuccessful runs, model training, or infrastructure depreciation. It is also not a future price for Astra. The estimate describes the token-equivalent inference cost of known successful outputs; it cannot establish the total cost per publishable result.

Two sets of evidence should become measurable over the next three to six months. External researchers can rebuild the ten Lean projects and assess novelty in the relevant literature. OpenAI can disclose the number of attempted problems, failure distribution, and compute used per attempt. The first test determines whether the ten results enter the scholarly record. The second is necessary before anyone can estimate how Astra performs when moving from selected successes to routine research work.

Sources:

Further reading:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.