722 Math Manuscripts and the Ten Details Behind Them

21人浏览 / 0人评论

A single minus sign just killed three mathematical proofs.

On October 7, OpenAI published a large batch of mathematical results produced by its internal models on GitHub. The first version of the catalog contained 722 manuscripts, grouped into 372 result families.

A day later, three manuscripts were gone.

One of them, a geometry paper dated September 18 about the algebraicity of Weil classes, carried a short withdrawal note: "The proof contains a sign error."

The proof needed the positive and negative counts of geometric intersections to cancel to zero before a later theorem could be invoked. An operation expected to add a positive term had actually added a negative one. In simplified arithmetic, the authors intended "−2 + 2 = 0" and instead produced "−2 − 2 = −4." The theorem's conditions were never met, and two other manuscripts borrowing the same construction lost their support.

A sign error in an AI geometry proof diagram, a glowing minus sign node breaking a chain of reasoning

This is not a story about failure alone.

On the same day the repository went live, combinatorialist Gil Kalai called the release "an amazing milestone for mathematics." Then he added the necessary caution: the proofs still have to be verified and digested by human mathematicians.

The debate over AI mathematics had already begun weeks earlier. In September, a claimed proof about critical percolation from Claude triggered widespread discussion. On September 11, Terence Tao and 24 other Fields Medalists signed an open statement acknowledging the progress of AI in mathematics while criticizing companies that treat problem-solving as a benchmark race and neglect understanding and the transmission of knowledge.

One side points to results that could reshape a field. The other side points to withdrawals that have already happened. So what, exactly, did these AI-generated math manuscripts deliver?

What OpenAI Handed to Mathematicians

The public repository contains paper PDFs together with their typesetting source files, so peers can inspect the arguments line by line. Some manuscripts include Lean proof artifacts — mathematical statements and reasoning written as strict code that a program can check.

Ten papers also ship with abridged reasoning traces. They show part of the exploration process, including attempts that did not work.

These manuscripts are effectively preprints: research texts released openly for reading, discussion, and checking. A preprint can contain major results, but its appearance says nothing about external peer review.

OpenAI says the overwhelming majority of results came from a single pipeline and a single internal model. If the major claims hold, that means one model system participated in research across many branches of mathematics.

The announcement also converted the average compute per result into a familiar unit: roughly three hours of thinking time on ChatGPT Pro.

What 722, 719, and 372 Actually Count

The first two numbers count manuscripts. The third counts groups. OpenAI calls each group a result family; the repository itself uses the plainer phrase "related papers."

Think of a family as a research folder. Inside it, papers divide the labor: one proves the main theorem, one fills a critical gap, another extends the conclusion to different conditions. Each is a separate manuscript, while the folder shares one family number.

A single paper can contain several conclusions, and a single folder can answer several different questions.

AI math manuscripts organized into result-family folders in a network, with one large family of fourteen papers

The largest folder, family 034, contains fourteen manuscripts in algebra and complex geometry, including both a general argument and a version specialized to four-dimensional objects.

A smaller example makes the logic clearer. Family 088 has only two papers, both studying a geometric indicator loosely connected to the "shadows" a shape casts in every direction. One paper asks which shapes minimize that indicator. The other supplies a counterexample showing that the presumed maximizer is not maximal at all. Same indicator, different questions, one family.

Grouping like this is common. Of the 372 families, 205 contain a single paper and 167 contain multiple papers. The multi-paper families are less than half the total, yet they hold 514 manuscripts — about 71 percent of the current collection.

Seventeen Fields, Seven Dominant Clusters

Across OpenAI's own subject classification, the 372 families span seventeen fields. The seven largest groups account for 227 families, about 61 percent. Theoretical computer science leads with 40 families, combinatorics follows with 37, and algebra with complex geometry contributes 36.

The fields ask different kinds of questions. Theoretical computer science measures how much computation an algorithm requires. Combinatorics studies how discrete objects can be arranged and counted. Algebra and complex geometry examine the properties of shapes defined by equations.

Headcounts tell two stories. Theoretical computer science contributes 40 families and 73 manuscripts, while probability and statistical mechanics contributes 29 families and 105 manuscripts. The first field has more folders; the second has more pages.

Five Shapes a "Breakthrough" Can Take

Based on a sampled reading of the math manuscripts, the contributions fall into five forms. This is not OpenAI's grading system, and one paper can combine several of them.

An answer to an open question. Pi can be approximated endlessly by fractions. The research asks how fast the error can drop, infinitely often, relative to the denominator. The new paper claims the "irrationality exponent" measuring that speed equals 2.

Wider applicability for an existing conclusion. Studies of particles interacting with electromagnetic fields often require the initial state to be small or specially symmetric. The new paper claims both restrictions can be removed, under the stated model and initial conditions, while long-term smooth solutions still exist.

Crossing a line presumed impassable. Longer integers usually demand more multiplication steps. The new paper claims that, under one fixed machine model, the number of steps can grow more slowly than the long-conjectured limit.

A premise shared by many conclusions. Numerous studies assume the Unique Games Conjecture, then prove limits on how well certain algorithmic tasks can be approximated. If the new paper holds, those studies no longer need to list it as an unproven assumption.

A new route to a known answer. One geometry counterexample paper openly states that "their result already gives a negative answer," then offers a direct construction. The added value is the proof path itself — whether it yields a clearer explanation or a reusable method remains to be seen.

The Three Withdrawals and a Twenty-Seven-Year-Old "If"

The withdrawn papers are the algebraicity of Weil classes on split abelian eightfolds, the algebraicity of Kuga–Satake correspondences for K3 surfaces, and the rational Hodge conjecture for products of K3 surfaces.

The latter two both involve K3 surfaces, and both use rewritten versions of the faulty construction from the first.

Crucially, a withdrawal is not a verdict that the statement is false. A gap in one argument does not disprove the proposition, nor does it rule out a different, valid proof.

All three belonged to family 032, which shrank from eight papers to five. The main Hodge paper for CM abelian varieties survived. CM abelian varieties are geometric objects with special arithmetic symmetry — and that surviving paper reaches back twenty-seven years.

In 1999, mathematician James Milne proved a conditional link: if one conjecture holds, another follows. What was missing was the first conjecture itself.

The first is the rational Hodge conjecture for all CM abelian varieties over the complex numbers. The second is the Tate conjecture for all abelian varieties over finite fields. One question lives among symmetric geometric objects; the other enters arithmetic systems with only finitely many elements.

The new math manuscript claims to prove the first. If it holds, Milne's already-proven bridge delivers the second, carrying the result beyond the objects it directly studies. On October 7, Milne added a note to his own paper catalog mentioning OpenAI's announcement and describing the consequence with "it seems" — a signal that he, too, sees the connection.

Why These Problems Suit an AI

OpenAI says the starting point for expanding its mathematics research was that "evaluations saturated" — standard problem sets no longer separated capable models from weaker ones. The company fed the model roughly 4,000 research questions, then filtered and grouped the output.

The suitable questions share a trait: even with no known answer, the task itself can be written precisely. A public task sheet lists the objects, the input conditions, the conclusion to prove, and the kind of counterexample that would settle it in the opposite direction.

Finding the route is left to the model.

An iterative proof-search loop as a circular maze, the path rerouting after hitting a dead end

One abridged trace records the design of a test meant to admit valid answers while rejecting invalid ones. The model discovered that one design exposed exploitable structure: invalid answers could slip through. Denser random perturbations blocked the exploit but lowered the probability that valid answers passed. The search continued — close the hole without breaking the guarantee.

That is what "finding a proof" becomes in practice: a repeatable loop of proposing an approach, locating where it fails, adjusting the route, and calling on theorems already proved by humans.

The favorable conditions are visible in hindsight: clear goals, existing research supplying reusable tools, and results that sometimes accept formal verification.

Paying Humans to Understand

Tools are not enough. Researchers need time to identify which steps are essential, which methods can be taught to students, and which ideas transfer to other problems.

On October 6, OpenAI announced funding for "workshops, conferences, and focused events" to help humans understand the important mathematical results produced by AI.

AGMAI, an advisory group on mathematics and artificial intelligence, had already sketched the options on September 29: working groups, summer schools, support for students or postdocs, and expository articles and books. It recommended that established nonprofit institutions choose the projects through mature processes, so understanding remains "community-led and community-driven."

For now, these remain advisory recommendations. OpenAI has made the funding commitment, but the October 6 announcement lists no budget, no named recipients, and no implementation dates, and it does not confirm that the proposed allocation mechanism is in place.

The math manuscripts are public. The verdict on them is not.

References

[1] Sharing AI progress in mathematics, OpenAI (https://openai.com/).

[2] OpenAI mathematics results public repository: project README, results catalog, and revision log, OpenAI, GitHub.

[3] Withdrawal note: Algebraicity of Weil classes on split abelian eightfolds, OpenAI.

[4] Withdrawal note: Algebraicity of Kuga–Satake Correspondences for K3 Surfaces, OpenAI.

[5] Withdrawal note: The rational Hodge conjecture for products of K3 surfaces, OpenAI.

[6] Gil Kalai, "Updates: Sharing AI progress on mathematics (amazing!); and my lecture plans," Combinatorics and more.

[7] Gil Kalai, "Amazing: There is no Percolation at the Critical Probability in all Dimensions," Combinatorics and more.

[8] "A Severe Misalignment of AI in Mathematics," joint statement initially signed by 25 Fields Medalists, Terence Tao's blog.

[9] Comparator project documentation, Lean Prover, GitHub.

[10] The irrationality exponent of π is 2, OpenAI.

[11] Integer multiplication below n log n, OpenAI.

[12] The Unique Games Theorem, OpenAI.

[13] Quasipolynomial Bounds for Arithmetic Progressions, OpenAI.

[14] The Quasi-Riemann Hypothesis: A Zero-Free Half-Plane Re(s)>7/8, OpenAI.

[15] The rational Hodge conjecture for CM abelian varieties, OpenAI.

[16] Global classical solutions of the three-dimensional relativistic Vlasov–Maxwell system, OpenAI.

[17] A product counterexample to the simplex maximum for projection-body volume, OpenAI.

[18] Subhash Khot, On the Unique Games Conjecture.

[19] Scott Aaronson, "The Mathocalypse" and comment thread (including Dana Moshkovitz's reading account), Shtetl-Optimized.

[20] J. S. Milne, Lefschetz Motives and the Tate Conjecture, 1999.

[21] J. S. Milne, "Mathematical Articles (with abstracts)," updated October 7, 2026, personal website.

[22] Ordinary NP-Hardness at the Basic Semidefinite Threshold, abridged reasoning trace, OpenAI.

[23] Mathematics and Artificial Intelligence advisory group: members and independence statement, AGMAI website.

[24] AGMAI, "On OpenAI's Release of Mathematical Results," October 6, 2026.

[25] AGMAI, "Responsible Release of AI-Generated Mathematics," September 29, 2026.

All illustrations are AI-generated.

全部评论