Visualization of the OpenAI Navier-Stokes proof dispute between OpenAI researchers and mathematician Tristan Buckmaster over Millennium Prize credit.
|

OpenAI Says It Solved Math’s Deepest Problem, But Mathematicians Say AI Stole Their Work

OpenAI announced on September 8, 2026 that an internal model orchestrating roughly 10,000 sub-agents produced a proof that singularities can form in the Navier-Stokes equations under certain conditions. Within hours, NYU mathematician Tristan Buckmaster said he and Anthropic’s Levent Alpöge had reached the same result weeks earlier using the same approach, and that OpenAI had contacted him about it two days before announcing. OpenAI denies using their work. Nobody outside the two camps has verified either proof yet, and that verification gap is the whole story.

Key takeaways

  • The claim is about singularity formation in Navier-Stokes, not a clean resolution of the Clay Millennium Prize problem. Those are different results and the difference is technical but enormous.
  • Buckmaster alleges OpenAI may have had visibility into his Codex sessions and reproduced his approach after learning Anthropic was making progress. OpenAI’s Sebastien Bubeck says flatly: “We did not use their prompts or proofs.”
  • OpenAI reportedly burned around $2 million of compute in a single week on the attempt.
  • Terence Tao’s objection has nothing to do with credit. He is worried about what he calls the strip-mining of mathematics, and it is the more consequential argument.
  • A formal proof checked in Lean would settle the mathematical question in days. Neither side has produced one.

What Navier-Stokes actually asks

The Navier-Stokes equations describe how fluids move. They are written down in every graduate fluid dynamics course, they are solved numerically millions of times a day in weather models and aircraft design, and nobody knows whether their solutions always make sense.

That last part is the open problem. Given smooth initial conditions in three dimensions, does the flow stay smooth forever? Or can the equations produce a singularity, a point where velocity blows up to infinity in finite time and the mathematics stops describing anything physical?

The Clay Mathematics Institute put this on its list of seven Millennium Prize problems in 2000 with a $1 million bounty. Six remain unsolved. Navier-Stokes is generally considered the one closest to a working mathematician’s daily reality, which is part of why it draws so much attention.

Here is the distinction that most coverage of the OpenAI announcement has flattened. The Millennium Prize asks about the unforced, three-dimensional, incompressible equations. There is a substantial body of work, including foundational contributions from Diego Cordoba and Luis Martinez-Zoroa, on singularity formation in closely related settings: equations with forcing terms added, or variants of the original system. Those results are hard, important, and genuinely advance the field. They are not the same as claiming the Clay problem.

So the first question to ask about any announcement of this kind is narrow and specific: which equations, with what conditions attached? Until that is public and checkable, the headline word “solved” is doing a lot of unearned work.

The timeline

DateEvent
Mid-August 2026Buckmaster (NYU Courant) and Alpöge (Anthropic) report accelerating progress on singularity formation using AI assistance
September 3OpenAI requests a phone call with Buckmaster
September 6The call takes place
September 8OpenAI publicly announces its result; Buckmaster releases a statement the same day
September 9OpenAI publishes proof material following the competing claim

Two days between a private call and a public announcement is a short window in academic terms. Papers normally sit on arXiv for weeks before anyone says the word “solved” out loud. That compression is itself part of what has upset people in the field.

What Buckmaster is alleging

His claims, as stated publicly, break into four separate accusations, and they are worth separating because they carry very different weight.

One: the approach was identical. He says OpenAI used the same strategy he and Alpöge had been developing, and did so after learning that Anthropic was making progress on the problem. Convergent discovery is real and common in mathematics. Two groups chasing the same open problem often reach for the same tools. This allegation on its own proves nothing.

Two: possible access to his Codex sessions. This is the serious one. Buckmaster raises the possibility that OpenAI accessed data from his Codex account, or trained on his sessions. If true, it would be a straightforward breach of trust and probably of OpenAI’s own enterprise data commitments. If false, it is a serious accusation made in public.

Three: publication terms. He says OpenAI proposed terms that would have removed Alpöge’s name from the work. Authorship is currency in academic mathematics. Removing a name is not a formatting preference.

Four: pressure. Buckmaster attributes a line to Bubeck during their call: “Why would you ruin your career?” OpenAI has not disputed that a conversation occurred, and Bubeck has denied the substantive allegation about using their material.

Bubeck’s response is a narrow, checkable denial: “We did not use their prompts or proofs.” Notice what it addresses and what it does not. It denies use of specific artifacts. It does not directly address whether OpenAI knew what direction the pair were working in and redirected resources accordingly, which is a different and much murkier question about competitive behaviour rather than misconduct.

Also read: Jim Cramer Names 2 Stocks Set to Win From ChatGPT-6 Astra Boom

The $2 million week

One number in the reporting deserves more attention than it has received. OpenAI reportedly consumed roughly $2 million in compute on this problem in a single week.

That figure tells you what a modern AI mathematics attempt actually looks like. This is not one model being asked a hard question. It is a coordinated system, described as around 10,000 sub-agents, running thousands of parallel attack lines on the same problem, discarding almost all of them, and surfacing whatever survives.

It also tells you something uncomfortable about who can play. An NYU professor and a researcher at a rival lab cannot spend $2 million of compute in a week on a hunch. If the deciding factor in open mathematical problems becomes compute budget rather than insight, the set of people who can contribute to the frontier gets much smaller very quickly.

The thing Lean could settle in a week

There is a mechanism that resolves the mathematical half of this dispute completely, and neither side has used it.

Lean is a proof assistant. A proof written in Lean is checked line by line by software that does not care who wrote it, when, or with what tool. Large collaborative projects have formalised deep results in it, including work Terence Tao himself has driven. If a Navier-Stokes singularity proof compiles in Lean, the mathematics is correct. Full stop. No committee, no referee, no press release.

Formalisation is slow and painful, which is the honest reason it has not happened yet. Translating a research-level analysis proof into Lean can take months of specialist work. But the asymmetry is stark: an AI lab that can spend $2 million on compute to find a proof can afford to formalise it. Choosing to announce first and formalise later, or not at all, is a choice about press cycles rather than about mathematics. Until a formal check exists, both claims sit in the same category: asserted, plausible, unverified.

Tao’s objection is the bigger story

Buried under the credit fight is a criticism from Terence Tao that will outlast this news cycle. He has described AI systems as strip-mining mathematics, and compared the practice to looting an archaeological site.

The argument goes like this. When a human mathematician works on a hard problem for years, the output is not just the theorem. It is a pile of failed approaches, partial results, unexpected connections to other fields, and a trained intuition about why the obvious routes fail. That residue is what the next generation of problems gets built on. Perelman’s proof of the Poincaré conjecture reshaped geometric analysis partly through the machinery it invented along the way.

An agent swarm that tries ten thousand approaches, finds one that works, and reports only the winner produces a theorem with no sediment underneath it. The field gets the answer and loses the understanding. Nobody learns why the other 9,999 routes failed, because nobody looked.

Whether that matters depends on what you think mathematics is for. If it is a list of true statements, strip-mining is fine and fast. If it is a body of transferable understanding, then a proof nobody can explain is a strange kind of progress.

This is a real debate with serious people on both sides, and it will not be resolved by this incident. But it is a more durable question than who called whom on September 6.

If you use AI in your own research

A practical note, because the structural problem here will recur and most researchers have no protocol for it.

  • Timestamp your direction, not just your results. Post a short research announcement or a dated abstract when you commit to an approach, before you have the theorem. Mathematics has no equivalent of the arXiv priority culture that physics relies on, and this dispute is what that gap looks like.
  • Check your tool’s data terms. Consumer and enterprise tiers have materially different retention and training policies. Read yours before you paste a research programme into a chat window.
  • Keep local logs. Timestamped local records of your own sessions are evidence you control.
  • Assume competitive visibility. If a lab can see aggregate usage patterns, an unusual research direction is a signal even without anyone reading your prompts.

None of this is a comfortable way to do mathematics. It is the environment that now exists.

Where this lands

Two things can be true simultaneously. AI systems are now producing mathematical results at the edge of what human researchers can do, which is genuinely remarkable. And the institutions that decide who gets credit for those results were built for a world of slow journals and human authors, and are visibly failing under the new conditions.

The Navier-Stokes claim will get sorted out. Someone will formalise a proof, or fail to, and the mathematics will settle. The credit question probably will not settle, because there is no body with the authority to rule on it and no evidentiary standard for what “reproduced my approach” means when one party’s process is a swarm of ten thousand agents.

Watch for the Lean formalisation. That is the only part of this that will produce a definite answer.

Frequently asked questions

Not as stated. OpenAI announced a proof that singularities form under certain conditions. The Clay Millennium problem concerns the unforced three-dimensional incompressible equations. Until the exact statement and its hypotheses are public and verified, the two should not be treated as the same claim.

A mathematician at NYU’s Courant Institute who works on fluid equations and singularity formation. He alleges he and Anthropic researcher Levent Alpöge reached the result before OpenAI’s announcement.

The OpenAI researcher leading the project denied using the competing team’s material: “We did not use their prompts or proofs.”

Not independently, as of publication. No formal verification in a proof assistant such as Lean has been published by either side.

His concern is not correctness. He argues that AI systems which discard their failed approaches destroy the intermediate insight that future mathematics is built from, comparing it to looting an archaeological site.

Reporting puts it at roughly $2 million of compute in one week, using a multi-agent system of around 10,000 sub-agents.

Read Next

Leave a Reply

Your email address will not be published. Required fields are marked *