Eighty-eight hours — one side: about ten thousand AI agents, running September 1–5, that OpenAI says made real progress on Navier–Stokes, one of the seven Clay Millennium problems, each worth a million dollars, open ninety years.

The other side is one year: the time Tristan Buckmaster, an NYU professor, and Levent Alpöge, a mathematician at Anthropic working in a personal capacity, spent on the same equations, publishing the day before OpenAI’s announcement. They did the work inside OpenAI’s Codex; every draft lived in a Codex session.

When Buckmaster asked whether the model that “finished” in 88 hours had been trained on that year, the answers came at two speeds. On access to their sessions: “the model did not look up user data.” On training: nothing. “I asked again, about training,” he wrote, “and I did not get an answer.”

The miracle story and the cheating story are writing themselves — and both miss it.

What OpenAI says happened

The Navier–Stokes equations describe how fluids move — air over a wing, water through a pipe, blood through a vessel, weather across a continent — the mathematics of turbulence. The prize question, in Terence Tao’s plain phrasing, relayed by The New York Times (2026): is there an initial configuration where the water splashes around and, at some point, all the energy implodes and sends the water particles off at infinite velocity? Do solutions stay smooth forever, or blow up in finite time?

On Tuesday, September 8, OpenAI’s blog said an unreleased internal model — trained from about August 28 and “significantly more capable,” in the company’s words, than its released GPT-6 Astra — had taken a swing with about ten thousand concurrent agents. They exchanged about three million messages and 130 billion output tokens on Navier–Stokes alone. Cost: roughly ten million dollars at OpenAI’s own pricing (BBC 2026); duration: 88 hours.

OpenAI says it resolved two of the four statements the Millennium Prize demanded, then adds: “We do not intend to claim the Millennium Prize for this result.” Commentators note the reported proof covers the forced equation — a fluid driven by an outside force — not necessarily the full problem as stated. Nothing is independently verified; Clay has accepted nothing; the model is unreleased, the run unreproducible. From the BBC to The Verge (2026), the note is skepticism — earned.

~1 year — Buckmaster & Alpöge (AI-assisted, machine-verified) ≈ 8,760 hours 88 hours — OpenAI’s ~10,000 agents (Sep 1–5) ≈ 1% of a year — a sliver at this scale The same 88 hours, zoomed in — September 1 to 5 Sep 1 Sep 2 Sep 3 Sep 4 Sep 5 ~3M messages exchanged · 130B output tokens · ≈$10M in compute
Both runs to scale: a year of human-plus-machine work versus 88 hours of agents — about one percent of a year. The strip zooms into September 1–5; the red tick at September 3 marks when Buckmaster says word of his progress reached OpenAI.

The day before

For about a year, Buckmaster and Alpöge worked these equations the modern way: OpenAI’s Codex for the heavy lifting, Astra alongside, Anthropic’s Claude in the mix. The result, published the day before OpenAI’s announcement (statement), was an AI-assisted proof of finite-time blow-up for several fluid equations under smooth forcing, machine-verified in Lean. Terence Tao praised it.

On September 3 — two days into OpenAI’s run — Buckmaster says he learned that “information about our progress had been passed to OpenAI,” and that OpenAI began work only afterward. He asked whether the model had been trained on, or given access to, their Codex sessions. First answer: “the model did not look up user data.” He asked again, about training. Nothing. He further alleges the company pressured him to drop Alpöge from the authorship and floated how its result would be announced.

Hear OpenAI out. Sébastien Bubeck said at a press briefing that the company did not see the pair’s work until it was public: “our proofs differ significantly and even the precise results proved are different.” The blog adds that “no specific user data was accessed in order to solve this problem” — then the sentence that turned this into a story: “while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

Buckmaster’s retort lands like a verdict: OpenAI is, he says, “openly admitting they used training data from a period after we found our result.”

A model can produce a genuine result without anyone opening a specific user’s file, and a company can still train on that user’s thinking without knowing which session it came from. That is not a contradiction. It is the problem.

From “machines can’t” to “whose machines?”

Yann LeCun has long insisted machines cannot do mathematics. Nobody asks whether the result is machine-made anymore; the mathematicians ask whose machine, trained on whose thinking, and who gets the credit. You only litigate the authorship of something real. The capability question was not answered. It was promoted.

The question nobody can answer

Put OpenAI’s two sentences side by side. “No specific user data was accessed in order to solve this problem” is about access. “We cannot rule out that de-identified data helped improve our models” is about training. Between them sits the question: did the model that ran September 1–5 get better because of what two mathematicians did in private sessions? Nobody has answered it. Not OpenAI. Not the mathematicians. Not the Clay Institute, which has no framework for machine claimants. The sport has no referee because nobody wrote the rules.

If you think your hardest thoughts inside someone else’s tool, you are Buckmaster; your reasoning is the feedstock. A model need not read your file to absorb your work — only to have been trained, in aggregate, on people like you. The old rules of attribution assumed a visible chain: preprint, seminar, citation. Machine learning breaks the chain; the transfer is invisible, statistical, unpointable — which is why it cannot be ruled out, and why “we did not look” will never again mean “we did not use.”

The quiet casualty

Mathematics was the most democratic of the sciences. The amateur with paper and pencil stood level with the professor: a proof’s only currency was public verification. What this episode privatizes is not mathematics but its means — a secret model, a ten-million-dollar run, training data from users’ private sessions. This is the first headline result in mathematics with a data moat. You cannot check the work without the model, and you cannot have the model without the money.

Tao has lamented companies using landmark open problems as marketing proof points. His deeper worry is incentives: when a lab can deploy massive on-demand compute the instant it catches a rumor, the rational response to a promising idea is to stop sharing it. He calls open problems “lighthouses”; lighthouses only work if ships can see them. Science/AAAS (2026) worries the episode will convince decision-makers that human mathematicians are obsolete, and young people that their passion has no future. Fortune (2026) frames it as two companies “racing one another to Armageddon” — a race whose absurdity peaks when the runners are on each other’s tracks: Anthropic’s mathematician spent a year inside OpenAI’s product, and if Buckmaster’s timeline holds, OpenAI’s agents spent 88 hours sprinting over the same territory.

Four rules while there is still time

None of these is exotic.

First, auditability. Point a model at a named open problem and its maker should certify training-data provenance: what it was trained on, and when the data stops. Not “we looked.” Certified.

Second, rumor protocols: labs should commit to not racing on whispers — Tao’s point, made into policy. Work you hear about in private stays off-limits until it is public.

Third, the Clay Institute should publish rules for machine claimants now — before the first credible claim. What counts as a resolution? What must be disclosed? Who referees a model nobody else can run?

Fourth, publish early — it just became the winning move. Buckmaster and Alpöge published first and own the moral high ground even though OpenAI’s model “finished faster.” In the new regime, first to publish beats first to prove, inverting the centuries-old norm of publishing only what is verified. That inversion is dangerous — and it is the antidote: against a data moat, the preprint is the only weapon that works.

So say the mathematics plainly first: unresolved. Two statements of four, on the forced equations — unverified, unclaimed, unaccepted. But that was never the story. OpenAI’s own sentence — “while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models” — is the first plain admission of a new kind of borrowing: no receipt, no trace, no remedy. You cannot cite an average. You cannot sue a weight.

That is the punch, and it is not that OpenAI cheated. It is that this kind of cheating cannot be shown — and the rules that would show it do not exist. Buckmaster asked twice. The first time, he got an answer. The second time, about training, he did not. If you think your hardest thoughts inside someone else’s tool, that silence is addressed to you.

Sources

  1. OpenAI — company blog post, 8 September 2026 (model and agent-run claims; all quoted statements are as cited in the article).
  2. T. Buckmaster — statement on the timeline and authorship dispute, 2026: cims.nyu.edu/~tristanb/statement.pdf.
  3. BBC News, 2026 — reporting including the ≈$10M compute estimate at OpenAI’s own pricing: bbc.com.
  4. The Verge, 2026 — coverage of the announcement: theverge.com.
  5. The New York Times, 2026 — Tao’s plain-language account of the Millennium question: nytimes.com.
  6. Fortune, 2026 — “racing one another to Armageddon”: fortune.com.
  7. Science / AAAS, 2026 — reporting on community and career worries: science.org.
  8. Clay Mathematics Institute — the seven Millennium Prize problems; rules and status: claymath.org.