AI Keeps Embarrassing Mathematicians. They're Done Staying Quiet.

Bottom line: In October 2025, OpenAI executives claimed GPT-5 had solved ten Erdős problems.

It had mostly found existing literature, and Google DeepMind's Demis Hassabis called it "embarrassing." A year later, the dispute has changed: mathematicians now object less to wrong answers than to the volume of unverifiable ones.

After OpenAI's October 6, 2026 release of several hundred AI-generated proofs, some researchers called for a boycott.

The core problem is the same one your team has with AI-generated pull requests: generating is cheap and verifying is expensive.

I once approved a 4,000-line pull request in eleven minutes because the tests were green and the diff looked plausible. Three weeks later it took us two days to untangle what it had quietly broken.

That's the story I kept thinking about while reading the mathematicians' open letters this fall.

Everyone is framing this as "AI versus mathematicians." It isn't. It's a verification bottleneck, and if you ship software, you already live inside one.

The Embarrassment That Started It

In October 2025, an OpenAI executive posted that GPT-5 had "solved" ten problems from Paul Erdős's famous collection. Within days, mathematicians pointed out what had actually happened.

The model had found existing solutions buried in the literature, not produced new ones.

The post was walked back. Hassabis's public reaction was three words: "This is embarrassing." Then-Meta chief AI scientist Yann LeCun said OpenAI had been "hoisted by their own" enthusiasts.

That should have been the end of it, a lesson about checking before you tweet. Instead it became the first chapter of a much longer story.

The Pattern: Real Progress, Oversold Packaging

Here's the part the skeptics skip. Some of the later results were real.

In January 2026, Terence Tao said Erdős problem #728 had been solved "more or less autonomously by AI." The commentary that followed was careful: the model solved a community interpretation of the problem, and similar techniques already existed in the literature.

Still, it was a genuine step.

In May 2026, OpenAI said a model had disproved Erdős's 1946 planar unit distance conjecture.

This time Thomas Bloom, who had publicly criticized the earlier claims, co-authored a companion paper validating it.

The result showed Erdős's proposed limit was too low, though it did not establish how fast the pairs actually grow.

That's what a healthy cycle looks like. A claim goes out, outside experts check it, and the part that holds up gets credit. The mathematicians weren't hostile to the result.

They were hostile to the process that came before it.

What Changed in September

Then the volume went up, and the dynamic broke.

Over the past few weeks, the conflict has moved from "that claim is wrong" to something more structural.

According to TechCrunch's September reporting, the feud with mathematicians was escalating before the latest release even landed.

The Leiden Declaration in June had already put over 150 mathematicians on record warning governments not to believe the hype.

Later came an open letter from 25 leading mathematicians, and a statement from Fields Medalists, Tao among them, warning of a "severe misalignment" between AI industry goals and the practice of mathematics.

Then on October 6, OpenAI published a very large batch of AI-generated proofs. Reports vary on the count, with some outlets citing 722 manuscripts, and the numbers are still being sorted out.

Futurism described the reaction as fury, and The Decoder reported that some mathematicians called for an OpenAI boycott.

I'll flag my own uncertainty here. Several of the aggregator sites covering this are low-quality, and I'd trust the specific counts less than the overall shape.

The shape is consistent across every credible source: a flood of output, and a profession asking who is going to read it.

It's Not About Being Wrong

This is the part that surprised me. The loudest objections aren't "these proofs are false."

Article illustration

The objections are closer to these:

That last one is the real issue. In infrastructure terms, you've built a producer that runs three orders of magnitude faster than the consumer, and you've put no backpressure between them.

Anyone who has watched a queue grow without bound knows how that ends.

Why Mathematicians Aren't Like the Rest of Us

Software has a escape hatch that mathematics doesn't. When AI-generated code is wrong, production tells you. Tests fail, pages fire, customers complain. The error eventually announces itself.

A mathematical proof has no production environment.

The only runtime is a human expert reading it carefully, and that's a scarce, unpaid resource that also happens to be the training pipeline for the next generation of mathematicians.

Flood that reviewer pool with plausible-looking manuscripts and you don't just waste time. You erode the thing that makes the field trustworthy.

A proof is only worth something because someone competent checked it and put their name on the check.

This is also why I think the "they're just gatekeeping" take is lazy. These people are defending a verification layer, not a turf.

The Reality Check

I don't want to oversell the mathematicians' side either. Some of this resistance is a profession reacting to a threat, and not every complaint is principled.

The unit distance result is real. Tao's comments on #728 are real.

Formal tools like Lean genuinely do reduce the verification burden for certain kinds of results, and in the long run they may be the right answer.

A world where proofs ship with machine-checkable certificates is better than the one we have now.

But "Lean-screened" isn't the same as "the statement proven is the statement that matters." Formalization catches logical errors and can still miss a misstated theorem.

If you've ever had a test suite pass against the wrong spec, you know exactly how that feels.

So both sides are partly right, and I think that's the honest position. The capability is real. The release process around it is the problem.

What You Should Take From This

You don't need to care about the unit distance conjecture to learn from this fight. The same dynamic is about to hit your codebase, your docs, and your security reviews.

Article illustration

1. Count your review capacity before you scale generation. If your team ships AI-written code faster than it can be reviewed, you aren't moving faster. You're accumulating unverified debt.

2. Demand reproducibility on claims. If a vendor says a model did something impressive, ask which model, which version, and whether you can run it.

"Trust us" is the wrong answer, and mathematicians just reminded the industry why.

3. Separate "generated" from "verified" in your own workflow. Label them. Track them. Don't let the first quietly become the second because the CI badge is green.

4. Put a human name on every check. Someone should be accountable for the sign-off, the same way a proof carries its referee.

5. Treat retractions as data. The October 2025 walk-back is the cheapest lesson the industry will get on what happens when you announce before you verify.

What I'm Watching

The Clay Mathematics Institute has not verified the Navier–Stokes claim that OpenAI floated in September, and I expect that story to keep moving.

The next few months will show whether labs adopt something like staged disclosure: named models, external verifiers before the announcement, and credit rules agreed in advance.

If they don't, I expect more letters, more boycotts, and a field that treats every AI announcement as presumptively noise. That would be a loss for everyone, because some of the results are good.

The cost of crying wolf is that people stop checking the real ones.

I keep coming back to that eleven-minute approval. The code wasn't the failure. My review was, and I'd scaled my trust faster than my attention.

So here's what I'm curious about: in your own work, where has AI output outrun your ability to verify it, and what did you do when you noticed?

Story Sources

YouTubeyoutube.com