An AI Disproved a Problem Mathematicians Sat On for 80 Years
A general reasoning model was handed a 1946 question from Paul Erdos and returned a 125 page counterexample. No hints, no step by step guidance, no training for the task.
Tuesday, August 4, 2026/5 min read

There is a difference between a machine that answers questions and a machine that finds something nobody knew. For most of the AI boom, the industry has been selling the first while gesturing at the second.
In May, that line moved. An OpenAI reasoning model was given the planar unit distance problem, a question Paul Erdos posed in 1946 about how many pairs of points in a plane can sit exactly one unit apart, and it produced a counterexample that disproves the conjecture.
It was not trained for mathematics. It was not walked through an approach. It was handed the statement and returned a complete proof running to 125 pages.
What the model actually found
The result is not a faster version of something a person had already sketched. According to OpenAI's account of the work, the model reached for point arrangements drawn from algebraic number theory, built by way of the Golod-Shafarevich criterion for infinite class field towers.
That construction beats the near linear square grid bound that mathematicians had long treated as essentially the best possible. The interesting part is not the improvement, it is the route. Nobody prompted it toward class field towers.
Why mathematicians are taking it seriously
Scepticism is the correct default with AI claims, and it usually survives contact with the details. Not here. Scientific American reported the reaction across the field as genuine astonishment rather than the usual polite interest.
Fields medallist Tim Gowers, writing in a companion paper, called the result a milestone in AI mathematics. That matters because Gowers has been unusually specific over the years about what machine assistance in mathematics would have to look like before it counted, and because he predicted something like this in 2000, imagining computers taking the routine checking while people chased the deeper ideas.
Writing in The Conversation, researchers made the point that lands hardest: a result of this kind would have earned a top journal and wide coverage if a person had produced it alone.
The caveat worth keeping
An analysis at Understanding AI argues the breakthrough played to the model's strengths rather than demonstrating general mathematical ability. That is a fair reading and it should temper the framing.
Disproving a conjecture means producing one counterexample. That is a search problem, enormous but bounded, and it is exactly the sort of thing a system with vast recall and tireless patience should be good at. Proving a theorem true is a different discipline, and nobody has shown that yet.
There is also the verification question. A 125 page argument has to be checked by people, and the checking is where a result becomes knowledge rather than a claim.
What this changes for everyone else
Most readers will never care about unit distances. They should care about the category shift.
Until now the honest description of these systems was that they compress and recombine what humans have already written. Producing a result that no human had produced, in a field where correctness is checkable and not a matter of taste, is a different claim about what the technology is.
Mathematics is the ideal proving ground precisely because it is unforgiving. A proof is right or it is not. There is no marketing department that can make a false one true.
What the problem actually asks
The unit distance problem is easy to state and brutal to answer, which is the Erdos signature. Place a number of points on a flat plane. Count how many pairs sit exactly one unit apart. As the number of points grows, how fast can that count of unit distances grow?
Erdos asked it in 1946 and the gap between the best known upper and lower bounds has stayed embarrassingly wide ever since. The natural guess is a square grid, and the near linear bound it produces was widely treated as close to the truth. The model showed it is not.
What makes the result unsettling is the distance between the question and the tools. Nothing about counting points at fixed distances suggests class field towers from algebraic number theory. That connection is the kind of thing mathematicians describe as taste, and taste was supposed to be the part machines lacked.
The verification problem nobody has solved
A 125 page proof is not self certifying. Somebody has to read it, and reading a long argument carefully is slow work that carries no professional reward if it turns out to be correct.
This is becoming the practical bottleneck. If systems can generate candidate proofs faster than humans can check them, the field acquires a queue of results that are probably true and not yet knowledge. Formal proof assistants are the obvious answer, but translating an informal argument into one is itself a substantial job.
The discipline has run into something like this before, with computer assisted proofs such as the four colour theorem, and the argument about what counts as understanding rather than verification has never fully closed.
Why mathematics was always going to be first
If machine research was going to arrive anywhere, this was the obvious field, and the reasons are structural rather than flattering to the technology.
Mathematics needs no laboratory, no ethics approval, no funding for materials and no waiting for an experiment to run. The entire discipline exists as text, which happens to be the format these systems consume and produce. A chemistry hypothesis has to be tested against physical reality. A proof only has to be tested against logic.
Correctness is also unusually decidable. In most fields a result is judged by a community over years, weighing methodology and replication. In mathematics an argument either holds or contains an error somebody can point at.
That combination makes it the cleanest possible testbed, and it also limits what the result implies about everything else. A system that produces valid proofs has not demonstrated that it can design a drug trial or interpret an ambiguous experiment.
What to watch next
Two things. First, whether the proof survives peer review intact, because a milestone that quietly acquires corrections is a smaller milestone. Second, whether this repeats on problems that reward construction rather than search.
If it does, the interesting question stops being whether machines can do research and becomes what mathematicians do with the time they get back.
Published in The Outspoken Digest
Newsletter
The Digest, in your inbox
One edition, sent when it is ready. No noise, and your address is never passed on.



