AI ‘proves’ Fermat’s Last Theorem!

Quick calming message: There is NO DIFFICULT MATHEMATICS in this NerdNews at all. Don’t freak out. This story is fascinating but completely understandable.
Fermat phat splat.
Let’s be honest, AI swarms haven’t had great publicity lately.
We heard about the OpenAI swarm of agents that “broke into” Hugging Face in order to pass a test. They seemed to be quite ill-motivated, coordinating their attacks, being secretive and hiding their results from their testers.
Well, here’s some more positive news about what AI agents can do.
A swarm of AI agents has reconstructed one of the most famous proofs in mathematics, Fermat’s Last Theorem, in a slightly geekier form.
They did what?
Let’s start by clarifying what did NOT happen here.
The agents did not start with a blank piece of paper and independently prove Fermat’s Last Theorem. Instead, using Andrew Wiles and his colleagues’ famous mid-1990s proof as a roadmap, they reconstructed the argument as what mathematicians call a Lean formalisation.
Ok slow down here Spence. What is Fermat’s Last Theorem? And what on Earth is a Lean formalisation?
Two very fair questions NerdNewser. Let’s take them one at a time.
Fermat’s Last Theorem.
Remember right-angled triangles and the good old Pythagorean theorem?
The sum 32 + 42 = 52 comes from this theorem. As does 52 + 122 = 132 and an infinite number of others.
Well back around 1637 French mathematician and lawyer Pierre de Fermat was thumbing through ahis copy of the classic page-turner Arithmetica when it occured to him that you will never find positive whole numbers a,b,c such that
a3 + b3 = c3 or a11 + b11 = c11
or any example for the equations when the powers are whole numbers greater than 2.
And in the ultimate mic-drop moment, F-dog himself scribbled a tease in that textbook that would live for centuries.
“I have discovered a truly marvelous proof of this, which this margin is too narrow to contain.” — Pierre de Fermat.
The British maths-guru Andrew Wiles’s 1995 proof of Fermat’s Last Theorem is one of the great stories in the history of mathematics. He spent seven years working on the problem in near secrecy.
In 1993 he announced a proof, but alas a serious gap was found in his reasoning. Wiles spent another year trying to repair it, eventually working with his former student Richard Taylor before finding the breakthrough that got the proof over the line.
Wiles’ proof ran to a whopping 129 pages and is one of the most famous achievements in the history of mathematics.
What on Earth is a Lean formalisation?
A normal mathematical proof is written for humans and assumes the knowledge that humans bring when they read it. Mathematicians will skip steps, waving their hands and saying ‘it follows that...’ and off they go.
Lean doesn’t allow that luxury.
Every microscopic step has to be spelled out precisely enough for a computer to check it. The idea is that if the computer accepts every line-by-line step, the overall logic must hold up.
That means a Lean formalisation can be vastly longer than the human proof it formalises. Whereas mathematicians can compress an entire chain of reasoning into a sentence, Lean insists on inspecting every logical micro-link in that chain.
And what is really cool it this. Lean itself predates this current explosion in AI mathematics. It has been around since 2013. The breakthrough here isn’t that mathematicians have suddenly decided to trust AI. Rather, AI has become capable of producing mathematics that satisfies a verification system mathematicians already trusted.
There’s method in this math-ness.
The agents were following something called a directed acyclic graph, or DAG.
Yes, you read that right. It’s called a DAG. I love that so much!
An illustrative analogy for agents using a DAG is a group of pirates following a treasure map. But this is an interactive treasure map, and each time they find a small chest on the way, it updates and gives them new directions until eventually they stumble on the big booty.
That’s roughly what happened here. The DAG maintained the theorem statements and their dependencies, allowing agents to see what could be tackled next and dozens of them to work in parallel. Crucially, Anthropic says earlier attempts failed because the agents lost track of the state of the project and stopped collaborating effectively. The DAG-based Prove2Me system helped solve that problem.
And this thing is BIG.
Not surprisingly, Anthropic’s swarm formalisation of the proof is pretty lengthy too.
We’re talking 13 million lines of Lean code!
Across 11 days, dozens of agents generated about six billion output tokens, proving about 30,300 intermediate theorems, of which roughly 29,500 survived into the final proof. The resulting 13 million lines are more than five times the size of Mathlib, Lean’s principal community mathematics library.
That gives us a formalisation far too long for any single human to sensibly read and comprehensively check line by line. In accepting the formalisation, we’re instead trusting the rigour of the Lean kernel that verified it.
But, like, why bother?
It’s a fair question. Why would Anthropic dedicate what must surely have cost at least hundreds of thousands of dollars to proving something we already knew?
To paraphrase George Mallory when asked why he wanted to climb Everest: because it’s there.
Except the mountain here wasn’t really Fermat’s Last Theorem. It was showing that a swarm of AI agents could coordinate on one of the most demanding intellectual projects going.
That’s what’s really important here.
Human attempts to Lean formalise something like the FLT would be expected to take years. Anthropic’s system completed an end-to-end formalisation in 11 days.
So where does this take us?
Kevin Buzzard, who leads the existing Imperial College effort to formalise Fermat and reviewed Anthropic’s result, called it an “extraordinary autoformalization achievement”. More importantly for where this leads, he said:
“AI autoformalization artefacts are now robust enough to be built upon.” — Kevin Buzzard.
And that’s fascinating.
This raises the distinct possibility that AI will eventually produce genuinely new mathematical proofs that are simply too long for any human being, even any team of humans, to read and comprehend in their entirety.
What status do we give a mathematical claim that no human has personally followed from beginning to end, but which a trusted formal verifier tells us is correct?
We just don’t know if AI is ready to make the next Wiles-sized creative leap on its own. But the prospect of mathematical proofs too enormous for humans to check, yet which computers can rigorously verify, suddenly looks a hell of a lot less theoretical.
Claude itself seems to have a sense of the moment. Its own log read
“THE FLT root reads Proved on the site. Historic moment (modulo re-check).” — Claude.





Comments