Fermat's Last Theorem: The Two Cases You Can Actually Prove
Here is a strange fact, and it’s worth sitting with why it’s strange before doing anything else. For , the equation has solutions everywhere — , pick any two positive integers and their sum is a third. For , it still has infinitely many solutions, and not just scattered ones: every primitive Pythagorean triple is generated by a clean two-parameter formula, so “solutions to ” is a completely understood, infinite, richly structured set. Then Fermat’s claim is that at this doesn’t just get harder — it goes to exactly zero, and stays at zero forever, for every integer exponent simultaneously. There’s no a priori reason the transition from “infinite family, fully classified” to “provably nothing, ever again” should be that abrupt. Proving a single equation has no solutions is one thing; proving that infinitely many equations, one for each exponent , all have no solutions, with a single stroke, is a different kind of claim, and it turns out to be enormously harder to actually establish than to believe.
Fermat wrote his claim in the margin of a copy of Diophantus’s Arithmetica around 1637, with a note — now famous, and worth being careful about — that he had “discovered a truly marvelous proof of this, which this margin is too narrow to contain.” That’s the standard paraphrase of the marginal note, not a verified verbatim quotation, and it should be treated as legend rather than historical fact: no such proof was ever found among his papers, and given what turned out to be required to actually prove the theorem, most historians of mathematics doubt Fermat had a correct general argument. What he certainly did have, and what he did leave a real proof for, is the case .
This post does two things honestly. First, it gives complete, checkable proofs of the cases and — genuine infinite-descent arguments, not sketches, that you could in principle verify line by line. Second, it explains why those methods, despite looking like they should generalize, provably don’t, and gives an accurate outline (explicitly not a derivation) of the real 20th-century strategy that finally closed the general case.
The precise claim, and a reduction
Fermat’s Last Theorem. For every integer , there are no positive integers with
The first real piece of progress, and one that’s genuinely worth deriving rather than citing, is that you don’t need to handle every exponent separately. You only need two cases: , and for an odd prime.
Claim. If the theorem holds for and for every odd prime , it holds for every .
Proof. Let .
- If , write with . Suppose had a positive integer solution. Then, setting , , ,
which is a positive integer solution to the equation. So if no such solution exists, none exists for either.
- If , then must have an odd prime factor: the only way it couldn’t is if were a power of , and since that power would have to be for , which is divisible by — excluded by assumption. So write for an odd prime and integer . The identical argument, with replaced by , shows a solution to would produce a solution with .
Either way, a solution for forces a solution for or for some odd prime .
So the entire theorem reduces to infinitely many prime exponents plus the single case . Fermat and Euler between them settled the two smallest, most tractable pieces of that reduction — completely elementarily, with one genuine subtlety that took another century to fully close.
The case : Fermat’s own descent
Fermat’s actual method here is infinite descent: assume a solution exists, and manufacture from it a strictly smaller solution to the same kind of equation. Since there’s no infinite strictly-decreasing sequence of positive integers, this is a contradiction, and no solution can exist in the first place. It’s worth proving something slightly stronger than needed, because the stronger statement is what the descent naturally produces:
Theorem. There are no positive integers with .
This immediately implies the case of Fermat’s theorem: a solution to is, in particular, a solution to with .
Proof. Suppose solutions exist, and pick one, , with as small as possible among all positive-integer solutions of .
Reduce to coprime . If a prime divided both and , then , which forces (compare exponents of on both sides), and dividing through by on the left and on the right produces a strictly smaller solution — contradicting minimality of . So .
First application of Pythagorean triples. Rewrite the equation as . Since , also , so this is a primitive Pythagorean triple. By the classification of primitive triples, exactly one of is even; relabel if necessary so is even and is odd. Then there exist coprime integers of opposite parity with
Second application, nested inside the first. Rearrange as
This is again a Pythagorean triple, and it’s primitive: any common prime factor of and would divide too (from ), hence divide , contradicting . Since is odd, must be the even leg, so there exist coprime integers of opposite parity — unrelated to the exponent or to any prime exponent discussed later, just local names for this one step — with
The perfect-square step. Now compute , so is a perfect square: . Moreover are pairwise coprime: by construction, and if a prime divided both and , it would divide and hence , contradicting (symmetrically for ). A product of pairwise coprime positive integers is a perfect square only if each factor individually is, so
for positive integers . Substituting into :
That’s a new solution of the exact same equation, .
The descent closes. Since forces , we get , so in particular . Then
so . We’ve produced a positive-integer solution with a strictly smaller third coordinate, contradicting the minimality of .
Every step here is elementary — nothing beyond the classification of Pythagorean triples and counting divisibility — but the structure (find a solution, extract a strictly smaller one, contradict minimality) is the entire idea, and it’s completely rigorous as stated.
The case : an elementary descent, plus one lemma
The case also falls to infinite descent, and — this may be surprising given how the last section ended — almost all of it is just as elementary as the proof: ordinary integer manipulation, parity, and one gcd case-split. There is exactly one place a ring bigger than is genuinely needed, and I’ve isolated it into a single, precisely stated lemma so the one real gap in this post is narrow and named, rather than spread through the whole argument.
It’s cleaner to work with the symmetric form of the equation. Showing has no solution in positive integers is equivalent to showing:
There are no nonzero pairwise-coprime integers with .
(Given positive with , take . Conversely, given such , they can’t all share a sign — the sum of three nonzero same-sign cubes is nonzero — and flipping the sign of all three if needed leaves exactly one negative, say , giving with everything positive.)
Setting up: parity, not divisibility by 3
Suppose with nonzero and pairwise coprime, and — among all such counterexamples — as small as possible, where is chosen to be the even one.
Exactly one of is even. If all three were odd, would be odd, not . If two were even, they’d share the factor , contradicting pairwise coprimality. So exactly one is even; relabel so it’s , meaning are both odd.
Since are odd, and are integers. They’re coprime: any common divisor divides and , and . They have opposite parity: is odd, so can’t both be even or both odd.
The key identity
Directly from the binomial expansion, . Since , :
is even, is odd. Whichever of is even, is odd (even+odd is odd, or odd+even is odd — either way). So the -adic valuation of the right side is exactly if , but at minimum if is odd. Since is even, has -adic valuation at least (as ); matching valuations forces , i.e. is the even one, and so is odd.
. Any odd prime dividing both and must divide ; since , , so . So the gcd (which is odd, since is odd) can only be or a power of — and one checks directly it’s never a higher power, so it’s or .
Case A:
The two factors on the right of are coprime, and their product is the perfect cube , so each is individually a perfect cube:
for integers . Since the gcd is , in particular (if then too, making the gcd at least ) — so , i.e. . And is odd, since is odd.
This is exactly the setup for the one lemma this proof needs:
Lemma. If are coprime integers and with , then there are integers with
(after possibly replacing by ).
Sketch. In the Eisenstein integers (, norm , a unique factorization domain with six units — the same ring used for the historical honesty note below), set . A direct computation gives , and since , this is an equation between elements of : , where is ‘s conjugate. One can check are coprime in using and (a common factor would have to divide both and , forcing it to divide and relate to in a way both hypotheses rule out). Since is a UFD and is a perfect cube with coprime factors, itself must be, up to one of the six units, a perfect cube: . Writing in the form (the -representation of , since ) and expanding recovers precisely the stated formulas for — provided the unit can be taken to be . Ruling out the other five units is the one piece of this lemma I’m not deriving here: the clean classical way to do it uses Gauss’s theory of composition of binary quadratic forms (specifically, that has class number ), which is real, substantial 19th-century algebra in its own right and a genuinely separate undertaking from anything else in this post. (honestly, a sketch — see the note above)
Given the lemma, . One checks , , are pairwise coprime (using , inherited from , and ), so each is individually a cube: , , for integers . Since , this gives
a new solution of the exact same equation .
Case B:
Here ; write . Then . Since and , we get and ; from these, and are coprime, so each is a cube:
(note here too, since ). Apply the lemma to (the same lemma, with relabeled as ):
Then , so ; writing gives , and the same coprimality argument as in Case A forces each factor to be a cube: , , . Since ,
again a new solution of the same equation.
The descent closes
Either case produces a genuinely new solution of , built out of , which are on the scale of — and are themselves on the scale of , which satisfy . Concretely, roughly , so , and the new solution’s entries are strictly smaller in absolute value than once . That contradicts the minimality of , completing the descent.
A historical honesty note. Euler’s actual 1770 argument worked with numbers of the form — that is, in , not — and at the key step (essentially the lemma above) needed unique factorization in that ring. But does not have unique factorization — for instance , and one can check , , are pairwise non-associate irreducibles (each has norm , and the only units are ), so this is a genuine double factorization into irreducibles. Euler’s argument had a real hole at exactly this point. The fix, understood over the following decades, is that is not the full ring of integers of — it sits inside as an index- subring, and does have unique factorization, which is exactly the ring the lemma above is proved in. This is exactly the kind of gap that’s easy for a popular account to paper over, and I’d rather name it than pretend Euler’s original version was airtight.
Why this doesn’t just keep going
Both proofs above have the same skeleton: factor the equation in a ring bigger than , use unique factorization to show coprime factors of a perfect power are themselves perfect powers, and produce a strictly smaller solution. It’s natural to guess that the general prime case falls the same way, working in for , using the factorization .
It doesn’t, and the reason is concrete rather than mysterious: is not always a unique factorization domain. Unlike the situation, this isn’t a matter of using the wrong ring — genuinely is the correct ring of integers for this problem — but for larger it can simply fail to have unique factorization (its class number can exceed ), and no amount of switching rings fixes that the way switching from to did.
Ernst Kummer, in the 1840s and 1850s, found a way to partially route around this: even without unique factorization into elements, one can develop a theory of “ideal numbers” (the direct ancestor of ideals in modern ring theory) that do factor uniquely. Using this machinery, he proved Fermat’s Last Theorem for every exponent in a class he called regular primes — primes not dividing the class number of (equivalently, by a theorem of his, not dividing the numerator of any of a certain finite list of Bernoulli numbers). This is a real, large chunk of the theorem, and most small primes turn out to be regular. But it’s honestly incomplete: irregular primes exist (the smallest is ), Kummer’s method says nothing about them, and it’s known that infinitely many irregular primes exist — while, remarkably, it’s still an open question today whether infinitely many regular primes exist. Descent-and-factorization approaches, pushed as far as 19th-century mathematics could push them, provably could not finish the job. Something structurally different was needed, and it took until the 1990s to arrive.
The real proof, sketched honestly
What follows is a sketch of the logical shape of the actual proof of Fermat’s Last Theorem — not a derivation. Executing it in full runs to hundreds of pages of published mathematics across several deep theories (elliptic curves, modular forms, Galois representations), built up over more than a decade by many people. I’m describing what each piece asserts and why it creates a contradiction, not reproducing the mathematics.
Suppose, for contradiction, that has a solution in positive integers for some prime (by the reduction proved earlier, together with the cases just proved in full, this is the only case left to rule out). Around 1984, Gerhard Frey had a striking idea: attach to this hypothetical solution the elliptic curve
Frey observed that such a curve, if it existed, would have to be extremely unusual — because can be taken pairwise coprime, turns out to be semistable (its bad reduction, at every prime where it has bad reduction, is the mildest possible kind — “multiplicative,” not “additive”) with a conductor that is unnaturally small and constrained, essentially just the product of the primes dividing . It looked, in Frey’s and others’ words at the time, too strangely well-behaved to be a genuine elliptic curve — but making “too strange to exist” into an actual contradiction required a precise conjecture linking elliptic curves to a completely different kind of object: modular forms.
The modularity theorem (the statement formerly known as the Taniyama–Shimura conjecture) says every elliptic curve over corresponds to a modular form — roughly, a highly symmetric analytic function on the upper half-plane — of a specific weight and level tied to the curve’s conductor. In 1986, Ken Ribet proved that if the Frey curve were modular, level-lowering arguments (building on work of Jean-Pierre Serre) would force it to correspond to a modular form of an impossibly small level — small enough that the relevant space of modular forms is provably empty. So modularity of the Frey curve, plus Ribet’s theorem, is already a contradiction — provided the Frey curve is modular in the first place.
That’s exactly what Andrew Wiles supplied. In a program culminating in 1994 (with a gap in the original 1993 announcement closed jointly with Richard Taylor, published in 1995), Wiles proved the modularity theorem for semistable elliptic curves over — precisely the class that Frey curves belong to. (The modularity theorem was later extended to all elliptic curves over , but the semistable case was already everything Fermat’s Last Theorem needed.)
Chain the three pieces together: a hypothetical solution produces a semistable Frey curve; Wiles’s theorem says that curve is modular; Ribet’s theorem says a modular semistable curve of this specific shape would require a modular form that doesn’t exist. Contradiction — no such can exist, for any prime . Combined with the elementary reduction to and odd primes, and the complete classical proofs of above, that closes Fermat’s Last Theorem for every exponent .
I want to end by being plain about the size of the gap between the two halves of this post. The and proofs above are complete: every step is either fully verified or explicitly flagged as the one spot worth double-checking, and nothing in them is beyond 19th-century tools applied carefully. The Frey–Ribet–Wiles argument, by contrast, is a sketch of a proof I have not carried out and could not carry out here — it depends on the theory of elliptic curves over , the theory of modular forms and the spaces they live in, and the Galois representations that connect the two, each of which is a substantial field with its own textbooks, let alone the specific hard theorems (Ribet’s level-lowering, Wiles’s modularity lifting) proved within them. Getting from “here’s roughly why this works” to an actual proof is exactly the distance those fields, and those specific results, are meant to cover — and it’s real work for another set of posts, not a paragraph.