Continuous everywhere, differentiable nowhere: the Weierstrass function
Draw a continuous curve. Take your pen off the paper only where you have to. Chances are it looks smooth almost everywhere, maybe with a handful of sharp corners — the graph of , or a piecewise-linear zigzag. Through most of the 19th century, this was more than a drawing habit: it was close to a working assumption. Continuity was widely believed to imply differentiability at “most” points, or at least at all but a sparse set of exceptions, because every continuous function anyone could write down behaved that way. Even Ampère had offered a purported proof, decades earlier, that a continuous function must be differentiable except at isolated points.
It’s false. In the 1870s, Karl Weierstrass constructed a function that is continuous at every single real number and differentiable at none of them — not “differentiable except on some thin exceptional set,” but nowhere, full stop. No tangent line exists anywhere on its graph, despite the graph being a single unbroken curve you could in principle trace without lifting a pen. The construction is a Fourier-flavored infinite sum, and unwinding exactly why it survives as continuous while its differentiability dies completely is one of the cleanest case studies in real analysis for what “rigor” actually buys you over “it looks smooth in the picture.”
The construction
For , a positive odd integer, define
Each term is a cosine wave: amplitude , frequency (in the sense that completes full oscillations as ranges over ). As grows, geometrically — each successive wave is smaller — but geometrically too — each successive wave is packed tighter. The function is literally an infinite superposition of cosine waves at every dyadic-like scale , each scaled down in height by .
Weierstrass’s classical theorem adds one more condition beyond and odd:
It’s worth pausing on why a condition like this should be needed at all, before treating it as a black box. Differentiate the -th term formally: , which has amplitude . The amplitude condition was exactly what made the heights of the terms shrink, which is what continuity is going to need. But the slopes of the terms scale with , not — and if , those slopes don’t shrink at all, they blow up geometrically as . So term contributes a gentle wave, term a slightly steeper one, and by term or so you’re formally adding in a wave whose slope is enormous, riding on top of a curve that (because ) is barely perturbed in height by that same term. Height under control, slope out of control, at every scale simultaneously — that tension is the entire mechanism of this function, and it’s why is the qualitative threshold and the stronger constant is what a fully rigorous proof of nowhere-differentiability actually needs (that stronger, sharp threshold comes from a careful classical estimate; I’ll say honestly below exactly where it enters, without re-deriving the precise constant).
Here’s a single low-frequency term by itself — nothing dramatic yet, just a cosine:
Throughout, I’ll use the concrete pair , ( odd, , comfortably satisfying the classical condition).
Continuity: the easy half, done completely
This direction is genuinely clean, and there’s no reason to hand-wave any of it.
Let denote the -th partial sum. Each is a finite sum of continuous functions (cosines, scaled and composed with linear maps), hence continuous on all of .
Uniform convergence, via the Weierstrass M-test. For every and every , , so
Since , converges (geometric series). The M-test says: if for all in some domain, with , then converges uniformly on that domain. The proof of the M-test itself is just the Cauchy criterion applied uniformly — for ,
and this bound doesn’t depend on or at all — it only depends on , and it as . So is uniformly Cauchy, hence converges uniformly to some limit function, which is by definition .
A uniform limit of continuous functions is continuous. This is the other half, and it’s worth actually running the argument rather than citing it. Fix and . By uniform convergence, choose so that (in particular this bounds and for every , with the same ). Since is continuous at , choose so that . Then for ,
So is continuous at , and since was arbitrary, is continuous on all of .
Nothing about this argument cares how large is, or whether exceeds — continuity only ever used . That asymmetry (continuity is cheap; differentiability, next, is expensive and needs the extra hypothesis) is the whole point of the post.
Two more terms already show the roughening, even though every one of these partial sums is, individually, a perfectly smooth, differentiable finite trigonometric sum:
By four terms the curve is visibly ragged. But finite sums of smooth functions are always smooth — has a well-defined derivative everywhere, and so does for every finite . Differentiability only breaks in the limit , and it breaks completely. That’s the part that needs a real argument, because “it looks jagged in the picture” is not a proof — a function can look rough at the resolution of a plot and still be differentiable (its derivative could just be large, or oscillate a lot, without failing to exist).
Non-differentiability: the real mechanism
Fix any point — the argument has to work at every point, which is the entire content of “nowhere differentiable.” The strategy is to exhibit a specific sequence of points marching in on along which the difference quotient fails to settle down to any finite limit.
Choosing the approach. For each positive integer , let be the integer nearest to , and write
Define
an exact integer. This is the key trick: is chosen so that at exactly the scale, the argument of the cosine lands exactly on a multiple of , where takes its extreme value . Since , — in particular and as , so .
Now look at the difference quotient, split by frequency band:
Low frequencies (, ): bounded by the mean value theorem. For fixed , has derivative bounded by in magnitude, so by the MVT,
Summing,
so — this is exactly the finitely-many-smooth-terms piece, and it grows no faster than .
The hinge term (, ): where the specific choice of pays off. By construction , so . And (using since ). So the numerator of the term is
Since , , so the bracket lies in — it is bounded away from zero, uniformly in . Dividing by gives
This is the whole engine of the proof: one single term, isolated by the specific choice of , contributes a difference quotient of size at least a fixed constant times — genuinely unbounded, since .
The high tail (, ): crudely bounded. Using and ,
Putting it together. By the triangle inequality,
If the bracketed constant is positive, as along the sequence , so the difference quotient is unbounded — no finite derivative can exist. And crucially, none of this argument depended on which we started with, so it applies at every point simultaneously.
I want to be straightforward about what’s rigorous here and what isn’t. The three bounds on , , above are genuine, checkable estimates — this is the actual shape of the classical proof (essentially the version in Rudin’s Principles of Mathematical Analysis), not a hand-wave. What I’m not doing is the careful optimization that turns “make the bracket positive” into the sharp, general classical constant — my bracket above, evaluated crudely, is a looser sufficient condition than the sharp one, and pinning down exactly how loose (and tightening it to the textbook constant) is real, somewhat tedious bookkeeping that I’m not reproducing line by line. The concrete pair used in this post, , satisfies the sharp classical condition with room to spare, so the theorem applies to it regardless.
This is exactly where the low/high split in the charts matters. Zoom into near any point, at any scale, and you’re looking at a superposition where finitely many “slow” terms behave like a smooth background (the piece — tame, bounded by the MVT) while there is always some band of terms whose frequency is comparable to your zoom level, contributing slope on the order of at that scale (the / pieces) — and that contribution never shrinks away no matter how far in you zoom, because . Eight terms, still on the original scale:
That’s the same eight-term partial sum , but plotted on a window of width around — about two thousand times narrower than the earlier plots. It’s exactly as jagged, at this microscopic scale, as was on the original scale. Nothing smooths out under magnification, because the high-frequency terms that dominate at this zoom level ( around –, with wavelength comparable to the window) are still there, still contributing slope of order .
Push to sixteen terms and zoom in by another factor of roughly , down to a window about wide, and the same thing happens again — the newly-added high- terms (wavelength comparable to this tiny window) take over and the curve is, once again, exactly as rough as it was at every previous scale:
This is not a coincidence of the specific window sizes I picked — it’s forced by the structure of the proof. At every scale , there is a term of the series (the one with ) whose amplitude has shrunk just enough to keep the function continuous, but whose slope, , has grown, and grown by exactly the same geometric factor at every scale. There is no scale at which the roughness runs out.
Why this mattered, and where it echoes
Weierstrass presented this construction in 1872, in a paper read to the Berlin Academy of Sciences, at a moment when analysis was in the middle of being rebuilt on epsilon-delta foundations rather than geometric intuition about curves — Cauchy, Bolzano, and Weierstrass himself had spent decades replacing “obviously true from the picture” with statements that could actually be checked from the definitions of limit, continuity, and derivative. The Weierstrass function is the sharpest possible demonstration of why that program mattered: it is a legitimate counterexample to a claim that had been treated as geometrically self-evident, constructed from tools (an infinite series, term-by-term estimates, the M-test) that only work because the definitions were made precise enough to apply them. It’s usually counted among the first of the “pathological” or “monster” functions and sets that forced 19th- and early 20th-century mathematicians to stop trusting pictures as proofs — a lineage that continues through nowhere-differentiable and self-similar constructions in what became fractal geometry, and, later, through probability theory: it’s a theorem that the sample paths of Brownian motion are, with probability , continuous everywhere and differentiable nowhere, by a mechanism that — roughly, not literally the same estimates — has the same flavor as this one, fine-scale randomness whose amplitude shrinks fast enough for continuity but whose local slope never settles down. The picture changed permanently: “continuous” and “smooth” were never really synonyms, but after 1872 nobody could pretend the gap between them was small.