Renormalization That ClicksBonus ③ / The central limit theorem was a renormalization group

Episode 1's 1/√N and Episode 4's universality were two faces of one theorem ── statistics has been doing renormalization for a hundred years

The central limit theorem was a renormalization group Treat "add them up and divide by \(\sqrt N\)" as a transformation and the Gaussian becomes its fixed point.
The central limit theorem is the claim that this fixed point is an attractor,
and skewness and kurtosis are irrelevant operators, dying off as \(N^{1-k/2}\).

Tools you'll need: Episode 1's \(1/\sqrt N\), Episode 4's fixed points and relevant/irrelevant, mean and variance The heart of this one: the \(k\)-th cumulant shrinks as \(N^{1-k/2}\)

Something occurred to me after finishing the main series: I wrote Episode 1's \(1/\sqrt N\) and Episode 4's universality as two separate stories. They are two faces of one theorem. Take "add up \(N\) independent variables and divide by \(\sqrt N\)" and regard it as a transformation from distribution to distribution ── it has exactly the same structure as Episode 4's block spins. Doing it twice equals doing it once for four (the composition rule); the Gaussian maps to itself (a fixed point); and the central limit theorem is the universality claim that everything flows there. Better still, the "quirks" of the original distribution ── skewness, kurtosis ── die off cleanly as \(N^{1-k/2}\): these are irrelevant operators. The whole renormalization group appears in a form a high-schooler can compute by hand.

01"Add and divide," seen as a transformation

Draw two independent samples from the same distribution (mean 0, variance 1), add them, divide by \(\sqrt2\).

A map from distributions to distributions
$$Z=\frac{X_1+X_2}{\sqrt2}\qquad\Longrightarrow\qquad \mathcal{T}[p](z)=\sqrt2\,(p*p)(\sqrt2\,z)$$

Here \(*\) is convolution (the operation that corresponds to adding). This \(\mathcal{T}\) is a map on the space of probability distributions. Episode 4 said "coarse-graining is a map from model to model" ── same shape. And the composition rule matches too: applying \(\mathcal{T}\) twice equals adding four and dividing by two.

Why divide by \(\sqrt2\)? Variance adds, so two samples give variance 2; dividing by \(\sqrt2\) brings it back to 1. That division is the same rescaling as shrinking the lattice in Episode 4. Coarse-graining always comes as a pair ── squash, then re-measure. The same is true here.

02The Gaussian is a fixed point

Feed the standard normal \(p(x)=\frac{1}{\sqrt{2\pi}}e^{-x^2/2}\) into \(\mathcal{T}\). A Gaussian convolved with a Gaussian is Gaussian, with variances adding to 2; divide by \(\sqrt2\) and the variance returns to 1 ── exactly the distribution you started with.

The fixed point equation
$$\mathcal{T}[p^*]=p^*\qquad\Longleftarrow\qquad p^*=\text{the standard normal}$$

Episode 4 defined a fixed point as "a model that maps to itself under coarse-graining." The Gaussian is a fixed point of this transformation. And the central limit theorem can be restated: this fixed point is an attractor for a wide range of starting points.

03Skewness and kurtosis are irrelevant operators

Here is the computable centrepiece. There are quantities called cumulants \(\kappa_k\) that measure a distribution's quirks (\(\kappa_1\) = mean, \(\kappa_2\) = variance, \(\kappa_3\) behind skewness, \(\kappa_4\) behind kurtosis). Cumulants have exactly one lovely property ── they simply add when you add independent variables.

The calculation ── how fast does a quirk die?

Adding \(N\) gives \(\kappa_k \to N\kappa_k\). Dividing by \(\sqrt N\) divides the \(k\)-th cumulant by \((\sqrt N)^k\):

$$\kappa_k^{(N)}=\frac{N\kappa_k}{N^{k/2}}=\kappa_k\,N^{\,1-k/2}$$

Tabulate it and the renormalization-group classification falls right out.

\(k\)quantity\(N\) dependencein RG language
1mean\(N^{+1/2}\) (grows)relevant (shift it and it diverges)
2variance\(N^{0}\) (fixed)marginal (it sets the scale)
3skewness\(N^{-1/2}\) (decays)irrelevant
4kurtosis\(N^{-1}\) (decays)irrelevant
\(k\ge3\)higher quirks\(N^{1-k/2}\)even more irrelevant

Episode 4 said "relevant directions are few, irrelevant ones overwhelmingly many." Here it is literally one relevant, one marginal, and all the rest irrelevant ── with the decay rates derivable by hand. This is the simplest fully computable instance of the renormalization group.

The information about "what shape the original distribution had" lives in \(\kappa_3,\kappa_4,\kappa_5,\dots\), and all of it dies as \(N^{1-k/2}\) ── so the origin is forgotten. The very phenomenon of Episode 4's water and magnet is happening with dice and coins.

04Are there other fixed points? ── heavy tails change everything

Everything so far quietly assumed the variance is finite (you can only divide by \(\sqrt N\) if variance is defined). What if it isn't?

The standard example is the Cauchy distribution \(p(x)=\frac{1}{\pi(1+x^2)}\). Its tail only falls as \(1/x^2\), so the variance integral diverges. And this distribution ── returns to itself if you add two and divide by \(2\), not by \(\sqrt2\).

There is more than one fixed point ── the stable family
$$Z=\frac{X_1+\cdots+X_N}{N^{1/\alpha}}\qquad(0<\alpha\le2)$$

Distributions that return to themselves under this scaling are called stable distributions, and they form a family of fixed points labelled by \(\alpha\) (the tail weight). \(\alpha=2\) is the Gaussian, \(\alpha=1\) the Cauchy. And what decides which fixed point you land on is only how the tail falls ── the shape near the middle is irrelevant, in every sense. In Episode 4's words:

stable distribution = fixed point, domain of attraction = universality class, tail exponent \(\alpha\) = the relevant parameter

Starting distributionTailDivide byFlows to
uniform, binomial, exponential, … (anything with finite variance)light\(\sqrt N\)Gaussian (\(\alpha=2\))
Cauchy\(1/x^{2}\)\(N\)Cauchy (\(\alpha=1\))
Pareto type (tail index 1.5)\(1/x^{2.5}\)\(N^{2/3}\)the \(\alpha=1.5\) stable law

The Gaussian looks "obvious" only because most quantities around us have finite variance. In heavy-tailed worlds ── earthquake magnitudes, city populations, financial crashes ── a different fixed point rules. That is why "just take the average and relax" fails there.

05Try it ── watch a distribution flow to a fixed point

The figure below actually convolves and rescales the distribution you pick, over and over (computed in your browser). The slider is the number added, \(N=2^k\).

Pick uniform, coin or exponential and within a few steps they all lie on top of the Gaussian (grey dashed) ── utterly different starting points, identical destination. The coin flip (±1) is especially good: watch a Gaussian rise out of two spikes. And pick Cauchy and it never becomes Gaussian. The shape doesn't change from beginning to end ── it is already sitting on a different fixed point.

Figure: the density after adding N=2ᵏ copies of the chosen distribution and dividing by N^(1/α) (convolution iterated in the browser). Grey dashed is the standard normal. Every finite-variance distribution flows to the Gaussian; Cauchy stays put on its own fixed point
current distribution standard normal (the Gaussian fixed point)
The central limit theorem breaks at a critical point The central limit theorem assumes independence. What if there are correlations? If the correlation length is finite, anything beyond it is effectively independent, so lumping them together still gives a Gaussian. But at a critical point the correlation length diverges (Episode 4). Then no "effectively independent chunk" can be formed and the theorem does not apply. That is why critical fluctuations are non-Gaussian and the exponents depart from mean-field values. Episode 4's \(\beta\approx0.326\) is not a simple number precisely because the CLT is broken there. Episode 1 (a world where the CLT works) and Episode 4 (a world where it breaks) join up here.

06Statistics has been coarse-graining for a hundred years

One more correspondence the main series failed to use.

Sufficient statistics ── the 1922 version of "keep this and discard the rest"

In 1922 Fisher introduced sufficient statistics. A function \(T(x)\) of the data is "sufficient" if, once you know \(T\), the raw data carries no further information about the parameter.
That is exactly what Episode 1 did ── temperature and pressure are sufficient statistics for the microstate. Translated into statistical language, "throw it away and the answer doesn't change" becomes "reduce to a sufficient statistic."

Maximum entropy ── statistical mechanics can be derived as inference

Jaynes (1957) argued: subject to the constraint that only the mean energy is known, choose the distribution that is maximally ignorant about everything else ── i.e. maximise entropy ── and the Gibbs distribution \(e^{-E/kT}\) comes out automatically.
So statistical mechanics can be derived as a procedure of inference, adding no new physical law. When Episode 1 said "entropy = the amount of information discarded," that was not a metaphor but the central claim of this lineage. Sufficient statistics, maximum entropy and exponential families are three faces of one mathematical structure.

◇ ◇ ◇
The honest line ── what is identical and what is not

Established: that the central limit theorem can be formulated as convergence to a fixed point of a transformation on distribution space (the probabilistic view of the renormalization group, Jona-Lasinio 1975 and others); the cumulant scaling \(\kappa_k^{(N)}=\kappa_k N^{1-k/2}\) (elementary and exact); that stable distributions are the fixed points of \(N^{1/\alpha}\) scaling and that domains of attraction are determined by the tail exponent (the Gnedenko–Kolmogorov classification, rigorous); Fisher's sufficient statistics and the factorisation theorem; Jaynes' derivation of the Gibbs distribution from maximum entropy; and that at a critical point the diverging correlation length breaks the CLT's premise so fluctuations become non-Gaussian ── all established mathematics and physics.

But do not over-identify: (1) The fixed points here correspond, in field-theory terms, to the Gaussian (free / mean-field) fixed point. Interacting fixed points like the Wilson–Fisher one ── the star of Episode 4, and where anomalous dimensions come from ── do not fit inside this simple CLT frame. The order matters: "the CLT is the simplest example of RG," not "RG is a restatement of the CLT." (2) The cumulant classification assumes the distribution is smooth enough that cumulants exist. (3) "Sufficient statistic = coarse-graining" is a correspondence once you fix the question (parameter estimation); change the question and what you should keep changes too (the general theory is Bonus ④).

Exercises
  1. Why divide by \(\sqrt N\)? What happens if you don't?
    See the answer
    Variance adds for independent sums, so \(N\) terms give \(N\) times the variance; dividing by \(\sqrt N\) restores it to 1. Without dividing, the distribution simply spreads forever and never converges. This is the same rescaling as re-expanding the lattice after squashing it in Episode 4 ── coarse-graining always needs "squash + re-measure."
  2. By what factor does skewness (\(k=3\)) shrink when \(N\) grows a hundredfold?
    See the answer
    \(N^{1-3/2}=N^{-1/2}\), so by \(\sqrt{100}=10\). Kurtosis (\(k=4\)) goes as \(N^{-1}\), so by 100. Higher-order quirks die faster ── the same structure as Episode 4's \((E/\Lambda)^{d-4}\).
  3. Why is the Cauchy distribution "another fixed point"?
    See the answer
    Because adding two and dividing by 2 (not \(\sqrt2\)) returns it to itself. Its tail falls as \(1/x^2\) and its variance is infinite, so the correct scaling exponent is \(\alpha=1\). Stable distributions form a family of fixed points labelled by \(\alpha\), and which one you flow to depends only on the tail exponent (= universality class).
  4. Why aren't fluctuations Gaussian at a critical point?
    See the answer
    The CLT presumes independence (or a finite correlation distance). At a critical point the correlation length diverges, so no "effectively independent chunk" exists, the premise fails, and you don't land on the Gaussian fixed point. Hence critical exponents deviate from the simple mean-field values.

Bonus ③ summaryEpisodes 1 and 4 were joined at the back by one theorem

Treat "add \(N\) independent variables and divide by \(\sqrt N\)" as a transformation on distribution space and the whole renormalization group appears ── the Gaussian is a fixed point, the central limit theorem is the claim that it attracts, and since the \(k\)-th cumulant scales as \(\kappa_k N^{1-k/2}\), the mean is relevant, the variance marginal, and skewness and everything above it irrelevant. Episode 4's classification, derivable by hand.

There is more than one fixed point. With infinite variance you divide by \(N^{1/\alpha}\) and a family of stable fixed points appears (\(\alpha=2\) Gaussian, \(\alpha=1\) Cauchy). What decides where you land is only the fall-off of the tail ── the middle of the distribution is irrelevant. That is a universality class. And when the correlation length diverges at a critical point, the CLT's premise breaks, which is exactly why non-trivial exponents like \(\beta\approx0.326\) appear. Episode 1 (where the CLT works) and Episode 4 (where it breaks) are joined by a single line.

This document is Bonus Episode ③ of the "Renormalization That Clicks" series, a reading piece for physics-loving high-schoolers and undergraduates. That the central limit theorem can be formulated as convergence to a fixed point of a transformation on distribution space (the probabilistic view of the renormalization group, Jona-Lasinio 1975 and others); the cumulant scaling \(\kappa_k^{(N)}=\kappa_kN^{1-k/2}\); that stable distributions are the fixed points of \(N^{1/\alpha}\) scaling with domains of attraction set by the tail exponent (the Gnedenko–Kolmogorov classification); Fisher's sufficient statistics and factorisation theorem (1922); Jaynes' derivation of the Gibbs distribution from maximum entropy (1957); and that the diverging correlation length at a critical point breaks the CLT's premise and makes fluctuations non-Gaussian ── all established mathematics and physics. That the fixed points appearing here correspond to the Gaussian (free / mean-field) fixed point of field theory and that interacting fixed points such as Wilson–Fisher lie outside this simple frame, and that "sufficient statistic = coarse-graining" holds once the question is fixed, are spelled out in the body's "honest line." The figure represents densities on a grid and iterates convolution and rescaling numerically, renormalising at each step to correct truncation loss from the finite window. ── To print, use your browser's "Print" and "Save as PDF" (in the print version the slider and answers are frozen and hidden). Related: Episode 1 / Episode 4 / Bonus ④ / Contents.

Print / make a PDF: ⌘+P (Ctrl+P on Windows). On screen, pick a starting distribution with the buttons and raise the number added with the slider to watch it flow to a fixed point. Only Cauchy fails to reach the Gaussian. "See the answer" opens each solution.