You don't have to track 10²³ molecules in a glass of water; temperature and pressure get the boiling time right ── why are you allowed to throw the rest away?
Say there are 180 mL of water in the glass in front of you. That's about 6×10²⁴ molecules. Write down each one's position (three numbers) and velocity (three numbers) and you have 3.6×10²⁵ numbers. Even at one byte each, that is many orders of magnitude beyond all the storage humanity owns. In principle, you cannot track it. And yet we can predict how many minutes it takes that glass to boil on a stove from just two numbers, temperature and pressure ── and we get it right. We threw away 10²⁴ numbers and kept 2 ── a compression ratio of 10⁻²⁴ ── and the answer didn't change. This miracle, "throw it away, same answer," is what this whole series is about. Episode 1 nails down, in one line, why you are allowed to throw things away.
First, let's be clear about how hopeless it is. Newtonian mechanics is a perfect recipe ── give it the initial position and velocity of every molecule and the future is determined; just solve the equations. The problem is the "give it" part: you need 10²⁵ numbers. You cannot gather them, cannot write them down, cannot solve them.
Thermodynamics, meanwhile, answers the same question about the same glass in three lines. Temperature, pressure, volume. Give it those three and out comes the boiling temperature, the expansion, the upper bound on the work an engine can extract. And it matches experiment. Of the 10²⁵ numbers, we discarded 10²⁵ and kept 3, and the prediction did not break. This operation is called coarse-graining.
This is not a compromise, not a "close enough." For macroscopic quantities the accuracy is in fact absurdly high. The reason is the next line.
Add up \(N\) scattered things and the scatter only grows like \(\sqrt{N}\), while the total grows like \(N\). So the relative scatter dies off as \(1/\sqrt{N}\). This is the central limit theorem in the form you'll actually use.
Roll one die and the result is anywhere from 1 to 6. Average 100 rolls and you cluster fairly tightly around 3.5; average 10,000 and you won't be off 3.5 by even 1%. The individuals are wild, but the average is sharp. That sharpness is what makes coarse-graining work.
With \(N\approx 6\times10^{24}\) molecules, the relative fluctuation is
$$\frac{1}{\sqrt{N}}\approx\frac{1}{\sqrt{6\times10^{24}}}\approx 4\times10^{-13}$$The pressure gauge's needle doesn't move until the 13th digit. The best pressure gauges in the world manage 6 or 7 digits, so the fluctuation is unobservable in principle. "Temperature" and "pressure" behave like real, solid quantities precisely because \(N\) is huge. The sharpness of macroscopic variables is underwritten by particle number.
Turn it around: if \(N\) is small, coarse-graining fails. A pollen grain jiggles as water molecules hit it (Brownian motion) ── that is a grain small enough that the collision count \(N\) is too low and \(1/\sqrt{N}\) became visible. Coarse-graining is not universal magic; it is a theorem with the condition that \(N\) be large. That condition is the subject of Episode 5.
| System | Number of parts N | Relative fluctuation 1/√N | Coarse-graining |
|---|---|---|---|
| A glass of water | 6×10²⁴ | 4×10⁻¹³ | works perfectly |
| A given protein inside one cell | ~10³ | 3% | fluctuations matter (life exploits this) |
| Collisions on a pollen grain | few enough to see | visible | breaks (Brownian motion) |
| A single electron | 1 | 100% | doesn't work at all |
Here the most mysterious quantity in thermodynamics gets a straight meaning. Entropy is precisely the amount of information you threw away.
Fix the temperature and pressure and there are still an enormous number of microscopic arrangements consistent with them (which molecule is where, moving how). Call that count \(W\):
\(W\) is "the number of microscopic arrangements that look the same macroscopically." So \(S\) is the amount of freedom left undetermined once the macrostate is fixed ── the information coarse-graining discarded. In bits it is \(S/(k\ln 2)\): roughly 10²⁵ bits for a mole of gas ── the amount dumped on the floor so that two numbers, temperature and pressure, could remain.
This reading is powerful. "Entropy increases" (the second law) becomes discarded information does not come back on its own. Milk stirred into coffee mixes and never spontaneously un-mixes, simply because the number of separated arrangements is overwhelmingly smaller than the number of mixed ones. That's all. Coarse-graining, entropy and the arrow of time are three names for one thing. (Why was it separated to begin with? See Bonus Episode ①.)
The figure below actually throws things away. The top and bottom rows are two completely different microscopic configurations, A and B (different random draws). On the left is the raw micro; on the right is the same thing averaged over blocks.
Drag the slider to grow the block size. Two things happen at once ── (1) the right-hand images smooth out (fluctuations die as \(1/\sqrt{N}\)); and (2) the right-hand A and B become indistinguishable. Two things that look nothing alike microscopically wear the same face once coarse-grained. What is being discarded is the information "which one was it?" That is a preview of Episode 4's universality.
The strangest by-product of coarse-graining is that concepts appear on the far side that were not there before.
A single water molecule has no "temperature." No pressure, no viscosity, no surface tension. These are quantities that only mean something once 10²⁴ molecules are gathered and the individuals discarded. So coarse-graining reduces information, and yet a new vocabulary is born in the reduction. That is why the world is a mille-feuille of independently closed layers.
particles → nuclei → atoms → molecules and chemistry → materials and heat → biology → mind → society.
A chemist predicts reactions without knowing about quarks; an engineer builds bridges without knowing about atoms. The laws of the floor below close without knowing the details of the floor above ── that is what it means for a layer to be a layer. The physicist Anderson called this "More is Different" (1972). Episode 4 makes the reason layers close exact, as the renormalization group.
That coarse-graining works, that relative fluctuations fall as \(1/\sqrt{N}\) (law of large numbers, central limit theorem), that Boltzmann's \(S=k\log W\) counts microstates, and that entropy can be read as information (Shannon, Jaynes) ── all of this is established physics and mathematics.
But you may not discard always. Coarse-graining is justified roughly when three things hold: (1) \(N\) is large (otherwise the fluctuations show); (2) the scales separate (fast microscopic motion and slow macroscopic change differ by orders of magnitude); (3) local equilibrium applies. At critical points, in turbulence and in chaos these fail and what you discarded comes back into the answer (Episode 5). Also, "which variables to keep" is not handed to you automatically by physical law ── there is an element of human choice depending on the question, which is why entropy carries a dependence on how you coarse-grained (a point pressed by Jaynes and others). This series takes the position that the renormalization group tells us even that choice.
Crush the 10²⁴ molecules in a glass of water down to two numbers and the prediction still lands. The reason is relative fluctuation ∝ 1/√N ── averages sharpen with count, and for a glass of water nothing moves until the 13th digit. That is why macroscopic variables behave like real quantities. The amount discarded has a name: entropy, \(S=k\log W\), and the second law is just "discarded information does not come back on its own."
As a by-product, concepts that did not exist before are born (temperature, pressure, viscosity) and the world becomes a mille-feuille of layers. But this only works while \(N\) is large and the scales separate ── Episode 5 counts the places where that fails. Next time we go to an extreme case: what does "discarding" mean for a single particle? We'll reread the neutron's 880-second lifetime as a story about phase, not probability.
Print / make a PDF: ⌘+P (Ctrl+P on Windows). On screen, moving the slider grows the block size: the coarse-grained image smooths out and the distinction between A and B vanishes. "Redraw the micro" regenerates the random numbers. "See the answer" opens each solution.