c·t = CONST, THAT CLICKS EPISODE 5 / Shortness and fit, priced in one currency
Zero description length — so how many bits is that zero worth?
What does
shortness cost?
Episode 4 found \(c\cdot t=\)const has zero dimensionless parameters — Occam at the limit.
But it loses on fit. Here both are converted into bits.
Episode 4 ended in a clean opposition: \(\Lambda\)CDM has 6 dimensionless parameters, \(c\cdot t=\text{const}\) has 0. The description length is overwhelmingly shorter. And yet the two \(H_0d_L/c\) curves differ by 0.21 magnitudes at \(z\simeq1.1\). Is the shorter one right, or the one that fits? This is not a matter of taste — information theory keeps a price list. This episode goes and reads it.
01Splitting description length in two
Imagine posting both a model and a data set to someone. Two things go in the envelope: the description of the model and the deviations from it. The best model is the one for which the total is shortest — that is the minimum description length (MDL) idea.
\(k\) is the number of parameters, \(N\) the number of data points
Holding one parameter lengthens the message by however many bits it takes to send its value — that is the right-hand term. Multiply that term by 2 and write it in natural logs and you get \(k\ln N\), which is exactly the BIC penalty. MDL and BIC were the same accounting in different units.
02The price list for one parameter
AIC (Akaike) — penalty \(2k\) in natural-log \(\chi^2\) units
$$\frac{2}{2\ln 2}=\frac{1}{\ln 2}=1.443\ \text{bits per parameter}\qquad(\text{independent of }N)$$BIC (Bayesian) — penalty \(k\ln N\)
$$\frac{\ln N}{2\ln 2}=\tfrac12\log_2 N\ \text{bits per parameter}\qquad(\text{rises with more data})$$| Data points \(N\) | AIC price | BIC price | BIC's \(\Delta\chi^2\) budget |
|---|---|---|---|
| 100 | 1.443 bit | 3.32 bit | 4.61 |
| 1000 | 1.443 bit | 4.98 bit | 6.91 |
| 1701 (Pantheon+ scale) | 1.443 bit | 5.37 bit | 7.44 |
| 10000 | 1.443 bit | 6.64 bit | 9.21 |
Read it like this: saving one parameter earns you 5.4 bits. So as long as you lose less than 5.4 bits on fit, the shorter model wins.
03Converting misfit into bits too
For Gaussian errors \(-2\ln L=\chi^2+\text{const}\), so
$$\text{bits lost to misfit}=\frac{\Delta\chi^2}{2\ln 2}=0.7213\,\Delta\chi^2$$The currencies now match. All that remains is to compute \(\Delta\chi^2\).
04Doing the accounts
Compare on the supernova distance modulus. In magnitudes, the difference between Episode 4's closed form \(H_0d_L/c=(1+z)\ln(1+z)\) and the \(\Lambda\)CDM numerical integral is:
| \(z\) | 0.05 | 0.1 | 0.3 | 0.5 | 1.0 | 1.5 | 2.0 |
|---|---|---|---|---|---|---|---|
| \(\Delta\mu\) [mag] | −0.027 | −0.052 | −0.126 | −0.171 | −0.213 | −0.206 | −0.181 |
A supernova's absolute magnitude is unknown, so any constant offset can simply be absorbed (this is the practical version of "\(H_0\) is dimensionful, so we do not count it"). What survives absorption is the difference in shape. Computing on a Pantheon+-like redshift distribution (median \(z=0.27\), 1701 supernovae, \(\sigma=0.15\) mag each):
Residual after absorbing the constant
$$\text{RMS}=0.053\ \text{mag}\qquad\Longrightarrow\qquad \Delta\chi^2=N\left(\frac{0.053}{0.15}\right)^2=213$$In bits
$$\text{cost of misfit}=\frac{213}{2\ln2}=154\ \text{bits}$$Against the saving
$$\text{one parameter}=5.4\ \text{bits}$$The bill for this episode
Earned: +5.4 bits Lost: −154 bits Net: −149 bits
The shortness bought by dropping one parameter is about 1/29 of what is paid on fit.
05From the other side — how well would it have to fit?
| Criterion | \(\Delta\chi^2\) budget | Residual RMS allowed | Actual |
|---|---|---|---|
| AIC | 2 | 5.1 mmag | 53 mmag |
| BIC | 7.44 | 9.9 mmag |
After averaging over 1701 supernovae, the two curves would have to agree to 10 millimagnitudes per supernova — one fifteenth of the intrinsic scatter (150 mmag). The actual discrepancy is 53 mmag. Short by an order of magnitude.
06So how many supernovae settle it?
| Count \(N\) | Misfit \(\Delta\chi^2\) | BIC budget \(\ln N\) | Winner |
|---|---|---|---|
| 10 | 1.25 | 2.30 | \(c\cdot t=\text{const}\) |
| ≈ 26 | 3.26 | 3.26 | a draw |
| 100 | 12.5 | 4.61 | \(\Lambda\)CDM |
| 1000 | 125 | 6.91 | \(\Lambda\)CDM |
With fewer than 26 supernovae, \(c\cdot t=\text{const}\) is the better model. This is not sour grapes — information theory really does say so. When data are scarce, the short theory is the correct choice. When accelerated expansion was found in 1998, the two teams had 42 and 16 supernovae. Back then this contest would have been close.
Figure: data points along the horizontal axis, bits up the vertical. The price of a parameter grows only as \(\log N\), while the cost of misfit grows as \(N\). So the two lines must cross, and past the crossing they never swap back.
Drag the discrepancy down and the crimson line slides right, taking the crossing with it. But they always cross — the slopes differ. Even at 1 mmag the crossing is only around \(N\simeq10^4\). Keep adding data and shortness must eventually lose.
07The reveal — Occam's razor has an expiry date
| How it grows | Why | |
|---|---|---|
| Price of a parameter | \(\propto\log N\) | more data demands more precision in the value you send — but only logarithmically more |
| Cost of misfit | \(\propto N\) | each point's residual must be re-sent, so it piles up in direct proportion |
The one line of this episode
The gain from shortness goes as \(\log N\); the loss from misfit goes as \(N\).
So with enough data the short model must lose — unless it is actually right.
Occam's razor is neither superstition nor aesthetics but a computable discount voucher. Its face value is merely \(\log N\), however, and it gets relatively cheaper as data accumulate. "Simpler theories are better" turned out to be a theorem with the proviso if the fits are comparable.
① The \(\Delta\chi^2=213\) of §04 is the expectation value on the assumption that \(\Lambda\)CDM (\(\Omega_m=0.315\)) is correct. It is not a fit to real data. The redshift distribution is a Pantheon+-like mock (median 0.27), not the real one, and \(\sigma=0.15\) mag is a representative intrinsic scatter. Including correlated systematics reduces the effective count and shrinks \(\Delta\chi^2\). Read it as an order-of-magnitude argument.
② The literature does not agree on the comparison with real data. Melia and collaborators argue that model-independent distance indicators favour \(R_h=ct\); other analyses (Shafer 2015 and others) report a strong preference for \(\Lambda\)CDM. This is a conditional calculation — "here is what follows if \(\Lambda\)CDM is right" — not an observational verdict.
③ The parameter counting is disputed too. For supernovae alone, \(H_0\) is degenerate with absolute magnitude, so it is effectively 1 parameter (\(\Omega_m\)) against 0. The "6" of Episode 4 is the standard \(\Lambda\)CDM basis including the CMB — a different arena. The 5.4 bits here is the price of one.
④ The \(\tfrac{k}{2}\log_2 N\) of MDL is asymptotic; exactly, there is a further term depending on the model's geometry (the volume of Fisher information). AIC and BIC derive from different premises (minimising prediction error vs maximising posterior probability), and which to use depends on the purpose — there is no single correct price list.
Exercises (solvable with this episode's formulas alone)
- At \(N=100\), what is the BIC price of one parameter in bits?
Show the answer
\(\tfrac12\log_2 100=3.32\) bits; in \(\Delta\chi^2\), \(\ln100=4.61\). Multiply the data by 17 (to 1701) and the price only goes 3.32 → 5.37 bits — because it is a logarithm. - How many bits is \(\Delta\chi^2=213\)?
Show the answer
\(213/(2\ln2)=154\) bits — about 29 times the 5.37 bits saved by dropping a parameter. A net loss of 149 bits. - Why is the AIC price independent of \(N\), and what does the difference from BIC reflect?
Show the answer
Because AIC's penalty \(2k\) contains no \(N\). AIC minimises prediction error: adding one parameter worsens the expected prediction error by 2 in \(\chi^2\). BIC maximises posterior probability: as \(N\) grows, "it fitted by luck" becomes less likely, so the fine gets heavier. Different purpose, different price list. - If the residual RMS were 20 mmag, up to how many supernovae would \(c\cdot t=\text{const}\) win?
Show the answer
Solve \(\ln N=N(0.020/0.15)^2=0.01778\,N\) numerically: \(N\simeq325\). Cutting the discrepancy by 2.7 multiplies the winnable count by about 12 — and still only reaches the low hundreds. \(\log N\) versus \(N\) was a losing contest from the start. - (Harder) Under what condition is "the shorter theory is better" correct?
Show the answer
As long as the difference in fit stays within about \(\log N\). A parameter costs only \(\tfrac12\log_2 N\) bits, so the moment the misfit exceeds that (\(\Delta\chi^2>\ln N\)) the ranking flips. And since misfit grows \(\propto N\), adding data must flip it eventually. Occam's razor is a theorem — but a theorem with an expiry date.
Summary — the price tag read 5.4 bits
"Is the shorter one right, or the one that fits?" Information theory answers with a price list. Description length splits into cost of misfit plus price of parameters, the latter being \(\tfrac{k}{2}\log_2 N\) bits — the same thing as BIC's \(k\ln N\) penalty. Under AIC it is 1.443 bits each, independent of \(N\); under BIC, \(\tfrac12\log_2 N\) — 5.37 bits at Pantheon+ scale (1701).
Misfit converts at \(\Delta\chi^2/(2\ln2)\) bits. Doing the accounts on the supernova distance modulus, the residual RMS after absorbing the absolute magnitude is 53 millimagnitudes, \(\Delta\chi^2=213\), i.e. 154 bits lost — about 29 times the 5.4 bits earned. A net loss of 149 bits. To break even the two would have to agree to 10 millimagnitudes per supernova, one fifteenth of the intrinsic scatter.
The interesting part came from varying the count: below 26 supernovae, \(c\cdot t=\text{const}\) is the better model. Not sour grapes — information theory says so. The 1998 discovery of accelerated expansion used 42 and 16 supernovae. When data are scarce, the short theory is the correct choice.
The reveal was in the slopes: gain from shortness \(\propto\log N\), loss from misfit \(\propto N\). The lines must cross and never swap back. Occam's razor is neither superstition nor aesthetics but a computable discount voucher — whose face value grows only logarithmically. "Simpler is better" was a theorem with the proviso if the fits are comparable.
This document is Episode 5 of "c·t = const, That Clicks", written for physics-minded high-school and university readers. The two-part MDL code \(-\log_2 L+\tfrac{k}{2}\log_2 N\), Akaike's criterion \(\mathrm{AIC}=-2\ln L+2k\), the Bayesian criterion \(\mathrm{BIC}=-2\ln L+k\ln N\), and \(-2\ln L=\chi^2+\text{const}\) for Gaussian errors are all standard. The conversion into bits (one parameter costing \(1/\ln2=1.443\) bits under AIC and \(\tfrac12\log_2 N\) under BIC, misfit costing \(\Delta\chi^2/(2\ln2)\) bits) is this document's own rewriting. The numbers in §04–§06 are computed here and are expectation values on the assumption that \(\Lambda\)CDM (\(\Omega_m=0.315\), \(\Omega_r=9.2\times10^{-5}\)) is correct — not fits to real data. The redshift distribution is a Pantheon+-like mock (\(N=1701\), median \(z=0.27\)) with \(\sigma=0.15\) mag per supernova (a representative intrinsic scatter); correlated systematics are not included, and including them lowers the effective count and \(\Delta\chi^2\). Comparisons of \(R_h=ct\) and \(\Lambda\)CDM on real data do not agree in the literature: Melia and collaborators argue in favour of \(R_h=ct\), while other analyses (Shafer 2015 and others) report a strong preference for \(\Lambda\)CDM. Using supernovae alone, \(H_0\) is degenerate with absolute magnitude, so the effective parameter count is 1 for \(\Lambda\)CDM against 0 for \(R_h=ct\); the "6" quoted in Episode 4 is the standard basis including the CMB. The \(\tfrac{k}{2}\log_2 N\) of MDL is asymptotic, with a further term depending on the volume of Fisher information. AIC and BIC rest on different premises (minimising prediction error vs maximising posterior probability) and no single criterion is uniquely correct. The 1998 discovery of accelerated expansion used 16 supernovae (High-z Supernova Search Team) and 42 (Supernova Cosmology Project). The academic standard remains the \(\Lambda\)CDM model including inflation. ── To make a PDF, use your browser's Print dialogue (sliders freeze and answers are hidden in the print version).