c·t = CONST, THAT CLICKS EPISODE 19 / Building the sorting procedure itself

Measured in bits, surprise falls into three clean strata

Is an identity really
not physics? This series has sorted coincidences again and again.
Here the criterion is stated, in the language of information theory.

What you need: one logarithmsurprise \(=-\log_2(\text{width}/\text{range})\)

This series has sorted "coincidences" again and again — Dirac's large numbers were an identity, being exactly at the Landauer limit an identity, \(\alpha+2\beta+\gamma=2\) an identity. Last episode's 1.96 fm was a coincidence. And \(\Omega/N=\ln2/3\pi^2(1+w)\) was physics. But is that sorting actually a procedure? Today we build it head on. The answer was on the information-theory side — measure the surprise in bits.

01What we have sorted so far

CoincidenceFromVerdict given
\(M/m_P=R_H/2\ell_P\) (Dirac's large numbers)prev. Extra 3identity
\(E/N=k_BT_H\ln2\) (Landauer)Episode 10identity
\(\alpha+2\beta+\gamma=2\) (critical exponents)Episode 14identity
\(\rho_\Lambda^{1/4}\) and \(m_\nu\) within a factor 22prev. Extra 5coincidence
Side of a bit's volume ≈ a nucleonEpisode 18coincidence
Koide's relation \(Q=2/3\)prev. Extra 4unexplained empirical formula
\(\Omega/N=\ln2/3\pi^2(1+w)\)Episode 1physics

Each verdict looked plausible, but the criterion was never stated. Today we state it.

02Measuring surprise in bits

How surprising a coincidence is can be measured as: how narrow a place did it land in, out of the range that was allowed beforehand?

Definition $$\text{surprise}=-\log_2\frac{\text{width it landed in}}{\text{range allowed beforehand}}\ \ [\text{bits}]$$

meaning

$$1\ \text{bit}=\text{one coin flip},\qquad 20\ \text{bits}=\text{one in a million}$$

This is exactly information theory's surprisal. And by this definition an identity comes out at precisely 0 bitsbecause, following from a definition, there was only ever one place for it to land.

03Actually measuring

CoincidenceSurpriseFeelClass
Dirac's large numbers0 bit──identity
Exactly at the Landauer limit0 bit──identity
\(\alpha+2\beta+\gamma=2\)0 bit──identity
\(\rho_\Lambda^{1/4}\) and \(m_\nu\)4.7 bit5 coin flipscoincidence
A bit's volume ≈ a nucleon7.4 bit7 coin flipscoincidence
Koide's relation \(Q=2/3\)15.7 bitone in 30,000empirical formula
CMB uniform across 9,600 patches\(1.6\times10^{5}\) bitanother worlda real problem

Conclusion of §03

As numbers, the strata separate cleanly — identities at 0, coincidences at a few bits, real problems at \(10^5\).
"The cosmological constant and the neutrino mass agree within a factor 22" is no more surprising than flipping five heads in a row.

This probably runs against intuition. Hearing that \(10^{-31}\) and \(10^{-30}\) — absurdly small numbers — are close sounds momentous, but counted logarithmically it is 1.3 orders out of 35. Last episode's 1.96 fm is the same: 0.36 orders out of 61 — 7.4 bits.

Only Koide's relation is surprising by orders The one entry at 15.7 bits is Koide's relation. Out of the range \([1/3,1]\) that \(Q\) can take, the measured value departs from \(2/3\) by only \(6.2\times10^{-6}\) — an agreement of one part in 30,000. Which is why, with no theoretical derivation at all, it has been taken seriously for over forty years. This table explains why certain coincidences alone get treated seriously.

Figure: the surprises of the coincidences handled so far, in bits. The slider changes how the prior range is chosensurprise depends on how much you consider could have happened. That is the weakest point of this measure.

35 orders
depends on the prior range (coincidences) independent of the prior range

Move the slider and only the top two bars stretch — "coincidence surprise" shifts by a few bits with the prior range. Identities (0 bits), Koide (whose \(Q\) range is mathematically fixed) and the CMB (whose patch count and precision are fixed by observation) do not move. So the former are "weak" coincidences and the latter "strong" ones.

◇ ◇ ◇

04The sorting procedure — count inputs and outputs

Apart from measuring bits, there is a more structural sort: what does the relation take in, and what does it put out?

ClassInputsOutputsFalsifiable?Surprise
Identity0 (follows from definitions)0no0 bit
Coincidence2 or more independent measurements0noa few bits
Physics\(n\) measurements\(m\) predictionsyeslarge

Apply it to Episode 1's \(\Omega/N=\ln2/3\pi^2(1+w)\): the input is one expansion law \(w\), the output is "operations per bit". It produces a quantity another measurement can reach, so it is physics. Whereas \(E/N=k_BT_H\ln2\) is the horizon's energy divided by the horizon's temperature and entropy — no input, no output.

05Rewritten as description length, it returns to Episode 5

This sorting is in fact what Episode 5 did. There we counted description length — a model is good if it lets you write the data shorter.

Conclusion of §05

A good relation is one that reduces the description length of the data.
An identity reduces it by 0 bits. A coincidence reduces it by a few, and has no mechanism to do the reducing.
Physics reduces it by orders — \(\Lambda\)CDM explains \(6.3\times10^6\) multipoles up to \(\ell\le2500\) with six parameters.

Episode 5 saw that "shortness buys only \(\log N\) while misfit costs \(N\)". Today the same balance is applied not to models but to coincidences. A coincidence demands explanation only when it could greatly reduce the description length.

06So is there nothing left in an identity?

This needs care. "An identity is not physics" does not mean "an identity is meaningless".

It is not a predictionit produces no quantity a new measurement can reach, so it is not a "discovery"
It is a consistency check\(E=T_HS\) holds because holography and thermodynamics do not contradict each other. If it failed, one of them would be wrong
Breaking it would be a major discoveryidentities are not unfalsifiable; they fail exactly when a premise of their derivation fails, and the failure means that premise has failed

So the accurate phrasing is: an identity is not a "prediction" but a "check". Every time this series has found one and written "not physics", it meant do not treat it as a mystery — not throw it away.

In Dirac's defence In 1937 Dirac did not know that \(M/m_P=R_H/2\ell_P\) is a consequence of the Friedmann equations (Friedmann cosmology was not yet established). With the knowledge of the time, it genuinely looked like a mystery. Recognising an identity is hindsight, and discovering that something "was an identity" is itself progress — one mystery fewer. Episode 7's "his prescription spins freely" was said with today's knowledge.
The honest line — what this episode assumes

① "Surprise in bits" shifts by a few bits with the choice of prior range. That is exactly what the figure's slider shows: the 4.7 bits for \(\rho_\Lambda^{1/4}\) and \(m_\nu\) assumes a 35-order range. Twenty orders gives 3.9 bits, sixty gives 5.5. The procedure can rank things but is not an absolute number — in Bayesian terms, it is the choice of prior.

② A log-uniform prior is assumed. Natural for a scale (a Jeffreys prior), but not the only choice.

③ The CMB value (\(1.6\times10^5\) bits) uses Episode 17's naive count of "17 independent bits × 9,600". Real fluctuations are structured by acoustic oscillations, so it is not an exact information content (Episode 17 ②). Read it as an order-of-magnitude comparison.

④ Koide's 15.7 bits takes \(Q\in[1/3,1]\) as the prior range. That is the mathematical range of \(Q\) for three positive masses, so the prior is barely arbitrary — which is part of what makes this coincidence "strong". But as Extra 4 of the previous series showed, the relation holds only for pole masses and the discrepancy grows 205-fold when run to \(M_Z\). Being surprising and being right are different things.

⑤ The "\(n\) inputs → \(m\) outputs" sort is this series' own. Philosophy of science offers many criteria — falsifiability, predictive power, unifying power — and this extracts only the part expressible in information terms.

Exercises (solvable with this episode's formulas alone)

  1. Why is an identity's surprise 0 bits?
    Show the answer
    Because, following from definitions, there was only one place for it to land. The "range allowed beforehand" is a single point, so \(-\log_2(1)=0\). There is nothing to be surprised by.
  2. Find the surprise that \(\rho_\Lambda^{1/4}\) and \(m_\nu\) agree within a factor 22, with a 35-order prior.
    Show the answer
    \(\log_{10}22.3=1.35\) orders; probability \(1.35/35=0.0385\); \(-\log_2 0.0385=\) 4.7 bits — five coin flips. Agreement between absurdly small numbers is, counted logarithmically, not much of a surprise.
  3. Why is Koide's relation alone so surprising?
    Show the answer
    Because \(Q\)'s range \([1/3,1]\) is mathematically fixed, yet the measurement departs from \(2/3\) by only \(6.2\times10^{-6}\). Probability \(1.85\times10^{-5}\), i.e. 15.7 bits = one in 30,000. With little arbitrariness in the prior, the coincidence is "strong" — which is why it is taken seriously despite having no derivation.
  4. Is "an identity is not physics" the same as "an identity is meaningless"?
    Show the answer
    No. An identity is not a prediction but a consistency check. \(E=T_HS\) holds because holography and thermodynamics do not contradict each other, and if it broke, one of the premises would be wrong. "Not physics" means do not treat it as a mystery, not discard it.
  5. (Harder) What is the greatest weakness of this measure?
    Show the answer
    Dependence on the prior range. As the slider shows, the \(\rho_\Lambda\)/\(m_\nu\) surprise is 3.9 bits at 20 orders and 5.5 at 60. In Bayesian terms it is the choice of prior: usable for ranking, not as an absolute number. Conversely, where the prior is fixed by mathematics or observation — Koide, the CMB — the weakness shrinks.

Summary — measured in bits, surprise falls into three strata

We turned this series' repeated sorting — identity, coincidence or physics — into a procedure, using information theory's surprisal: \(-\log_2(\text{width}/\text{prior range})\).

Measured, the strata separate cleanly. Identities at 0 bits (Dirac's large numbers, the Landauer limit, \(\alpha+2\beta+\gamma=2\), Episode 18's \(N/(R/\ell_P)^3\)). Coincidences at a few bits — the factor-22 agreement of \(\rho_\Lambda^{1/4}\) and \(m_\nu\) is five coin flips (4.7 bits), last episode's 1.96 fm is 7.4. Agreement between absurdly small numbers is, logarithmically, not much of a surprise. And real problems at \(10^5\) bits (CMB uniformity). The one entry at 15.7 bits is Koide's relation, which is why it has been taken seriously without a derivation.

We built a structural sort too — identities take 0 inputs and give 0 outputs; coincidences take 2 or more and give 0; physics takes \(n\) and gives \(m\) predictions. It is the same balance as Episode 5's description length: a good relation reduces the description length of the data, and \(\Lambda\)CDM explains \(6.3\times10^6\) multipoles with six parameters.

One point of care at the end. "An identity is not physics" does not mean "meaningless" — an identity is a check, not a prediction, and breaking it would mean a premise had failed. Dirac in 1937 did not know the Friedmann equations, so with the knowledge of the time it genuinely looked like a mystery. Discovering that something was an identity is itself progress — one mystery fewer.

This document is Episode 19 of "c·t = const, That Clicks", written for physics-minded high-school and university readers. Surprisal \(-\log_2 p\) is a standard information-theoretic quantity. The surprise values given here (4.7 bits for \(\rho_\Lambda^{1/4}\) and \(m_\nu\), 7.4 for a bit's volume against a nucleon, 15.7 for Koide's relation, \(1.6\times10^5\) for CMB uniformity) are computed here (kenshou/calc23.py). They depend on the choice of prior range: the 4.7 bits assumes a 35-order log-uniform (Jeffreys) prior, becoming 3.9 at twenty orders and 5.5 at sixty. The procedure ranks; it is not an absolute number. The CMB's \(1.6\times10^5\) bits uses Episode 17's naive count of "17 independent bits × 9,600"; real fluctuations are structured by acoustic oscillations, so it is not an exact information content. Koide's 15.7 bits takes \(Q\in[1/3,1]\) (the mathematical range for three positive masses) as the prior, with \(Q=0.66666051\) departing from \(2/3\) by \(6.2\times10^{-6}\) — though the relation has no theoretical derivation, holds only for pole masses, and its discrepancy grows 205-fold when run to \(M_Z\) (Extra 4 of the previous series). Being surprising and being right are different things. The "\(n\) inputs → \(m\) outputs" sort is this series' own, extracting from the philosophy of science (falsifiability, predictive power, unifying power) only what can be written in information terms. The remark that Dirac (1937) did not know \(M/m_P=R_H/2\ell_P\) to be a consequence of the Friedmann equations is a point about historical context. The six \(\Lambda\)CDM parameters and the \(6.3\times10^6\) \(a_{\ell m}\) modes up to \(\ell\le2500\) are indicative. Linear expansion (\(c\cdot t=\)const, \(R_h=ct\)) is a minority model under examination. The academic standard remains the \(\Lambda\)CDM model including inflation. ── To make a PDF, use your browser's Print dialogue (sliders freeze and answers are hidden in the print version).

Print / PDF: ⌘+P (Ctrl+P on Windows). On screen, the slider changes the prior range and only the coincidence bars move. "Show the answer" opens each solution.