Starlight

Every distance is measured with the last one

No single method reaches from a planet to a distant galaxy. The ladder is built rung by rung, each calibrated on the one below, and the errors multiply all the way up.

Assumes Parallax and Magnitudes.

There is no instrument that measures the distance to a galaxy. There is no instrument that measures the distance to a star. What exists is a sequence of methods, each of which works over a limited range and each of which is calibrated by comparison with the one below it.

That structure has a name — the cosmic distance ladder — and the metaphor is exact in a way metaphors usually are not. A ladder is climbed one rung at a time, each rung bears on the one below, and if a lower rung moves then everything above it moves with it.

The consequence is arithmetic and unavoidable. Fractional errors compound as the ladder is climbed, so a distance to a remote galaxy carries, inside it, the uncertainty of a radar echo off Venus.

The fractional distance error, rung by rung. Cumulative fractional uncertainty in a measured distance against the distance itself, both on logarithmic axes. Each rung of the ladder is calibrated against the one below it, so its scatter adds in quadrature to everything already inherited, and the total can only rise: radar to the planets 10⁻⁵%, parallax 1.0%, main-sequence fitting 5.1%, Cepheids 6.5%, Type Ia supernovae 8.2%.
Fig. 1 The cumulative fractional uncertainty against distance. Each rung adds its own scatter in quadrature to everything inherited, so the curve is a staircase that only ever rises: from a part in ten million inside the solar system, through 1% at a kiloparsec, to 8% at the reach of Type Ia supernovae.

The bottom rung is not astronomical at all

The ladder begins with a measurement made with a radio transmitter.

Bouncing radar off Venus, Mercury or Mars gives the distance to that planet as a round-trip light travel time, and the speed of light is a defined constant, which is what makes a time into a distance. The result is a distance in metres, to a few hundred metres out of 101110^{11} — a fractional precision of a few parts in 10910^{9}.

That single measurement fixes the entire scale of the solar system, because Kepler’s third law gives every planet’s semi-major axis in units of the Earth’s, and the radar echo supplies the conversion factor. Before radar, this rung was the weakest of the whole ladder: the astronomical unit was known to about 0.1% in 1900, from transits of Venus and from parallaxes of Eros, and everything above it inherited that.

Since 2012 the astronomical unit is defined as exactly 149,597,870,700 metres, so it no longer carries an uncertainty at all. What carries uncertainty now is the relationship between the AU and any particular planet’s orbit, and that is known to parts in 101110^{11}.

The bottom rung, in other words, is not an astronomical measurement and is not a source of error. That is unusual and it is worth noticing, because every rung above it is both.

Parallax: the only rung that is geometry

The second rung is the only one on the whole ladder that measures a distance without assuming anything about the object.

Trigonometric parallax uses the Earth’s orbit as a baseline and measures the annual shift of a nearby star against the background. The definition of the parsec makes the arithmetic trivial — d=1/pd = 1/p with dd in parsecs and pp in arcseconds — and the method requires no knowledge whatever of the star’s brightness, composition or nature.

That is why parallax is the calibrating rung. Every method above it is a method for inferring an object’s intrinsic brightness from something else observable, and every such method needs a sample of objects whose distances are known independently in order to be calibrated. Parallax is what supplies that sample.

The reach has grown by four orders of magnitude in two centuries. Bessel measured 61 Cygni at 0.314″ in 1838 — a distance of 3.2 parsecs, with an error of about 10%. Hipparcos, in 1997, reached about 1 milliarcsecond, giving useful distances to a few hundred parsecs. Gaia measures parallaxes for 1.5 billion stars at 20 to 300 microarcseconds, giving 1% distances out to about 2 kiloparsecs and usable ones to 10. Even Gaia does not reach far. Two kiloparsecs is a twentieth of the way across the Milky Way and a ten-thousandth of the way to the Virgo cluster. The ladder has to keep climbing.

Standard candles: the trade the ladder makes

Every rung above parallax works the same way. Find a class of object whose intrinsic luminosity can be inferred from something that does not depend on distance — a period, a spectral type, a light-curve shape — and then compare that inferred luminosity with the observed brightness. The distance modulus converts the ratio into a distance.

The trade is explicit: reach is bought with assumption. Parallax assumes nothing and reaches 2 kpc. Cepheids assume a period–luminosity relation and reach 30 Mpc. Type Ia supernovae assume a light-curve-shape relation and reach a thousand times further. Each step outward is a step away from geometry. The rungs in the modern ladder, with their approximate reach and their own scatter:

Main-sequence fitting — the HR diagram of a star cluster is slid vertically until its main sequence overlies a calibrated one, and the shift is the distance modulus. Reach 50 kpc, scatter 5%, and it depends on the cluster’s metallicity and reddening being right.

Cepheid variables — the period–luminosity relation. Reach 30 Mpc with the Hubble and JWST, scatter 4% once calibrated.

The tip of the red giant branch — the sharp upper limit to the luminosity of red giants, set by the helium flash, which is nearly independent of a star’s mass and metallicity. Reach 20 Mpc, scatter 5%, and it is the main independent cross-check on Cepheids.

Type Ia supernovae — a white dwarf detonating near the Chandrasekhar mass, so the energy released is nearly always the same — set by a counting rule rather than by any accident of the star. After correcting for the correlation between peak brightness and decline rate, the scatter is 5% in distance — on objects visible at a redshift of 2.

The rung that is one star at a time

There is a step on the ladder that is easy to skip past and is worth a paragraph on its own, because it is the only one that measures a distance to a single object without a candle.

If a star’s angular diameter can be measured and its surface brightness inferred from its spectrum, its physical diameter follows, and comparing the two gives the distance. That is the principle behind detached eclipsing binaries: two stars that eclipse each other give their radii in absolute units from the eclipse durations and their velocities, and give their surface brightnesses from their colours. The distance follows with no candle and no calibration against a lower rung.

The Large Magellanic Cloud’s distance is now fixed this way, from twenty late-type eclipsing binaries, to 1%: 49.59±0.5449.59 \pm 0.54 kpc. That single number carries much of the modern ladder, because the LMC is where the Cepheid period–luminosity relation is anchored.

It is worth noticing what has happened there. A rung that used to be built on Cepheids calibrated by parallax is now built on binaries calibrated by nothing at all, and the ladder has quietly grown a second foot. Redundancy at the bottom is the only structural defence a chain of calibrations has.

Why the errors multiply

Each rung’s calibration is a comparison against the rung below, so its zero point carries the lower rung’s uncertainty. The errors are independent, so they add in quadrature:

σtotal2(k)=σtotal2(k1)+σrung2(k).\sigma_{\text{total}}^2(k) = \sigma_{\text{total}}^2(k-1) + \sigma_{\text{rung}}^2(k).

Adding in quadrature is merciful compared with adding directly — five rungs at 5% each give 11%, not 25% — but it has a property that catches people out: the total is dominated by the largest single term. A rung at 8% and four rungs at 2% give 8.5%. Improving the four small ones to zero would change almost nothing. That is why effort concentrates where it does. When the Hubble constant was uncertain at the factor-of-two level in the 1970s, the dominant term was the Cepheid zero point, and it was worth a Key Project to fix it. Now the dominant terms are the Cepheid–supernova calibration and the supernova standardisation, each around 1.5%, and the parallax rung beneath them has been pushed to 0.3% by Gaia — beyond the point where it matters.

The other property that catches people out is that a revision low on the ladder moves everything above it, by the same fractional amount, all at once. When Walter Baade discovered in 1952 that there were two distinct classes of Cepheid with different period–luminosity relations, and that the calibration had used one class and the measurements the other, every extragalactic distance doubled overnight. The Andromeda galaxy went from 250 kiloparsecs to 500; the universe’s age went from 1.8 billion years — embarrassingly less than the age of the Earth — to something respectable. Nothing above the Cepheid rung had been remeasured. One rung moved and the whole ladder moved with it.

What was actually measured

The current state of the ladder is best read through the disagreement at the top of it, because a disagreement is the only thing that tests a measurement chain rather than merely reporting it.

The SH0ES programme measures the Hubble constant by climbing: Gaia parallaxes and detached eclipsing binaries in the Large Magellanic Cloud calibrate Cepheids; Cepheids in 42 galaxies that have also hosted Type Ia supernovae calibrate the supernovae; the supernovae reach into the Hubble flow. The 2022 result is 73.04±1.0473.04 \pm 1.04 km/s/Mpc — a 1.4% measurement, built from three rungs each calibrated on the last.

The Planck satellite measures the same constant without a ladder at all, by fitting a cosmological model to the pattern of temperature fluctuations in the microwave background. Its result is 67.4±0.567.4 \pm 0.5.

The two disagree by about 5σ. That is the Hubble tension, and the reason it is interesting rather than merely annoying is precisely the ladder structure: a systematic error anywhere on any rung would produce exactly this, and so would new physics, and the two are hard to tell apart. Enormous effort has gone into finding a rung that is wrong. The tip of the red giant branch, used instead of Cepheids, gives about 70 — between the two, and with its own disputes about how the tip is measured. The honest summary is that the ladder now has systematic uncertainties smaller than the discrepancy it is being asked to resolve, which is a position no distance measurement in astronomy has ever been in before, and nobody knows yet whether the answer is a subtle calibration error or a missing ingredient in cosmology.

The fractional error on an occurrence rate. Cumulative fractional uncertainty in an occurrence rate against the number counted, both on logarithmic axes. Each factor of an occurrence rate is measured separately and multiplies the last, so its fractional error adds in quadrature to everything already inherited: geometric probability 2.0%, detection completeness 15.1%, reliability 25.1%.
Fig. 2 The identical structure in a subject that has nothing to do with distance. An exoplanet occurrence rate is a count divided by a geometric probability, a completeness and a reliability — three multiplicative corrections, each measured separately, each with its own fractional error, and the total can only rise. The shape of the argument is the ladder’s exactly: what is quoted at the top is a product of things measured at the bottom, and no amount of care in the last step recovers what was lost in the first.

The generalisation: chained calibration everywhere

The distance ladder is the most visible example of a structure that appears wherever a quantity cannot be measured directly across its whole range.

Geological time is dated by a chain: radiometric dates on volcanic ash beds calibrate biostratigraphic zones, which calibrate the relative sequence, which is what most rock is actually dated by. A revision to a decay constant moves the entire Phanerozoic.

Radiocarbon dating is calibrated against tree rings, which are calibrated against counted varves and coral, and the calibration curve has structure that no single measurement could reveal.

Temperature scales below 1 K are chained through a sequence of fixed points and interpolating thermometers, each valid over a limited range and each calibrated in the overlap with the last.

All of them share the three properties of the astronomical ladder: reach is bought with assumption, errors compound in quadrature, and a revision at the bottom propagates all the way up without anything above being remeasured. The third is the one that makes such a structure feel unsatisfying, and it is also the one that makes it possible to improve everything at once by improving one thing.

The fractional distance error, rung by rung. Cumulative fractional uncertainty in a measured distance against the distance itself, both on logarithmic axes. Each rung of the ladder is calibrated against the one below it, so its scatter adds in quadrature to everything already inherited, and the total can only rise: radar to the planets 10⁻⁵%, parallax 1.0%, Cepheids 6.1%, Type Ia supernovae 7.9%.
Fig. 3 The same chain with one rung removed. Dropping main-sequence fitting and calibrating Cepheids directly on parallax shortens the ladder by a link and lowers the total — but only because the link removed was contributing scatter, not because fewer rungs is better in itself. What a rung costs is exactly its own fractional scatter added in quadrature, so a rung that is more precise than the ones beneath it is nearly free and one that is worse dominates everything above it.
The fractional distance error, rung by rung. Cumulative fractional uncertainty in a measured distance against the distance itself, both on logarithmic axes. Each rung of the ladder is calibrated against the one below it, so its scatter adds in quadrature to everything already inherited, and the total can only rise: parallax 1.0%, Cepheids 4.1%, Type Ia supernovae 6.5%.
Fig. 4 The same ladder with two rungs removed — parallax straight to Cepheids to supernovae, which is the modern chain now that Gaia’s parallaxes reach the Cepheids directly. Two fewer quadrature additions and the error at the top falls accordingly. Removing a rung is worth more than improving one, because each rung contributes its whole scatter to everything above it and no amount of care on a middle rung recovers what it costs merely by existing.

What a rung costs to add

A new rung is worth having only if it reaches further than the one below and overlaps it enough to be calibrated, and those two requirements pull against each other.

The overlap has to contain enough objects to fix a zero point. Cepheids and Type Ia supernovae overlap in exactly the galaxies that have hosted both — a set of 42 in the current calibration, built up over thirty years, because a supernova has to happen in a galaxy near enough for the Hubble Space Telescope to resolve its Cepheids. That number is the binding constraint on the whole Hubble-constant measurement: with 42 anchors and 5% per supernova, the statistical floor is 0.8%, and every additional anchor moves it as the square root.

The tip of the red giant branch was adopted so quickly in the 2010s for the same reason from the other side: it is measurable in every nearby galaxy, including elliptical ones that have no Cepheids at all, so its overlap with the supernova rung is much larger than the Cepheid overlap and it can be calibrated against a different population entirely.

A rung with no overlap is not a rung. Several proposed distance indicators — some of them physically excellent — have never joined the ladder because the range over which they work contains too few objects whose distances are known some other way.

The fractional distance error, rung by rung. Cumulative fractional uncertainty in a measured distance against the distance itself, both on logarithmic axes. Each rung of the ladder is calibrated against the one below it, so its scatter adds in quadrature to everything already inherited, and the total can only rise: parallax 1.0%, Cepheids 4.1%, Type Ia supernovae 6.5%.
Fig. 5 The three rungs that actually carry the Hubble constant, with the solar-system links removed. Nothing below parallax matters for an extragalactic distance — the astronomical unit is known to metres and contributes nothing — so the ladder that sets H0H_0 is three links long and its total is 6.5 per cent. That is the number the local determination quotes, and every attempt to improve it is an attempt on one of these three.
The fractional distance error, rung by rung. Cumulative fractional uncertainty in a measured distance against the distance itself, both on logarithmic axes. Each rung of the ladder is calibrated against the one below it, so its scatter adds in quadrature to everything already inherited, and the total can only rise: radar to the planets 10⁻⁵%, parallax 1.0%, main-sequence fitting 5.1%, Cepheids 6.5%, Type Ia supernovae 8.2%.
Fig. 6 The same five rungs starting from a tenth of a milliparsec rather than a hundredth — the radar rung shortened, and nothing else changed. The curve above it is identical, which is the point: the bottom of the ladder contributes nothing to the error at the top, because radar to the planets is good to a part in ten million and adding that in quadrature to five per cent changes nothing at all. A chain is as strong as its weakest link and this chain’s weakest links are all near the top.

Where the model stops

The rungs are not independent measurements of the same thing. Quoting five methods that agree is not five confirmations if four of them are calibrated on the fifth. Genuine independence — Planck, or gravitational-wave standard sirens, which measure a distance from a waveform and need no ladder at all — is rare and valuable for that reason.

Every rung above parallax assumes uniformity. A standard candle is standard only to the extent that its composition allows — the objects in a distant galaxy are like the objects nearby, and the distant ones formed from gas that had been through fewer generations of stars. Metallicity dependence is the standing worry on every rung above the second.

Dust is not an error, it is a bias — and it is measured by the reddening it produces. Extinction always makes things look fainter and therefore always makes them look further away. It has to be measured and removed, and residual extinction produces a one-sided error that no amount of averaging fixes.

The figures show a staircase, and the real thing is a graph. The hero figure draws the rungs in sequence, as though each were used over one range and handed over cleanly to the next. In practice the rungs overlap heavily, several are used simultaneously with weights, and the calibration is a global fit rather than a chain of handovers. The staircase is the argument’s structure, not the procedure’s.

The fractional error on a siren distance. Cumulative fractional uncertainty in a siren distance against the distance itself, both on logarithmic axes. A siren has one rung: the amplitude of the wave gives the distance directly, and nothing beneath it is calibrated on anything: gravitational-wave amplitude 15.0%.
Fig. 7 And what a ladder with one rung looks like. A gravitational-wave siren gives its own distance from the amplitude of the wave, calibrated on nothing — so the chain is one link long and its error is one link’s error, 15 per cent for a single well-measured event. That is far worse than the three-rung ladder’s 6.5 per cent today and it improves as the root of the number of events, which is the whole argument for building more detectors rather than refining more rungs.

A distance with no ladder under it

One measurement mentioned in passing above deserves its own section, because it is the only genuinely new rung to have appeared in a century and it does not sit on anything.

When two compact objects spiral together, the waveform they emit carries an amplitude that depends on the distance and a frequency evolution that depends on the masses. Because the theory relating the two is known exactly, the observed amplitude and frequency evolution together give the luminosity distance directly — in metres, with no calibration against any other object and no assumption about intrinsic brightness.

That is why such events are called standard sirens rather than standard candles: nothing is assumed to be standard. The distance falls out of the physics of the source.

Two things stand between the principle and a cosmological measurement. The amplitude also depends on the orbit’s inclination — a face-on system looks nearer than an edge-on one of the same distance — so the two are degenerate, and breaking the degeneracy needs either a network of detectors or additional information about the source. And the redshift, which is the other half of any expansion measurement, is not carried by the waveform at all; it requires identifying the host galaxy, which requires seeing the event some other way.

One event so far has supplied both: a neutron-star merger in 2017 with an electromagnetic counterpart bright enough to locate its host. The Hubble constant that followed had an uncertainty of about ten per cent, which is far worse than either of the values it sits between and is the point — it is a measurement that shares no systematic with either.

The fractional distance error, rung by rung. Cumulative fractional uncertainty in a measured distance against the distance itself, both on logarithmic axes. Each rung of the ladder is calibrated against the one below it, so its scatter adds in quadrature to everything already inherited, and the total can only rise: radar to the planets 10⁻⁵%, parallax 0.5%, main-sequence fitting 5.0%, Cepheids 5.4%, Type Ia supernovae 7.4%.
Fig. 8 The same five rungs with two of them improved — parallax halved to half a per cent and the Cepheid rung halved to two. The top of the ladder falls, and it falls by less than either improvement: the errors add in quadrature, so halving one contributor of several buys a fraction of a halving overall. This is the arithmetic that makes the Hubble tension expensive, because every route to a smaller error at the top requires improvements at several rungs at once, and each of those is a separate decade of work.

The error accounting above is one way of drawing the ladder. The other is simply to draw its reach, and the two together are the whole of what a distance scale is.

The distance ladder, and its overlaps. The reach of each distance technique on a logarithmic scale in parsecs. Each rung is calibrated where it overlaps the one below it, so an error low on the ladder propagates all the way to the top.
Fig. 9 The reach of each technique on a logarithmic scale in parsecs, with the overlaps between them. Every rung is calibrated against the one below it over the range where both work, so the width of each overlap is what limits how well the calibration can be done — and two of the overlaps are narrower than anybody would design.

The ladder from here

Later rungs on this anchor: parallax as it is actually done, with the Gaia astrometric solution and its parallax zero-point offset. Main-sequence fitting and the metallicity problem. The tip of the red giant branch, and why the helium flash makes it sharp. Type Ia supernovae and the Phillips relation. Surface brightness fluctuations, the Tully–Fisher relation, and the other rungs this essay skipped. Standard sirens, and a distance measured from a waveform. The Hubble tension, followed through every rung.

There is a nice irony in the bottom of the ladder. The one rung that is not astronomical, not uncertain, and not calibrated against anything — the radar echo — is the rung that was the worst for the whole of history before 1961. Determining the astronomical unit was the great observational problem of the eighteenth and nineteenth centuries, prosecuted by sending expeditions to the ends of the earth to watch Venus cross the Sun, and it was solved in the end by pointing a radio transmitter at a planet and waiting four and a half minutes.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 39 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Astronomical unit (AU)Cluster turnoffDistance ladderDistance modulusGiant branchMetallicityParsecPhotometrySupernovaTrigonometric parallax