Cosmology

The same constant, measured twice, five sigma apart

The distance ladder gives an expansion rate of about 73 kilometres per second per megaparsec. The microwave background gives 67.4. Both quote errors near one per cent, both have been rebuilt from scratch by rival teams, and the gap between them has grown as the measurements have improved.

Assumes Hubble constant and Distance ladder.

A disagreement between two measurements is ordinary. A disagreement that survives twenty years of both sides being rebuilt, by people trying hard to find their own mistakes, and that grows rather than shrinks as the error bars come down, is not.

H₀: nine determinations in two families. Published determinations of H₀, each with its quoted one-sigma interval, sorted into two families — measured locally, calibrated by a ladder, against inferred from z ≈ 1100 through a model. The shaded band behind each family is that family's inverse-variance weighted mean: 72.66 ± 0.75 across 5 of them, against 67.40 ± 0.41 across 4. The difference is 5.26 ± 0.85 km/s/Mpc, which is 6.2 standard deviations, computed here from the quoted errors alone. That number is an upper bound on the significance rather than the significance: the determinations within each family share calibrations, samples and in two cases the same supernovae, so they are not independent, and a correlated pair combines to something wider than the formula used here gives. What the figure does establish is that the split is not one discrepant measurement against a consensus — it is two internally consistent groups, and the grouping is by method rather than by result.
Fig. 1 Nine published determinations of the expansion rate, with their quoted one-sigma intervals, sorted not by method or by date but by what each one has to assume. The five in the upper group measure a distance and a redshift for objects in the local flow and divide: they assume nothing about the contents of the universe and everything about the calibration of a ladder. The four below measure an angle on a map from z1100z \approx 1100 and infer H0H_0 through a six-parameter model: they assume nothing about any ladder and everything about the model. The weighted means are 72.7 ± 0.7 and 67.4 ± 0.4, and the difference computed from the quoted errors alone is 6.2 standard deviations. That number is an upper bound rather than the significance, because the determinations within a family share calibrations and in two cases share supernovae — but the shape of the plot is the finding, and the shape is two internally consistent groups split by method.

The two things being compared are not the same measurement

It is tempting to describe this as two measurements of one number that disagree. That is not quite what it is, and the difference matters for every proposed resolution.

The local route measures H0H_0 more or less directly. It finds objects whose intrinsic brightness is known, measures how bright they look, converts that to a distance, measures a redshift, and divides. Nothing in that chain requires knowing what the universe is made of or how old it is. It is a measurement in the sense that a survey is a measurement.

The early-universe route measures nothing of the kind. It measures an angle — the angular scale of the acoustic peaks in the microwave background — and that angle is a ratio of a length to a distance. Extracting H0H_0 from it requires knowing the length, which requires the physics of the pre-recombination plasma, and requires converting the distance into an expansion rate, which requires the whole expansion history in between. H0H_0 comes out of the early route as a derived parameter of a fit, not as a reading.

So the disagreement is not “two rulers give different answers”. It is “a ruler and a model give different answers”, and the resolutions divide accordingly: something is wrong with the ruler, something is wrong with the model, or something is wrong in between.

H₀: nine determinations in two families. Published determinations of H₀, each with its quoted one-sigma interval, sorted into two families — measured locally, calibrated by a ladder, against inferred from z ≈ 1100 through a model. The shaded band behind each family is that family's inverse-variance weighted mean: 72.658 ± 0.749 across 5 of them, against 67.396 ± 0.408 across 4. The difference is 5.262 ± 0.853 km/s/Mpc, which is 6.2 standard deviations, computed here from the quoted errors alone. That number is an upper bound on the significance rather than the significance: the determinations within each family share calibrations, samples and in two cases the same supernovae, so they are not independent, and a correlated pair combines to something wider than the formula used here gives. What the figure does establish is that the split is not one discrepant measurement against a consensus — it is two internally consistent groups, and the grouping is by method rather than by result.
Fig. 2 The same nine determinations quoted to three decimals, which is more precision than any of them possesses and is worth drawing once. Nothing about the picture changes — the values, the intervals and the grouping are identical — and the extra digits are noise from the arithmetic rather than information. Set against the one-decimal version further down, the pair bracket the honest number of figures: two, which is where the separation between the families is visible and the digits are still meant.
G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — torsion balance, in one form or another, against beam balance, pendulum, atom interferometry. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67435 ± 0.00004 across 11 of them, against 6.67343 ± 0.00009 across 3. The difference is 0.00092 ± 0.00010 10⁻¹¹ m³ kg⁻¹ s⁻², which is 9.3 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 3 The same kind of disagreement in a quantity nobody suspects of cosmology. Determinations of the gravitational constant each quote an uncertainty of a few parts in a hundred thousand and disagree with one another by ten times that — so the scatter between experiments is far larger than any of them admits. That is what an unrecognised systematic looks like from outside, and it is the null hypothesis every claimed tension has to be tested against before it is a discovery.

What the local route rests on

Two of the three geometric anchors deserve naming, because they are the reason the bottom rung is no longer the weak one. Gaia has measured parallaxes for Milky Way Cepheids to tens of microarcseconds, which is a direct trigonometric distance to the calibrators themselves. And NGC 4258 carries a disc of water masers orbiting the black hole at its centre — spots whose positions, line-of-sight velocities and centripetal accelerations are all measured, which over-determines the geometry and gives a distance of 7.58 Mpc good to 1.5 per cent with no photometry in it at all. A galaxy at that distance is far enough to contain Cepheids observable with the same instruments used on the distant ones.

The middle rung is the one that carries the most weight and attracts the most argument. The team that produces the highest-weight local number has spent two decades on exactly those effects: observing the calibrating and the distant Cepheids with the same instrument and the same filters so that the calibration errors cancel, measuring the metallicity dependence rather than assuming it, and resolving crowded fields with the James Webb Space Telescope to check whether unresolved companions were inflating the brightnesses. The last of those was the most-cited candidate explanation for the tension, and the check found the crowding correction to be small. An independent local route replaces Cepheids with the tip of the red giant branch. In an old stellar population every star that reaches the top of the giant branch ignites helium at very nearly the same core mass and therefore at very nearly the same luminosity, so the giant branch on a colour–magnitude diagram has a sharp upper edge, and the edge is a standard candle. It has two advantages over a Cepheid: the stars are old, so they live in the smooth outskirts of a galaxy rather than in crowded star-forming regions, and the metallicity dependence is much weaker in the infrared. It returns 69.8, between the two camps and consistent with either at about two sigma, and whether that is a third result or a low local one is genuinely unsettled — which is itself informative, because the two Cepheid-free local methods do not agree with each other as well as either agrees with its own error bar.

What the early route rests on

The peak position is an angle: θ=0.0104110±0.0000031\theta_* = 0.0104110 \pm 0.0000031 radians, which is the best-measured quantity in cosmology. It is the ratio of the sound horizon at last scattering, a length, to the comoving distance to last scattering. That is where the assumption enters, and it enters in one specific place: the sound horizon rsr_s. If rsr_s is smaller than ΛCDM says, the same measured angle implies a smaller distance to last scattering, which implies a larger H0H_0. Everything else about the microwave background can be held fixed. The tension can be phrased entirely as a disagreement about the length of a ruler nobody can measure directly, and a reduction of about 7 per cent in rsr_s would remove it.

There is also a hybrid route worth naming, because it is the one that shows the early number is not just a Planck result. Take the baryon acoustic scale from galaxy surveys, calibrate its absolute length with the baryon density from deuterium rather than from the microwave background, and run the ladder downward instead of upward. That “inverse distance ladder” uses no microwave-background map at all, and returns 67.4.

H₀: nine determinations in two families. Published determinations of H₀, each with its quoted one-sigma interval, sorted into two families — measured locally, calibrated by a ladder, against inferred from z ≈ 1100 through a model. The shaded band behind each family is that family's inverse-variance weighted mean: 72.7 ± 0.7 across 5 of them, against 67.4 ± 0.4 across 4. The difference is 5.3 ± 0.9 km/s/Mpc, which is 6.2 standard deviations, computed here from the quoted errors alone. That number is an upper bound on the significance rather than the significance: the determinations within each family share calibrations, samples and in two cases the same supernovae, so they are not independent, and a correlated pair combines to something wider than the formula used here gives. What the figure does establish is that the split is not one discrepant measurement against a consensus — it is two internally consistent groups, and the grouping is by method rather than by result.
Fig. 4 The same comparison quoted to one decimal rather than two, which is how the disagreement usually reaches a headline. Rounding hides it: at one decimal the two combined values look almost adjacent, and at two the separation is unmistakable. That is a small thing worth naming, because a tension is the ratio of a difference to an uncertainty, and any rounding that touches either changes the number of sigma being claimed.

What would have to be wrong

The candidate resolutions fall into three groups, and it is worth stating what each costs.

Something is wrong with the local ladder. This was the majority view for a decade and it has weakened. The specific candidates — Cepheid crowding, metallicity, the supernova standardisation, the possibility that the local volume is inside an underdense region and so expanding slightly fast — have each been tested and each comes out too small by a factor of several. The local void in particular is bounded by the galaxy counts and by the supernova sample’s own uniformity: a void deep enough to explain the gap would show up as a redshift-dependent H0H_0 within 200 Mpc, and it does not.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — one laboratory, four determinations, against every other laboratory. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67425 ± 0.00005 across 4 of them, against 6.67414 ± 0.00005 across 10. The difference is 0.00012 ± 0.00007 10⁻¹¹ m³ kg⁻¹ s⁻², which is 1.6 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 5 The null hypothesis drawn as a laboratory would draw it. The same fourteen determinations of GG sorted by which group made them rather than by method: the four from one laboratory agree with the other ten to 1.6 sigma, where the method split reports 9.3. That is what an unrecognised systematic looks like when the partition is changed — the significance moves by a factor of six on unchanged data — and it is the reason a tension has to be shown to be robust against re-grouping before it is a discovery. The Hubble determinations have been subjected to exactly that test and the split survives it.

Something is wrong with the model between recombination and now. This is the hardest to arrange. The late universe is heavily constrained by supernovae, by baryon acoustic oscillations at several redshifts, and by lensing, and those constrain the shape of the expansion history very well while being nearly blind to its overall scale. Modifications that raise H0H_0 tend to break one of them.

Something is wrong with the model before recombination. This is where most current proposals live, precisely because it is the least constrained epoch and because shrinking rsr_s is the cleanest way to move H0H_0. Adding a component that behaves like dark energy for a brief window around matter–radiation equality and then dilutes away faster than radiation — “early dark energy” — raises H(z)H(z) just when the sound horizon is being laid down and shortens it. It works, at the cost of a new field with a tuned mass and a tuned turn-on time, and it leaves fingerprints in the higher acoustic peaks that the current data mildly disfavour.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — quoted better than 100 ppm, against quoted worse than 100 ppm. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67420 ± 0.00004 across 11 of them, against 6.67301 ± 0.00048 across 3. The difference is 0.00120 ± 0.00048 10⁻¹¹ m³ kg⁻¹ s⁻², which is 2.5 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 6 And the other diagnostic worth borrowing. Sorting the GG determinations by their own quoted precision puts the eleven most confident in one family, and they span five hundred parts per million while quoting intervals between twelve and a hundred and thirty. That pattern — the tightest error bars scattering furthest in units of themselves — is the signature of a systematic nobody has found, and it is exactly what the Hubble determinations do not do. Within each Hubble family the measurements are consistent with their own errors; it is only across the families that they are not.

A fourth possibility deserves stating plainly because it is not a proposal: that one of the two families has an unmodelled systematic that nobody has thought of. That has been the resolution of most historical tensions of this kind, and the base rate for it is not low. It is also the possibility that cannot be argued for, only found, which is why the honest summary of the field’s position is that the tension is real, that it is unexplained, and that the explanation is more likely to be dull than interesting.

What the picture cannot show

The hero figure treats nine measurements as nine independent facts, and they are not. Three of the five local determinations use the same supernova sample; two use the same geometric anchor. Correlated measurements combine to something wider than the inverse-variance formula gives, so the drawn 6.2 sigma is an upper bound. The published estimates that account for the correlations land between four and five, which is still a great deal.

No figure here can show a systematic error, because a systematic error that anyone could draw would already have been fixed. Every point on the hero plot has an error bar containing the systematics its authors could think of. The disagreement is, by construction, about the ones they could not.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — torsion balance, time-of-swing, against every other technique, torsion included. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67400 ± 0.00007 across 6 of them, against 6.67428 ± 0.00004 across 8. The difference is -0.00028 ± 0.00008 10⁻¹¹ m³ kg⁻¹ s⁻², which is 3.4 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 7 A third partition of the same fourteen: time-of-swing torsion balances against every other technique, torsion included. Three point four sigma, from measurements using the same instrument and differing only in how a period is read. A method split that survives restricting to one method is not a method split, and the lesson transfers directly — the Hubble tension’s own grouping is by what each determination assumes rather than by what apparatus it used, and that is a stronger partition precisely because it is not a property of the hardware.

And the plot cannot show what would settle it. What would settle it is a measurement that is neither of these: a local determination with no ladder in it at all, or an early-universe determination with no microwave-background map. Both exist in prototype. Gravitational-wave standard sirens read a luminosity distance straight off a waveform, with no calibration whatever; the single event with an electromagnetic counterpart so far gave 708+1270^{+12}_{-8}, which excludes nothing and is the right kind of number. Strongly lensed quasars measure a time delay against a lens model and currently give 73, on the local side, with the lens mass profile as their dominant uncertainty.

The history is short and it goes the wrong way

For most of the twentieth century the argument about H0H_0 was a factor of two: Allan Sandage’s school held out for 50 and Gérard de Vaucouleurs’s for 100, and the dispute lasted thirty years. The Hubble Space Telescope Key Project was built to end it and did, reporting 72±872 \pm 8 in 2001 — a result that was celebrated precisely because it agreed with everything within its errors.

The tension is what happened when that eight became a one. Planck’s first cosmological results in 2013 gave 67.3; the local number had by then tightened to 73.8. Every subsequent improvement on either side has kept the central values and shrunk the intervals, which is the opposite of what a statistical fluctuation does and the opposite of what an unresolved systematic usually does when the measurement is rebuilt.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — before 2005, against 2005 and after. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67427 ± 0.00008 across 5 of them, against 6.67418 ± 0.00004 across 9. The difference is 0.00010 ± 0.00009 10⁻¹¹ m³ kg⁻¹ s⁻², which is 1.0 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 8 And the historical control. Sorting the GG determinations before and after 2005 gives one standard deviation — the weakest of the four partitions — so two centuries of improvement have shrunk the quoted intervals by three orders of magnitude and left the disagreement between laboratories where it was. That is what a systematic looks like when a measurement is rebuilt: the central values stay and the errors fall, which is the same behaviour the previous paragraph attributes to the Hubble tension. The difference between the two cases is that here the scatter is within each family and there it is between them.

The instructive comparison is with the age crisis of the 1990s, which had the same shape: a local measurement and a model-based expectation, irreconcilable, with a decade of argument about which side had the error. The resolution turned out to be neither — it was a missing component of the universe. That is the historical precedent people have in mind, and it is why the tension is taken seriously rather than filed under “somebody’s calibration”.

The number is an inverse time

It helps to remember what the quantity being disputed actually is, because its units are the reason a nine per cent disagreement is felt everywhere.

The Hubble constant has dimensions of one over time. Written in the units the field uses — kilometres per second per megaparsec — that is disguised, but converting gives a Hubble time of 1/H01/H_0: 14.5 billion years at 67.4, and 13.4 at 73. It is also, up to a factor, an inverse length: the Hubble distance c/H0c/H_0 is 4.45 gigaparsecs at the lower value and 4.11 at the higher.

Every distance inferred from a redshift is proportional to 1/H01/H_0. So are every luminosity computed from such a distance, every mass computed from that luminosity, and every volume in every number density. That is why catalogues quote quantities with an explicit factor of h=H0/100h = H_0/100 attached — a galaxy’s absolute magnitude comes with a +5logh+5\log h, a cluster’s mass with an h1h^{-1}, a density with an h2h^2 — so that the number can be rescaled when the constant moves.

The convention exists because the constant has moved before, by a factor of two, and a generation of published numbers would otherwise have had to be recomputed. It survives as a small piece of bookkeeping that records how uncertain this one number used to be.

What it does to the age

The consequence that presses hardest is the age of the universe, because the age is one of the very few things that can be checked against an object rather than against a model.

In a flat model with the matter density fixed by the microwave background, the age is close to 0.96/H00.96/H_0. At 67.4 that gives 13.8 billion years. At 73 it gives about 12.8 — a full billion years younger.

The check is the oldest stars. Globular clusters are dated from the turn-off of their main sequences, and the oldest come out at 12.5 to 13 billion years with an uncertainty of several hundred million; white-dwarf cooling in the Galactic disc and radioactive dating of the oldest halo stars give consistent figures. Those stars cannot have formed before the universe existed, and they cannot have formed immediately either — structure formation needs a few hundred million years.

At 13.8 billion years there is comfortable room. At 12.8 there is very little, and the margin becomes comparable to the uncertainty in the stellar ages themselves.

That does not refute the higher value: stellar ages are not precise enough to exclude it, and the age also depends on the matter density and the dark-energy equation of state, both of which a resolution might move. But it is the one place where the tension touches an observation that involves no cosmology at all, and it is the same argument that produced the age crisis of the 1990s — a local expansion rate implying a universe younger than the stars in it, resolved by a component nobody had expected.

The parallel is not exact, since the ages then were in conflict with the model rather than with each other, and the missing component was found by a different measurement entirely. What carries over is the shape of the eventual answer: the disagreement was real, both measurements were sound, and what was wrong was a term nobody had put in the model.

It is also the reason the age is quoted so much less often than the expansion rate. The two carry the same information in a flat model with a fixed matter density, and only one of them can be measured to a part in a hundred.

One further consequence of the units is worth stating because it disposes of a common misreading. A Hubble time is not the age of the universe; it is the age the universe would have if it had always expanded at the present rate. The true age depends on the whole history, and in the current model the deceleration of the matter-dominated era and the acceleration since happen to nearly cancel — which is why the factor relating the two is 0.96 rather than something far from one. That near-cancellation is a coincidence of the epoch it is being observed from, not a general feature, and it would not hold in a model with a different dark-energy equation of state.

The generalisation

The pattern here recurs whenever a quantity is reachable by a direct measurement and by an inference, and it has a name in this collection already: a number and a model are not the same kind of thing, and the error bar on the second one is the model.

A galaxy’s dynamical mass against its luminous mass is the same comparison, and it resolved in favour of the model being incomplete. A planet’s mass from a radial velocity is a lower bound until a transit removes the inclination, and there the direct measurement wins outright. The lithium abundance from big-bang nucleosynthesis disagrees with the halo stars by a factor of three and has done so for twenty-five years without resolving.

What distinguishes them is whether the disagreement has anywhere to hide. The Hubble tension has three places — the ladder, the late expansion, the early expansion — and the whole current effort is to close two of them so that the third has to answer.

Where the ladder goes next

The obvious next question is what the early route is actually measuring, since so much of this rests on the acoustic scale. That takes three essays: the microwave background’s spectrum, its peaks, and the surface it comes from.

Later rungs on this anchor: standard sirens and what a hundred of them would settle; time-delay cosmography and the mass-sheet degeneracy that limits it; the S8S_8 tension, which is a second and smaller disagreement of the same shape about the clumpiness rather than the expansion rate; whether H0H_0 can be measured without any distance at all, from the ages of the oldest objects; and the specific early-dark-energy models, and the higher-peak measurements that will decide between them.

What links here

The 8 of 16 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Distance ladderEarly dark energyHubble constantHubble tensionInverse distance ladderPeriod luminosity relationSound horizonSystematic errorTime-delay cosmographyTip of the red giant branch