A length in centimetres, measured against an angle
Assumes Sunyaev zeldovich, Clusters and Distance ladder.
The first rung of this anchor was about a survey property: the decrement a cluster imprints on the microwave background is a fraction of a background that does not dim, so a cluster at redshift one is as easy to find as one at redshift a tenth. That makes the effect a finding tool. It is also, and quite separately, a ruler — and the ruler does not work by being found more easily. It works because the same cluster offers a second measurement of the same electrons with a different power of the density in it.
Everything about a cluster of galaxies that can be observed is an integral down a line of sight. Nobody has ever measured the depth of one. The two integrals available happen to weight the gas differently, and two differently weighted integrals of one unknown quantity along one unknown path are two equations in two unknowns.
Two integrals, two powers, one gas
The Compton parameter is the optical depth to electron scattering multiplied by the fractional energy each scattering hands over, integrated along the line of sight:
It is linear in the electron density because a scattering needs one electron.
The X-ray brightness is thermal bremsstrahlung, which needs an electron and an ion at the same place at the same time, so its emissivity is quadratic:
The factor of is the ordinary cosmological dimming of a surface brightness, and it is the reason the X-ray half of this pairing runs out of reach long before the Compton half does. Both integrals need the gas temperature, which is measured from the shape of the X-ray spectrum and enters the two expressions differently — linearly in one, as a square root in the other.
The temperature is where the two halves are joined, and it is worth dwelling on because it is the one input neither integral can do without. The Compton parameter is an integral of pressure, so it carries to the first power; the bremsstrahlung emissivity carries , because the emitted power per collision rises with the electron speed while the time each electron spends near an ion falls. Neither is a strong dependence, which is fortunate — a ten per cent error in the temperature moves the recovered distance by about fifteen per cent rather than by a factor. But the temperature is measured from the shape of the X-ray continuum, and for a 10 keV cluster the exponential cut-off sits at photon energies where the effective area of every X-ray telescope ever flown is falling steeply. Cluster temperatures measured by two satellites have historically disagreed by ten to twenty per cent, and that disagreement propagates here undiluted.
The gas is modelled as an isothermal β model, , which is not a physical derivation of anything — it is a two-parameter fitting form that has described the outer parts of cluster X-ray images since the 1970s and is used because it integrates in closed form. Everything below inherits whatever that form gets wrong. Real clusters are not isothermal either: the temperature typically peaks around a tenth of the virial radius and falls by a factor of two by the edge, and a cool core drops it sharply in the middle. Fitting a single temperature to such a cluster returns something between the two, weighted towards wherever the emission is brightest — which is the core, because the emission is quadratic. The pressure integral, which is not, is then being evaluated at a temperature characteristic of the wrong part of the cluster.
Solving for the length
Along the central line of sight the two integrals are
where is a dimensionless number that depends on the shape parameter and on nothing else. The two unknowns are the central density , in electrons per cubic centimetre, and the path scale , in centimetres. The first integral gives the product ; the second gives . Their ratio is , and dividing the first by that gives .
That is the whole of it, and the striking thing is what has just happened dimensionally. Two measurements, neither of which contains a length, have produced a length in centimetres. It works because the second one contains a square: a quantity that appears once and a quantity that appears twice cannot both be absorbed into one unknown, so the pair carries one more piece of information than either alone.
The angular core radius comes from the same X-ray image that gave the surface brightness. A physical length divided by the angle it subtends is an angular-diameter distance, and that is the output.
Collecting the dependences shows how the errors will propagate. The distance goes as , times a temperature factor and the two shape constants, so the Compton parameter enters squared and everything else linearly. A five per cent error in the decrement is a ten per cent error in the distance. That squaring is the reason the method waited two decades between being proposed and being useful: the idea is from 1978, and a decrement of a hundred microkelvin against a two-point-seven kelvin sky, measured well enough to square, needed interferometers that did not exist.
The round trip is the test
Nothing in the last two figures is an illustration of an argument made elsewhere. The cluster is specified in physical units — a central density, a core radius in megaparsecs, a temperature, a redshift — its two observables are computed from that state, and the observables alone are handed to an inversion that is told nothing else. What comes back has to be the cluster.
A round trip that closes proves less than it seems and more than nothing. It does not show that the method works on the sky; it shows that the arithmetic is the arithmetic claimed, that no exponent has been mistyped, and that the two shape constants have not been swapped — which is exactly the class of error that a plausible-looking figure hides. What it cannot test is any assumption the forward and the backward calculation share, and both of them assume the gas is smooth and the cluster is round.
There is a whole second category the round trip is blind to, which is contamination of either observable by something that is not the cluster. A radio-loud galaxy in the cluster — and clusters are where radio-loud galaxies are — emits at exactly the frequencies where the decrement is deepest, and fills it in. An under-measured decrement is squared before it reaches the distance, so a source contributing ten per cent of the signal takes nineteen per cent off the answer, in the direction of a larger Hubble constant. The remedy is to observe at several frequencies and subtract a fitted power law, and the residual after that subtraction is one of the larger error terms in every published determination.
What the method assumes
Those two assumptions are the entire difficulty, and they are not small.
Real intracluster gas is not smooth. It has cool cores, infalling subclusters, and structure below whatever resolution the X-ray telescope has. Because the emission goes as , a given mean density with structure in it emits more than the same mean density without, by the clumping factor . The inversion takes that extra brightness at face value and concludes the gas is denser than it is, so the path length comes out short and the distance comes out small.
Real clusters are also not round. A cluster elongated along the line of sight presents more path than its angular width suggests, so the inferred distance is too large; one flattened along the line of sight does the reverse.
The bias in the distance is exactly the elongation divided by the clumping, which is worth stating because it is the rare case where a systematic has a closed form. Both enter the two integrals in the same simple way — clumping multiplies the quadratic one, elongation multiplies both — and the algebra carries them straight through.
Both defects are invisible in the data being fitted. A clumped cluster and a smooth denser one produce identical X-ray images; an elongated cluster and a spherical more distant one produce identical maps in both wavebands. Nothing in the observation distinguishes them, which is what separates this from an ordinary systematic that better data would reduce.
Orientation and clumping behave completely differently in a sample, and the difference is the reason the method has the reputation it has. Elongation has no preferred sign: over many clusters observed in random orientations it averages away, leaving a scatter of ten to twenty per cent per object that shrinks as the sample grows. Clumping has a sign. Gas is lumpy or it is smooth, never anti-lumpy, so always, and averaging a hundred clusters produces a hundred-times-more-precise estimate of a biased number.
Why it reaches where a ladder does not
Against those difficulties sits a structural advantage that no rung-based method has. This measurement is absolute. There is no calibration step, no anchor galaxy, no zero point handed down from a parallax; the distance is a length in centimetres derived from atomic physics and divided by an angle.
There is a second thing an absolute distance can do that a relative one cannot, and it is a consistency test rather than a measurement. Luminosity distance and angular-diameter distance are related by in any metric theory in which photons travel on null geodesics and are conserved in number, whatever the expansion history is. Supernovae give the first and this method gives the second, for objects at the same redshifts, so the pair tests an identity that is prior to cosmology. A failure would mean photons going missing along the way — absorbed by dust, or converted into something. The test currently confirms the identity to about ten per cent, which is not a strong constraint and is a constraint on something nothing else constrains at all.
Reaching redshift one directly matters more than it sounds, because the angular-diameter distance is not monotonic. It rises, turns over near redshift 1.6 in the standard model, and falls; a measurement of a length at redshift 0.9 therefore constrains the expansion history in a way no local calibration can. The comparison is with a distance measured with a stopwatch, where a time delay between lensed images gives an absolute length in the same spirit, and with a distance with no ladder under it, where a gravitational-wave chirp carries its own amplitude calibration. Three methods, three unrelated pieces of physics, and one property in common: none of them is a rung.
What was actually measured
The clusters in the figures above are constructions. Setting them beside the observational history is the more sobering exercise.
The measurement is: a radio or millimetre map of the decrement, at a resolution good enough to fit a profile; an X-ray image, giving the surface brightness and the angular core radius; and an X-ray spectrum, giving the temperature. Then a joint fit of a β model to both, with the temperature profile assumed flat, the geometry assumed spherical and the gas assumed smooth.
The chronology is a useful corrective to how simple the algebra looks. The distance argument was published in 1978, within a year of the effect’s own confirmation. The first credible detection of a decrement took most of the following decade and consisted of single-dish observations that were disputed for years afterwards, because a hundred-microkelvin dip in a two-point-seven-kelvin sky is indistinguishable from a fault in a receiver unless something else moves. What settled it was interferometry: an interferometer resolves out the smooth sky and responds only to structure of the angular size the cluster has, which turns an absolute-calibration problem into a differential one. The first sample of distances worth quoting appeared in the late 1990s, twenty years after the method.
The advice the whole programme converged on was to use relaxed clusters and to avoid their outskirts, and both halves follow from the two biases. A relaxed cluster — no double X-ray peak, no offset between the gas and the galaxies, a smooth image — is one that has not recently merged, and a merger is what produces both large clumping factors and shapes far from spherical. The outskirts are where simulations put clumping factors of 1.3 and above, because that is where infalling gas has not yet been mixed in; restricting the fit to the inner regions is a way of trading signal for a smaller bias. Neither step measures either defect. Both reduce a number nobody can observe by an amount nobody can quote.
Published results from that programme span a range wider than any of their quoted uncertainties. A sample of eighteen clusters in 2002 returned 60 km s⁻¹ Mpc⁻¹ with a stated error of four; a sample of thirty-eight in 2006 returned 76.9 with a stated error of about four. Those two are more than three of their own error bars apart and they overlap in objects. The difference is not in the data but in the treatment — which clusters were judged relaxed, whether a cool core was excised, what temperature profile was assumed, and how much clumping was corrected for.
Two further items are always in those error budgets and never dominant. The relativistic correction to the spectral shape matters because a 10 keV electron is moving at a fifth of the speed of light: the distortion is no longer the simple non-relativistic form, the 217 GHz null moves upward by a few gigahertz, and the decrement at the frequencies actually observed is a few per cent shallower than the textbook expression. And the photons doing the scattering are the best-characterised light in astronomy, which is one input in this whole chain that contributes no error at all.
That spread is the honest state of the method as a competitor in the argument about the Hubble constant. It brackets both contenders. Its value now is not as an arbiter of that disagreement but as a check on cluster physics: with the constant taken from elsewhere, the same equations run backwards give the clumping factor, and that is a measurement of gas structure that nothing else provides. The same inversion has switched from being a distance to being a diagnostic, which is a common fate for a method whose systematics grow slower than its rivals’ precision.
The generalisation
The structure worth extracting is that two measurements of one quantity at different powers are not two measurements. They are a measurement and a scale.
If an observable goes as and another as along the same path, the pair separates the density from the path length, because no single rescaling of an unknown can satisfy both. The same trick recurs whenever a system offers a linear and a nonlinear probe of itself. A star’s radius follows from combining a flux with an angular diameter; a cluster weighed three ways is over-determined for the same reason, with the lensing mass linear in the potential and the X-ray mass depending on its gradient.
The corollary is the warning that this essay’s second half is entirely about. The separation is only as good as the assumption that the two observables sample the same thing. The instant the quadratic probe sees structure the linear one averages over, the extra information the square provided becomes an extra error, with a sign. That is not a defect peculiar to clusters — it is what always happens when a method’s power comes from a nonlinearity, because a nonlinear average is not the average of the nonlinear.
Where the ladder goes next
The next rung stays with the same gas and changes which motion is measured. Scattering off electrons that are moving bodily along the line of sight shifts the background’s spectrum in a different way from scattering off hot random motion, with a different frequency dependence — so the same map, observed at the right frequencies, separates a cluster’s peculiar velocity from its temperature. That gives a velocity for an object whose redshift says nothing about its motion, and it is the only way anything has ever been measured moving with respect to the microwave background other than the local group itself.
Further rungs on this anchor: the relativistic corrections to the spectral shape, which matter at the per-cent level for the hottest clusters and shift the 217 GHz null; the integrated signal as a mass proxy and the scatter in that relation; the effect’s use to count clusters as a function of redshift, which is a measurement of structure growth rather than of distance; and the kinematic effect’s statistical detection in pairs of clusters, which works where the individual measurement does not.
About the same objects
Not linked from either essay — found by the objects both name.
- The baryons that are not in the galaxies beta model · bremsstrahlung · intracluster medium · the sunyaev–zel'dovich effect
- A velocity that has the colour of the sky intracluster medium · the sunyaev–zel'dovich effect · thomson scattering
- Too few clusters, or a scale that reads light angular-diameter distance · the sunyaev–zel'dovich effect
What links here
Essays that link to this one from their own argument.
- A null that moves with the temperature cosmology
The objects this essay names
Each one links to every other essay that touches it.
Angular-diameter distanceBeta modelBremsstrahlungCompton parameterElectron densityGas clumpingHubble constantIntracluster mediumLine of sight integralThe Sunyaev–Zel'dovich effectThomson scatteringX-ray surface brightness