Cosmology

Too few clusters, or a scale that reads light

A catalogue selected on the microwave shadow is, past redshift one half, very nearly a catalogue of everything above a fixed mass — so its count by redshift is the growth of structure read almost directly. Almost, because the mass behind the threshold comes from a calibration, and a scale that reads twenty per cent light is indistinguishable from a universe with less in it.

Assumes Sunyaev zeldovich, Large-scale structure and Clusters.

A cluster of galaxies is the largest object gravity has finished building, and how many of them exist at each epoch is a direct record of how fast building has gone. The count above a fixed mass depends on the amplitude of density fluctuations through an exponential, so a small change in the amplitude is a large change in the count — which is what makes clusters worth counting, and what makes every error in the count’s mass scale an error of the same exponential size.

The shadow a cluster casts on the microwave background does not fade with distance, and that property was presented as a way of finding clusters to high redshift. It does something more specific for a survey. Combined with how clusters of fixed mass change with cosmic time, it makes the detection threshold nearly a threshold in mass — so that a catalogue ordered by redshift is a catalogue of the same kind of object at every epoch, and its counts are the growth of structure with little else folded in.

That would be the cleanest measurement in cosmology if the mass at the threshold were known. It is inferred, through a calibration against masses that are themselves known to be biased by an amount that has resisted measurement for more than a decade. The counts are exact; the scale they are read against is not; and the arithmetic of which one moved is the subject here.

1020 clusters or 372, and the survey cannot say which parameter moved. The number of clusters per unit redshift a survey finds above an integrated Compton signal of 8·10⁻⁵ arcmin² over 6 per cent of the sky, computed from the Sheth–Tormen halo count grown by the linear growth factor, the comoving volume in each redshift slice, and the calibrated relation between signal and mass. Each curve is one pair of assumptions: the amplitude of structure, σ₈, and the hydrostatic mass bias, 1 − b, which says how far the X-ray masses the relation was calibrated on fall below the true masses. With σ₈ = 0.811 and 1 − b = 0.8 the survey finds 1020; σ₈ = 0.811 with 1 − b = 0.6 gives 372; σ₈ = 0.75 with 1 − b = 0.8 gives 548. The counts fall at low redshift because there is little volume, and at high redshift because massive halos have not yet formed, and the peak sits near z = 0.24. Lowering 1 − b pushes the threshold onto more massive and rarer halos; lowering σ₈ makes every halo rarer. The two lower curves differ in total by 47 per cent and in normalised shape by at most 2 per cent of the peak. A survey that assumed 1 − b = 0.8 would read the 372 clusters of the curve with 1 − b = 0.6 as σ₈ = 0.716, with a redshift distribution that differs from it by at most 12 per cent of the peak — which is the only handle the survey has on the difference, and it is smaller than the counting noise in any redshift bin holding fewer than about 67 clusters. A total count cannot distinguish a universe with less structure from a survey that has misjudged its masses, and it is that degeneracy, not the counting, that has been argued about since the first large catalogue.
Fig. 1 The number of clusters per unit redshift above an integrated Compton signal of 8·10⁻⁵ arcmin² over 6 per cent of the sky with a 1.2-arcminute beam — roughly the shape of a deep ground-based millimetre survey. With σ8\sigma_8 = 0.811 and a hydrostatic mass bias of 1 − b = 0.8 the survey finds 1020; with the mass bias at 0.6 instead it finds 372; with σ8\sigma_8 = 0.75 it finds 548. A survey assuming 1 − b = 0.8 would read the 372 as σ8\sigma_8 = 0.716, and the redshift distributions of those two universes differ by at most 12 per cent of the peak.

Why a count above a mass is a measurement of growth

The number density of dark-matter haloes of mass MM at redshift zz is set, to good accuracy, by one variable: the height of the collapse threshold in units of the density fluctuations on the scale of that mass,

ν=δcσ(M)D(z)\nu = \frac{\delta_c}{\sigma(M)\,D(z)}

where σ(M)\sigma(M) is the root-mean-square fluctuation smoothed on the mass scale today, normalised by the familiar σ8\sigma_8, and D(z)D(z) is the linear growth factor, which is one today and smaller in the past. For a massive cluster ν is two to four. The fraction of mass in haloes above a threshold of that height falls as eν2/2e^{-\nu^2/2} once ν is well above one, and that is where the leverage comes from. At ν = 3 a five per cent change in σ8\sigma_8 changes ν by five per cent, the exponent by ten per cent of 4.5, and the count by more than a third.

The same exponential makes the count at high redshift a measurement of D(z)D(z). The number of clusters above 101510^{15} solar masses at redshift one, compared with the number today, is set by how much the growth factor has changed in between — and the growth factor is set by the competition between gravity and whatever is driving the acceleration. That is why cluster counts became one of the independent lines of evidence, beside supernovae, that the growth of structure slowed at late times. A catalogue with good masses and good redshifts constrains the density of matter and the equation of state of dark energy through a quantity that no distance measurement touches directly.

The halo count itself, the part a survey is compared against, is not measured. It comes from the linear power spectrum and a multiplicity function calibrated on simulations, and it is accurate to perhaps ten or twenty per cent in normalisation at cluster masses. That uncertainty is real and it enters every figure below identically. What the figures isolate is the part that is not about simulations at all: what a survey’s threshold means as a mass.

A threshold that becomes a mass

A survey is not selected on mass. It is selected on an observable, and the observable at the threshold corresponds to a different mass at every redshift.

For a Sunyaev–Zel’dovich survey the observable is the integrated signal YY, and the calibrated relation between signal and mass has the form

E(z)2/3DA2Y500    [(1b)M500]1.79E(z)^{-2/3}\, D_A^2\, Y_{500} \;\propto\; \left[(1-b)\,M_{500}\right]^{1.79}

The angular-diameter distance is there because YY is an integral over solid angle: a cluster twice as far away covers a quarter of the solid angle at the same Compton parameter. The expansion rate E(z)E(z) is there because a cluster of given mass that formed earlier formed from a denser universe, is more compact and hotter, and has more thermal energy per unit mass. The factor 1b1-b is the mass bias, and it is deferred to the next section.

Beyond z = 0.5 a Compton threshold is a mass threshold. The smallest cluster mass M₅₀₀ a survey can detect, against redshift, for two Sunyaev–Zel'dovich surveys with integrated-signal thresholds of 0.001 arcmin² and 8·10⁻⁵ arcmin², and an X-ray survey with a flux limit of 3·10⁻¹² erg cm⁻² s⁻¹. The Compton limit uses the calibrated relation between integrated signal and mass with a hydrostatic mass bias of 1 − b = 0.80; the X-ray limit uses a representative luminosity–mass scaling and no K-correction. At low redshift both limits rise quickly, because a nearby cluster of any mass is bright. Past about z = 0.5 they part company. The integrated signal falls as the square of the angular-diameter distance, which stops growing and turns over near z = 1.6, and rises as the two-thirds power of the expansion rate, because a cluster of given mass at high redshift is denser and hotter; the two nearly cancel, and the limiting mass falls by 30 per cent between z = 0.5 and 1.8, for a 1.2-arcminute beam whose noise grows with a cluster's angular size once the cluster is resolved. The X-ray limit has the luminosity distance in the same place, and nothing that turns over, and climbs by a factor of 2.8 over the same range. A catalogue selected above a nearly constant mass is a catalogue in which the number of objects at each redshift is the growth of structure, read almost directly — provided the mass that sits behind the threshold is the mass the relation says it is.
Fig. 2 The smallest detectable mass against redshift for two Sunyaev–Zel’dovich surveys — a shallow one at 10⁻³ arcmin² and a deep one at 8·10⁻⁵ — and for an X-ray survey with a flux limit of 3×1012 erg cm2 s13\times10^{-12}\ \mathrm{erg\ cm^{-2}\ s^{-1}}. At low redshift every limit rises steeply, because a nearby cluster of any mass is bright. Beyond z = 0.5 the Compton limits stop climbing: the shallow survey’s limiting mass actually falls by 30 per cent between z = 0.5 and 1.8, while the X-ray limit, with a luminosity distance and nothing that turns over, climbs by a factor of 2.8. The beam matters at low redshift, where a resolved cluster spreads its signal over more noise.

Past redshift one half the two factors nearly cancel. The angular-diameter distance grows slowly, reaches a maximum near z = 1.6, and then shrinks, because the universe was smaller when the light left; the expansion rate keeps rising. Their combination in the relation above changes little, and a survey whose noise is set by a beam that does not resolve distant clusters sees a limiting mass that changes by a quarter or less from redshift one half to two — against a factor of nearly three for the X-ray survey over the same range.

An X-ray survey has the luminosity distance in the same place, and the luminosity distance grows without limit. Its threshold mass climbs steadily, so its catalogue at redshift one is a catalogue of a rarer class of object than at redshift one tenth, and the evolution of the count mixes growth with selection. Every X-ray cluster cosmology has had to model that mixing; a millimetre survey has much less of it.

Where the mass scale comes from

Nothing in a millimetre survey measures mass. The relation above has a normalisation, and the normalisation was set by comparing YY with masses measured some other way for a subsample of well-observed clusters — historically, X-ray masses from hydrostatic equilibrium.

A hydrostatic mass assumes that the gas pressure alone holds the gas up against gravity. It does not: clusters are still assembling, their gas carries turbulent and bulk motions, and those motions supply part of the support. A mass computed from thermal pressure alone is therefore too low, by the fraction of support that is non-thermal. That fraction is written bb, and the masses the relation was calibrated on are (1b)(1-b) times the true ones. Simulations of cluster formation put 1b1-b between about 0.7 and 0.9, with 0.8 the conventional choice, and the choice is not a measurement.

The consequence for a survey is exactly a shift of its threshold. If the calibration masses were twenty per cent low, then the survey’s threshold, quoted in calibration masses, corresponds to true masses twenty-five per cent higher than stated. Every halo in the prediction has to be heavier to be counted, and heavier haloes are exponentially rarer.

Beyond z = 0.5 a Compton threshold is a mass threshold. The smallest cluster mass M₅₀₀ a survey can detect, against redshift, for a Sunyaev–Zel'dovich survey with an integrated-signal threshold of 8·10⁻⁵ arcmin², and an X-ray survey with a flux limit of 3·10⁻¹³ erg cm⁻² s⁻¹. The Compton limit uses the calibrated relation between integrated signal and mass with a hydrostatic mass bias of 1 − b = 0.60; the X-ray limit uses a representative luminosity–mass scaling and no K-correction. At low redshift both limits rise quickly, because a nearby cluster of any mass is bright. Past about z = 0.5 they part company. The integrated signal falls as the square of the angular-diameter distance, which stops growing and turns over near z = 1.6, and rises as the two-thirds power of the expansion rate, because a cluster of given mass at high redshift is denser and hotter; the two nearly cancel, and the limiting mass falls by 23 per cent between z = 0.5 and 1.8, for a 1.2-arcminute beam whose noise grows with a cluster's angular size once the cluster is resolved. The X-ray limit has the luminosity distance in the same place, and nothing that turns over, and climbs by a factor of 2.8 over the same range. A catalogue selected above a nearly constant mass is a catalogue in which the number of objects at each redshift is the growth of structure, read almost directly — provided the mass that sits behind the threshold is the mass the relation says it is.
Fig. 3 The deep survey’s limiting mass drawn for a mass bias of 1 − b = 0.60 rather than 0.80, beside an X-ray survey ten times deeper than the one in the previous figure. The shape of the Compton limit is unchanged — it falls by 23 per cent between z = 0.5 and 1.8 — and the whole curve sits a third higher, because the same observed signal now corresponds to a larger true mass. Nothing in the survey data distinguishes this curve from the previous one. The X-ray limit still climbs by a factor of 2.8 whatever its depth, because depth moves a flux limit up or down and does not change its shape.

What forty per cent would mean for the gas

A mass bias is a statement about motion, and it can be turned into a speed. If a fraction of a cluster’s support comes from random bulk motions rather than thermal pressure, the ratio of the two pressures is roughly the ratio of the squared velocity dispersions. The thermal one is fixed by the temperature: for gas at 8 keV with a mean particle mass of 0.6 proton masses, the one-dimensional thermal speed is about 1,130 km/s.

A conventional bias of 1b=0.81-b = 0.8 means the non-thermal pressure is a quarter of the thermal, which needs turbulent or bulk motions of about half the thermal speed — some 570 km/s in each direction, throughout the region where the mass is measured. The bias that would reconcile the counts with the microwave background, 1b=0.61-b = 0.6, needs non-thermal pressure two-thirds of the thermal and motions of about 920 km/s in each direction — a three-dimensional speed of 1,600 km/s, above the gas’s own sound speed of about 1,460. A cluster supported that way would be in a state of supersonic turbulence over hundreds of kiloparsecs, which would heat itself and would not look relaxed in any X-ray image.

The only direct measurement of such motions was made in 2016, when an X-ray microcalorimeter resolved the widths of iron lines in the core of the Perseus cluster before its spacecraft was lost. The line-of-sight velocity dispersion was about 160 km/s — a non-thermal pressure of a few per cent of the thermal, in a core. That does not settle the question, because the bias that matters is at large radii, where the gas is still falling in and where simulations put most of the non-thermal support. But it is the right kind of measurement: a Doppler width, which is a mass bias in kilometres per second, and its successor instruments are now measuring the same widths in more clusters and further out.

One lever with two names

Put the two together and the degeneracy is immediate. Lowering σ8\sigma_8 makes every halo rarer. Lowering 1b1-b moves the threshold onto heavier haloes, which are rarer. Both reduce the count, both reduce it more at high mass than at low, and both reduce it more at high redshift than at low, because in each case the effect acts through ν and the exponential is steeper where ν is larger.

That is why the hero figure’s two lower curves are so hard to tell apart. A deep survey that finds 372 clusters where σ8\sigma_8 = 0.811 and 1b=0.81-b = 0.8 predicted 1,020 can conclude that σ8\sigma_8 is 0.716 with the calibration taken at face value, or that σ8\sigma_8 is 0.811 and the calibration masses were forty per cent low. The totals are identical by construction. The redshift distributions differ by at most twelve per cent of the peak — and a redshift bin would need several dozen clusters in it before a difference of that size rose above its own counting noise, which a catalogue of 372 spread over redshift has in only a few bins.

1020 clusters or 646, and the survey cannot say which parameter moved. The number of clusters per unit redshift a survey finds above an integrated Compton signal of 8·10⁻⁵ arcmin² over 6 per cent of the sky, computed from the Sheth–Tormen halo count grown by the linear growth factor, the comoving volume in each redshift slice, and the calibrated relation between signal and mass. Each curve is one pair of assumptions: the amplitude of structure, σ₈, and the hydrostatic mass bias, 1 − b, which says how far the X-ray masses the relation was calibrated on fall below the true masses. With σ₈ = 0.811 and 1 − b = 0.8 the survey finds 1020; σ₈ = 0.811 with 1 − b = 0.7 gives 646; σ₈ = 0.78 with 1 − b = 0.8 gives 753. The counts fall at low redshift because there is little volume, and at high redshift because massive halos have not yet formed, and the peak sits near z = 0.24. Lowering 1 − b pushes the threshold onto more massive and rarer halos; lowering σ₈ makes every halo rarer. The two lower curves differ in total by 17 per cent and in normalised shape by at most 1 per cent of the peak. A survey that assumed 1 − b = 0.8 would read the 646 clusters of the curve with 1 − b = 0.7 as σ₈ = 0.765, with a redshift distribution that differs from it by at most 6 per cent of the peak — which is the only handle the survey has on the difference, and it is smaller than the counting noise in any redshift bin holding fewer than about 319 clusters. A total count cannot distinguish a universe with less structure from a survey that has misjudged its masses, and it is that degeneracy, not the counting, that has been argued about since the first large catalogue.
Fig. 4 The same survey with more modest departures: a mass bias of 0.7 instead of 0.8, and a σ8\sigma_8 of 0.78 instead of 0.811. The first reduces the count from 1020 to 646, the second to 753. A survey that assumed 1 − b = 0.8 would read 646 clusters as σ8\sigma_8 = 0.765, with a redshift distribution that differs by at most 6 per cent of the peak — a difference that needs more than 300 clusters in a single bin to see. A twelve per cent error in the mass scale is a six per cent error in σ8\sigma_8, which is larger than the whole of the disagreement this measurement became famous for.

The conversion rate is worth carrying. On these figures a mass scale wrong by a fraction ε moves the inferred σ8\sigma_8 by roughly half of ε. The quantity is the whole reason cluster cosmology is a calibration problem: a statistically perfect count, with an error of one per cent in σ8\sigma_8 from Poisson noise, is limited to an accuracy of five per cent if its masses are known to ten.

The disagreement that followed

The first large cosmological sample of Sunyaev–Zel’dovich clusters came from the Planck satellite in 2013: a couple of hundred objects, selected across most of the sky, with a relation between signal and mass calibrated on hydrostatic X-ray masses and a mass bias of 1b=0.81-b = 0.8 adopted from simulations. The count implied σ8\sigma_8 ≈ 0.77. The same satellite’s map of the primary microwave background, extrapolated forward with the standard model, implied σ8\sigma_8 ≈ 0.83.

The difference between those two numbers is small in σ8\sigma_8 and large in clusters: the primary background predicted roughly twice as many as were found. Either the growth of structure has been slower than the standard model predicts — which would mean new physics, perhaps massive neutrinos removing structure by streaming out of it — or the mass bias is larger than assumed. Asking what mass bias would reconcile them gave 1b0.61-b \approx 0.6: hydrostatic masses forty per cent low, which is outside what simulations suggested and inside what nobody could exclude.

The resolution had to be a mass measurement that assumes nothing about the gas. Weak lensing is that measurement — the shear of background galaxies responds to total mass, hydrostatic or not — and several programmes measured the bias directly by comparing lensing masses with hydrostatic ones for the same clusters. They disagreed with one another. One found 1b1-b around 0.7, another near 0.78, a third close to 0.95, with quoted errors that did not all overlap. Each lensing analysis has its own calibration of galaxy shapes, its own treatment of the redshifts of the background galaxies, and its own sample of clusters, and each of those is a mass scale in its own right.

The situation since has been gradual rather than decisive. Counts calibrated with lensing masses from ground-based surveys have moved closer to the primary background’s value; the degeneracy has been narrowed, not broken; and a similar and still partly open difference in the amplitude of structure measured by lensing surveys that do not use clusters at all has made it harder to decide whether the clusters were ever pointing at anything physical.

Wide and shallow, or deep and narrow

The two survey designs sample the degeneracy differently, and the difference is visible in where their clusters sit.

876 clusters or 335, and the survey cannot say which parameter moved. The number of clusters per unit redshift a survey finds above an integrated Compton signal of 0.001 arcmin² over 65 per cent of the sky, computed from the Sheth–Tormen halo count grown by the linear growth factor, the comoving volume in each redshift slice, and the calibrated relation between signal and mass. Each curve is one pair of assumptions: the amplitude of structure, σ₈, and the hydrostatic mass bias, 1 − b, which says how far the X-ray masses the relation was calibrated on fall below the true masses. With σ₈ = 0.811 and 1 − b = 0.8 the survey finds 876; σ₈ = 0.811 with 1 − b = 0.6 gives 335; σ₈ = 0.75 with 1 − b = 0.8 gives 477. The counts fall at low redshift because there is little volume, and at high redshift because massive halos have not yet formed, and the peak sits near z = 0.10. Lowering 1 − b pushes the threshold onto more massive and rarer halos; lowering σ₈ makes every halo rarer. The two lower curves differ in total by 42 per cent and in normalised shape by at most 1 per cent of the peak. A survey that assumed 1 − b = 0.8 would read the 335 clusters of the curve with 1 − b = 0.6 as σ₈ = 0.718, with a redshift distribution that differs from it by at most 13 per cent of the peak — which is the only handle the survey has on the difference, and it is smaller than the counting noise in any redshift bin holding fewer than about 58 clusters. A total count cannot distinguish a universe with less structure from a survey that has misjudged its masses, and it is that degeneracy, not the counting, that has been argued about since the first large catalogue.
Fig. 5 A shallow survey of 65 per cent of the sky with a 7-arcminute beam and a threshold of 10⁻³ arcmin², roughly the shape of an all-sky satellite catalogue. With σ8\sigma_8 = 0.811 and 1 − b = 0.8 it finds 876 clusters, peaking near z = 0.10; with the mass bias at 0.6 it finds 335, and with σ8\sigma_8 = 0.75 it finds 477. A survey assuming 1 − b = 0.8 reads the 335 as σ8\sigma_8 = 0.718. The catalogue is dominated by massive, nearby clusters, for which the halo count is least certain and the leverage on growth is smallest.

An all-sky survey at low resolution finds the most massive clusters in the local universe and few beyond redshift one half. Its counts are dominated by objects at very high ν, where the exponential is steepest and the sensitivity to the mass scale is largest, and it has almost no redshift range to exploit the small difference in shape. A deep survey over a few per cent of the sky reaches lower masses at every redshift out to two, where the growth factor has changed most, and it has the redshift lever. Neither breaks the degeneracy alone; together, with different mass ranges and different redshift ranges, they constrain it from two sides. The same logic runs through comparing a halo count with a galaxy count: two samples that respond differently to the same unknown are worth more than twice one sample.

The other lever is internal. Clusters cluster, and how strongly they do depends on their mass through the same ν that sets their abundance; a catalogue that measures its own spatial correlation function is weighing its members by a second route. For the catalogue sizes drawn here that route is weak, but it improves faster than the count itself as catalogues grow, because the number of pairs grows as the square of the number of clusters.

What the figures leave out

The threshold here is sharp, and no survey’s is. A cluster of given mass has a signal scattered by about ten to twenty per cent around the mean relation, and since there are many more clusters below any threshold than above it, scatter moves more clusters up across the threshold than down. The effect is the Eddington bias that inflates every steep count, and at a threshold where the mass function is steep it adds tens of per cent to the count. Modern analyses model the scatter jointly with everything else; these figures do not.

The halo masses in the count are total masses at the mean density, and the relation uses M500M_{500}, the mass inside the radius at which the enclosed density is five hundred times critical. The figures convert with a single factor of 0.65, which is right to about fifteen per cent for typical cluster concentrations and wrong in detail, since concentration depends on mass and redshift. Because the factor multiplies every curve identically it moves the totals and not the comparisons, which is the only reason it is tolerable.

The beam model is a caricature. A matched filter’s noise depends on the cluster’s profile, on the noise spectrum of the map, and on the primary background it is being extracted from, and real selection functions are computed by injecting simulated clusters into real maps. And the relation’s normalisation carries the relativistic correction, which lowers the Compton signal of the hottest clusters by several per cent if ignored — a small mass bias of its own, concentrated exactly where the count is most sensitive.

A count above a steep threshold measures the threshold

A count against a threshold measures the threshold as much as it measures the population, and when the population falls exponentially above the threshold, it measures the threshold almost entirely. The same structure appears in every census on a steep distribution: a luminosity function’s faint end is a statement about completeness, a planet occurrence rate is a statement about a detection efficiency, and a cluster count is a statement about a mass scale.

What makes the cluster case unusual is that the threshold’s calibration is a physical assumption rather than an instrumental one. An efficiency can be measured by injection. A hydrostatic bias is the degree to which assembling structures are out of equilibrium, which is itself something structure formation predicts — so the count and its calibration are two answers to one question, and a survey that fits both has to be careful not to find that each has explained the other.

Still open: whether the scale can be set without any gas at all

The route out is a mass measured by something that responds to total mass and has no scaling relation of its own. Lensing of background galaxies is one, and its calibration is still disputed at the level that matters. Lensing of the microwave background itself by each cluster is another: the background behind every cluster is a source at a precisely known distance with a precisely known statistical structure, and its deflection by the cluster’s mass needs no galaxy shapes and no redshifts of anything behind. Stacked over hundreds of clusters it has already given average masses to within ten or twenty per cent, and the question it can answer — whether the scale reads light — is the one the counts have been waiting on since the first catalogue.

About the same objects

Not linked from either essay — found by the objects both name.

The objects this essay names

Each one links to every other essay that touches it.

Angular-diameter distanceCluster mass functionDegeneracyHalo mass functionHydrostatic biasHydrostatic massIntegrated ySelection functionSigma 8Structure growthThe Sunyaev–Zel'dovich effectWeak lensing