Galaxies

Two parameters that lensing measures as one

A weak-lensing survey measures how much the shapes of distant galaxies are distorted by the matter in front of them, and that distortion depends on how much matter there is and on how clumpily it is arranged. The two enter as a product. What the survey determines is one number, and the disagreement between surveys and the microwave background is stated in that number because neither measures either quantity alone.

Assumes Weak lensing and Large-scale structure.

Light from a distant galaxy passes through everything between it and the observer, and the matter along the way deflects it. The deflection is not uniform, so the galaxy’s image is sheared — stretched slightly in one direction — by about one per cent.

That distortion is far smaller than a galaxy’s own ellipticity, so it is measured statistically: average the shapes of a million galaxies, and their intrinsic ellipticities cancel while the coherent shear does not. The result is a measurement of how much structure the light passed through.

Two bands 2.1 standard deviations apart, and neither is a point. The plane of the matter density against the amplitude of matter fluctuations, with the constraints from a weak-lensing survey and from the microwave background drawn as bands. Lensing measures the shear produced by structure along the line of sight, and that shear depends on how much matter there is and on how clumpy it is in a fixed combination — more matter arranged less clumpily gives the same signal. The locus is a power law of exponent one half, and the combination it fixes is written S₈. The two bands are separated by 2.1 standard deviations, and whatever that separation is, it is a statement about the growth of structure between recombination and now rather than about either measurement's precision. Neither band alone determines either quantity, which is why the disagreement is quoted in the combination rather than in the parameters.
Fig. 1 What that measurement constrains. The shear signal depends on how much matter there is and on how clumpy it is, and the two enter in a fixed combination — more matter arranged less clumpily gives the same distortion. The locus is a power law and the combination it pins down is conventionally written S₈. The band from the microwave background sits above the band from lensing, by two to three standard deviations.

A one-per-cent distortion needs a million galaxies to see, and this essay is about what the million galaxies determine once they have been averaged — which is one number rather than the two the number is usually reported as.

Where the combination comes from

The shear at a given angular scale is an integral along the line of sight of the matter power spectrum, weighted by a geometric factor.

Two things control its amplitude. The matter density Ωm\Omega_{\rm m} sets how much matter there is to do the deflecting and also how the geometry works out. And σ8\sigma_8 — the root-mean-square fluctuation of the matter density in spheres of eight megaparsecs — sets how clumpy that matter is.

Doubling the density with half the clumpiness leaves the shear almost unchanged over the range of scales a survey measures well. The exact trade depends on the survey’s depth and angular range, and it is close to σ8Ωm\sigma_8\sqrt{\Omega_{\rm m}}, which is why that combination is given a name and quoted as the result.

That is not a defect of the analysis. It is what the observable is: a single number characterising an amplitude, measured across a range of scales too narrow to separate two parameters that both control it.

It is worth being precise about why the two parameters are not separable from the shear alone. The lensing signal is an integral of the matter power spectrum over the scales the survey resolves, weighted by geometry. Changing Ωm\Omega_{\rm m} changes the geometry, the amount of matter, and the shape of the power spectrum — and the shape change is the part that could break the degeneracy, because it alters the signal’s dependence on angular scale. A survey covering a decade in angle sees only a little of that shape change, so it recovers one amplitude and almost nothing about the shape. A survey covering three decades would do better, and no survey does, because the small scales are contaminated by baryonic physics and the large ones by the survey’s own area.

One more property of the combination is worth noticing because it explains the exponent. The shear signal scales roughly as the matter density times the fluctuation amplitude at the scales the survey reaches, and the geometry contributes a further weak dependence on the density. Working the two through gives an exponent near one half — and the exponent is not universal: it depends on the survey’s depth and on the range of angular scales fitted, so different surveys quote S8S_8 with slightly different definitions. Comparing them requires either using the same definition or projecting each posterior onto a common one, and both are done. Every distance is measured with the last one has an analogue here: every amplitude is quoted in somebody else’s combination.

Why the disagreement is stated in the combination

The microwave background measures the same two quantities in a completely different way: it observes the fluctuations at recombination and computes forward, using the standard model of how structure grows under gravity, to predict what σ8\sigma_8 should be now.

Comparing the two is therefore a test of the growth of structure over thirteen billion years. And they disagree, by two to three standard deviations, with the lensing surveys finding less structure than the microwave background predicts.

The disagreement is stated in S8S_8 rather than in σ8\sigma_8 for the reason above: neither measurement determines σ8\sigma_8 on its own, and quoting the comparison in a quantity that neither measures would be misleading about where the tension lies.

Two bands 3.6 standard deviations apart, and neither is a point. The plane of the matter density against the amplitude of matter fluctuations, with the constraints from a weak-lensing survey and from the microwave background drawn as bands. Lensing measures the shear produced by structure along the line of sight, and that shear depends on how much matter there is and on how clumpy it is in a fixed combination — more matter arranged less clumpily gives the same signal. The locus is a power law of exponent one half, and the combination it fixes is written S₈. The two bands are separated by 3.6 standard deviations, and whatever that separation is, it is a statement about the growth of structure between recombination and now rather than about either measurement's precision. Neither band alone determines either quantity, which is why the disagreement is quoted in the combination rather than in the parameters.
Fig. 2 The same comparison with the tighter error bars a modern survey achieves. The separation grows to about three standard deviations, which is the regime in which a result is interesting and not conclusive — precisely where a systematic in either measurement would be doing its most damage. The history of the field is that such tensions have resolved as systematics about as often as they have persisted, and the effort has gone into the systematics accordingly.

There is a further reason the comparison is worth taking seriously, and it is about what the two measurements have in common: almost nothing. One observes the shapes of galaxies at redshifts below one, in the optical, with an eight-metre telescope; the other observes temperature fluctuations at redshift eleven hundred, in the millimetre, from space. They share no instrument, no calibration and no astrophysics. The one thing they do share is the model used to connect them, and if the model is right they must agree. That is the ideal configuration for a test and it is the reason a two-sigma disagreement between them is discussed more than a five-sigma disagreement between two similar experiments would be.

Be explicit about what “structure growth” means as a measurement, because that phrasing does a lot of work. The microwave background fixes the amplitude of the fluctuations at recombination directly. Between then and now they grow under gravity at a rate set by the expansion history and by whether gravity behaves as general relativity says. Lensing measures the amplitude now. So the comparison is a measurement of the integrated growth over the interval, and a disagreement points at something in the interval — a different expansion history, a massive neutrino suppressing small-scale growth, or a modification of gravity. It does not point at either endpoint, which is why the discussion is about the model rather than about the instruments.

What tomography adds

The degeneracy is broken, partly, by measuring the shear as a function of the source galaxies’ redshift.

Galaxies at different distances have their light passing through different amounts of structure and different geometry, so the shear signal’s dependence on source redshift carries information about how the two parameters combine. Splitting a survey into redshift bins and measuring the shear correlation between and within them — tomography — turns one number into a matrix of numbers, and the matrix constrains the parameters separately.

It does so weakly. The degeneracy is narrowed rather than removed, because the redshift dependence of the two parameters’ effects is similar over the range a survey reaches. What tomography does much better is constrain the growth rate, which is the quantity the tension is really about.

A mass profile from shapes, with no dynamics and no light in it. Mean tangential ellipticity in 8 logarithmic annuli, from a simulated catalogue of 21206 background galaxies at 30 per square arcminute behind a singular isothermal sphere of Einstein radius 14″. Each point is an average over 125 to 7939 galaxies and its error bar is 0.3/√N and nothing else — no galaxy in the sample was measurably distorted. The curve is the reduced shear γ/(1 − κ) the lens was built with, not a fit: the points sit χ²/N = 0.76 from it. The measurement is a projected mass profile obtained with no assumption whatever about the lens's dynamical state, which is the one thing neither a velocity dispersion nor an X-ray temperature can offer, and it is why a merging cluster can be weighed at all. The open points are the cross component, the same shapes rotated by 45°: χ²/N = 1.36 from zero, and if they were not, the mass map would be a picture of the telescope.
Fig. 3 The measurement itself, at the level a single lens produces: the shear profile around a mass concentration, falling with radius. Everything in this essay is an average of profiles like this one over millions of foreground structures and millions of background galaxies, and the amplitude of that average is the number in dispute. The individual profile is dominated by noise from the sources’ own shapes, which is why the technique needs a million galaxies rather than a thousand.
Two bands 0.9 standard deviations apart, and neither is a point. The plane of the matter density against the amplitude of matter fluctuations, with the constraints from a weak-lensing survey and from the microwave background drawn as bands. Lensing measures the shear produced by structure along the line of sight, and that shear depends on how much matter there is and on how clumpy it is in a fixed combination — more matter arranged less clumpily gives the same signal. The locus is a power law of exponent one half, and the combination it fixes is written S₈. The two bands are separated by 0.9 standard deviations, and whatever that separation is, it is a statement about the growth of structure between recombination and now rather than about either measurement's precision. Neither band alone determines either quantity, which is why the disagreement is quoted in the combination rather than in the parameters.
Fig. 4 The same comparison as it stood a decade ago, with a lensing error bar twice as large. The two bands overlap comfortably and there is no tension at all. Nothing about the physics changed between then and now; the survey got bigger. That progression — a disagreement that appears as errors shrink — is the ordinary way a systematic reveals itself and also the ordinary way a real effect does, and distinguishing the two is what the next generation of surveys is for.

There is a further consequence of measuring an amplitude rather than a shape that decides how the tension can be resolved. Because the observable is a single amplitude, anything that lowers it lowers it uniformly — a calibration bias in the shapes, a systematic in the source redshifts, or a real suppression of small-scale power. The three are degenerate in the shear signal alone and are separated only by their different behaviour across angular scale and across redshift bin, which is exactly the information tomography adds. So tomography’s real value is not in breaking the parameter degeneracy but in distinguishing a physical suppression from an instrumental one, which is a different and more useful job.

What was actually measured

There are three, and the third is why the tension is taken seriously.

Consistent results from independent surveys. Three large lensing surveys, with different telescopes, different filters, different shape-measurement algorithms and different redshift calibrations, report S8S_8 values agreeing with each other and sitting below the microwave background’s prediction. Agreement between independent instruments is the main argument that the low value is not one survey’s systematic.

Consistent results from other probes of the same epoch. Galaxy clustering combined with lensing, cluster abundances, and redshift-space distortions all measure structure growth at low redshift, and most of them sit on the low side. That is a weaker argument than it looks — several share calibration ingredients — and it is the reason the tension is quoted at two to three sigma rather than higher.

And the systematics have been hunted hard. The three that could produce a spurious low value are all under active control: shape measurement biases, calibrated by injecting simulated galaxies into the real images; photometric redshifts, calibrated against spectroscopic subsamples and cross-correlations; and intrinsic alignments, the tendency of physically nearby galaxies to be aligned by the same tidal field they are being lensed by, which is modelled and marginalised over. Each is capable of the required shift and none has been shown to produce it.

A per-cent shear costs 900 galaxies, and there is no cheaper route to it. The scatter of the mean ellipticity against the number of galaxies averaged, measured from 60–400 independent realisations at each N rather than quoted. The line through the points has a slope of −0.499 in the log plane against the −0.5 that averaging independent shapes must give, and the intercept is the intrinsic dispersion 0.3. Everything about a lensing survey is on this plot: a one-per-cent shear at a signal-to-noise of one needs 900 galaxies, at ten it needs a hundred times as many, and the only quantities an observer controls are the depth — galaxies per square arcminute — and the area. No improvement in the telescope changes the slope, because the noise is the galaxies' own shapes and not the instrument's; that is the difference between this measurement and almost every other one in the collection, where a bigger mirror helps. It is also why the systematic floor matters so much: below about 10⁻³ the PSF's own anisotropy stops being negligible, and averaging more galaxies makes it no smaller.
Fig. 5 Why the measurement needs so many galaxies. Each source’s intrinsic ellipticity is about thirty per cent and the shear signal is one per cent, so the signal-to-noise per galaxy is a thirtieth and the number needed to reach a per cent measurement of the shear is of order a million. That arithmetic sets everything about how such surveys are designed: the area, the depth, the seeing requirement and the number of visits all follow from it.
A coherent one-per-cent distortion, invisible on every galaxy in the picture. 150 background galaxies behind a lens of Einstein radius 14″, each drawn at its own ellipticity: an intrinsic shape with a dispersion of 0.3 per component, plus the reduced shear the lens adds. The strongest shear on any galaxy here is 0.035, one part in 8 of the intrinsic scatter, so no object in this field is measurably distorted and the tangential alignment cannot be seen by eye at all. Averaged over these 150, the mean tangential ellipticity is 0.0088 ± 0.0245 against the 0.0130 the lens model predicts — consistent with the lens and equally consistent with nothing, because 150 galaxies buy a precision of 0.024 and the signal is 0.013. Detecting it at five sigma takes about 13,275 of them, which is not a picture anybody can draw. The cross component — every shape rotated by 45°, which gravitational lensing cannot produce — averages −0.0016 ± 0.0245, consistent with nothing, and that null is what separates a mass from a badly figured optic. The signal is not in any galaxy. It is in the sum, and the whole design of a lensing survey follows from that.
Fig. 6 What is actually observed before any of the averaging: a field of galaxy shapes, each with its own intrinsic ellipticity, carrying a coherent distortion of about a per cent. The pattern is there and it is invisible to the eye; extracting it is a matter of measuring a million shapes to a precision far better than any one of them matters. Everything in this essay concerns what happens after that extraction, and every systematic in the extraction — the point spread function, the noise bias, the blending of overlapping galaxies — enters as a multiplicative error on the amplitude in dispute.

The amplitude has a multiplier in front of it

There is one systematic that acts on the disputed number the way a gain acts on a voltage, and it is worth separating from the others because nothing about the cosmology can absorb it.

A galaxy’s observed shape is its true shape convolved with the telescope’s point spread function and then sampled by pixels with noise in them. Deconvolving that is an estimation problem with a bias, and the bias is multiplicative: the recovered shear is the true shear times a factor near one plus an offset. The factor comes from the noise — a shape measured from a noisy image is systematically rounder or more elongated depending on the estimator — and from the point spread function’s own ellipticity leaking into the galaxy’s.

A one per cent error in that factor is a one per cent error in S8S_8, which is most of the tension. So the calibration of the multiplicative bias has to be better than a per cent, on a measurement whose per-galaxy signal-to-noise is a thirtieth.

Two approaches have been made to work. The first is simulation: generate images of galaxies with known shears, pass them through a model of the telescope and the pipeline, and measure what comes out. That works to the extent the simulated galaxies resemble real ones — and the resemblance that matters is in the faint, small, blended objects that dominate the sample, which are precisely the ones least well characterised. The second is to calibrate on the data themselves: artificially shear each observed image by a small known amount, re-run the measurement, and use the response to that shear as the calibration. That removes the dependence on a galaxy population model and replaces it with a dependence on the noise being correctly propagated.

Blending is the term that neither handles cleanly. In a deep survey a large fraction of objects overlap another object on the sky, and an overlapping pair measured as one object has a shape that belongs to neither. The fraction rises with depth, so the deepest surveys — the ones with the most galaxies and therefore the smallest statistical errors — have the largest blending correction. That is the wrong way round for a measurement approaching a per cent, and it is the main reason the next generation’s error budget is dominated by a term that was negligible in the first.

None of these is exotic physics; they are the ordinary difficulties of measuring a shape. What makes them decisive is the arithmetic at the top of this essay: a one per cent signal, averaged over a million objects, compared against a prediction at the two per cent level. An error budget added in quadrature is the right frame for the whole exercise, and here the terms that matter are the ones that multiply rather than the ones that add.

Where the picture stops

There are three, and the second is the one most likely to be the answer.

Baryons are not dark matter. The predictions are computed from simulations of gravity alone, and gas physics — feedback from supernovae and from active nuclei — pushes matter out of haloes and suppresses the power spectrum on small scales by several per cent. That suppression is in the direction of the observed tension, its size is uncertain by a factor, and it is currently marginalised over with a parameterised model rather than predicted.

The redshift calibration is the leading candidate. The shear signal depends on the geometry through the source distances, so a systematic error in the mean redshift of a source bin translates directly into the inferred amplitude. The required shift is small — a few hundredths in mean redshift — and it is at the edge of what the calibration can exclude.

And the comparison is model-dependent on both sides. The microwave background’s prediction for σ8\sigma_8 today is a computation through thirteen billion years of the standard model. A disagreement is a statement about that model as much as about either measurement, and the interesting possibilities — a different growth rate, a neutrino mass, an interaction in the dark sector — are all changes to what happens between the two epochs rather than to either observation.

One more deserves stating because it is the one that will decide the outcome. Both sides of the comparison are being improved: the lensing surveys are growing by an order of magnitude in area, and the microwave background’s prediction depends on parameters that other experiments are pinning down. If the tension is a systematic in the lensing, it will shrink as the systematics are controlled; if it is real, it will grow as the errors shrink. That is a clean test, it is a few years away, and it is the reason the field’s response has been to build larger surveys rather than to propose new physics. A map that is not of positions is one of the other probes of the same epoch, and its independent agreement or disagreement is part of the same test.

Why naming the combination is the honest thing to do

The underlying point deserves stating on its own because it is a matter of scientific hygiene rather than of technique.

When a measurement constrains a combination, the combination is the result. Quoting a value for σ8\sigma_8 from a lensing survey requires marginalising over Ωm\Omega_{\rm m} with a prior, and the answer then depends on the prior — so two surveys with identical data and different priors report different σ8\sigma_8 and identical S8S_8. Reporting the combination is what makes the comparison between experiments meaningful.

That discipline is not universal and its absence causes real confusion. A radial-velocity orbit measures a mass times a sine and the convention is to quote the product, which is right; a transit measures a radius ratio and the convention is to quote the radius, which imports the star’s error into every planet; a tidal rate measures a Love number over a quality factor and both conventions are in use.

The rule that emerges is simple: quote what was measured, and quote separately what was assumed to turn it into something else. A reader can undo an assumption that is stated and cannot undo one that is not.

A parting observation about how such a combination gets its name. S8S_8 is defined with an exponent of one half and a pivot at Ωm=0.3\Omega_{\rm m} = 0.3 because those choices make the combination close to what a particular generation of surveys measured best. A survey with different depth measures a slightly different combination, and quoting its result as S8S_8 involves a small projection along the degeneracy direction. That is harmless while the surveys are similar and it is a real source of confusion when comparing across generations — the same symbol denotes slightly different quantities in papers a decade apart. The honest presentation is the full posterior in the plane, and the single number is a summary of it whose definition has to be carried alongside. The mass that is not the light is measured by all of these methods, and comparing their answers requires knowing which combination each one actually determined.

Where the ladder goes next

Later rungs on this ladder start with intrinsic alignment: why physically associated galaxies are aligned before any lensing happens, how large the contamination is, and why it has the opposite sign for the two kinds of correlation. The rung after that is the baryonic suppression — what feedback does to the small-scale power spectrum, and why a measurement of cosmology now requires a model of how galaxies blow gas out of their own haloes.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Cosmic shearDegeneracyIntrinsic alignmentMatter densityPhotometric redshiftSigma 8Structure growthSystematic errorTomographyWeak lensing