Cosmology

The shape a galaxy was given before it was lensed

Weak lensing assumes that galaxies point in random directions, so that a coherent stretch can only have been put there by the light's journey. The tidal field that set every galaxy's spin also stretched its shape — towards the very mass that lenses the galaxies behind it — and the resulting error is not noise but a signal of the opposite sign.

Assumes Tidal torque theory, Weak lensing and Large-scale structure.

A weak-lensing survey measures a distortion of about one per cent in the shapes of galaxies whose own shapes scatter by thirty per cent. It can do that only because of a single assumption: that the thirty per cent points in random directions. Average a million galaxies and the random part cancels, leaving the coherent stretch that the light picked up on its way past the intervening matter. That is the whole logic of measuring mass from shapes, and every other part of the method — the point-spread function, the redshift calibration, the shear estimator — is a refinement of it.

The assumption is false, and it is false for a reason that has already been met elsewhere in this collection. A galaxy’s spin was applied by the tidal field of its neighbours while it was still an expanding patch of matter. The same field that exerted the torque also pulled the patch out of round, and it did so in a direction set by the surrounding mass. A galaxy’s intrinsic shape is therefore not random. It carries a small, coherent imprint of the large-scale structure it formed in — and large-scale structure is exactly what a lensing survey is trying to weigh.

Stretched towards the mass, sheared across it. Two rings of galaxies around the same concentration of mass. Left, galaxies physically beside it: its tidal field stretches each one along the line to the centre, so their long axes point radially and their mean tangential ellipticity is −0.30. Right, galaxies far behind it: their light is deflected past the mass and each image is sheared tangentially, with mean tangential ellipticity +0.30. The two patterns are perpendicular. A weak-lensing survey correlates the shapes of foreground and background galaxies to measure the shear, and when a foreground galaxy was aligned by the very structure that lenses the background one, their product is negative — −0.090 here — and subtracts from the signal rather than adding noise to it. What the picture exaggerates is the size: real intrinsic alignments are a per cent or two in the mean, buried under a random shape scatter of about 0.3, and are found only statistically.
Fig. 1 The two patterns drawn at an exaggerated thirty per cent so that the geometry is visible. Galaxies near a mass are stretched along the line towards it; galaxies far behind it are sheared across that line. In a real survey both effects are a per cent or two in the mean and are recovered only by averaging, but their relative orientation is exactly as drawn — perpendicular — and that is the fact everything below follows from.

A tide pulls towards the mass and lensing shears across it

The contamination has a definite sign, and the sign is the first thing worth understanding, because it is what makes the error dangerous rather than merely large.

A tidal field is the difference between the pull on the near side of a body and the pull on its far side. For a body beside a concentrated mass, that difference stretches the body along the line to the mass and squeezes it across that line — the same geometry that raises two ocean bulges on the Earth, and that pulls a satellite apart altogether when it comes close enough. A protogalaxy forming beside a cluster, or inside a filament feeding one, is stretched towards it. When it collapses and its stars settle, the long axis of the resulting galaxy remembers that direction.

Lensing does the opposite. Light from a galaxy far behind the cluster is deflected past it, and a small image is sheared tangentially: stretched around the mass rather than towards it. The pattern of weak shear around a cluster is a set of tiny arcs concentric with the cluster — the faint, statistical version of the giant arcs that strong lensing draws close to the core — and it is that tangential alignment which the survey detects.

So a foreground galaxy aligned by some structure and a background galaxy lensed by the same structure are aligned perpendicular to each other. Their shapes are correlated, and the correlation is negative. A survey that correlates foreground and background shapes to measure the shear finds less signal than the lensing alone would give — and reads the deficit as less matter.

This cross term is called the gravitational–intrinsic correlation, GI, and it was pointed out only in 2004, several years after the simpler contamination had been recognised and modelled. The simpler one is intrinsic–intrinsic, II: two galaxies that formed near each other were stretched by the same field and are aligned with each other, adding a positive correlation that has nothing to do with lensing at all. II needs the two galaxies to be physically close; GI needs one galaxy to sit in front of the other and near the mass that lenses it. They are different effects with different geometry, and a survey has to model both.

Where along the line of sight the damage is done

The size of the GI term for any pair of galaxy samples is an overlap integral, and drawing its two ingredients shows why the contamination is distributed so unevenly across a survey.

The galaxy that lenses and the galaxy that is aligned. Three curves against redshift, each scaled to a peak of one. The foreground galaxies' distribution, a photometric-redshift bin centred at z = 0.4 with width 0.1; the lensing efficiency of a background bin at z = 1, which is zero at the observer and at the sources and peaks at z = 0.51, roughly halfway in distance; and their product, which is the integrand of the gravitational–intrinsic term. A foreground galaxy aligned by some structure contaminates the background's shear exactly to the extent that the same structure lenses the background, and the product says where that happens. The foreground bin overlaps the background's lensing kernel 6.3 times as strongly as it overlaps its own, because a source a long way behind is lensed most by what sits halfway to it — which is why the contamination is worst for the most widely separated pairs of bins rather than the closest. What the curves cannot show is the sign: the overlap is positive and the term it produces is negative, because stretched and sheared shapes are perpendicular.
Fig. 2 A foreground sample at redshift 0.4 against the lensing efficiency of sources at redshift 1. The efficiency is zero at the observer and zero at the sources and peaks near z = 0.5 — roughly halfway in distance — because a lens deflects most usefully when it sits between the two. The foreground sample falls almost entirely under that peak, so almost every structure that aligns a foreground galaxy also lenses the background ones. The dashed curve is the foreground sample’s own lensing efficiency, which peaks at lower redshift and barely overlaps the galaxies at all.

A lensing survey does not use a single pool of galaxies. It divides them by photometric redshift into several bins and measures the shear correlation between every pair of bins, a technique called tomography, because the way the signal grows with source distance is what separates the amount of matter from the rate at which structure has grown. Each pair of bins gives its own spectrum, and each spectrum has its own contamination.

The foreground galaxies’ distribution and the background’s lensing kernel have to overlap for GI to exist. That overlap is largest when the foreground bin sits about halfway to the background bin — which is where the background’s lensing kernel peaks — and it is small when the two bins are close together, since a galaxy is hardly lensed by structure at its own distance. The integral has a factor of 6.3 in it, in the case drawn: the foreground galaxies sit 6.3 times more heavily under the background’s lensing kernel than under their own.

That asymmetry is the whole story of the next figure.

Two errors that move in opposite directions

Evaluating the three spectra — the true lensing signal GG and the two contaminants — for a background bin at redshift 1 and a foreground bin sliding towards it gives a picture that is at first sight paradoxical.

Two contaminants that move in opposite directions. The two intrinsic-alignment terms as fractions of the true shear–shear signal, for background sources in a bin at z = 1 cross-correlated with a foreground bin whose centre runs from 0.2 up to the background's own redshift, bins 0.1 wide, in the linear alignment model with amplitude A = 1 and a matter spectrum approximated as a power law of slope −1.5. The gravitational–intrinsic term is negative throughout and is 1.28 times the signal for the widest separation, falling to 0.022 when the bins coincide. The intrinsic–intrinsic term needs the two galaxies to be physically near each other, so it lives only where the two redshift distributions overlap: below a ten-thousandth of the signal at the widest separation, 0.0135 when the bins coincide. II never overtakes GI for these bins. Neither is noise: both are coherent, both scale with the same tidal field, and a survey that ignored them would read the first as too little matter and the second as too much. The widely separated pairs look the most contaminated because their true signal is small — the foreground galaxies are barely lensed by anything — and not because the alignment is larger there.
Fig. 3 The two contaminants as fractions of the true lensing correlation, for bins 0.1 wide in redshift, in the linear alignment model at unit amplitude. GI is negative at every separation and exceeds the signal itself when the foreground bin is at 0.2; it falls to two per cent when the bins coincide. II is positive and is confined to the right-hand side, where the two redshift distributions overlap, reaching 1.4 per cent. The matter power spectrum is approximated as a power law in wavenumber, which sets how much weight nearby distances receive; the ratios drawn do not depend on its normalisation at all.

The pair with the widest separation, foreground at 0.2 against background at 1, is contaminated by 128 per cent: the GI term is larger than the lensing signal it is meant to be a correction to. That is not because the alignment is stronger there. It is because the true signal is small. The lensing correlation between two bins is limited by the nearer bin’s lensing, and galaxies at redshift 0.2 are barely lensed by anything — there is very little matter in front of them. The contamination, which depends only on the far bin’s lensing of the near bin’s aligners, has no such limit.

At the other end, when the two bins coincide, GI falls to about two per cent and II rises to about one and a half. The two contaminants trade places, and a survey that fitted only one of them would fit it wrongly in half of its spectra.

Neither of these is a small correction to a precise measurement. The widely separated bin pairs carry little signal and are down-weighted for that reason, so their enormous fractional contamination does less damage than the figure suggests. But the pairs that carry most of the constraint — adjacent and overlapping bins at the high-redshift end — are contaminated at the several-per-cent level, which is comparable to the entire statistical error of a modern survey.

A photometric redshift is a spread, and the spread lets II in

Real redshift bins are not 0.1 wide. A photometric redshift is estimated from a handful of broadband colours rather than from a spectrum, and its scatter is a few per cent in 1 + z, with a tail of outright failures. The bins that result have long wings, and the wings overlap.

Two contaminants that move in opposite directions. The two intrinsic-alignment terms as fractions of the true shear–shear signal, for background sources in a bin at z = 1 cross-correlated with a foreground bin whose centre runs from 0.2 up to the background's own redshift, bins 0.25 wide, in the linear alignment model with amplitude A = 1 and a matter spectrum approximated as a power law of slope −1.5. The gravitational–intrinsic term is negative throughout and is 0.51 times the signal for the widest separation, falling to 0.050 when the bins coincide. The intrinsic–intrinsic term needs the two galaxies to be physically near each other, so it lives only where the two redshift distributions overlap: 3.3e-3 of the signal at the widest separation, 0.0058 when the bins coincide. II never overtakes GI for these bins. Neither is noise: both are coherent, both scale with the same tidal field, and a survey that ignored them would read the first as too little matter and the second as too much. The widely separated pairs look the most contaminated because their true signal is small — the foreground galaxies are barely lensed by anything — and not because the alignment is larger there.
Fig. 4 The same calculation with bins 0.25 wide, closer to what a photometric survey actually achieves. II no longer vanishes at wide separation, because the wings of the two distributions now overlap even when their centres do not; it is a third of a per cent at the widest separation. GI at the widest separation has fallen from 128 per cent to 51, because the broad foreground bin has moved some of its galaxies out from under the peak of the background kernel. Neither contaminant can be removed by choosing which pairs of bins to use, since every pair now contains some of each.

This matters because the most obvious defence against II — use only the cross-correlations between bins that do not overlap, and throw away the auto-correlations — works only if the bins really do not overlap. With photometric redshifts they always do. A galaxy assigned to redshift 0.4 may be at 0.9, and if it is, it sits beside the galaxies of the far bin and shares their tidal field. The II term leaks into every cross-spectrum in proportion to the overlap of the tails, which is set by the quality of the photometric redshifts rather than by anything about the alignment.

That is the reason the two leading systematics of cosmic shear — the redshift calibration and the intrinsic alignments — cannot be treated separately. An error in the width of a redshift distribution changes the II contamination as well as the lensing kernel, and the two effects are fitted jointly or not at all.

Why an elliptical is aligned and a spiral is not

The model behind every figure so far is called linear alignment. It says that a galaxy’s intrinsic ellipticity is proportional to the tidal field at the place and time it formed, projected onto the sky, with one free constant. It works for a galaxy whose shape is set by being stretched: a pressure-supported elliptical, whose stars were given their orbits in a collapse that the tidal field distorted. For such a galaxy the alignment is linear in the field.

A disc galaxy is different. Its shape on the sky is set by the orientation of the disc, and the orientation is set by the spin, and the spin is a product — the tidal tensor contracted against the patch’s own inertia tensor. Both of those were generated by the same random density field. The disc’s alignment is therefore quadratic in the field, and a quadratic function of a Gaussian field decorrelates much faster than the field itself.

An alignment that is squared is gone by twenty megaparsecs. How the correlation between two galaxies' alignments falls with their separation, for a density correlation function that is a power law of slope −1.8 with a correlation length of 5 Mpc/h, both curves set equal at one megaparsec. A galaxy whose shape is stretched by the tidal field carries an alignment linear in that field, and the correlation between two such galaxies falls as the field's own correlation does. A disc whose orientation follows its spin carries an alignment quadratic in the field, because the torque is a product of the field with the patch's shape, and for a Gaussian field the covariance of two squared quantities is twice the square of the covariance — checked here by drawing sixty thousand correlated pairs, which give 0.491 against the exact 0.500. So the quadratic alignment falls as the square: at twenty megaparsecs it is 4.6e-3 of the linear one. That is the prediction that separates the two populations, and it is what the surveys find — pressure-supported red galaxies are aligned on large scales and blue discs show no detected alignment at all. What the curves leave out is the mixture: the scales a lensing survey uses are larger than a few megaparsecs, where only the linear term survives.
Fig. 5 How fast the two kinds of alignment decorrelate with separation, for a density correlation function of the usual power-law form, both normalised to agree at one megaparsec. The stretched shapes follow the correlation function itself; the spun discs follow its square, which for a Gaussian field is exact — the covariance of two squares is twice the square of the covariance, checked by drawing correlated pairs. By twenty megaparsecs the quadratic alignment is about half a per cent of the linear one, and on the scales a lensing survey uses it has effectively gone.

The prediction is that red, pressure-supported galaxies should show intrinsic alignments on large scales and blue discs should not, and that is what the measurements find. Samples of luminous red galaxies with spectroscopic redshifts show a clear alignment of their shapes with the surrounding density field, extending tens of megaparsecs, with an amplitude several times the reference value and rising with luminosity. Samples of blue emission-line galaxies show no detected alignment at all on those scales, with upper limits well below the red galaxies’ signal.

This is one of the cleaner confirmations the tidal picture has. The linear and quadratic mechanisms were written down for different reasons — one to explain shapes, the other to explain spins — and they predict a difference in how two galaxy populations should behave that nobody put in by hand. The difference is observed, and it is in the direction and roughly of the size the two exponents require.

It also means that the contamination of a lensing survey depends on what its galaxies are. A deep survey is dominated by faint blue galaxies at high redshift, which are the least aligned population, and its effective amplitude is small. A survey selected for bright red galaxies would be contaminated several times more strongly by the same tidal field.

Measured where lensing cannot reach

The amplitude of the alignment cannot be read off a lensing survey, because there it is entangled with the signal it contaminates. It is measured instead where lensing is negligible: in nearby galaxies with spectroscopic redshifts, where the light has passed through too little matter to be sheared by any detectable amount and where the redshifts are precise enough to know which galaxies are physically close.

The statistic is a cross-correlation between the positions of galaxies and the shapes of their neighbours. If shapes are aligned by the density field, a galaxy’s long axis should point, on average, towards where the other galaxies are — the same radial pattern as the left half of the opening figure, but built up statistically from pairs rather than drawn around one mass. The signal is a function of separation, and its shape and amplitude are compared with the linear model’s prediction. The redshifts matter twice: once to select pairs that are genuinely close in three dimensions, and once to correct for the stretching of a redshift map by the peculiar velocities of the very flows the tidal field drives, which moves galaxies along the line of sight in a way correlated with their surroundings.

The constant in the model was normalised originally from a photographic survey of low-redshift galaxy shapes published in 2002, and every amplitude quoted since has been expressed relative to that reference. The luminous red galaxies come out at several times it; the blue galaxies at a value consistent with zero. What no spectroscopic measurement can supply is the amplitude for the faint galaxies at redshift one that dominate a lensing survey, because those galaxies are too faint to have spectra in the numbers required. Their alignment is extrapolated — in luminosity, in colour and in redshift — from populations they do not belong to.

A contamination that can change sign

For a single broad sample of source galaxies the two contaminants combine into one correction, and the way they combine depends on the alignment amplitude in a way that makes the size of the error hard to guess from the physics alone.

A contamination with a zero in it. The shear–shear spectrum of a single broad source sample, centred at z = 0.8 with width 0.3, as a fractional error when intrinsic alignments are ignored, against the alignment amplitude A of the linear model. The gravitational–intrinsic term is linear in A and negative, −9.9% of the signal at A = 1; the intrinsic–intrinsic term is quadratic and positive, 0.79% at A = 1. Their sum is a parabola through the origin that dips and then rises, and it returns to zero at A = 12.6: a sample that aligned that strongly would show no net contamination at all, while being aligned more strongly than any sample has been measured to be. The figure's point is that the sign of the net error is not fixed by the physics but by the sample — a population of discs with A near zero is clean, a mixed sample near one is biased low by 9.1%, and a sample of luminous red galaxies with A = 5 is biased by −30%, with the positive term already cancelling 40% of the negative one. It cannot show which amplitude a survey has; that is fitted from the data alongside the cosmology.
Fig. 6 The fractional error in the shear spectrum of a single source sample centred at redshift 0.8, if the alignments are ignored, against the alignment amplitude. GI is linear in the amplitude and negative, about ten per cent of the signal at unit amplitude. II is quadratic and positive, less than one per cent at unit amplitude but twenty-five times that at an amplitude of five. The net correction is a parabola: most negative somewhere in the middle, and back to zero at an amplitude of about thirteen, which no measured population has approached.

At unit amplitude, roughly what a mixed sample of red and blue galaxies might carry, the spectrum is underestimated by nine per cent. Since the shear spectrum scales roughly as the square of the amplitude of the matter fluctuations, a nine per cent deficit in the spectrum reads as a deficit of four or five per cent in the inferred clumpiness of matter. That is the size of the discrepancy between lensing surveys and the microwave background that has been debated for a decade, and the coincidence of scale is the reason intrinsic alignments are always among the first suspects in that argument.

At an amplitude of five, the value measured for luminous red galaxies, the net error is thirty per cent, and the positive II term has already cancelled two-fifths of the negative GI. At larger amplitudes still it would cancel all of it. There is no amplitude at which the sign of the net error can be stated without also knowing the redshift distribution, since the relative size of the two terms depends on how much of the sample overlaps itself.

Current surveys do not assume an amplitude. They fit it simultaneously with the cosmology, as a nuisance parameter with a prior, and in the most recent analyses they fit a second parameter as well — the amplitude of the quadratic, torque-driven term, which the literature calls tidal torquing and which is the same mechanism that spun the halo in the first place. The fitted amplitudes for the mixed samples of recent surveys have come out small, a fraction of one, with uncertainties comparable to their values. That is consistent with a faint blue sample and it is also consistent with the model being incomplete, and the data do not yet distinguish the two readings.

Separating the two by their geometry

There is a way out that uses nothing but the redshift dependence, and it is worth stating because it is the one defence that does not depend on the alignment model at all.

The lensing signal between two bins and the GI contamination between them depend on the redshifts in different ways. Lensing grows with the distance of the source behind the lens according to a known geometric factor; GI depends on where the foreground galaxies are, not on how far behind them the sources sit. So for a fixed foreground bin, the lensing correlation with a sequence of background bins rises in a predictable way as the background recedes, while the GI part follows the lensing kernel of the background evaluated at one redshift. Given enough bins with good enough redshifts, the two can be separated by fitting that difference in shape — a technique called self-calibration — or the GI part can be nulled outright by weighting the background bins so that their combined lensing kernel vanishes at the foreground’s redshift.

Both work in principle and both pay for it. Nulling throws away the part of the signal that shares the contaminant’s geometry, which is a large part, and both methods require the redshift distributions to be known far better than photometric redshifts currently know them. The method is limited by exactly the systematic the previous section showed to be entangled with the alignments: the widths of the bins.

Where the linear model stops

Everything drawn here uses the tidal field on large scales, where perturbation theory holds. Most of the constraining power of a lensing survey is on smaller scales, where it does not.

Inside a single dark-matter halo the alignment is of a different kind. Satellite galaxies orbiting a massive central galaxy are stretched radially towards it, like the foreground ring in the opening figure, and the central galaxy’s own long axis tends to point along the major axis of its halo and hence towards the nearest filament. Neither of those is described by the linear model’s single constant, and both are measured to be real in groups and clusters. Modelling them requires a picture of how galaxies occupy halos — a halo model with alignment terms — and the parameters of that model are fitted from the same data they are supposed to correct.

There is also a question the model assumes away: when the alignment was imprinted. Linear alignment evaluates the tidal field at the epoch of galaxy formation and assumes the shape is frozen after that. A galaxy that merged, or whose shape was re-stretched by the field at a later time, carries a different alignment from the one the model assigns it, and simulations that follow galaxy shapes through their histories find that the alignment strength evolves.

Finally there is the pure measurement problem. An intrinsic alignment is measured from galaxy shapes, exactly like the lensing signal, and it is contaminated by the same imperfect correction for the telescope’s own blurring. A residual point-spread-function error with a coherent pattern across the sky mimics an alignment as easily as it mimics a shear.

What the picture cannot show

The opening figure draws the alignment at thirty per cent so that its direction can be seen, and that is the figure’s one real distortion. Measured alignments are a per cent or two in the mean ellipticity, a small bias on a distribution whose width is thirty times larger, and they have been found only statistically — from correlations over hundreds of thousands of galaxies with spectroscopic redshifts. No individual galaxy is known to be tidally aligned, and none could be: for any one object, the stretch it was given is indistinguishable from the shape it happens to have.

The figures of contamination are also model outputs. Their shape — which pairs of bins are worst, which term dominates where, the parabola with its zero — follows from geometry and the Gaussian statistics of the density field and is robust. Their amplitude depends on a constant fitted from one survey of nearby galaxies in 2002, rescaled by a free parameter that current data barely constrain. The honest statement is that the direction and the redshift dependence of the error are known and its size is not.

Still open: whether the torque term is in the data

The one piece of this that connects back directly to the theory of spin is the quadratic term. It is what the disc galaxies ought to carry, and it has been added to the model because the physics says it must exist. Whether it has been detected is another matter. The surveys that fit it report values within one or two standard deviations of zero, and not yet a consistent picture from one analysis to the next.

What would settle it is a measurement of spin directions rather than shapes: a large sample of disc galaxies whose rotation has been measured kinematically, so that the spin axis is known rather than inferred from the projected ellipse. The alignment of those spins with the surrounding tidal field is what tidal torque theory predicts directly, what that alignment can and cannot reveal about the field is a question in its own right, and at the moment the answer is limited less by theory than by the number of galaxies with measured rotation.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Correlation functionCosmic shearGalaxy biasIntrinsic alignmentLarge-scale structureLinear growthPhotometric redshiftShape noiseSpin alignmentTidal tensorTidal torque theory