Galaxies

Weighed by the light that bends past it

Every other mass in this collection is measured from something orbiting, which requires the system to have settled down. A gravitational lens weighs whatever is in the way with no such assumption — the light does not care whether the mass is in equilibrium.

Assumes Microlensing and Dark matter.

Every mass in this collection so far has been weighed by watching something go round it. A planet’s mass comes from its moons, a star’s from its companion, a galaxy’s from its rotation, an elliptical’s from its dispersion. All of those require the system to be in a steady state: an orbit that has been completed, or a velocity distribution that has settled.

Light passing a mass is deflected by it, and the deflection does not require anything to have settled. That is the whole reason gravitational lensing matters as a measurement rather than as a spectacle.

Why an Einstein radius is a mass. The geometry, drawn at an angle some ten thousand times larger than the real one so that anything is visible at all. Light from a source directly behind a lens reaches the observer along every path that passes the lens at the same distance, so the image is a ring rather than a point. The ring's angular radius is θ_E = √(4GM/c² · D_ls/D_l D_s), which for a lens of 1.0×10¹² M☉ at these distances is 2.52 arcseconds and encloses 10 kpc at the lens. Rearranged, it is a mass in terms of an angle and three distances — and the mass so obtained is inside a cylinder rather than a sphere, and assumes nothing whatever about the lens being in equilibrium, which is the assumption every other weighing in this collection makes.
Fig. 1 The geometry, drawn at an angle some ten thousand times larger than the real one. Light from a source directly behind a lens reaches the observer along every path that passes the lens at the same distance, so the image is a ring. Its angular radius is θ_E = √(4GM/c² · D_ls/D_l D_s), which for a 10¹² solar-mass lens at these distances is about a second of arc and encloses 91 kiloparsecs at the lens.

The deflection, and the factor of two

A ray passing a mass MM at impact parameter bb is deflected by

α=4GMc2b.\alpha = \frac{4GM}{c^2 b}.

The expression is worth pausing on, because it is one of the very few places in this collection where a Newtonian calculation gives an answer that is exactly half right.

Treating light as a stream of fast particles on hyperbolic orbits gives 2GM/c2b2GM/c^2b — the calculation Soldner did in 1801 and Einstein repeated in 1911. General relativity doubles it, because space is curved as well as time and both contribute equally for a photon.

The bending angle, and the factor of two. Deflection against impact parameter for a 1.0×10¹² M☉ lens, on logarithmic axes. The relativistic value is 4GM/c²b and the Newtonian calculation — a corpuscle of light on a hyperbolic orbit — gives exactly half of it. The two differ by a constant factor at every impact parameter, so no measurement of the shape of the curve can separate them and only an absolute measurement can, which is what the 1919 eclipse was. At the Einstein radius of this configuration, 10 kpc, the deflection is 4.0 arcseconds.
Fig. 2 Deflection against impact parameter, on logarithmic axes. The two predictions differ by a constant factor at every impact parameter, so no measurement of the shape of the curve can separate them — only an absolute measurement can. That is what the 1919 eclipse expedition was: a measurement of the deflection of starlight grazing the Sun, at 1.75 arcseconds against the Newtonian 0.87.

The Einstein radius, which is a mass

Put a source directly behind a lens. By symmetry, light leaving the source in every direction that passes the lens at the right distance arrives at the observer, and the image is a ring.

The condition that fixes its radius is that the deflection exactly compensates the offset, and working it through in the small-angle limit gives

θE=4GMc2DlsDlDs.\theta_E = \sqrt{\frac{4GM}{c^2}\frac{D_{ls}}{D_l D_s}}.

Rearranged, that is a mass in terms of an angle and three distances. Nothing else is needed: no velocities, no orbits, no assumption that the lens has settled into anything.

The mass so obtained has two features that must be stated with it. It is the mass inside a cylinder of radius θEDl\theta_E D_l along the line of sight, not inside a sphere — lensing is sensitive to the projected surface density, so everything in front of and behind the lens on the same line contributes. And it is exact only for a circular lens; a real galaxy is elliptical, and the ring becomes arcs.

Two images, always

A source not exactly behind the lens produces two images rather than a ring, and where they are is worth knowing because it is the observable in every real case.

In units of the Einstein radius, the lens equation for a point mass is β=θ1/θ\beta = \theta - 1/\theta, whose solutions are

θ±=12(β±β2+4).\theta_\pm = \tfrac12\left(\beta \pm \sqrt{\beta^2+4}\right).

Two images, always, and where they sit. The positions of the two images of a point source against the source's true offset, both in units of the Einstein radius, from θ = ½(β ± √(β²+4)). The images always straddle the lens: one outside the Einstein radius and one inside it, on opposite sides. As the source approaches alignment the two converge on the ring from either side and both brighten without limit; as it moves away the outer image tends to the source's own position and the inner one collapses onto the lens, faint and usually unobservable. Nothing about this depends on what the lens is made of.
Fig. 3 The two image positions against the source’s true offset, both in Einstein radii. The images always straddle the lens: one outside the Einstein radius, one inside it, on opposite sides. As the source approaches alignment they converge on the ring from either side and both brighten without limit; as it moves away, the outer image tends to the source’s own position while the inner one collapses onto the lens and fades. Nothing here depends on what the lens is made of.

Two consequences follow that are used constantly.

The total magnification is never less than one. Lensing conserves surface brightness — which is the distance-independent quantity — while increasing the solid angle, so it always adds flux. That is why lensing clusters are used as telescopes: a foreground cluster magnifies the galaxies behind it by factors of ten or more, reaching sources that no instrument could otherwise detect.

The images have different arrival times. They travel different path lengths and pass through different depths of the gravitational potential, so a source that varies is seen to vary in one image before the other, by days to years. Measuring that delay gives a length, and a length divided by a redshift is a Hubble constant — a route to the expansion rate that bypasses the distance ladder entirely.

A delay of 81 days, and a sheet nobody can see that moves H₀ to 82.4. Above: the arrival-time surface of a lensed source, along the line through the lens. The curve is the Fermat potential in days — the geometric cost of taking a longer path, minus the gravitational cost of climbing out of the potential — and the images sit at its stationary points, at -1.20″ and 2.04″, which for an isothermal sphere is β ± θ_E. Fermat's principle is doing all of the work here: light does not take the shortest path or the quickest one, it takes every stationary one, and the number of images is the number of stationary points. The vertical distance between the two is 81 days, and it is measurable — the source is a quasar, quasars vary, and the same wiggle appears in one image and then the other. That single number carries an absolute distance: the delay is D_Δt/c times a dimensionless function of the lens model, and D_Δt goes as 1/H₀, so a monitoring campaign gives the Hubble constant with no rung of any ladder beneath it. Below: the two light curves, shifted by exactly that delay. The dashed curve is the second image with the delay removed, and the agreement is the measurement. What the picture also shows is the reason the answer keeps moving. The second arrival-time curve is the same lens with a uniform sheet of convergence added and the source moved to compensate: every image sits at the same place, every flux ratio is the same, every image shape is the same, and the delay is λ = 0.85 times as long. A lens model fitted to positions alone cannot see the sheet, and inferring H₀ from the same delay under it gives 82.4 instead of 70 — a 15 per cent shift with no observable attached. Breaking it needs a mass measured some other way: the velocity dispersion of the deflector, or a count of everything else along the line of sight.
Fig. 4 The delay, and the reason it has not settled the expansion rate. Two images of one varying source arrive 81 days apart for this configuration, and that interval is a length divided by the speed of light — so a measured delay and a modelled potential give H0H_0 with no rung of the distance ladder underneath. What the modelled potential cannot see is a sheet of mass spread smoothly across the whole field: adding one rescales every image position and every magnification together, leaves the observables untouched, and moves the inferred H0H_0 from 67 to 82. That degeneracy is the method’s whole error budget, and it is broken only by measuring the lens galaxy’s own velocity dispersion — which is a different physics with a different set of assumptions.

The same equations, one scale down

This collection has already met lensing, in a form where the images are never resolved.

A magnification, and a spike inside it. The brightness of a background star as a foreground one passes in front of it. The smooth curve is exact: a point mass magnifies a point source by (u² + 2)/(u√(u² + 4)), which peaks at 5.1 for this track. The spike is the planet, of mass ratio 0.004, lensing one of the two images. Its duration is the Einstein time scaled by √q — about 1.9 days against 30 days — so the whole planetary signal is a few hours in an event lasting a month, and it never repeats. The deviation is drawn in the approximation that the planet lenses the image in isolation; a real caustic crossing has structure this smooths over, and its true height is set by the source's size rather than by the geometry.
Fig. 5 Microlensing: the same lens equation with a lens of a solar mass rather than 10¹², so the Einstein radius is a milliarcsecond and the two images cannot be separated. What is observed is their combined brightness as the alignment changes, and the mass enters through the duration of the event rather than through an angle. The deviation partway through this curve is a planet — the same mathematics, one further factor of a thousand down in mass.

The continuity is exact and worth spelling out. A stellar-mass lens gives an Einstein radius of about a milliarcsecond, so the images are unresolvable and the observable is a light curve. A galaxy-mass lens gives about an arcsecond, so the images are separated and the observable is their configuration. A cluster gives ten to thirty arcseconds, so the images are arcs stretched around the cluster’s core.

Three regimes, three names, one equation, and a factor of 101510^{15} in mass between the ends.

Why an Einstein radius is a mass. The geometry, drawn at an angle some ten thousand times larger than the real one so that anything is visible at all. Light from a source directly behind a lens reaches the observer along every path that passes the lens at the same distance, so the image is a ring rather than a point. The ring's angular radius is θ_E = √(4GM/c² · D_ls/D_l D_s), which for a lens of 100.0×10¹² M☉ at these distances is 23.78 arcseconds and encloses 104 kpc at the lens. Rearranged, it is a mass in terms of an angle and three distances — and the mass so obtained is inside a cylinder rather than a sphere, and assumes nothing whatever about the lens being in equilibrium, which is the assumption every other weighing in this collection makes.
Fig. 6 The top of that range: a cluster of 10¹⁴ solar masses at the same sort of distance. The Einstein radius scales as the square root of the mass, so a hundredfold heavier lens has a ten times larger ring — tens of arcseconds rather than one, which is why cluster arcs were seen and recognised in the early 1980s while galaxy-scale lenses needed better resolution. The enclosed radius is now well over a hundred kiloparsecs, and the mass measured is a substantial fraction of the whole cluster’s.

The critical density, which decides whether a lens is strong

There is a threshold in the subject, and it is the cleanest way to say what “strong” lensing means.

The deflection at a given radius depends on the mass inside that radius, projected. Writing the projected mass as a surface density Σ\Sigma and asking when the deflection is large enough to produce multiple images gives a critical value:

Σcrit=c24πGDsDlDls.\Sigma_{\text{crit}} = \frac{c^2}{4\pi G}\frac{D_s}{D_l D_{ls}}.

A lens with Σ>Σcrit\Sigma > \Sigma_{\text{crit}} somewhere produces multiple images and arcs; one below it everywhere produces a slight distortion and nothing more. The ratio Σ/Σcrit\Sigma/\Sigma_{\text{crit}} is called the convergence, and it is the natural variable of the whole subject.

The numbers are worth having. For a source at cosmological distance and a lens halfway to it, the critical surface density is of order a few thousand solar masses per square parsec. The centre of a massive elliptical exceeds it; the centre of a rich cluster exceeds it over a region tens of kiloparsecs across; a spiral galaxy’s disc, seen face-on, does not.

That single comparison explains the census of known lenses: they are ellipticals and clusters, and nothing else, because those are the only structures dense enough in projection. It also explains why lensing tells so little about the outskirts of anything — the convergence falls away, the images become a distortion too small to see in any individual galaxy, and only the statistics of thousands of background shapes recover the signal.

Why this weighing is different from all the others

The point cannot be made too plainly: a lensing mass requires no assumption about the dynamical state of the lens.

A rotation curve assumes circular orbits. A velocity dispersion assumes virial equilibrium. An X-ray mass assumes hydrostatic equilibrium. Each of those is a statement about the system having had time to settle, and each is wrong for a system that has recently been disturbed — which, for clusters, is most of them.

Lensing assumes only that the deflection formula is right and that the distances are known. That independence is what makes the agreement between lensing masses and dynamical masses meaningful: two measurements that share no assumptions agreeing is evidence, where two measurements that share an assumption agreeing is not.

The sharpest use of that independence is a system in which the two disagree by construction. In the Bullet Cluster — two clusters caught shortly after passing through each other — the X-ray gas has been slowed by the collision and sits between the two galaxy concentrations, while the lensing mass follows the galaxies. The mass is not where most of the ordinary matter is. It is difficult to produce that separation with any modification of the force law, because a modified force still tracks the matter that is there.

What lensing has actually established about galaxy masses

Applied systematically, strong lensing has produced a few results that are worth separating from the machinery.

The total mass profile of a massive elliptical is very close to isothermal. Combining an Einstein radius with the galaxy’s own velocity dispersion constrains the profile between the two radii, and the answer is ρr2\rho \propto r^{-2} to within a few per cent, over hundreds of lenses. That is not what either the stars alone or the dark halo alone would give — the stellar profile is steeper and the halo’s is shallower — and it means the two conspire to produce a power law. Why they should is an unsolved problem with its own name: the bulge–halo conspiracy.

The dark fraction inside the Einstein radius is about a half. Which is a strong constraint on the stellar mass-to-light ratio, and therefore on the initial mass function, from a direction that has nothing to do with counting stars.

And substructure shows up in flux ratios. A smooth lens predicts the relative brightnesses of its images; the observed ratios frequently disagree, and the disagreement is what a clump of mass near one image path would produce. That is the most direct evidence available that dark matter is not perfectly smooth on small scales, and it is one of the few observations capable of distinguishing between candidate particles.

The observation behind the number

Strong lenses are found rather than predicted, and finding them is a search for coincidences.

The first was Q0957+561 in 1979: two quasars 6 arcseconds apart with identical spectra and identical redshifts, which is not a plausible pair of objects and is a very plausible pair of images. The lensing galaxy was found afterwards.

Since then the search has been industrialised. Spectroscopic surveys find galaxies whose spectra contain emission lines at two redshifts — the foreground galaxy’s own, and a background source’s — and imaging then reveals the arcs. Automated searches of wide imaging surveys now find them by shape.

The measurement itself is a modelling exercise. The observables are the positions, shapes and brightnesses of the images; the model is a mass distribution; and the fit adjusts the model until the predicted images match. What comes out is tightly constrained at the Einstein radius and much less so elsewhere, which is the honest summary of what a strong lens measures: one number, the projected mass inside one radius, to a few per cent.

One more thing lensing measures that nothing else can, and it is worth naming before the limitations.

Because the deflection depends on the projected mass and not on what that mass is doing, a lens measures the total column along the line of sight — including matter that is not bound to the lens at all. For a single system that is a nuisance, and it is a leading contributor to the scatter in lensing masses. Averaged over thousands of lines of sight it becomes the signal itself: the statistical distortion of background shapes maps the mass of the whole cosmic web, filaments and all, and not merely the mass of the objects bright enough to have been catalogued.

The transformation the images cannot see

A lens measurement looks like a clean geometric statement, and it has one degeneracy that no amount of imaging removes.

Take a lens model that reproduces the observed image positions. Now add a uniform sheet of mass across the whole field and shrink the model’s own mass profile by a compensating factor, while rescaling the unlensed source. The image positions are unchanged. The flux ratios between images are unchanged. The shapes of the arcs are unchanged.

Everything observable about the configuration is identical, and the total mass inside the Einstein radius is not — the two models differ by whatever the sheet contributes.

That is the mass-sheet degeneracy, and it is exact rather than approximate. It survives because the lensing observables depend on the derivatives of the deflection, and the added sheet’s contribution is a term the observables cannot distinguish from a change of scale in the source, which is itself unobservable since the unlensed source is never seen.

What the degeneracy does not touch is the mass inside the Einstein radius when the sheet is known to be zero — which is why the standard result quoted for a lens is the enclosed mass under an assumption about the environment, and why a lens sitting in a cluster is harder to interpret than one in the field.

Breaking it requires an observable that responds to the total mass rather than to its derivatives. Three are used. Stellar kinematics of the lens galaxy measure the depth of its potential well independently, which fixes the scale. A known source size — a supernova of known luminosity, say, behind the lens — fixes the magnification and therefore the sheet. And time delays between images depend on the potential itself rather than on its gradient, so they respond to the transformation, which is why they can break it and also why a measurement of the expansion rate from lensing carries the degeneracy as its leading systematic.

A degeneracy that is exact cannot be reduced by better data of the same kind, and every improvement in lensing masses over the last decade has come from adding a different kind of measurement rather than from sharper images.

The same is true of the other degeneracies in lens modelling, which are approximate rather than exact and which behave better: adding data of the same kind does reduce them, and the practical limit on a well-observed lens is now the number of independently identifiable features in the source rather than the resolution of the images.

That is a comfortable position for a technique to be in, and it is worth distinguishing from the exact case above, where no amount of source structure helps at all.

The practical result is that a lens with a single featureless arc is modelled with several free parameters and constrained by few, while one with a lumpy source and four images is over-constrained and can be tested rather than merely fitted.

What the picture cannot show

The mass-sheet degeneracy. Adding a uniform sheet of surface density to a lens and rescaling the source changes no observable image position or shape. So a strong lens does not determine the mass profile uniquely, and breaking the degeneracy requires something else — a velocity dispersion for the lens galaxy, or weak lensing further out, or a second source at a different redshift.

The distances are cosmological. Every mass here is proportional to a ratio of angular-diameter distances, which are computed from redshifts through a cosmological model. A lensing mass is therefore never a purely local measurement, and the time-delay route to the Hubble constant is the same dependence read in the other direction.

And the geometry drawn here is not to scale, by an enormous factor. The bending angles are arcseconds; the figure draws them at tens of degrees. Nothing about the real configuration would be visible otherwise, and the caption says so, but the distortion hides how extraordinary the measurement is: a mass of 101210^{12} suns, ranged over dozens of kiloparsecs, announces itself as a one-arcsecond kink in the path of light that has travelled for billions of years.

A cluster weighed three ways, and its stars weighed once. A cluster of galaxies with a velocity dispersion of 1000 km/s, gas at 8 keV and a strong-lensing Einstein radius of 25 arcseconds, each turned into a mass inside 1.5 Mpc by its own relation and nothing else: 7.0, 8.9 and 15.2 × 10¹⁴ M☉. The three assume, respectively, that the galaxies are in equilibrium, that the gas is, and nothing whatever — so their agreement to within a factor of 2.2 is not three restatements of one assumption. The lensing bar is the loosest of the three and is drawn that way deliberately: it measures the mass inside a cylinder of radius 109 kpc, 1.11×10¹⁴ M☉, and carrying that out to 1.5 Mpc as though it were a sphere overstates it. The stars are 2.9 per cent of it.
Fig. 7 The comparison the independence buys. Three masses for one cluster, each from its own relation: the galaxies’ speeds, the gas temperature, and the bending of light. The first two assume equilibrium and the third assumes nothing about it, and they agree to within a factor of two — which is the check that no single method can perform on itself. The stars are about a per cent of the total, and that number is the one Zwicky reported in 1933 and nobody acted on for forty years.

The generalisation

The move is one that recurs whenever a measurement is hard: find a probe that shares none of the assumptions of the existing method, even if it is much harder to obtain.

The value of such a probe is not that it is better. A lensing mass is less precise than a rotation curve and vastly more expensive to obtain. Its value is that it fails differently — the systematic errors of a lensing mass have nothing in common with those of a dynamical one, so where the two agree, both are confirmed in a way that neither could confirm itself.

The same logic runs through the three routes to a cluster’s mass, through the comparison of transit and radial-velocity masses for a planet, and through every case in this collection where a number is quoted twice. The habit worth taking from it is to ask, of any agreement between measurements, what they have in common — because an agreement between two applications of the same assumption is not a check on anything.

One more reading covers the range of source offsets a survey actually finds.

Two images, always, and where they sit. The positions of the two images of a point source against the source's true offset, both in units of the Einstein radius, from θ = ½(β ± √(β²+4)). The images always straddle the lens: one outside the Einstein radius and one inside it, on opposite sides. As the source approaches alignment the two converge on the ring from either side and both brighten without limit; as it moves away the outer image tends to the source's own position and the inner one collapses onto the lens, faint and usually unobservable. Nothing about this depends on what the lens is made of.
Fig. 8 Image positions for source offsets from nearly aligned to well outside the Einstein radius. There are always exactly two images, their separation barely changes, and their brightness ratio runs over orders of magnitude — so what a survey selects on is the ratio rather than the separation.

Where the ladder goes next

The next rung is weak lensing: the same deflection, far too small to produce multiple images, measured statistically as a slight preferential alignment in the shapes of thousands of background galaxies. It maps a mass distribution rather than measuring one number, and it reaches radii where no other method works.

Later rungs on this anchor: the mass-sheet degeneracy and how it is broken; time-delay cosmography and the Hubble constant it gives; cluster lenses used as telescopes for the earliest galaxies; substructure in lenses as a probe of dark matter that is not smooth; the Bullet Cluster and what a collision separates; and lensing of the microwave background, which is the same effect measured against the most distant source there is.

What this makes readable

Essays that name this one as a prerequisite.

What links here

The 8 of 12 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

ConvergenceCritical surface densityDark matterDeflection angleEinstein radiusGravitational lensingImage multiplicityMagnificationThe mass-sheet degeneracyTime-delay