A distance measured with a stopwatch
Assumes Strong lensing, Hubble constant and Quasars.
The rung below this one turned an image geometry into a mass. The Einstein radius is an angle, the angle plus three distances is a mass inside a cylinder, and the inference assumes nothing about the lens being in equilibrium — which makes it the cleanest weighing in this collection.
The same geometry contains a second measurement, and it is of a different kind entirely. Light reaching an observer from a lensed source has taken several paths, and those paths are not the same length. If the source varies, the variation appears in one image and then in the other, and the interval between them is measured with a clock.
That interval is proportional to a distance. And because every distance in cosmology scales as , a quasar that flickers gives the Hubble constant in one step, with nothing calibrated beneath it.
Why the images are where they are
The arrival time along a path with image position and source position is
the first term being the extra geometric path length of a bent ray and the second the Shapiro delay from passing through the lens’s potential. The combination in brackets is the Fermat potential.
Fermat’s principle then does all the work. Light does not take the shortest path or the quickest one; it takes every path along which the arrival time is stationary. So the images are at the stationary points of , the number of images is the number of stationary points, and their types are fixed: minima, saddle points and — for a lens with a smooth core — a maximum, usually too demagnified to see.
Two consequences follow immediately.
A time delay is a length. The height difference between two stationary points is an interval, and the prefactor is with
which has dimensions of length and scales as while its dependence on the other cosmological parameters is weak.
And the delay is dominated by the potential, not by the path length. For a typical galaxy lens the two terms are comparable and partially cancel, so the delay is a difference of two larger quantities — which is why the lens model matters as much as the measured delay.
What has to be measured
Three things, and they are independent problems with independent failure modes.
The delay. The source has to vary, the variation has to be non-periodic, and the images have to be monitored for years. Quasar variability obliges: it is stochastic, of order tenths of a magnitude on timescales of months, and the same wiggle appears in each image. The lens model. Image positions, flux ratios and — where the host galaxy is lensed into an extended arc — thousands of surface-brightness pixels are fitted with a parameterised mass distribution. The delay depends on the difference of the Fermat potential between image positions, so what is needed is not the total mass but the shape of the potential between the images, and the constraint is strong exactly where the images are and weak elsewhere.
And the line of sight. Everything between the observer and the source contributes convergence and shear. A group of galaxies near the lens, or a chance overdensity halfway to the source, adds a mass sheet that the lens model cannot see.
The degeneracy
The mass-sheet transformation is the reason time-delay cosmography is difficult, and it is exact rather than approximate.
Take a lens model with convergence and replace it with
adding a uniform sheet of convergence and scaling the rest. Simultaneously rescale the source position, . Then:
- every image sits at the same place;
- every flux ratio is unchanged;
- every image shape is unchanged;
- and every time delay is multiplied by .
Since is inferred from the delay divided by the model’s predicted Fermat-potential difference, the inferred scales by . A ten per cent sheet is a ten per cent shift in the Hubble constant, and there is nothing in the imaging that can detect it.
The source’s true brightness is rescaled too, and that is the only observable that changes — but a quasar’s intrinsic luminosity is not known to a factor of two, let alone ten per cent.
Breaking it
Two things constrain the sheet, and both come from outside the lensing data.
Stellar kinematics of the deflector. The velocity dispersion of the stars in the lens galaxy measures the depth of its potential well by a route that has no lensing in it. A model with more mass in the smooth sheet and less in the galaxy predicts a lower dispersion, so a measured dispersion selects . Counting the line of sight. A photometric survey of the field, compared against similar random fields, estimates the external convergence statistically; the ray-tracing of a cosmological simulation converts a galaxy count into a distribution for . The result is a probability distribution of a few per cent width, not a number.
Neither breaks the degeneracy completely, and the residual is the dominant systematic. Published time-delay measurements of have moved by more than their quoted error bars as the treatment of the sheet changed — most visibly when a set of results near 74 km s⁻¹ Mpc⁻¹, obtained with a particular family of parameterised profiles, moved down by several units and gained much larger error bars once the profile family was made more flexible.
The other degeneracy, which is not exact
The mass sheet is the famous one because it is exact. There is a second, approximate degeneracy that in practice does more damage, and it concerns the radial profile of the lens rather than anything added to it.
Write the deflector’s convergence as a power law, , so that is isothermal. The Fermat-potential difference between two images depends on directly, and steeply: for a typical two-image configuration a change of 0.1 in the slope moves the inferred Hubble constant by several per cent. Since the observed range of slopes across elliptical galaxies is about 0.2 wide, the slope alone is worth a ten per cent systematic if it is assumed rather than measured.
What constrains it is the extended image of the source’s host galaxy. A quasar gives four points; its host, lensed into arcs or a ring, gives thousands of surface-brightness pixels spanning a range of radii, and fitting those determines the slope where the arcs are. That is why systems with a bright, well-resolved ring are worth far more than systems with four bright points and nothing else, and why the published sample is chosen on the host rather than on the quasar.
The trap is that a slope measured from arcs at one radius is being used to compute a potential difference between images at slightly different radii, so the constraint is local and the extrapolation is short but not zero. And the slope and the mass sheet are themselves correlated: a family of models with a free slope can absorb part of a sheet, and a family with a fixed slope cannot, which is exactly why results obtained with a rigid profile family had smaller error bars and moved when the family was widened.
A small error bar obtained by assuming a shape is not a measurement of anything but the shape. The history of this method’s published values is a history of that lesson being applied.
What was actually measured
Some half-dozen to a dozen lensed quasars have delays good enough to use, and the campaigns behind them are long.
RXJ1131−1231 is the standard example: a quadruply imaged quasar at redshift 0.658 behind a galaxy at 0.295, with delays measured from years of nightly monitoring on metre-class telescopes. The longest delay is around 90 days and is measured to about 1.5 per cent. Its Einstein radius is 1.6 arcseconds; the host galaxy is lensed into a nearly complete ring, which supplies thousands of constraints on the potential; and the deflector’s velocity dispersion has been measured spatially resolved.
For an individual system the uncertainty is six to eight per cent, of which roughly half is the delay and half the model. Combining several systems reaches two to three per cent — which puts the method in the same class as the two established ones, and lands it in the middle of a disagreement.
The arithmetic, on one system
It is worth doing once, because the chain is short enough to follow and the places where a model enters are then visible.
A lens at and a source at , in a fiducial cosmology, have angular-diameter distances of roughly 900, 1420 and 780 megaparsecs for , and . The time-delay distance is then
at the fiducial . For an isothermal lens the Fermat-potential difference between the two images of a source offset by is , and with and that is steradians — dimensionless, and tiny.
Multiply: the delay comes out around eighty days. Measure eighty-five instead, and falls by six per cent; measure seventy-five, and it rises by the same.
Every one of the three distances scales as and their combination scales as too, which is the whole reason the method works: the ratios that remain depend on the matter density and the equation of state only weakly, so a measured delay is very nearly a measurement of one number.
What the arithmetic hides is that — the source’s true, unlensed position — is not observable. It is an output of the lens model, and it is exactly the quantity the mass-sheet transformation rescales.
Measuring the delay is its own problem
The delay was listed above as one of three measurements, and it is worth saying what makes it hard, because it is not photometric precision.
Two light curves of the same quasar are not translates of one another. Each image is independently microlensed by the stars in the deflector, which adds a slow, smooth, uncorrelated trend to each — of the same order as the intrinsic variability and on a similar timescale. So the estimator has to find the shift that best aligns two curves which differ by more than noise, and the extra difference is not white.
Every technique for this makes an assumption about the microlensing. Fitting a spline to the intrinsic variability and separate low-order polynomials to each image’s extrinsic trend is the standard approach; a dispersion-minimisation estimator that never models the intrinsic curve at all is the standard alternative. On the same photometry the two have disagreed by more than their own uncertainties, and the difference propagates directly into .
The response has been to test the estimators on simulated curves whose true delay is known, generated blind and analysed by teams who were not told the answer. Those exercises found real biases — several methods returned delays systematically short, and several reported uncertainties too small by a factor of two — and the published delays now come from methods that survived them.
The same discipline was extended to the cosmology. Recent analyses have been performed blinded, with the inferred value of hidden from the people making the modelling choices until the analysis was frozen, precisely because the answer sits inside a contested five-sigma disagreement and every modelling choice has a known direction of effect. In a measurement whose systematic is a choice of model family, blinding is not a courtesy; it is the only thing standing between an estimate and the value the analyst expects.
It is worth being clear about what blinding can and cannot protect. It removes the pull of a known answer on choices made after the data are in hand — which family of profiles to fit, which pulsar of the line-of-sight galaxies to include as an explicit perturber, where to cut a sample. It does nothing about a choice made before, and nothing at all about a systematic shared by every analysis in the field, which is what a common software pipeline or a common convergence prescription supplies. The measurements that moved most when the profile family was widened were blinded measurements.
The remaining defence is the one the rest of this collection keeps recommending: arrange for the same quantity to be measured by a route with different weaknesses. Within lensing that means a lensed supernova, whose unlensed brightness is known and which therefore breaks the mass sheet outright rather than constraining it. Outside lensing it means the ladder and the microwave background, which is the disagreement this method was brought in to arbitrate — so the method cannot be validated against the thing it is meant to settle, and its credibility has to be built entirely out of internal checks. That is an uncomfortable position and it is the honest description of where time-delay cosmography currently stands.
Where the model stops
Microlensing by stars in the deflector perturbs each image’s brightness independently on timescales of months — the same physics as a fold in a light curve, acting on a quasar rather than a star. That adds uncorrelated variability to each light curve and is the main reason delay measurements need years rather than months.
The source is not a point. Different parts of an accretion disc are lensed slightly differently and vary on different timescales, so a delay measured in the optical and one measured in X-rays need not be the same. The difference is small and is not zero.
The lens is not smooth. Substructure — dwarf satellites of the deflector — perturbs flux ratios strongly, which is one of the few probes of mass that emits nothing and delays weakly, which is fortunate; but the same substructure biases the recovered profile slope if it is not modelled.
And the picture cannot show the arrival-time surface. It is a construction. What exists observationally is two or four points of light and two or four light curves, and everything about the surface between them is inferred from a model that is constrained where the images are and assumed everywhere else.
Three of the construction’s ingredients are worth reading at other values, because each is fixed by a different observation.
The generalisation
The measurement’s power and its weakness come from the same source: it converts a time into a length using nothing but the speed of light.
That is the most direct conversion available anywhere in astronomy, and it appears wherever a geometry can be timed. A reverberation lag is a radius, , for a source nothing can resolve. A round-trip light time is a range, to a metre, for a spacecraft nothing can image. A supernova’s light echo dates the geometry of the dust around it. In each case a clock replaces a ruler, and no calibration stands between them.
The weakness is equally general. A timing measurement gives a path difference, and a path difference depends on the geometry along the whole path. Where the geometry is known — a spacecraft’s, a shell of dust — the conversion is clean. Where the geometry is the thing being modelled, as here, the clock’s precision is not the limiting factor, and the answer is only as good as the description of the space the light crossed.
And the deflection law itself at a smaller mass, which is the regime the same physics is tested in.
Where this ladder goes next
Later rungs on this anchor: lensed supernovae, whose delays are measured from a light curve of known shape rather than from a stochastic one and whose first example was found in 2014; the standardisable lensed Type Ia, which breaks the mass-sheet degeneracy outright because its unlensed brightness is known; time-delay measurements at radio wavelengths, where microlensing is absent because the source is larger; spatially resolved kinematics of deflectors from integral-field spectroscopy, which is the current front on the degeneracy; and the double source plane systems, in which two sources at different redshifts behind one lens give a ratio of distances that constrains the cosmology with no in it at all.
About the same objects
Not linked from either essay — found by the objects both name.
- A one-per-cent distortion, and a million galaxies to see it convergence · the mass-sheet degeneracy
What links here
Essays that link to this one from their own argument.
The objects this essay names
Each one links to every other essay that touches it.
ConvergenceFermat-potentialHubble constantLens modelLine of sight structureThe mass-sheet degeneracyQuasar variabilityStationary pointStrong lensingTime-delayTime-delay distanceVelocity dispersion