A residual that is somebody else's velocity
Assumes Dark energy and Hubble constant.
A galaxy’s redshift has two contributions: the expansion of space between here and there, and the galaxy’s own motion through space at its own location. The first is what cosmology wants. The second is a few hundred kilometres a second and is not distinguishable from the first in any single measurement.
Far away the expansion dominates and the motion is negligible. Nearby it is not — and “nearby” extends further than intuition suggests.
The shift is not a Doppler shift — the cosmological part of a redshift is not a velocity at all — and yet the contaminant is, which is one of the reasons the two are so easily confused. What is measured is a single wavelength ratio containing both, and separating them requires a distance measured some other way.
The two components behave differently
The peculiar velocity field has two parts and they matter for different reasons.
A random dispersion of about 250 kilometres a second per galaxy, from the small-scale gravitational field. It is uncorrelated between widely separated galaxies, so averaging a sample reduces it as the square root of the number — it behaves like a measurement error and it is treated as one.
A coherent bulk flow of a few hundred kilometres a second, shared over volumes tens of megaparsecs across, produced by the gravitational pull of large-scale structure. This one does not average away over a sample within the flow, because every galaxy in it moves the same way.
That distinction decides everything about how the contamination is handled. The dispersion is a noise term to be added in quadrature; the flow is a systematic to be modelled or cut away.
It is worth putting a number on the coherence length, because it decides how large a survey has to be. The velocity field’s correlation length is set by the scale of the structures producing it, which is tens of megaparsecs — so galaxies within about thirty megaparsecs of each other share most of their peculiar velocity. A sample confined to a sphere of that radius therefore contains one independent measurement of the flow, not a thousand, however many galaxies it holds.
How a peculiar velocity is measured at all
Everything above treats the peculiar velocity field as though it were known. It is not measured directly, and the way it is measured explains why its amplitude is still disputed.
A galaxy’s peculiar velocity is the difference between its observed recession and the recession the expansion alone would give at its distance. So it requires a distance measured without using the redshift — a line width that is a distance for a spiral, the fundamental plane for an elliptical, a supernova where one has been seen. Those indicators carry a fractional distance error of fifteen to twenty per cent, and the recession velocity they are subtracted from grows with distance. So the absolute error on a peculiar velocity grows with distance too: twenty per cent of the recession at fifty megaparsecs is seven hundred kilometres a second, already larger than the signal.
That is the whole difficulty in one line. The quantity being measured is constant with distance and its error grows linearly, so peculiar velocities are measurable individually only within about thirty megaparsecs and statistically only within about a hundred. The bulk flows reported at two hundred megaparsecs and beyond are averages over samples whose individual errors are several times the claimed signal, and whether such an average is trustworthy depends entirely on whether the distance indicator has a distance-dependent bias — which is precisely the thing a fifteen-per-cent scatter makes hard to rule out.
There is a second circularity worth naming. The distance at which a galaxy’s peculiar velocity is evaluated is often the redshift distance, because that is the only distance available for most of the sample — so the correction is applied at a position that the correction itself would move. In practice the field is solved iteratively, or fitted in redshift space with the mapping built into the model. Neither is wrong, and both mean the recovered velocity field is not independent of the cosmology assumed while recovering it.
Why it sets a lower cut-off
Measuring the Hubble constant means fitting a slope to distance against velocity, and the fit’s precision is limited by the scatter. Nearby the scatter is peculiar velocities and far away it is the distance indicator’s own error, so there is an optimum range.
The usual choice is to discard everything below about forty megaparsecs. At that distance the recession is three thousand kilometres a second and a three-hundred-kilometre flow is ten per cent — which is already several times the precision being sought, and it is a systematic because the local flow has a definite direction and the nearby sample does not cover the sky uniformly.
Convert the numbers into the units the observation is actually made in, because the effect looks negligible in redshift and is not. Three hundred kilometres a second is a redshift of one part in a thousand. A supernova’s distance is good to about seven per cent, which at forty megaparsecs is three megaparsecs, and the recession changes by two hundred kilometres a second across that. So the peculiar velocity and the distance error are comparable at the cut — which is why the cut sits there rather than anywhere else, and why arguments about it are arguments about which of two errors one would rather carry.
The cut is a judgement rather than a derivation, and moving it changes the answer at the level of a per cent or so. That is small compared with the disagreement between the early- and late-universe determinations, and it is not negligible compared with the quoted uncertainties.
There is a second reason the cut cannot simply be pushed outward. The distance-ladder measurement of the Hubble constant works by calibrating a secondary indicator — supernovae — on galaxies close enough to contain a primary one, and the primary indicators reach only a few tens of megaparsecs. So the calibration sample and the peculiar-velocity-contaminated region are the same region, and the cut that removes the contamination removes the calibrators. The standard resolution is to calibrate on the nearby galaxies and fit on the distant ones, with the peculiar velocities corrected rather than cut in the calibration step — which puts the correction, and its uncertainty, into the calibration rather than into the fit.
What is done about it
Three approaches, and all three are in use.
Cut. Discard the nearby objects. Simple, defensible, and it discards the calibrators.
Correct. Use a map of the local density field, derived from a galaxy redshift survey, to predict each galaxy’s peculiar velocity and subtract it. That works — the predicted and observed velocities correlate well — and it imports the survey’s own systematics and an assumed relation between galaxies and mass.
Marginalise. Include the peculiar velocity field as a set of nuisance parameters with a prior from theory, and integrate over it. That is the most honest and it widens the error bar by the amount the flow is uncertain, which is the point.
There is a fourth approach that avoids the problem rather than treating it, and it is worth naming because it is what a modern analysis does: fit the whole diagram at once. Rather than measuring a Hubble constant from a nearby sample and a deceleration from a distant one, fit a cosmological model to the entire redshift range simultaneously, with the peculiar-velocity covariance included in the likelihood. That treats the correlated errors as what they are — a covariance between nearby objects — and it propagates them correctly into every parameter rather than into one. It is more work, it produces a slightly larger error bar, and it is the only version that is defensible when the errors are shared.
What was actually measured
There are three, and the second is the one that constrains the systematic.
The bulk flow of the local volume. Surveys of peculiar velocities out to a hundred megaparsecs find a coherent flow of a few hundred kilometres a second, in a direction consistent with the pull of the large structures beyond. The amplitude agrees, within uncertainties, with what the standard model predicts from the observed density field — which is a non-trivial check and also a weak one, because the predictions have a large cosmic variance.
The dependence of the fitted Hubble constant on the cut. Recomputing the local measurement with cuts from twenty to eighty megaparsecs changes the answer by about a per cent, with no systematic trend beyond about forty. That stability is the main argument that the local determination is not being driven by the flow, and it is quoted in every analysis.
And the convergence of the flow. The peculiar velocity field’s amplitude falls with the scale averaged over, and by two hundred megaparsecs it is a few tens of kilometres a second — under one per cent of the recession. Beyond that the Hubble flow is clean, and the remaining question for a distance-ladder measurement is entirely about the calibration rather than about the motions.
Could the motions explain the tension?
The question is unavoidable, because the local determination of the expansion rate disagrees with the early-universe one by more than its stated error, and the local one is the one contaminated by motions. So it is worth asking what a peculiar velocity would have to do to produce the gap.
The gap is about eight per cent. To manufacture it, the local calibration volume would have to be moving with respect to the more distant universe at eight per cent of its own recession — which at forty megaparsecs is two hundred and forty kilometres a second, and at a hundred megaparsecs is six hundred. That is not an absurd number; it is roughly the amplitude of the flows actually observed. What makes the explanation fail is the sign structure: a flow shifts the fitted rate up in one direction on the sky and down in the opposite one, so a sample covering both hemispheres averages it away, and the residual after averaging goes as the sample’s dipole asymmetry rather than as the flow itself.
The stronger version of the objection is a local void. If the volume within a few hundred megaparsecs were underdense by ten or twenty per cent, everything inside it would be accelerating outward relative to the mean, and the locally measured rate would exceed the global one by roughly the amount required — with no dipole, because a void is monopolar. That hypothesis is testable and has been tested: an underdensity of the required depth and extent would show in galaxy counts, in the number of clusters within the volume, and in the supernova residuals as a break at the void’s edge. The counts do show a mild local underdensity, of a few per cent over a couple of hundred megaparsecs, and a few per cent is not eight.
So the honest position is that the motions cannot currently account for the discrepancy and cannot be ruled out as contributing a part of it. That is a weaker statement than either side of the argument usually makes, and it is what the measurements support.
Where the picture stops
Three limits stand out, and the second is a genuine open question.
The correction depends on a density field that is itself a measurement. Predicting peculiar velocities requires a map of the mass, which is inferred from a map of the galaxies through a bias parameter that has to be fitted. So the correction carries a parameter degenerate with the amplitude of the flow, and the amplitude of the flow is what the correction is for.
Some measurements of the bulk flow are larger than the model allows. Several analyses have reported flows of five to nine hundred kilometres a second extending to a hundred megaparsecs and beyond, which is above the standard model’s prediction. Other analyses of overlapping data find smaller values consistent with it. The disagreement is between methods rather than between datasets, and it is unresolved.
And the observer is inside the flow. Everything is measured relative to a frame that is itself moving, and the correction to the microwave background frame is made using the dipole. That correction is exact for the solar system’s motion and does not remove the motion of the whole local volume with respect to the more distant universe — which is what a bulk flow is.
A fourth sits alongside them, and it concerns the sample’s coverage of the sky. A bulk flow has a direction, so its effect on a sample depends on where the sample is: a set of galaxies spread uniformly over the sky sees the flow as a dipole in the residuals, which averages to zero in the fitted slope; a set concentrated in one hemisphere sees a net offset, which does not. Since supernova samples are built from surveys with uneven sky coverage — avoiding the Galactic plane, favouring one hemisphere’s observing seasons — the flow’s contribution to the fitted Hubble constant depends on the sample’s geometry in a way that has to be computed for each compilation. The way a survey covers the sky is part of its data, and here it enters through the direction of a velocity field.
Why a correlated error is a different animal
The general point is worth pulling out because it decides how a sample size helps.
A random error shrinks as the square root of the sample; a correlated one does not shrink at all until the sample is larger than the correlation length. Peculiar velocities are correlated over tens of megaparsecs, so a survey of a thousand galaxies within thirty megaparsecs has the statistical power of a handful of independent measurements. Adding more galaxies inside the same volume buys almost nothing.
That is a general feature of any measurement contaminated by large-scale structure, and it is the same statement as cosmic variance: there are only so many independent volumes within reach, and no amount of observing increases the number. The optical depth’s cosmic-variance floor is the same limitation at the largest scales, and the parallax zero point is the same structure with a different cause — an error common to every object, which a million objects share equally.
The practical rule is short and is often ignored: before increasing a sample, ask whether the dominant error is shared. If it is, the observing time is better spent on a second, independent volume or a second method than on more objects in the same place.
End on what the contamination is worth, since this essay has treated it as an obstacle. The peculiar velocity field is not noise from the point of view of cosmology — it is a measurement of the gravitational field of large-scale structure, and comparing observed velocities against those predicted from the observed galaxy distribution measures how strongly matter clusters relative to galaxies and how fast structure grows. That comparison is one of the few tests of gravity on cosmological scales that does not go through the microwave background, and it uses exactly the quantity a Hubble-constant measurement is trying to discard. An expansion that was supposed to be slowing was found in the same diagram, and the same scatter that hid it at low redshift is a different experiment’s signal.
Where the ladder goes next
The rung above this one is the velocity field itself: how a map of galaxy positions is turned into a prediction of their motions, what the bias parameter is, and why the comparison between predicted and observed velocities is a cosmological measurement in its own right. Above it again sits the high-redshift end — where peculiar velocities are irrelevant and everything depends on whether the candles have been standardised the same way at both ends of the diagram.
What links here
Essays that link to this one from their own argument.
The objects this essay names
Each one links to every other essay that touches it.
Bulk flowCorrelated errorCosmic varianceHubble diagramLow redshift cutPeculiar velocityRedshiftSelection effectSupernova cosmologySystematic error