A blur undone with the field that made it
Assumes Baryon acoustic oscillations, Large-scale structure and Perturbations.
A ruler measured along and across states the problem in one paragraph and then moves on. Galaxies move by of order ten megaparsecs over the age of the universe, which is a few per cent of the acoustic scale, so the feature is blurred by that amount rather than shifted by it — and to second order it is shifted, by about a third of a per cent, which is comparable with a modern survey’s statistical error — and which no amount of survey volume removes.
Both effects are fixed by the same operation, and the operation is a good deal more surprising than the problem.
Why the feature blurs rather than moves
A pair of galaxies separated by the acoustic scale is not a rod. It is two members of an excess probability, and each of them has been moved since recombination by the gravitational pull of everything around it.
If both moved the same way, the separation would be unchanged and nothing would happen. They do not, because the field that moves them varies on scales smaller than 150 megaparsecs — so the two displacements are drawn from a distribution with a substantial variance and only partial correlation.
What matters is the pairwise dispersion: the mean squared difference between the displacements of two points that far apart. For the real universe it is about eight megaparsecs, and convolving a feature of width twelve with a Gaussian of width eight gives one of width fifteen. The peak is lower, broader, and in the same place.
A symmetric blur does not move a symmetric peak — the same reason a homogeneous universe above a hundred megaparsecs can carry a sharp feature at all, which is why the leading effect is a loss of precision rather than a bias. The bias enters at second order, because the displacement field is correlated with the density: a pair at the acoustic separation is very slightly more likely to be pulled together than apart, so the peak creeps inward by about a third of a per cent at the present epoch.
The estimate, and why it is nearly self-consistent
The move that makes reconstruction possible is that the displacement is not an unknown. It is a function of the density field, and the density field is what a survey measures.
To leading order — the Zel’dovich approximation, which is the first term of Lagrangian perturbation theory — the displacement satisfies
so solving one Poisson equation on the observed density gives the displacement that produced it. Move every galaxy by and the field is returned approximately to its initial state, with the acoustic feature approximately as sharp as it was laid down.
The circularity is the interesting part. The density used to estimate the displacement is the displaced density, not the initial one — so the estimate is of the wrong thing, by an amount second order in the displacement. That is the same order as the effect being corrected, which ought to be fatal and is not: the correction removes the first-order blur completely and the second-order residual is second order in a quantity that was already small.
There is a second approximation and it is the one that limits the result in practice. The relation between density and displacement is linear, and the density field is not linear on small scales — so the estimate has to be smoothed, typically on ten to fifteen megaparsecs, before it is used. Whatever the smoothing removes is displacement that does not get undone.
The optimum nobody can compute
The smoothing scale is therefore a trade with a minimum, and the two sides of it are of different kinds.
Too coarse and the displacement is underestimated, the blur is only partly removed, and the gain in precision is small.
Too fine and the estimate is contaminated by shot noise, so galaxies are moved by amounts that have nothing to do with any real flow — which adds its own dispersion and can leave the feature worse than it started.
Where the minimum sits depends on the tracer’s number density, which varies across a survey, and on the galaxy bias, which has to be assumed. The scale used is chosen by running the whole procedure on simulated catalogues that match the survey’s density and selection, and picking the value that works — the same feed-the-pipeline-a-known-answer discipline that every step of this measurement rests on. The parameter is tuned on mocks rather than derived, and the resulting measurement inherits whatever the mocks got wrong about the small scales.
That is an uncomfortable dependency and it is bounded in a useful way. Reconstruction cannot move the peak’s true position, because it is a deterministic operation applied to the data and its effect on the estimator is measured in the same simulations. What a badly chosen smoothing costs is precision and a small residual bias, both of which are quantified by running the pipeline on mocks with the peak in a known place.
What the sharpening is worth
A number is available for the gain, and it is the reason the procedure is universal rather than optional.
The precision with which a feature’s position can be measured goes roughly as its width divided by the square root of the number of independent modes contributing to it. Narrowing the feature from fifteen megaparsecs to nine therefore improves the measurement by something like a third, at no cost in observing time.
That is equivalent to increasing the survey’s volume by a factor of two. For an instrument that took a decade to build and five years to observe, a factor of two in effective volume from a post-processing step is an enormous return, and it is the reason reconstruction went from a proposal to standard practice within a few years of being demonstrated.
The gain is not uniform across a survey. It depends on the tracer density, because the displacement estimate is noisier where galaxies are sparse, so the improvement is largest in the densest samples and smallest at the survey’s faint limit. Published analyses quote the improvement per sample, and it runs from nearly nothing for the sparsest quasar catalogues to about forty per cent for the densest galaxy ones.
A procedure whose benefit scales with the density of the data is a procedure that rewards depth over area, which is the opposite of what the volume argument says about the raw measurement — and the two together are why survey design is a genuine optimisation rather than a rule.
What it does to the shift
Sharpening is the advertised benefit and the removal of the second-order shift is the more important one.
The peak’s inward creep comes from the correlation between the displacement field and the density field — the same correlation reconstruction exploits. Undoing the displacement therefore undoes most of the shift, and what is left is well under a tenth of a per cent.
That matters because the shift is a bias and the blur is a variance. A blurred feature is measured less precisely and a shifted one is measured wrongly, and adding more survey volume fixes the first and not the second. As the surveys have grown, the statistical error on the acoustic scale has fallen through the size of the unreconstructed shift — so the correction that was optional in the first measurements is not optional now.
The other thing it removes
Reconstruction was invented for the blur and it fixes something else as a side effect, and the side effect is worth as much.
Every galaxy’s radial coordinate comes from a redshift, and a redshift contains the galaxy’s own motion. On large scales those motions are coherent — everything falls towards the overdensities — so structures are compressed along the line of sight by a factor that depends on the growth rate. That compression is a distortion of the same statistic the acoustic scale is measured in, and separating it from the geometric distortion is the central difficulty of measuring a ruler along and across.
The displacement field that reconstruction estimates is the same field that produces those velocities. So the standard procedure moves each galaxy by the estimated displacement and removes the estimated line-of-sight component of its velocity, which undoes most of the linear redshift-space distortion at the same time.
The result is a reconstructed field that is closer to isotropic, in which the acoustic feature is nearly a sphere again, and in which the geometric anisotropy the Alcock–Paczyński test is looking for is a larger fraction of what remains.
One operation, estimated from one field, fixing three things: a blur, a bias, and a distortion. They are three consequences of one displacement, which is why one correction addresses all of them, and it is the cleanest example in observational cosmology of a nuisance being removed by modelling rather than by marginalising over it.
What was actually measured
The input is a catalogue of positions and redshifts and a model of the survey’s selection, and the output is the same catalogue with every galaxy moved.
Four ingredients go into that move and none is observed.
The galaxy bias. The density field the displacement is computed from is the galaxy density, and galaxies are not matter. Dividing by a bias factor is the first step, and the bias is fitted from the same data at larger scales.
The growth rate. Removing the velocity component needs the ratio of the growth rate to the bias, which is fitted alongside.
The survey’s boundary. The Poisson equation has to be solved in a region with edges, and the displacement near an edge depends on the density outside it. The standard treatment fills the exterior with a random catalogue, which is a choice with a systematic attached.
And the smoothing scale, as above.
What is published is a distance, and the check that the reconstruction has not invented it is the same one every acoustic measurement rests on: the whole pipeline is run on simulated catalogues whose true acoustic scale is known, and the recovered value is compared with it. The recovered value agrees to well under a per cent, and that agreement is the measurement’s actual warranty.
A field that is not a field
There is a conceptual objection to the whole procedure that is worth taking seriously, because the answer to it says what reconstruction is actually doing.
Reconstruction moves galaxies. Galaxies are objects with histories, and a galaxy that has been moved to where its progenitor material was is not in a place it was ever in — the material that became it was spread over a region tens of times its own size, and “where it came from” is an average rather than a position.
So the reconstructed catalogue is not a picture of the early universe. It is a catalogue whose two-point statistic is closer to the early universe’s than the observed one was, which is a much weaker and entirely sufficient claim.
That distinction matters twice. It says why the procedure can be applied without worrying about what the objects are: nothing in it is a statement about an individual galaxy, and the errors made about individual galaxies average away in a statistic over millions of pairs. And it says where the procedure stops being useful: any measurement that is not a two-point statistic — a count of clusters, a distribution of voids, a measurement of a galaxy’s environment — cannot use a reconstructed catalogue at all, because those are statements about objects.
An operation valid for one statistic is not valid for a catalogue, and the reconstructed positions are properly thought of as an intermediate quantity inside an estimator rather than as data. Published analyses treat them that way: the reconstructed catalogue is generated, the correlation function is measured from it, and the catalogue is discarded.
Where the model stops
The figures have no noise. Every curve here is an idealisation in which a finer smoothing is always better, and the whole practical difficulty is that it is not. What the figures establish is the shape of the effect and the direction of the cure, not the size of the gain.
The Zel’dovich approximation is first order. Second-order Lagrangian perturbation theory gives a better displacement and is used, and it does not converge everywhere either — inside collapsed structures no perturbative displacement exists, because the mapping from initial to final position is no longer one-to-one.
Shell crossing is the real boundary. Once two streams of matter have passed through each other, the position a galaxy came from is not a function of where it is now. Reconstruction cannot undo that, and no refinement of the algorithm will.
And the acoustic scale itself is not reconstructed. What is recovered is the sharpness of the feature in the observed statistic. The comoving length the feature represents was fixed at recombination, and nothing in this essay measures it — it is computed from the pre-recombination physics and imported, which is what makes it a standard length rather than a measured one.
The order in which a survey is analysed
Reconstruction sits in the middle of a pipeline whose order is not obvious, and the order carries information about what the procedure is.
First the catalogue is built and its selection function estimated, because the density field cannot be computed without knowing where the survey looked. Then a random catalogue is generated with the same selection, because the density contrast is defined against it. Then the density field is smoothed, divided by a bias, and the displacement solved for. Then every galaxy and every random point is moved — the randoms by the displacement alone and the galaxies by the displacement plus the velocity correction — and the correlation function is computed from the moved catalogues.
Moving the randoms is the step that looks wrong and is essential. The correlation function is a comparison between the galaxy distribution and a distribution with no clustering in it, and if the galaxies are moved and the randoms are not, the comparison acquires an artefact with the shape of the displacement field. The standard treatment moves both, in different ways, and which points get which treatment is the part of the recipe that took longest to settle.
A procedure applied to data has to be applied to whatever the data are compared against, which is obvious once stated and was not obvious when the technique was new — the first implementations moved only the galaxies, and the resulting bias was found in simulations rather than predicted.
The generalisation
The structure worth carrying is that a corruption whose cause is measurable is a corruption that can be removed rather than merely characterised, and that the two are different in kind.
Most systematics are handled by modelling their effect and marginalising over their parameters, which costs precision in proportion to how badly the parameters are known. This one is handled by measuring the cause — the displacement field, from the density field it produced — and undoing it. What remains is not an uncertainty in a nuisance parameter but a residual of the operation itself, and it is second order.
The question to ask of any systematic is whether its cause leaves a signature in the same data. If it does, the correction can be deterministic; if it does not, the best available is a prior. Here the flows that blurred the feature are the flows that built the density field, so the data contain their own repair — and the reason that works is not cleverness but that one field is responsible for both.
There is a second reading and it is about circularity. Estimating a correction from the corrupted data looks like a mistake and is not, because the error made by doing so is of higher order than the correction itself. A self-referential estimate is legitimate when its own error is smaller than what it fixes, and establishing that is a calculation rather than an intuition — which is why the procedure was proposed, tested on simulations, and adopted, in that order.
Still open: the tracer beyond the reach of galaxies
What comes next changes what is being counted. Beyond redshift one a galaxy survey runs out of galaxies bright enough to get redshifts for, and the only tracer left is the absorption in the spectra of quasars further away still. That makes a three-dimensional map out of one-dimensional pencils, with a sampling so anisotropic that the balance between the two acoustic observables inverts.
Beside it lies the void–galaxy cross-correlation, in which the same geometric test is applied to underdense regions rather than to the feature — where the distortion is larger, the velocities are outflows rather than infalls, and the reconstruction described here has to be done in the opposite sense.
About the same objects
Not linked from either essay — found by the objects both name.
- A map stretched by the thing it measures correlation function · redshift-space distortion
- A map that is not of positions correlation function · redshift-space distortion
What links here
Essays that link to this one from their own argument.
- A ruler read in absorption cosmology
The objects this essay names
Each one links to every other essay that touches it.
Baryon acoustic oscillationsCorrelation functionDisplacement fieldNon-linear dampingPerturbation theoryReconstructionRedshift-space distortionSmoothing scaleStandard rulerZeldovich approximation