Cosmology

A forest with no continuum left

Measuring how much neutral hydrogen sits between here and a distant quasar means measuring the fraction of its light that survives, which means knowing how much light there was. At high redshift nothing survives at the wavelengths that would show it, so the level is extrapolated across the region being measured — and the optical depth is the logarithm of a number divided by a guess.

Assumes Reionisation and Line formation.

Light from a distant quasar passes through the intergalactic medium, and every parcel of neutral hydrogen along the way absorbs at its own redshifted Lyman-alpha wavelength. The result is a forest of absorption lines blueward of the quasar’s own emission, and the amount of light removed measures how much neutral gas there is.

The measurement is a fraction: the observed flux divided by what the quasar would have emitted. The numerator is measured well. The denominator is not observed at all.

A 5 per cent continuum error, and an optical depth wrong by 1.1. The mean transmitted flux of the Lyman-alpha forest against redshift, with a 5 per cent uncertainty in the quasar continuum drawn as a band. The continuum is not observed: at these redshifts every part of the spectrum blueward of the emission line is absorbed, so the level has to be extrapolated from the red side across a region where the quasar's own spectrum has structure. A fractional error in that level is a fractional error in the flux, and since the optical depth is minus the logarithm of the flux, the resulting error in the optical depth is the fractional error divided by the flux — which grows without bound as the forest goes black. At z = 2 it is 0.06; at z = 6.2 it is 1.1. That is why measurements of when reionisation ended are quoted as limits rather than values above about redshift six.
Fig. 1 What that does. The mean transmitted flux falls steeply with redshift as the forest thickens, and a five per cent uncertainty in the continuum is a five per cent uncertainty in the flux — which becomes an error in the optical depth of the fractional error divided by the flux. At redshift three that is a few hundredths; at redshift six the flux is a few per cent and the same continuum error is an optical depth wrong by more than one.

The difficulty is not a small correction on an otherwise clean measurement. Above redshift five it is the dominant error, it grows without bound as the forest saturates, and it is the reason a quantity that could in principle be measured to a per cent is quoted as a limit.

Why the level is not observable

At low redshift the forest is thin and there are stretches of spectrum between the lines where the quasar’s own continuum shows through. Drawing a curve through those stretches is straightforward and the result is good to a per cent or two.

At high redshift there are no such stretches. By redshift five the mean transmission is a few per cent and the absorption is continuous rather than discrete — every pixel is absorbed, and the highest points of the spectrum are not the continuum but the least-absorbed parts of the forest.

So the continuum has to be extrapolated from the red side of the Lyman-alpha emission line, where there is no forest, across the emission line itself, into the region being measured. That extrapolation spans a region where the quasar’s own spectrum has real structure: the broad emission line, the smaller lines around it, and a continuum slope that varies from object to object.

The extrapolation is the same operation as predicting a magnitude in a band the source was never observed in, and it fails in the same way: the prediction is only as good as the template’s applicability to this object, and the objects that most need it are the ones the template describes worst.

The standard method is to build a set of principal components from low-redshift quasars whose full spectra are visible, fit the components to the red side of a high-redshift quasar, and predict the blue side. The prediction is good to five to ten per cent, and that number is the systematic on every high-redshift forest measurement.

There is a second difficulty that compounds it and is worth naming separately: quasars are not a homogeneous population. Their continuum slopes, emission-line strengths and the relative prominence of their broad lines vary substantially from object to object, and the variation correlates with luminosity — which correlates with redshift in any flux-limited sample. So the template built from low-redshift quasars may not describe high-redshift ones, and the mismatch would be a systematic rather than a scatter, in a direction that depends on how the population evolves. That is the hardest part of the error to bound, because bounding it requires knowing the evolution the measurement is being used to study.

Why the error grows as the forest darkens

The optical depth is minus the logarithm of the transmitted fraction, so an error in the fraction propagates as

δτ=δFF,\delta\tau = \frac{\delta F}{F},

and FF is falling exponentially with redshift. A fixed fractional error in the continuum is a fixed fractional error in FF, so the error in τ\tau grows as one over the flux.

At redshift three, with F0.7F\approx0.7, a five per cent continuum error is δτ0.07\delta\tau\approx0.07 against τ0.35\tau\approx0.35 — twenty per cent, unpleasant and survivable. At redshift six, with F0.03F\approx0.03, the same error is δτ1.7\delta\tau\approx1.7 against τ3.5\tau\approx3.5 — half the measurement.

Beyond that the flux is consistent with zero and the optical depth has no upper bound at all. That is the reason measurements above redshift six are quoted as lower limits on the optical depth rather than as values.

A 10 per cent continuum error, and an optical depth wrong by 3.6. The mean transmitted flux of the Lyman-alpha forest against redshift, with a 10 per cent uncertainty in the quasar continuum drawn as a band. The continuum is not observed: at these redshifts every part of the spectrum blueward of the emission line is absorbed, so the level has to be extrapolated from the red side across a region where the quasar's own spectrum has structure. A fractional error in that level is a fractional error in the flux, and since the optical depth is minus the logarithm of the flux, the resulting error in the optical depth is the fractional error divided by the flux — which grows without bound as the forest goes black. At z = 2 it is 0.11; at z = 6.5 it is 3.6. That is why measurements of when reionisation ended are quoted as limits rather than values above about redshift six.
Fig. 2 The same relation with a ten per cent continuum uncertainty, which is realistic for the faintest and most distant quasars where the red side is poorly measured. Every conclusion is doubled in severity: the optical depth is uncertain by a tenth at redshift three and by more than three at redshift six. The technique’s precision is set entirely by how well a quasar’s unabsorbed spectrum can be predicted from the part of it that is not absorbed.
A 3 per cent continuum error, and an optical depth wrong by 0.3. The mean transmitted flux of the Lyman-alpha forest against redshift, with a 3 per cent uncertainty in the quasar continuum drawn as a band. The continuum is not observed: at these redshifts every part of the spectrum blueward of the emission line is absorbed, so the level has to be extrapolated from the red side across a region where the quasar's own spectrum has structure. A fractional error in that level is a fractional error in the flux, and since the optical depth is minus the logarithm of the flux, the resulting error in the optical depth is the fractional error divided by the flux — which grows without bound as the forest goes black. At z = 2 it is 0.03; at z = 5.5 it is 0.3. That is why measurements of when reionisation ended are quoted as limits rather than values above about redshift six.
Fig. 3 The regime where the technique is a measurement rather than a limit: below redshift five and a half, with a three per cent continuum for a bright, well-observed quasar. There the transmitted flux is tens of per cent, the optical depth is uncertain by a few hundredths, and the evolution of the ultraviolet background is measurable. Everything quantitative that the forest has contributed to cosmology comes from this part of the diagram, and it is the part where the continuum is nearly visible.

What the measurement is trying to say

The reason the optical depth matters is that it constrains the neutral fraction of the intergalactic medium, and the neutral fraction is what reionisation is about.

The relation is unhelpfully steep in the other direction. Because the Lyman-alpha cross-section is enormous, an intergalactic medium with a neutral fraction of even 10410^{-4} is completely opaque — so a saturated forest says only that the neutral fraction exceeds about that, which is a very weak statement. A trough proves the forest survived rather than proving reionisation had not happened.

So the measurement’s useful regime is narrow: below about redshift five and a half the forest is transmitting and the optical depth is a measurement, and above about six it is saturated and only limits are available. The interesting epoch is exactly at the boundary.

A neutral fraction of 2.4·10⁻⁶ is already opaque. The Gunn–Peterson optical depth against the neutral fraction of the intergalactic medium, at z = 3, 5, 6.3, for Ω_b = 0.0493 and h = 0.674. Note the range of the vertical axis. A fully neutral medium at z = 6.3 gives τ = 4.1·10⁵, which is not absorption but extinction of everything; the medium reaches τ = 1 — the point at which it stops transmitting most of the light — at neutral fractions of 6.1·10⁻⁶ at z = 3, 3.3·10⁻⁶ at z = 5, 2.4·10⁻⁶ at z = 6.3. That is why the argument runs from the flux that survives rather than from the flux that does not. A spectrum showing any transmission at all between Lyman α and Lyman β is a measurement that the medium is ionised to better than one part in 164,727, and no fit to any absorption line is needed to establish it.
Fig. 4 The relation the optical depth is being used to invert: the Gunn–Peterson optical depth against the neutral fraction, which is so steep that the forest saturates at a neutral fraction of one part in ten thousand. The steepness is the reason the technique is sensitive to the very end of reionisation and blind to everything before it, and it is also the reason a continuum error at the saturated end costs so much — the inversion divides by a derivative that has gone to nothing.

There is a further use of the forest that is much less sensitive to the continuum and is worth naming because it is now the main one: the flux power spectrum. The fluctuations of the transmitted flux along a line of sight trace the density fluctuations of the intergalactic gas, and their power spectrum measures the matter power spectrum on scales far smaller than any galaxy survey reaches. A multiplicative continuum error changes the mean flux and largely divides out of the fluctuations’ shape, so the power spectrum is a far more robust observable than the mean. It is the forest’s main contribution to cosmology today — a constraint on the neutrino mass and on the small-scale power — and it is robust for exactly the reason the mean flux is not.

There is a third robust statistic worth naming, and it is the one that turned the saturated regime from useless into informative: the length distribution of dark gaps. Where the flux is consistent with zero the optical depth is unmeasurable, and how far a dark stretch extends is still measurable — it is a length on the sky and needs no photometric calibration at all. Long dark gaps require large regions of enhanced neutral fraction or suppressed ionising background, and their distribution constrains the patchiness directly. That statistic uses only where the flux crosses a threshold, so a continuum error shifts the threshold slightly and changes the gap lengths hardly at all. A trough that proves the forest survived becomes a measurement rather than a limit when it is the trough’s length being measured.

The continuum comes back through the normalisation

The flux power spectrum is described above as robust because a multiplicative continuum error divides out of the fluctuations’ shape. That is true of the shape and it is not the whole story, and the way the error returns is worth following because it is how a systematic survives a statistic designed to remove it.

The forest’s power spectrum is not interpreted directly. It is compared against hydrodynamic simulations of the intergalactic medium, and those simulations have a free parameter that nothing predicts: the intensity of the ultraviolet background, which sets how ionised the gas is. The standard procedure is to fix it by rescaling the simulated optical depths until the simulated mean flux matches the observed one.

So the observed mean flux — the quantity the continuum error lives in — enters the interpretation as the normalisation of the model the power spectrum is compared against. A continuum error that changed the mean flux by five per cent changes the rescaling, which changes the simulated fluctuation amplitude, which changes the inferred amplitude of the matter power spectrum. The statistic is robust and the inference is not.

The size of the leak is modest and it is not negligible against what the forest is used to constrain. A neutrino mass of a tenth of an electron volt suppresses small-scale power by a few per cent, which is the same order as the mean-flux systematic — so the forest’s neutrino limit carries a continuum error inside it, arriving three steps from where it was introduced. An amplitude and a depth that arrive multiplied is the same structure in the microwave background: a nuisance parameter that multiplies the quantity of interest and is fixed by a separate, weaker measurement.

Two further leaks travel with it. The gas temperature enters the small-scale power through thermal broadening, and it is constrained largely by the shape of the same power spectrum, so temperature and amplitude trade against each other. And the quasar sample is flux-limited, so the objects measured at high redshift are the most luminous ones — which have systematically different continua and larger proximity zones than the population as a whole. A sample brighter than the population it came from is not usually thought of as a spectroscopic problem, and here it decides which template applies.

The general lesson is that a statistic immune to a systematic is not the same as an inference immune to it. The immunity has to hold all the way through to the number being quoted, and a model normalised on the contaminated quantity gives it a way back in.

What was actually measured

There are three, and the third is the one that changed the picture.

The evolution of the mean flux. Measured over redshifts two to five with continuum errors under five per cent, the mean transmitted flux falls smoothly and is well described by a slowly evolving photoionisation rate. That is the regime where the method works, and it constrains the intensity of the ultraviolet background over three billion years.

The first complete troughs. Quasars beyond redshift six show stretches with no detected flux at all, extending over hundreds of megaparsecs. That was initially read as the end of reionisation, and it is consistent with a neutral fraction anywhere above 10410^{-4} — which includes fully ionised.

And the scatter between sight lines. Different quasars at the same redshift show very different amounts of transmission, far more than density fluctuations alone predict. That scatter is now the primary evidence about the end of reionisation: it indicates that the ultraviolet background is patchy, which is what a process ending at different times in different places produces. The scatter is a ratio between sight lines, so the continuum errors partly cancel — which is why it has become the preferred statistic.

Bubbles that meet at z = 5.3, and a scattering depth of 0.047. The fraction of the volume of the universe filled by ionised bubbles, integrated from redshift 20 down to 4.5. The equation has two terms and no others: photons escaping from young galaxies open new volume, at a rate taken from the measured cosmic star formation history with an escape fraction of 0.2; recombinations inside the bubbles close it again, on a timescale that is one over the density times the recombination coefficient times a clumping factor of 3. Early on the density is high and recombination wins almost everything; the curve is nearly flat. As the universe expands the recombination time lengthens as the cube of one plus the redshift while the star formation rate is still rising, the balance tips, and the filling factor runs to one in under half a billion years. It reaches unity at redshift 5.31, which is overlap — the moment the bubbles meet and the last neutral walls between them disappear. The same integration gives an electron-scattering optical depth of 0.0465 for the microwave background, against the 0.054 that is measured, and that agreement is the check: the two observations constrain the same history from opposite ends, one fixing when it finished and the other how long it took.
Fig. 5 The picture the scatter supports: ionised bubbles growing around sources and merging, with the last neutral regions surviving between them. The observable consequence is that two lines of sight at the same redshift pass through different amounts of neutral gas, so their transmitted fluxes differ — and the variance between sight lines carries information that the mean does not, and carries it in a statistic much less sensitive to the continuum.
A Thomson depth of 0.054 puts the midpoint at z = 7.7. The neutral fraction of the intergalactic medium against redshift. The heavy curve is an ionisation history whose only constraint is the microwave background's Thomson optical depth of 0.054 — the fraction of CMB photons scattered on the way to us, measured from the polarisation at large angular scales and having nothing whatever to do with quasars. Requiring the integral ∫ σ_T n_e c dt/(1+z) to reproduce that number fixes the midpoint at z = 7.72. The forest says the same thing by an unrelated route: troughs appear in quasar spectra below about z = 6, and the Gunn–Peterson depth says a trough needs a neutral fraction above roughly 10⁻⁴, which is the shaded band. Two measurements — one of scattered photons across the whole sky, one of transmitted photons along a handful of sight lines — agree on when the medium was ionised, and neither could have been predicted from the other.
Fig. 6 The history the measurements are constraining: the neutral fraction against redshift, with the forest’s sensitivity confined to the very end of it. The forest saturates at a neutral fraction of one part in ten thousand, which on this diagram is essentially the bottom axis — so the entire epoch during which the universe went from neutral to ionised is invisible to it, and only the last moments are measurable. The technique that gave reionisation its name constrains almost none of it.

One more thing worth recording, because it is a rare case of the problem being solved by better data rather than by a better statistic. The templates used to predict a quasar’s unabsorbed spectrum are built from low-redshift quasars, and their quality is limited by how many such quasars have been observed with the same instruments and the same signal-to-noise. Large spectroscopic surveys have now measured hundreds of thousands, which has reduced the prediction error from ten per cent to five and, more usefully, has made the error’s distribution measurable — so an analysis can now marginalise over a realistic continuum uncertainty rather than adopting a single prediction. That is a modest improvement in a number and a substantial improvement in what can be claimed from it.

Where the picture stops

Three of them, and the second is the one that motivates the alternatives.

The continuum error is correlated across a spectrum. A principal-component prediction that is wrong is wrong smoothly, so the error does not average down over a spectrum’s length. It behaves like a multiplicative offset applied to the whole forest of one quasar, and combining many quasars reduces it only if the prediction errors are independent between objects — which they partly are and partly are not, since they share a template.

Saturation destroys the information rather than degrading it. Where the flux is consistent with zero there is nothing to correct. No improvement in the continuum recovers a measurement from a saturated region, and the only routes into the fully neutral epoch are different observables entirely — the damping wing of a quasar’s own absorption profile, the twenty-one centimetre line, or the microwave background’s optical depth.

And the quasar’s own neighbourhood is not typical. A quasar ionises its surroundings, so the forest immediately blueward of its emission line is more transmitting than average. That proximity zone has to be excluded, and how far to exclude depends on the quasar’s luminosity and lifetime, which are not known independently.

A fourth belongs with them, and it concerns the resolution of the spectrum. Everything above treats the transmitted flux as a smooth quantity, and it is not: at high resolution the forest resolves into individual lines with cores and wings, and the mean flux measured at low resolution is an average over structure the instrument did not resolve. Since the relation between flux and optical depth is nonlinear, an average of the flux is not the flux at the average optical depth — so a low-resolution measurement and a high-resolution one of the same sight line give different mean optical depths, by an amount that depends on the distribution of line strengths. That is a second denominator problem, arriving at the level of the instrument rather than the source.

Why a denominator nobody measured is a familiar problem

The general shape is one this collection has met in a different subfield, and the parallel is exact.

An equivalent width is an area measured against a continuum that is drawn rather than observed, and in a crowded spectrum the drawn level is systematically too low. Here the same operation is performed across an entire spectral region rather than around a line, and the level is extrapolated from outside rather than interpolated from within — which is worse, because an extrapolation has no local check.

The remedies are the same in both cases and they are worth stating together. Work differentially, comparing objects that share the systematic: the sight-line-to-sight-line scatter here, and differential abundance analysis there. Move to a statistic that does not need the level: the flux power spectrum, whose shape is nearly independent of a multiplicative continuum error. And find a different observable that measures the same physics without the denominator, which for reionisation means the damping wing and the twenty-one centimetre line.

What none of the remedies does is measure the continuum. It is not measurable, and every technique that has made progress here has made it by not needing to.

A final observation about how the field’s framing changed. For thirty years the question was “when did the forest first appear”, and the answer was pursued by finding higher-redshift quasars and measuring their troughs. The troughs arrived, they were consistent with everything above a neutral fraction of 10410^{-4}, and the question stopped being answerable that way. What replaced it was a set of questions about variance — how much sight lines differ, how large the ionised bubbles were, how patchy the background was — all of which are answerable with the same data and none of which needs the continuum to better than the scatter being measured. Reionisation ends when the walls meet, and measuring the meeting turned out to require statistics of the differences rather than a measurement of the mean.

It is worth noting one more thing the continuum problem has cost, since it is quantifiable. Published values of the mean optical depth at redshift five from different groups differ by twenty to thirty per cent, with individually quoted uncertainties of ten. The difference is almost entirely continuum methodology — which template set, which region of the red side was fitted, how the proximity zone was excluded. That spread between analyses is the field’s own measurement of the systematic, and it is larger than any single analysis’s error bar. The disagreement between analyses is the honest uncertainty wherever a reference has to be constructed rather than observed, and it is worth looking for before believing an error bar.

Where the ladder goes next

The rung above this one is the damping wing: the Lorentzian absorption profile a neutral medium imprints redward of a quasar’s Lyman-alpha line, which is unsaturated even when the forest is black and which is currently the best probe of the neutral fraction above redshift six. The rung beyond it is the twenty-one centimetre line — an emission or absorption signal from the neutral gas itself, requiring no background source and no continuum at all.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Continuum placementThe Gunn–Peterson troughThe Lyman-α forestMean transmitted fluxNeutral fractionOptical depthPrincipal component analysisQuasar spectrumSaturationUpper limit