Starlight

A continuum that was never observed

An equivalent width is an area measured relative to the continuum, and the continuum is not in the data. It is drawn — a curve through the highest points of the spectrum — and in any spectrum with many weak lines those highest points are already below the true continuum, because the weak lines have eaten the gaps.

Assumes Line formation and Spectra.

The quantity that turns a spectral line into an abundance is its equivalent width: the width of a completely black rectangle that would remove the same amount of light. It is an area, it is independent of the instrument’s resolution, and it is the input to every classical abundance analysis.

An area has to be measured against something, and the something is the continuum — the flux the star would emit with no lines at all. That flux is not observed. It is inferred, by drawing a smooth curve through the parts of the spectrum that appear to have no lines in them.

A continuum drawn 4.3 per cent below the real one. A short stretch of spectrum with one strong line in it and 150 weak ones scattered across the same interval. The upper dashed line is the true continuum — the flux the star would emit with no lines at all — and it is not observable. The lower one is what a fit through the highest points of the spectrum returns, which is 4.3 per cent lower, because the weak lines have depressed the gaps between the strong ones. Measuring the strong line's equivalent width against the apparent continuum instead of the real one makes it 11.2 per cent too small. The error has a sign, it is worse in spectra with more lines, and it therefore correlates with metallicity — which is exactly the quantity being measured.
Fig. 1 Why that inference is biased. The spectrum drawn has one strong line and a hundred and fifty weak ones scattered across the same interval. The upper dashed line is the true continuum and it is not observable. The lower one is what a fit through the highest points returns, several per cent below — because the weak lines have depressed the gaps between the strong ones. Measuring the strong line against the apparent continuum makes its equivalent width too small.

The problem is old and it is not solved. It is the reason an equivalent width is not simply a line’s depth, the reason the most careful abundance work is done differentially rather than absolutely, and the reason two competent groups analysing the same spectrum of the same star routinely differ by a tenth of a dex.

The error has a sign

This is not a random error. Weak lines only ever remove flux, so the apparent continuum is only ever below the true one, so every equivalent width measured against it is only ever too small, so every abundance derived from it is only ever too low.

A systematic with a fixed sign is in some ways easier to handle than a random one — it can be corrected if its size is known — and in one important way it is worse: it does not average away over lines, over stars or over observers.

That last clause deserves emphasis. Two astronomers analysing the same spectrum with the same software will place the continuum slightly differently, and their answers will differ; that difference is a random error and it is what the quoted uncertainty usually reflects. But both will place it too low, by nearly the same amount, because the physical cause is in the spectrum rather than in either of them. The scatter between analysts is therefore a measurement of the random part and says nothing at all about the systematic — a distinction that is easy to state and is routinely lost when a published uncertainty is derived from the spread of independent reductions.

The size depends on how many weak lines there are, which depends on the temperature, the surface gravity and above all the metallicity. That last dependence is the awkward one.

The error grows with the quantity being measured. How far the apparent continuum sits below the true one, against the number of weak lines in the same stretch of spectrum. The relation is steep and it has an uncomfortable property: the number of lines is set by the metallicity, so a spectrum of a metal-rich star has a more depressed continuum than one of a metal-poor star, and the resulting error in the abundance therefore depends on the abundance. That is a systematic that cannot be calibrated out with a single number and does not cancel in a differential comparison unless the two stars have closely similar metallicities. It also means that the classical technique works best where it is least needed: for a metal-poor halo star the gaps really are empty, and for a metal-rich giant the continuum is a fiction.
Fig. 2 How far the apparent continuum sits below the true one, against the number of weak lines in the same stretch of spectrum. The number of lines is set by the metallicity, so a metal-rich star has a more depressed continuum than a metal-poor one, and the resulting error in the abundance depends on the abundance. The method works best where it is least needed: for a metal-poor halo star the gaps really are empty, and for a metal-rich giant the continuum is a fiction.

It is worth putting a number on it. A continuum placed three per cent low makes a weak line’s equivalent width about three per cent small, and on the linear part of the curve of growth that is a three per cent abundance error — 0.013 dex, which is small. The same continuum error applied to a saturated line, where the equivalent width depends on the abundance only as a square root of a logarithm, corresponds to an abundance error several times larger. And the error is common to every line measured in that region, so it does not reduce when many lines are averaged. Averaging twenty lines reduces the random scatter by a factor of four and leaves the continuum error exactly where it was.

Where the depressed regions are

The depression is not uniform across a spectrum, because lines are not uniformly distributed. They cluster where the atomic physics puts them: the blue is far more crowded than the red, because most of the strong transitions of the iron-peak elements lie there.

The consequence is that the continuum error is wavelength-dependent, and a wavelength-dependent multiplicative error in a spectrum is a reddening. A star analysed with a blue-heavy line list acquires a continuum systematically more depressed in the blue than in the red, and any quantity derived by comparing blue and red — a temperature from a colour, a spectral slope — inherits it.

A continuum drawn 0.0 per cent below the real one. A short stretch of spectrum with one strong line in it and 40 weak ones scattered across the same interval. The upper dashed line is the true continuum — the flux the star would emit with no lines at all — and it is not observable. The lower one is what a fit through the highest points of the spectrum returns, which is 0.0 per cent lower, because the weak lines have depressed the gaps between the strong ones. Measuring the strong line's equivalent width against the apparent continuum instead of the real one makes it 0.0 per cent too small. The error has a sign, it is worse in spectra with more lines, and it therefore correlates with metallicity — which is exactly the quantity being measured.
Fig. 3 The same construction in an uncrowded region. Forty weak lines instead of a hundred and fifty leaves genuine gaps, the fit lands within a per cent of the true continuum, and the equivalent width is nearly right. This is why the classical method works at all, and it is why every practitioner’s first instinct is to choose clean regions — a selection that is correct, that biases the sample of usable lines, and that is not always possible in the parts of the spectrum where the elements of interest have their lines.

There is a second geometric consequence, and it bites hardest in exactly the objects people most want to measure. A rapidly rotating star has broadened lines, so its weak lines overlap more and its continuum is depressed further; a cool star has more lines and molecular bands; a metal-rich star has more lines still. So the continuum error is worst for cool, metal-rich, rotating stars, and those are the ones whose abundances are of most interest for Galactic chemical evolution and for planet-host statistics. The method is at its worst where the science is.

What resolution does and does not fix

A natural response is to observe at higher resolution: if the weak lines can be resolved, the gaps between them can be found.

This helps and it does not solve the problem. Higher resolution reveals the individual weak lines and shows where the true gaps are, so the continuum can be placed better. But the number of lines per unit wavelength is a property of the star, not of the spectrograph, and beyond some density there are no gaps at any resolution — the lines overlap intrinsically, because they are broadened by thermal motion and by pressure to widths comparable to their spacing.

That limit is reached in the blue for metal-rich stars, and it is reached everywhere for cool giants, whose spectra are covered in molecular bands that have no continuum anywhere at optical wavelengths. For those objects the quantity that is fitted is not a continuum but a pseudo-continuum — an agreed reference level defined by a model rather than by the data — and every abundance is quoted relative to it.

The curve of growth. The equivalent width of an absorption line against the number of absorbers along the sight line, both logarithmic, computed by integrating a Voigt profile with damping parameter 0.005. Three regimes: the width grows in proportion to the abundance while the line is weak, then almost not at all for two decades once the core saturates, then as the square root once the damping wings dominate. A line measured in the middle stretch carries almost no information about the abundance, and most strong lines in a stellar spectrum are there.
Fig. 4 The relation the equivalent width is being fed into: how a line’s strength grows with the number of absorbing atoms. On the linear part the equivalent width is proportional to the abundance, so a five per cent error in the width is a five per cent error in the abundance. On the flat part the line is saturated and the same five per cent error corresponds to a much larger abundance error, because the curve is nearly horizontal. So a continuum error costs least on weak lines and most on strong ones — which is the reverse of the intuition that strong lines are better measured.
A continuum drawn 14.7 per cent below the real one. A short stretch of spectrum with one strong line in it and 300 weak ones scattered across the same interval. The upper dashed line is the true continuum — the flux the star would emit with no lines at all — and it is not observable. The lower one is what a fit through the highest points of the spectrum returns, which is 14.7 per cent lower, because the weak lines have depressed the gaps between the strong ones. Measuring the strong line's equivalent width against the apparent continuum instead of the real one makes it 40.6 per cent too small. The error has a sign, it is worse in spectra with more lines, and it therefore correlates with metallicity — which is exactly the quantity being measured.
Fig. 5 A crowded region: three hundred weak lines in the same interval, which is what the blue part of a metal-rich giant’s spectrum looks like. There are no gaps at all, the fitted level is far below the true continuum, and the strong line’s equivalent width comes out badly wrong. Nothing about the measurement is careless. There is simply no observational definition of the continuum available in this spectrum, and any number quoted for it is a number from a model.

There is a third response that is worth naming because it is what most large surveys now do: stop measuring equivalent widths at all. If the analysis is going to fit a synthetic spectrum to the observation anyway, then the continuum can be one of the fitted parameters — a low-order polynomial multiplying the synthetic spectrum, adjusted along with the abundances. That does not make the continuum observable; it makes it a nuisance parameter, marginalised over rather than measured, and the uncertainty it contributes appears honestly in the error bar instead of silently in the answer. The price is that the fit is only as good as the line list, because a missing line in the synthetic spectrum is absorbed into the fitted continuum and reappears as an abundance error somewhere else.

What was actually measured

Three lines of evidence establish the size of the effect.

Synthetic spectra with known input. Compute a spectrum from a model atmosphere with a known abundance, degrade it to a real instrument’s resolution and signal-to-noise, and hand it to the standard analysis. The recovered abundance is low, by an amount that grows with metallicity, and the size matches the continuum depression computed from the same line list.

The solar analysis, done twice. The Sun’s photospheric abundances have been measured for a century, and the modern values differ from the older ones by amounts that are large in astrophysical terms — the oxygen abundance fell by about 0.2 dex between 1998 and 2005. Several things changed at once, of which the three-dimensional model atmosphere was the most discussed, and the treatment of the continuum in crowded regions was another.

Differential analyses. Comparing two stars of nearly identical temperature, gravity and metallicity, line by line, and quoting only the difference in abundance, removes the continuum error almost exactly — because the two spectra have the same lines in the same places and the same depression. Such analyses reach precisions of a few thousandths of a dex, forty times better than absolute analyses of the same stars. That gap between the differential and absolute precision is the most direct measurement there is of how large the shared systematics are.

The error grows with the quantity being measured. How far the apparent continuum sits below the true one, against the number of weak lines in the same stretch of spectrum. The relation is steep and it has an uncomfortable property: the number of lines is set by the metallicity, so a spectrum of a metal-rich star has a more depressed continuum than one of a metal-poor star, and the resulting error in the abundance therefore depends on the abundance. That is a systematic that cannot be calibrated out with a single number and does not cancel in a differential comparison unless the two stars have closely similar metallicities. It also means that the classical technique works best where it is least needed: for a metal-poor halo star the gaps really are empty, and for a metal-rich giant the continuum is a fiction.
Fig. 6 The same relation for deeper weak lines, which is what a cooler star or a lower-resolution spectrograph produces. Everything shifts upward: at the same line count the depression is larger, so the effect depends on the instrument as well as on the star, and a comparison between two stars observed with different spectrographs carries a differential continuum error that no amount of care with either one removes.
Hydrogen is half ionised by 9,555 K. Ionisation fractions from the Saha equation at an electron pressure of 20 N/m², against temperature. Each species' stages sum to one at every point, which is checked rather than assumed. Hydrogen crosses half-ionised at 9,555 K and is 99% ionised by 12,527; calcium, whose first ionisation potential is 6.113 eV against hydrogen's 13.598, is already 99% singly ionised at 6,000 K, where hydrogen is 100.00% neutral. Those two curves are the reason the same photosphere shows strong Ca II and weak Balmer at one temperature and the reverse at another — with the abundances fixed and only the temperature moving. The partition functions are held at their ground-state weights, which is the usual first approximation and shifts these curves by a few hundred kelvin rather than by their shape.
Fig. 7 Why the continuum error propagates further than it looks. An abundance analysis does not stop at one line: it requires that the same element give the same abundance from lines of different strengths and from two ionisation stages, and those requirements are what fix the temperature and the gravity. A continuum error that changes the equivalent widths of weak and strong lines differently therefore changes the inferred temperature as well as the abundance, and the temperature then changes every other abundance in the analysis. One systematic in one measurement moves the whole solution.

There is also a fourth strand of evidence that is quieter and in some ways the most convincing. Abundances derived from lines in the crowded blue and from lines of the same element in the sparse red should agree, and for metal-rich stars analysed classically they do not: the blue lines give systematically lower abundances, by an amount that grows with metallicity. That is precisely the signature the continuum argument predicts, it is measurable within a single spectrum with no external reference, and it is available for free in any analysis that uses lines across a wide wavelength range. It is the crater-count trend of this subject — plot the answer against something it should not depend on, and see whether it does.

One more piece of evidence deserves a mention because it comes from outside spectroscopy altogether. Abundances derived from stellar spectra can be compared against abundances derived from the emission lines of the interstellar gas the stars formed out of, and against the abundances measured in meteorites for the Sun. Where those comparisons are possible they agree to a few tenths of a dex and not better, and the disagreements have the direction the continuum argument predicts. It is a weak constraint, because the other methods have systematics of their own, and it is the only external check the subject has.

Where the picture stops

There are three, and the third is where the field went.

The definition of the continuum is model-dependent for cool stars. For a spectrum with no line-free regions there is no observational definition at all, and different groups’ pseudo-continua differ. Abundances quoted for such stars are on a scale set by a choice, and comparing two of them requires knowing that the choice was the same.

Signal-to-noise interacts with the bias. Placing a continuum by eye or by fitting an upper envelope favours the high noise excursions as well as the true gaps, so at low signal-to-noise the continuum is placed too high and the bias reverses sign. There is a signal-to-noise at which the two effects cancel, it depends on the line density, and nobody observes there on purpose.

And the whole framework treats lines one at a time. The modern alternative is spectrum synthesis: compute a full synthetic spectrum including every line, and fit it to the observation over a whole region, with the continuum as one of the fitted parameters. That removes the need to identify a continuum by hand and replaces it with a need for a complete and accurate line list — which is a different problem with a different set of errors, and one that has been the limiting factor for the last two decades.

A fourth deserves stating because it is often assumed away. The whole construction supposes there is a continuum in the physical sense — a level the star would radiate with no lines — and for a real atmosphere that is itself an idealisation. The opacity that sets the continuum is not line-free: it includes the negative hydrogen ion, bound-free transitions, and the accumulated wings of vast numbers of distant lines, which is why an opacity is a mean dominated by where the gaps are. The “true continuum” used above is the flux computed with the line list switched off, which is a well-defined quantity in a model and not a well-defined quantity in a star.

Why the reference is the hard part

The recurring shape is one this collection meets in every field, and it is worth stating in its general form.

A measurement of a difference requires a reference, and the reference is usually harder to establish than the difference. An equivalent width is the difference between the observed flux and the continuum; the observed flux is measured to a fraction of a per cent and the continuum is drawn by hand. A radial velocity is the difference between an observed wavelength and a laboratory one; the observed wavelength is measured to a part in 10910^9 and the laboratory value is often known to a part in 10710^7. A solar radius from an eclipse is the difference between two edges, one of which is a mountain range.

Composition is read from what is missing, and what is missing is defined against what would have been there — which is the one part of the picture no telescope records.

In each case the effort naturally goes to the half being measured, because that is the half the instrument is pointed at, and the accuracy is set by the half nobody is looking at.

The escape, where it is available, is differential measurement: compare two objects that share the reference, and the reference cancels. That is why differential abundance analysis reaches thousandths of a dex while absolute analysis reaches tenths, and why the modern practice for the most demanding work — chemical tagging, the search for planet-hosting signatures, the abundance differences between binary companions — is entirely differential. It buys two orders of magnitude and it costs the ability to say what the abundance actually is.

A closing observation, and it is the one that makes the subject tractable rather than hopeless. Every difficulty in this essay is shared between spectra taken the same way of similar stars. The continuum depression depends on the line density, the line density depends on the star’s parameters, and two stars with the same parameters have the same depression. So the error is enormous in absolute terms and nearly zero in a comparison, and a field that needed absolute abundances would be stuck where it was in 1970 while a field that needs differences has advanced by two orders of magnitude. What was learned was not how to measure the continuum. It was which questions do not require it — and the second thing a spectrum says turned out to be answerable without ever answering the first.

A last practical note, because it is the part a reader can act on. Any published abundance carries three things that should be stated and often are not: which wavelength regions the lines came from, whether the analysis was absolute or differential, and how the continuum was placed. Two of those three decide the systematic, and none of them appears in the error bar. A tenth of a dex is the honest uncertainty on an absolute abundance from a classical analysis of a solar-type star, and a few thousandths is the honest uncertainty on a differential one between two stars of nearly identical parameters. The two numbers describe different measurements of different things, and the same word is used for both.

One practical consequence deserves stating for anybody comparing abundances across the literature. Because the depression depends on how crowded the spectrum is, it depends on the star’s temperature and metallicity — a cool, metal-rich star has far more weak lines per unit wavelength than a hot, metal-poor one, so the continuum is depressed more and the abundances are biased more. The bias is therefore not a constant offset that cancels in a comparison; it is a function of the very parameters being compared. That is why an abundance trend with temperature or with metallicity is the most suspect kind of result in this subject, and why the standard defence is a differential analysis against a star of nearly identical parameters, where the crowding is the same and the depression cancels. The defence works, it restricts the comparison to stars that resemble each other, and it is the reason the most precise abundances in the literature are for stars chosen to be solar twins rather than for stars chosen to be interesting.

It is also worth remembering that the whole difficulty is invisible in the data: a spectrum with a depressed continuum looks exactly like a spectrum with a correct one, because the depression is smooth and the eye has nothing to compare it against. That is what makes it a systematic rather than a source of scatter, and it is why the size of the effect is estimated from synthetic spectra rather than measured from real ones.

Where the ladder goes next

The next rung is spectrum synthesis: fitting a computed spectrum to the whole observed region rather than measuring lines one at a time, and what the completeness of a line list has to be for that to work. The rung after that is the pair of quantities the continuum placement is entangled with — the temperature and the gravity that no single diagnostic separates, both of which shift when the continuum moves.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Abundance analysisContinuum normalisationCurve of growthDifferential analysisEquivalent widthLine blanketingMetallicityPseudo continuumSpectral resolutionSystematic error