Starlight

A velocity that is an average of lines that disagree

A radial velocity measured to a metre a second is not measured from a line. It is the position of the peak of a cross-correlation against a mask of thousands of lines, and those lines do not agree with each other by hundreds of metres a second — because each one forms at a different depth in an atmosphere that is boiling.

Assumes The Doppler effect and Spectra.

A radial velocity of one metre a second is a Doppler shift of three parts in a thousand million. Across a spectral line five kilometres a second wide, it is a displacement of a five-thousandth of the line’s own width, and no single line can be centred that well.

The precision comes from the number of lines. A spectrograph records thousands of absorption features at once, and if each one’s centre is measured independently the errors are independent while the shift is common, so averaging them reduces the uncertainty as the square root of the count. That is the whole of how the metre-per-second regime is reached.

4000 lines averaged into one profile, and 2.3 m/s out of it. A cross-correlation function: the average absorption profile obtained by shifting a mask of 4000 line positions across a spectrum and summing what falls under it. The faint curves behind are individual lines, each with its own depth, its own width and its own small offset; the heavy curve is what averaging them produces. The velocity is the position of the peak, and its precision is the width divided by the contrast, the signal-to-noise and the square root of the number of lines — 2.3 metres a second here. Nothing about this construction is a measurement of any one line. It is a measurement of where a weighted average of thousands of them sits, and the weights are a choice: a mask built for one spectral type applied to another weights the disagreement between the lines differently, and moves the peak.
Fig. 1 How the averaging is actually done. A mask — a list of line positions and weights — is slid across the spectrum and the flux under it is summed at each velocity; the result is an average absorption profile whose peak is the velocity. The faint curves are individual lines, each with its own depth, width and small offset, and the heavy curve is what averaging four thousand of them produces. The precision is the width divided by the contrast, the signal-to-noise and the square root of the line count.

The construction is old — Griffin built the first cross-correlation spectrometer as a physical mask on a photographic plate in 1967, and the modern version does the same arithmetic in software — and it is the reason a line shift became a speedometer good enough to find planets rather than merely good enough to sort stars into moving groups.

What the average is an average of

The construction is a matched filter and it is optimal in the usual sense: given a set of lines with known positions and a common shift, the maximum-likelihood estimate of the shift is a weighted sum, and the cross-correlation is that sum evaluated on a grid.

The weights matter, and they are a choice. A mask built for a solar-type star lists the lines a solar-type star has, weighted by their depths; applied to a cooler star it includes lines that are not there and omits lines that are, and the resulting velocity is a differently weighted average. Two pipelines using different masks on the same spectrum return velocities that differ by tens of metres a second — a constant offset for a given star, and therefore harmless for a planet search and fatal for anything absolute.

There is a more basic point hiding in that. The cross-correlation function is not a line. It is a construction, its shape depends on the mask, and quantities read off it — its width, its depth, its asymmetry — are properties of the mask-plus-star combination rather than of the star.

400 lines averaged into one profile, and 31.3 m/s out of it. A cross-correlation function: the average absorption profile obtained by shifting a mask of 400 line positions across a spectrum and summing what falls under it. The faint curves behind are individual lines, each with its own depth, its own width and its own small offset; the heavy curve is what averaging them produces. The velocity is the position of the peak, and its precision is the width divided by the contrast, the signal-to-noise and the square root of the number of lines — 31.3 metres a second here. Nothing about this construction is a measurement of any one line. It is a measurement of where a weighted average of thousands of them sits, and the weights are a choice: a mask built for one spectral type applied to another weights the disagreement between the lines differently, and moves the peak.
Fig. 2 The same construction with ten times fewer lines and a third of the signal-to-noise, which is a small telescope observing a faint star rather than a large one observing a bright one. The precision falls to tens of metres a second, and it falls exactly as the square root of the line count and the first power of the signal-to-noise. Neither factor is negotiable, and the difference between them is why a planet search is a competition for aperture rather than for cleverness.

A second consequence of the same point is that the CCF’s width is not the star’s rotational broadening, and its depth is not a line depth. Both are used as diagnostics — the width to estimate the projected rotation, the depth to estimate the metallicity — and both need calibrating against stars whose properties are known independently, mask by mask. A quantity read off a construction has to be calibrated against the construction.

It is worth being explicit about what the mask contributes and what it does not, because the distinction decides which errors matter. The mask’s positions have to be right or the correlation is smeared and the precision falls, but a small error in a position is a constant offset in the velocity and cancels in a differential series. The mask’s weights are what determine which lines dominate the average, and because the lines genuinely disagree, the weights determine the answer. So the requirement on the positions is a precision requirement and the requirement on the weights is a stability requirement, and they are not the same kind of thing at all.

The lines do not agree

Measure each line’s velocity separately and the answers are spread over hundreds of metres a second. That spread is not noise. It repeats from night to night, it correlates with line strength, and its cause is that different lines are formed at different heights in the star’s atmosphere.

Lines that disagree by 350 metres a second, in order. The velocity measured from each of 90 individual absorption lines, against how deep the line is. The scatter is not random: a strong line is opaque high in the atmosphere, where the convective overturning has died away, while a weak line is transparent down to the granules themselves, which are rising and hot and therefore blueshifted. The result is a trend of 318 metres a second across the range of line strengths, with only 78 of genuine scatter about it. A cross-correlation averages over this, so its answer depends on which lines the mask contains and how they are weighted — and when the star's activity changes, the convective pattern changes and the whole trend shifts, which a single averaged velocity reports as a change in the star's motion.
Fig. 3 The velocity from each of ninety individual lines, against how deep the line is. A weak line is transparent all the way down to the granules, which are hot, rising and therefore blueshifted; a strong line is opaque high up, where the convective motion has died away and the gas is nearly at rest. The result is a trend of hundreds of metres a second across the range of line strengths, with only tens of metres a second of genuine scatter about it.

The star’s surface is a boiling fluid. Bright granules cover most of the area and are rising; dark intergranular lanes cover less and are sinking faster. The area-weighted, brightness-weighted average is a net blueshift of a few hundred metres a second, and how much of it a given line sees depends on how deep the line’s photons come from.

This is the convective blueshift, and it is the reason the absolute radial velocity of the Sun measured spectroscopically differs from the one measured by radar ranging to the planets. The difference is about 300 metres a second and it is not an error in either measurement.

The mechanism is worth spelling out because its size is not obvious. Granules rise at about a kilometre a second and cover roughly two thirds of the surface; the intergranular lanes sink at about two kilometres a second and cover the rest. Mass is conserved, so the two flows balance in mass — and they do not balance in light, because the rising gas is hotter and therefore brighter. A brightness-weighted average of the two flows is therefore not zero: it favours the rising material, and the net is a blueshift of a few hundred metres a second. The effect exists for exactly the same reason that the part of a star that boils is visible as granulation at all, and it would vanish if the granules and the lanes were equally bright.

Lines that disagree by 350 metres a second, in order. The velocity measured from each of 90 individual absorption lines, against how deep the line is. The scatter is not random: a strong line is opaque high in the atmosphere, where the convective overturning has died away, while a weak line is transparent down to the granules themselves, which are rising and hot and therefore blueshifted. The result is a trend of 328 metres a second across the range of line strengths, with only 78 of genuine scatter about it. A cross-correlation averages over this, so its answer depends on which lines the mask contains and how they are weighted — and when the star's activity changes, the convective pattern changes and the whole trend shifts, which a single averaged velocity reports as a change in the star's motion.
Fig. 4 The same lines with the convective pattern slightly altered, as it is when the star’s magnetic activity rises: magnetic fields suppress convection, the blueshift is reduced, and every line moves — the weak ones most. A cross-correlation averages over this and reports a change of a few metres a second in the star’s velocity. Nothing about the star’s motion has changed. This is the floor that stops the technique reaching the centimetre-per-second regime an Earth-mass planet requires.

The size of the effect is worth stating against the precision being sought. The spread between the weakest and strongest lines is three to four hundred metres a second. The instrument’s stability is a few centimetres a second. The signal from an Earth-mass planet in a habitable orbit is nine centimetres a second. So the disagreement between the lines is four thousand times the signal, and the only thing that makes the measurement possible at all is that the disagreement is nearly constant — it is a fixed offset that cancels in a differential series. Everything difficult about the technique is a consequence of that “nearly”.

There is also a hard floor from a direction nobody chose: the star’s own oscillations. A solar-type star rings in acoustic modes with periods of minutes and amplitudes of tens of centimetres a second each, and thousands of modes are excited at once. Their sum is a velocity signal of a metre or two, coherent over minutes and incoherent over hours, which is removed by the simple expedient of integrating for longer than a few oscillation periods. That works, it is cheap, and it is the one component of the stellar noise budget with a clean solution — which is worth knowing, because it means an exposure time that looks generously long is usually being set by this rather than by photons.

Three things the average hides

The activity signal. Because the trend with line depth shifts when the convective pattern changes, and because the magnetic field that changes it varies with the star’s rotation and its cycle, a velocity series contains a signal at the rotation period and another at the cycle period. Both look exactly like planets.

The mask dependence of the answer. A velocity is the weighted mean of a set of numbers that genuinely differ. Change the weights and the mean changes. Two masks that both give excellent internal precision give different velocities, and the difference varies with the star’s activity, so it does not even cancel as a constant offset.

The information in the disagreement. The spread across lines is not a nuisance to be averaged away — it is a measurement of the atmosphere. Line-by-line velocities, tracked over time, separate a genuine Doppler shift, which moves every line by the same fraction, from an activity signal, which moves lines differently according to their formation depth. That is the current frontier of the technique and it is only available because the individual measurements were kept.

A spot that reads as -31 m/s, and gives itself away in the shape. Left: a rotationally broadened absorption line from a star at v sin i = 6 km/s, drawn three ways — undisturbed, shifted rigidly by 50 m/s as a planet would shift it, and with a dark spot covering 2.0 per cent of the disc at 0.45 of the projected radius. Right: the bisector of each, the midpoint of the line at every depth, magnified. A rigid shift moves the bisector sideways and changes its span by 0.07 m/s, which is not zero only because the bisector is read by interpolation on a finite grid: a translation cannot change a shape. The spot moves the line's centroid by -31 m/s, which a velocity pipeline reports as a companion, and bends the bisector by 96 m/s, 1297 times the interpolation floor. That difference is the test, and it is the reason a radial-velocity detection is confirmed by a bisector that did not move rather than by a period that repeated.
Fig. 5 The classical version of the same diagnostic, applied to the averaged profile rather than to individual lines: the bisector, which traces the midpoint of the line at each depth. A true Doppler shift translates the whole profile and leaves the bisector’s shape unchanged; a spot rotating across the disc distorts the profile and bends the bisector. The two are distinguishable in principle and the distinction requires signal-to-noise the individual lines do not have — which is why the line-by-line approach, which uses the same information differently, has displaced it.

There is a fourth thing, and it is the one that has caused the most trouble in practice: the average is a single number, and it invites being treated as a measurement with a formal error. The formal error from the cross-correlation is the photon-limited precision, which for a bright star is centimetres a second, and it is quoted alongside a value whose true uncertainty is metres. Series of velocities with error bars ten times too small are common in the literature, and fitting a Keplerian orbit to such a series produces a planet with an implausibly small formal uncertainty and an implausibly good chi-squared — which is the signature to look for. A survey’s own selection is the first thing to model and the second is its error bars.

What was actually measured

Three measurements pin the argument down, and they were made in three different ways.

The Sun’s convective blueshift, measured against the planets. The solar spectrum’s lines are blueshifted relative to their laboratory wavelengths by an amount that depends on line strength, running to about 400 metres a second for the weakest lines and towards zero for the strongest. The comparison is possible only for the Sun, because only for the Sun is there an independent absolute velocity — from radar ranging and spacecraft tracking — to compare against.

The centre-to-limb variation. Observing the Sun as a resolved disc rather than as a star, the blueshift is largest at the centre of the disc, where the line of sight looks straight down into rising granules, and falls towards the limb, where the view is across them. The variation is measured and it is exactly what a model of granulation predicts, which is the strongest confirmation that the effect is convective rather than instrumental.

The activity correlation in other stars. For stars other than the Sun the test is statistical: velocity residuals correlate with activity indicators measured from the same spectra — the emission cores of the calcium H and K lines, chiefly — and the correlation has the sign and roughly the size the suppression of convection predicts. It is the basis of every attempt to remove activity from a velocity series.

A true peak at 0.3103 and a false one at 0.6897, from the same data. The Lomb–Scargle periodogram of 162 simulated observations taken over 419 nights from one site, of a star carrying a 4.2-unit sinusoid at 0.3103 cycles per day under 3 units of Gaussian noise per point. The injected signal is recovered at 0.3103 cycles per day, within the 0.0024 resolution element the baseline allows. The second peak, at 0.6899, is 87 per cent as tall and corresponds to nothing: it is 1 − f, the signal reflected in the one-cycle-per-day spike of the sampling. Neither peak is more real than the other in this picture — deciding between them needs a second site at a different longitude, or a run long enough for the seasonal window to separate them. The dashed line is the power a pure-noise series would exceed once in 1000 trials, computed from 586 independent frequencies rather than from the 4691 grid points searched; using the grid count instead would put the line 2.08 higher and reject a real detection.
Fig. 6 Where the consequence shows up. A velocity series contains a planet’s signal and the star’s own, and both are periodic — the planet at its orbital period, the star at its rotation period and its harmonics. Nothing in a periodogram distinguishes them, and several published planets have been withdrawn after being identified as rotation signals. The distinction has to come from information the periodogram does not use, which is what the line-by-line velocities supply.
A 28.7 km/s correction, and a 12.5 m/s planet underneath it. Two years of radial velocities of a star at ecliptic latitude 12°, orbited by a companion whose reflex semi-amplitude is 12.5 m/s — the Sun's own, from Jupiter. The upper panel is what the spectrograph measures: the Earth's motion about the barycentre of the solar system, amplitude 28.72 km/s, which is V⊕ cos β to a fraction of a per cent. The planet is in that curve and is 2,297 times smaller than it, which is a line thinner than the stroke it is drawn with. The lower panel is the same data after the correction, and the correction is not a fit: it is computed from an ephemeris, the observatory's position on a rotating deformable Earth, and the star's own coordinates and proper motion. To leave a centimetre a second it has to be right to one part in 2.9·10⁶ — the light-travel time across the Earth's orbit, the relativistic terms, and the fact that the star moves are all inside that budget. What remains is the planet, at 4333 days, and a scatter of 1.2 m/s that is the star rather than the instrument.
Fig. 7 The correction that has to be applied before any of this is visible: the observer’s own motion, thirty kilometres a second of it, removed to a part in three million so that a metre per second survives. That correction is exact, calculable and not the difficulty. What this essay is about is what is left after it — a residual velocity that is an average over an atmosphere, at the level of a few metres a second, and which no ephemeris can remove.

One more measurement belongs on this list because it closes the loop between the Sun and the stars. Instruments now observe the Sun as a star — feeding sunlight into the same spectrograph a planet search uses, through a small integrating telescope — and produce a solar velocity series with the same pipeline, the same mask and the same systematics as a stellar one. The result is a velocity that varies by a few metres a second with the solar rotation and the solar cycle, on a star whose planets are known exactly. It is the only calibration of activity-induced velocity that exists, and every technique for removing activity is now tested against it before being applied anywhere else.

Where the picture stops

Three limits stand out, and the third is where the subject is now.

Tellurics. The Earth’s own atmosphere imprints absorption lines on every ground-based spectrum, and they do not move with the star — they move with the observatory, which is to say they stay put while everything else shifts through the year. A mask that includes a telluric line by accident acquires a spurious annual signal, and the water-vapour lines change strength with the weather, so the contamination is not even constant.

The mask is not the star. Everything above treats the lines as a fixed list with fixed weights. A physically better approach uses a high signal-to-noise template built from the star’s own spectra, which automatically has the right lines with the right weights. That improves the precision and does nothing about the convection, because the template contains the convective pattern averaged over whatever epochs went into it.

And the floor is astrophysical rather than instrumental. Modern spectrographs are stable to a few centimetres a second, verified against laser frequency combs. The stars are not: granulation, oscillation and magnetic activity contribute somewhere between fifty centimetres and several metres a second for a quiet solar-type star. Detecting an Earth around a Sun requires a signal of nine centimetres a second, which is an order of magnitude below the star’s own noise, and every proposal for getting there is a proposal for using the structure of that noise rather than for averaging it down.

9000 lines averaged into one profile, and 0.6 m/s out of it. A cross-correlation function: the average absorption profile obtained by shifting a mask of 9000 line positions across a spectrum and summing what falls under it. The faint curves behind are individual lines, each with its own depth, its own width and its own small offset; the heavy curve is what averaging them produces. The velocity is the position of the peak, and its precision is the width divided by the contrast, the signal-to-noise and the square root of the number of lines — 0.6 metres a second here. Nothing about this construction is a measurement of any one line. It is a measurement of where a weighted average of thousands of them sits, and the weights are a choice: a mask built for one spectral type applied to another weights the disagreement between the lines differently, and moves the peak.
Fig. 8 The other end of the range: a very bright star observed with a large telescope and a mask of nine thousand lines. The photon-limited precision is well under a metre a second, and it is a number that has stopped meaning anything on its own — the star’s own surface contributes several times that, and the quantity limiting the measurement is no longer in this figure. That transition, from an instrument-limited to a star-limited technique, happened around 2005 and is the reason the subject’s effort moved from spectrographs to stellar atmospheres.

Why the disagreement is the interesting part

The recurring shape here is worth naming, because it is the opposite of the usual advice.

Averaging is the standard response to scatter, and it is correct when the scatter is noise. Here the scatter is not noise: it is a set of genuinely different measurements of genuinely different things, and averaging them produces a number that is precise, reproducible and not quite a measurement of anything. The average of a thousand lines’ velocities is not the star’s velocity; it is a weighted average of the velocities of a thousand different layers of a moving atmosphere, and the weighting is an accident of which lines the mask contains.

That is exactly the situation in which the residuals carry more information than the mean. The trend of velocity with line depth is a direct measurement of the convective structure. Its change over time is a measurement of the magnetic activity. Neither is available from the averaged velocity, and both were being discarded for thirty years because the average was what the instrument was built to produce.

There is a second, more uncomfortable observation. For thirty years the internal precision of these measurements improved by three orders of magnitude while the accuracy — the agreement between what was measured and the star’s actual motion — improved hardly at all, because the limiting term was never the one being worked on. Every gain in the spectrograph was real and every gain went into a residual dominated by something else. That is a common shape in instrumental astronomy and it is only visible from outside the programme: from inside, each improvement is measured against the previous instrument and looks like progress. The diagnostic is the one this collection keeps returning to — plot the answer against something it should not depend on, and see whether it does.

The general form: when a measurement is an average over things that differ, keep the things. The spread of line widths in a spectrum is another instance, and so is the disagreement between distance indicators. In each case the summary statistic was designed when the individual measurements were too noisy to be worth keeping, and in each case they no longer are.

One practical consequence deserves stating on its own, because it changes what a velocity series is for. If the average velocity contains a stellar signal at the rotation period, then a planet at or near that period is undetectable by this technique, and a planet at half or twice it is at risk of being confused with a harmonic. For a solar-type star that removes a band of orbital periods around twenty to thirty days from the searchable space — which happens to include the habitable zone of stars a little cooler than the Sun. The technique’s blind spot and the region of most interest overlap, and they overlap for a reason that has nothing to do with either: the star’s rotation and its habitable orbit are set by different physics and land in the same place by accident. A velocity measured from the shape of a line rather than its position is one of the few routes that does not share the blind spot.

It is worth closing by naming what changed. The technique’s first thirty years treated the star as a point with a velocity, and the spectrum as a noisy measurement of it. That model was adequate while the errors were tens of metres a second and it stopped being adequate at one. What replaced it is a model in which the star is an extended, structured, time-varying object whose light carries several velocities at once, and the quantity wanted — the motion of the centre of mass — is one component of it. That is a much harder measurement and it is the same one an occulted planet’s shadow crossing a rotating line is making from the other direction: both are attempts to read a spatially resolved property out of light that arrives unresolved.

Where the ladder goes next

The obvious next rung is the template: how a star’s own spectrum, built up over many epochs, replaces a mask, and why that improves the precision without touching the systematic. The rung beyond it is the deconvolution of the two — how a velocity series is separated into the part that moves every line equally, which is a Doppler shift, and the part that does not, which is the star’s own surface.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Convective blueshiftCross-correlationGranulationLine bisectorLine maskRadial velocityStellar jitterSystematic errorTelluric linesWeighted mean