Starlight

A period found in the gaps

Astronomical time series are sampled when the sky is dark and clear and the target is up, which is a schedule with a spectrum of its own. That spectrum is convolved with the real one, so a single sinusoid produces several peaks — and the tallest is not always the true one.

Assumes Photon noise, The Doppler effect and Variable stars.

Almost every period quoted in astronomy was extracted from a series of measurements taken at whatever times the sky allowed. The star was above the horizon for six hours in November and eleven in February; it was cloudy for four nights running; the telescope was allocated to somebody else in March; and for three months of the year the target was on the far side of the Sun.

That schedule is not a nuisance to be averaged away. It has a spectrum, the data’s spectrum is the true one convolved with it, and every feature of the schedule is copied onto every real signal in the data.

A true peak at 0.3103 and a false one at 0.6897, from the same data. The Lomb–Scargle periodogram of 162 simulated observations taken over 419 nights from one site, of a star carrying a 4.2-unit sinusoid at 0.3103 cycles per day under 3 units of Gaussian noise per point. The injected signal is recovered at 0.3103 cycles per day, within the 0.0024 resolution element the baseline allows. The second peak, at 0.6899, is 87 per cent as tall and corresponds to nothing: it is 1 − f, the signal reflected in the one-cycle-per-day spike of the sampling. Neither peak is more real than the other in this picture — deciding between them needs a second site at a different longitude, or a run long enough for the seasonal window to separate them. The dashed line is the power a pure-noise series would exceed once in 1000 trials, computed from 586 independent frequencies rather than from the 4691 grid points searched; using the grid count instead would put the line 2.08 higher and reject a real detection.
Fig. 1 The consequence, on simulated data with one sinusoid in it. The injected signal is recovered at 0.310 cycles per day, and a second peak two-thirds as tall stands at 0.690 — which is 1f1 - f, and corresponds to nothing at all. Neither peak is more real than the other in this picture. Deciding between them takes either a second telescope at a different longitude or a run long enough for the seasonal structure to separate them.

The window function

Write the sampling as a set of times {tj}\{t_j\} and define

W(ν)=1N2je2πiνtj2.W(\nu) = \frac{1}{N^2}\left|\sum_j e^{-2\pi i \nu t_j}\right|^2.

This is the spectral window: the power spectrum of the observing schedule, containing no data whatever. It always has a peak at zero frequency, and every other peak in it is a periodicity in when the telescope was pointed.

The sampling has a spectrum of its own, and it spikes at one cycle per day. The spectral window |Σ exp(−2πiνt)|²/N² of the same 162 observation times, with no data in it at all — this is a picture of when the telescope was pointed, not of what it saw. The spike at 0.9999 cycles per day reaches 0.82 of the zero-frequency value, because observations are taken at night and nights are one day apart. Around it sit sidelobes about a cycle per year away, at 0.9970, from the months each year in which the target is not up. A periodogram of real data is this function convolved with the true spectrum, which is the whole reason a single sinusoid produces more than one peak: every feature here is copied to every real frequency. The remedy is not a better statistic but a better window — a second telescope at a different longitude fills the nightly gap, and its window has no spike at one.
Fig. 2 The window of the same simulated campaign. The spike at exactly one cycle per day is there because observations are taken at night and nights are one day apart; the structure immediately around it comes from the months each year when the target is not up. A periodogram of real data is this function convolved with the true spectrum, so every feature here is copied to every real frequency. The remedy is not a better statistic but a better window.

The convolution statement is exact in the limit of a large number of samples and approximately true otherwise, and it is the single most useful fact in the subject. A real signal at frequency ff produces peaks at ff, at f±1d1f \pm 1\,\text{d}^{-1}, at f±1yr1f \pm 1\,\text{yr}^{-1}, and at ff plus any other frequency in the window — with heights in proportion to the window’s own peaks.

There is a subtlety in the position of that daily spike which repays attention. Weather and the working day repeat on the solar day; a target’s visibility repeats on the sidereal day, which is three minutes fifty-six seconds shorter. A long campaign therefore has two closely spaced spikes near one cycle per day rather than one, separated by about 1/3651/365 of a cycle per day — and since the resolution of a periodogram is 1/T1/T, a run longer than a year resolves them. The pair is not a curiosity: it is the reason a multi-year baseline breaks aliases that a single season cannot, and the reason the alias structure of a survey changes character once its baseline passes a year.

The sampling has a spectrum of its own, and it spikes at one cycle per day. The spectral window |Σ exp(−2πiνt)|²/N² of the same 458 observation times, with no data in it at all — this is a picture of when the telescope was pointed, not of what it saw. The spike at 1.0000 cycles per day reaches 0.81 of the zero-frequency value, because observations are taken at night and nights are one day apart. Around it sit sidelobes about a cycle per year away, at 1.0028, from the months each year in which the target is not up. A periodogram of real data is this function convolved with the true spectrum, which is the whole reason a single sinusoid produces more than one peak: every feature here is copied to every real frequency. The remedy is not a better statistic but a better window — a second telescope at a different longitude fills the nightly gap, and its window has no spike at one.
Fig. 3 The same window over a run nearly three times as long. The resolution element is 1/T1/T, so tripling the baseline divides the width of every spike by three — and the structure that was a single blur near one cycle per day in the shorter run has resolved. This is what buying a longer baseline actually buys: not a taller peak, since the window’s height is normalised, but a narrower one, and therefore a smaller region of frequency space in which an alias can hide a real signal. The seasonal sidelobes at one cycle per year narrow in the same proportion, which is why the same campaign run for four years rather than one has an alias problem of a different character rather than a smaller one.

The sinusoid at 1f1 - f has a physical interpretation that makes the ambiguity feel less like an artefact. Between two observations one day apart, a wave at 0.31 cycles per day advances 0.31 of a turn; a wave at 0.69 cycles per day advances 0.69 of a turn, which is 0.31 of a turn backwards. Sampling once a night cannot tell those apart.

Two periods, 3.22 days and 1.45, through the same observations. Twelve days of the same simulated run, with the best-fitting sinusoid at the injected 0.3103 cycles per day and at its one-day alias 0.6897 drawn together. The two curves differ by a root-mean-square of 4.49 units between the samples and 1.42 at them — a factor of 3.2 — and that is the definition of an alias rather than a coincidence: sampling once a night cannot distinguish a wave that advances 0.31 of a cycle between observations from one that advances the same fraction backwards, and 1 − f is exactly that wave. The discrimination that exists comes from the scatter in the observation times, a few hours either way as the target rises and sets, and from the length of the run; both are why the periodogram prefers one peak slightly over the other rather than not at all. A dataset does not contain a period. It contains a set of periods it cannot rule out, and the width of that set is a property of the schedule.
Fig. 4 The two, drawn through the same points. Twelve days of the same run, with the best-fitting sinusoid at each of the two candidate frequencies: they agree to a root-mean-square of a fraction of a unit at the samples and disagree by several times that between them. What little discrimination exists comes from the scatter in the observation times — a few hours either way as the target rises and sets — which is why the true peak is slightly taller and not decisively so.

The statistic

For evenly sampled data the natural tool is the discrete Fourier transform. For unevenly sampled data it is not, because the transform of an irregular series has no clean statistical interpretation.

The way out, due to Lomb and Scargle, is to abandon the transform and fit. At each trial frequency, fit a sinusoid Acosωt+BsinωtA\cos\omega t + B\sin\omega t to the data by least squares, and record how much of the variance it removed. That is a well-posed question for any sampling whatever.

The algebra then collapses into something that looks like a transform:

P(ω)=12σ2[(jyjcosω(tjτ))2jcos2ω(tjτ)+(jyjsinω(tjτ))2jsin2ω(tjτ)],P(\omega) = \frac{1}{2\sigma^2}\left[\frac{\left(\sum_j y_j\cos\omega(t_j-\tau)\right)^2}{\sum_j\cos^2\omega(t_j-\tau)} + \frac{\left(\sum_j y_j\sin\omega(t_j-\tau)\right)^2}{\sum_j\sin^2\omega(t_j-\tau)}\right],

with τ\tau chosen so that the sine and cosine sums are orthogonal. That offset is the whole difference between this and a naive transform: it makes the statistic invariant under a shift of the time origin, which a naive transform of unevenly spaced data is not.

Two refinements matter in practice.

Fit a floating mean. The classical form above assumes the mean has been subtracted correctly, and for uneven sampling the sample mean is not the signal’s mean — if a sinusoid was observed preferentially near its maximum, subtracting the sample mean removes part of the signal. Fitting the offset as a third free parameter fixes it, and the resulting statistic is the one any modern implementation uses.

Weight by the errors. Points with different uncertainties should not contribute equally, and a weighted least squares is no harder than an unweighted one.

What a false alarm probability actually says

Under pure Gaussian noise, the normalised power at any single frequency is exponentially distributed with unit mean. The probability of exceeding some level zz at one frequency is eze^{-z}, and the probability of the highest of MM independent frequencies exceeding it is 1(1ez)M1 - (1-e^{-z})^M.

Everything difficult is in the word independent.

Two other things have to be right before the arithmetic means anything, and both are about timekeeping rather than statistics. The times must be on a uniform scale — six kinds of second are in circulation and the differences between them reach tens of seconds — and they must be referred to the solar system’s barycentre rather than to the observatory, since the Earth’s own orbital motion introduces a light-travel variation of up to 16.6 minutes over a year. An uncorrected series carries an annual signal that nothing in the sky put there.

The resolution of a periodogram is 1/T1/T, where TT is the baseline — two sinusoids closer together than that cannot be told apart by the data. So the number of independent frequencies in a search up to fmaxf_{\max} is about TfmaxT f_{\max}, and it does not change when the grid is made finer. Evaluating on a grid ten times denser finds the peaks more precisely and adds no independence at all.

A true peak at 0.3103 and a false one at 0.6897, from the same data. The Lomb–Scargle periodogram of 162 simulated observations taken over 419 nights from one site, of a star carrying a 4.2-unit sinusoid at 0.3103 cycles per day under 3 units of Gaussian noise per point. The injected signal is recovered at 0.3103 cycles per day, within the 0.0024 resolution element the baseline allows. The second peak, at 0.6899, is 87 per cent as tall and corresponds to nothing: it is 1 − f, the signal reflected in the one-cycle-per-day spike of the sampling. Neither peak is more real than the other in this picture — deciding between them needs a second site at a different longitude, or a run long enough for the seasonal window to separate them. The dashed line is the power a pure-noise series would exceed once in 100 trials, computed from 586 independent frequencies rather than from the 4691 grid points searched; using the grid count instead would put the line 2.08 higher and reject a real detection.
Fig. 5 The same periodogram with the false-alarm line drawn at one per cent rather than one in a thousand — the only thing changed. The line moves down, more of the noise reaches it, and the peak’s significance is a different number for data that are identical. That is the whole content of a false-alarm probability: it is a threshold somebody chose, converted into a height through an assumed noise distribution and an assumed count of independent frequencies. Neither peak in this drawing moved. What moved is a line describing how surprised the reader should be, and the two things are worth keeping separate, because only one of them is a measurement.

Quoting a false-alarm probability computed from the number of grid points rather than the number of independent frequencies is a common and consequential error, and it goes in the conservative direction — it makes real detections look marginal — which is why it survives.

Red noise, which is the real problem

Stars are not constant, and they are not randomly variable either. Granulation, oscillations, spots rotating in and out of view, and long-term activity cycles all produce variability that is correlated in time — power that rises towards low frequencies rather than being flat.

For such noise the exponential distribution is wrong, and it is wrong in the dangerous direction: correlated noise produces high peaks far more often than white noise does, so a false-alarm probability computed under the white assumption can be too small by orders of magnitude. The practical responses are all empirical. Estimate the noise’s own power spectrum and use it in the significance calculation. Or bootstrap: shuffle the measured values among the observed times, recompute the highest peak, repeat ten thousand times, and read the false-alarm level off the resulting distribution. Shuffling preserves the sampling and destroys any real signal, so it measures exactly the thing wanted — how tall a peak this schedule produces from nothing.

A true peak at 0.3103 and a false one at 0.6897, from the same data. The Lomb–Scargle periodogram of 162 simulated observations taken over 419 nights from one site, of a star carrying a 2-unit sinusoid at 0.3103 cycles per day under 3 units of Gaussian noise per point. The injected signal is recovered at 0.3100 cycles per day, within the 0.0024 resolution element the baseline allows. The second peak, at 0.6899, is 91 per cent as tall and corresponds to nothing: it is 1 − f, the signal reflected in the one-cycle-per-day spike of the sampling. Neither peak is more real than the other in this picture — deciding between them needs a second site at a different longitude, or a run long enough for the seasonal window to separate them. The dashed line is the power a pure-noise series would exceed once in 1000 trials, computed from 586 independent frequencies rather than from the 4691 grid points searched; using the grid count instead would put the line 2.08 higher and reject a real detection.
Fig. 6 The same star and the same schedule with the signal cut to two thirds of the noise per point, which is a more typical detection than the hero’s. Both peaks survive — the sampling is what puts them there, and a weaker signal is aliased just as faithfully as a strong one — but the gap between the true peak and its reflection has narrowed, and the forest of noise peaks has climbed towards both. This is the regime in which the alias problem does its damage, because the discriminant described above is the difference in height between the two candidates, and that difference shrinks with the signal while the reflection does not.

The Nyquist frequency, and why it is not one

Even sampling has a hard ceiling: a series sampled every Δt\Delta t cannot distinguish frequencies above 1/2Δt1/2\Delta t, because every higher frequency has a lower one that agrees with it at every sample. That is the Nyquist limit, and it is a theorem.

Uneven sampling does not have one — or rather, the limit is set by the precision of the times rather than by their spacing. Nights are not exactly 24 hours apart, observations within a night are scattered by hours, and each of those irregularities breaks a degeneracy that regular sampling would leave intact.

The consequence is one of the pleasant surprises in the subject. A star observed once a night can yield a period of a few hours, provided the times are recorded well enough and the run is long enough. Kepler’s long-cadence photometry, sampled every 29.4 minutes, routinely detects oscillations well above its nominal Nyquist frequency, because the spacecraft’s own barycentric motion modulates the sampling by a few seconds over a year and that is enough to separate a super-Nyquist signal from its reflection.

The sampling has a spectrum of its own, and it spikes at one cycle per day. The spectral window |Σ exp(−2πiνt)|²/N² of the same 712 observation times, with no data in it at all — this is a picture of when the telescope was pointed, not of what it saw. The spike at 0.9999 cycles per day reaches 0.81 of the zero-frequency value, because observations are taken at night and nights are one day apart. Around it sit sidelobes about a cycle per year away, at 0.9967, from the months each year in which the target is not up. A periodogram of real data is this function convolved with the true spectrum, which is the whole reason a single sinusoid produces more than one peak: every feature here is copied to every real frequency. The remedy is not a better statistic but a better window — a second telescope at a different longitude fills the nightly gap, and its window has no spike at one.
Fig. 7 The same campaign with four observations a night instead of one, and the daily spike is still there. That is the point worth taking from it: the spike comes from the gaps between nights, not from the rate within them, so quadrupling the data does not remove it. What the extra points do is fill in the intra-night structure, which breaks the degeneracy the last section was about — several hours of coverage each night constrains a signal’s phase within the night, and no amount of once-a-night sampling can. The remedy for aliasing is a different window rather than more of the same one, and this is the cheapest version of a different window there is.

The limit is set by the irregularity, not by the average rate, which is the exact opposite of the intuition regular sampling trains.

Where the periods come from

The statistic is used far outside the case it was written for, and the applications are the reason it earns a place here.

The factor of two

There is a second ambiguity that has nothing to do with the window and that catches people more often, because it survives even perfect sampling.

A periodogram fits a sinusoid. A signal that is periodic but not sinusoidal has power at its fundamental and at its harmonics, and how that power is distributed depends on the waveform. For some shapes the tallest peak is not at the true period at all.

The standard case is an eclipsing binary with two similar components. Its light curve has two minima per orbit — one when each star passes in front of the other — and if the two eclipses are of similar depth the dominant Fourier component is at twice the orbital frequency. A period search returns half the true period, and folding on that half produces a perfectly clean single-eclipse light curve with no hint that anything is wrong.

The discriminant is the difference between the two minima. Two stars of different temperature produce eclipses of different depth, and folding on the true period separates them while folding on half superimposes them. So the test is to fold on twice the recovered period and look for an alternation: if the odd and even minima differ, the longer period is the right one.

That works when the components differ and fails when they are nearly identical, which is exactly the case that produces the ambiguity in the first place. For those systems the period is settled by radial velocities — which show one full cycle per orbit rather than two, because a velocity has a sign and a depth does not.

The same trap appears wherever a signal has two similar features per cycle. A spotted star with two active longitudes on opposite hemispheres gives a rotation period too short by a factor of two; a pulsator with a strong first overtone can be catalogued at the overtone rather than the fundamental. A period search returns the strongest periodicity in the data, and the strongest periodicity is not always the period, which is a distinction the statistic itself cannot draw.

Two periods, 1.61 days and 2.63, through the same observations. Twelve days of the same simulated run, with the best-fitting sinusoid at the injected 0.62 cycles per day and at its one-day alias 0.3800 drawn together. The two curves differ by a root-mean-square of 3.85 units between the samples and 1.25 at them — a factor of 3.1 — and that is the definition of an alias rather than a coincidence: sampling once a night cannot distinguish a wave that advances 0.62 of a cycle between observations from one that advances the same fraction backwards, and 1 − f is exactly that wave. The discrimination that exists comes from the scatter in the observation times, a few hours either way as the target rises and sets, and from the length of the run; both are why the periodogram prefers one peak slightly over the other rather than not at all. A dataset does not contain a period. It contains a set of periods it cannot rule out, and the width of that set is a property of the schedule.
Fig. 8 The two-candidate picture at twice the hero’s frequency — 1.61 days against its one-day alias at 2.63 — which is where the ambiguity is at its most dangerous. Both periods are entirely plausible for the things a survey is looking for, so neither can be dismissed on physical grounds the way a period of a hundred days or of two hours could be. The two curves here disagree between the samples by three times what they disagree at them, which is the entire discriminating power available from one site: a factor of three in a residual, against a noise level that is a fair fraction of the signal.

A famous ambiguity

The alias problem is not academic, and the clearest cases come from radial-velocity searches for planets.

A survey observing a star from one site, on nights when it is up, produces exactly the window drawn above. A planet with a period near one year is therefore accompanied by a peak at the alias of one year — which is to say near infinity, or near zero frequency, where the star’s own long-term activity also lives. A planet with a period near a day is accompanied by a peak near infinity in the other direction.

The awkward middle case is a period of a few days to a few weeks, where the daily alias lands in a region where other planets plausibly are. More than one announced planet has turned out to be the alias of a signal at a different period, and more than one has turned out to be an alias of the star’s rotation. The corrections have been made by the same means each time: more data with a different window, from another longitude or another instrument, which moves the aliases without moving the truth.

The general test is that simple. A real signal sits at the same frequency in every window; an alias moves. Nothing about the height of a peak, the elegance of a fold or the smallness of a false-alarm probability substitutes for it.

Looking for the second signal

Most interesting time series contain more than one periodicity, and the standard procedure for finding them compounds every difficulty above.

The method is prewhitening: find the strongest peak, fit and subtract the corresponding sinusoid, recompute the periodogram of the residuals, and repeat until nothing significant remains. It is simple, it is what almost everybody does, and it has a failure mode that is worth naming.

If the first peak selected is an alias rather than the true frequency, the sinusoid subtracted is the wrong one. It removes some of the real signal — the alias and the truth agree at the sample times, which is why they were confusable — and it adds structure between the samples that was never there. The residuals then contain an artefact at the true frequency and at the differences between the two, and the second iteration finds one of those and reports it as a discovery.

The literature contains several multi-planet systems that have been reinterpreted this way, in which a claimed second signal turned out to be an artefact of having subtracted a first at an aliased frequency. The correction each time has been the same: fit all the signals simultaneously rather than sequentially, so that the model is compared against the data as a whole and no intermediate residual is ever formed.

Simultaneous fitting is more expensive and it removes the failure mode entirely, because a wrong frequency for one component is penalised by the fit to the others rather than being frozen into the residuals. It also makes the significance question honest: the quantity to compare is the likelihood of a two-signal model against a one-signal model, evaluated with the extra parameters counted, rather than the height of a peak in a residual spectrum whose noise properties nobody has characterised.

Subtracting a model to look for what is left assumes the model was right, and in a subject where the strongest peak is regularly at the wrong frequency that assumption does most of its damage on the second iteration rather than the first.

When the model is not a sinusoid

The Lomb–Scargle statistic is the least-squares fit of a sinusoid. When the signal has a different shape, fitting a sinusoid to it wastes most of the signal.

A transit is the extreme case: a flat line with a brief box-shaped dip, whose Fourier content is spread across many harmonics. Fitting a sinusoid captures a fraction of it, and the standard tool is instead a box-fitting statistic — the same idea, with the model replaced by a box of adjustable width, depth and phase. The general principle is worth stating because it is easy to miss: a periodogram is not a neutral measurement of periodicity. It is a matched filter for one particular waveform, evaluated at many frequencies, and anything shaped differently is detected inefficiently or not at all.

That has a consequence for what a survey finds. The set of objects a search recovers is the set whose signals resemble its model, at frequencies its window does not spoil, above a threshold set by noise it has assumed to be white — and every survey therefore draws a different sky. The detection efficiency of a period search is a function of period, amplitude and shape, and it is measured the same way a transit survey’s is: by injecting synthetic signals into the real time series and counting how many come back. A completeness curve for a periodogram is as necessary as one for a photometric survey, and is quoted far less often.

Where this ladder goes next

This rung has established what the statistic is, why the aliases exist, and what an independent frequency means.

The rung above is the treatment of correlated noise properly — modelling the star’s own variability as a Gaussian process with a kernel fitted alongside the signal, which is now standard practice in radial-velocity work and which changes both the recovered amplitudes and their uncertainties.

Beside it lies the multi-site and space-based answer to aliasing: a window function with no daily spike, obtained either by observing from several longitudes at once or by leaving the ground entirely.

And below it, the habit this rung is really about: a measurement inherits the structure of the schedule that produced it. Before believing a peak, look at the window; before quoting a significance, ask what the noise actually is; and before publishing a period, fold on it.

What this makes readable

Essays that name this one as a prerequisite.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

AliasBox least squaresFalse-alarm probabilityIndependent frequenciesThe Lomb–Scargle periodogramNyquist frequencyPhase foldingRed noiseSpectral leakageTime seriesUneven samplingWindow function