Starlight

The tallest peak in nothing at all

A periodogram of pure noise has peaks in it, and the tallest is not small. How tall it has to be before it means something depends on how many frequencies were searched and on what the noise actually is — and astronomical noise is almost never the white noise the standard formula assumes.

Assumes Periodograms and Photon noise.

A periodogram turns a time series into a spectrum of power against frequency, and the eye is drawn to its highest peak. The question that has to be answered before that peak means anything is how tall a peak noise alone would have produced.

The answer is not “small”. Searching many frequencies means taking the maximum of many random numbers, and the maximum of many draws sits far out in the tail of the distribution any one of them came from. A search over fifty thousand frequencies in pure noise produces a peak that would be a five-sigma detection if it had been the only frequency examined.

A peak worth 10.8 in a narrow search is worth nothing in a wide one. The probability that noise alone produces a peak at least as tall as a given power, for searches over four different numbers of independent frequencies. A single frequency examined in isolation gives a one-per-cent chance at a power of 4.6; searching fifty thousand frequencies for the same one-per-cent chance requires 15.4. The threshold rises as the logarithm of the width of the search, which is why the penalty is survivable — but it is a penalty, it is often not applied, and the number of independent frequencies in an unevenly sampled time series is not the number of frequencies on the grid. Overestimating that count is conservative and underestimating it is not, which is the one asymmetry worth remembering.
Fig. 1 The probability that noise alone produces a peak at least as tall as a given power, for four widths of search. A single frequency examined in isolation gives a one-per-cent chance at a power of 4.6; searching fifty thousand frequencies for the same one-per-cent chance requires 15.4. The threshold rises as the logarithm of the width of the search, which is why the penalty is survivable — and it is a penalty, and it is often not applied.

The effect has names in several fields — the look-elsewhere effect in particle physics, multiple testing in statistics, the false-alarm probability in time-series astronomy — and the fact that it was named separately three times says something about how easy it is to meet without recognising.

Where the formula comes from

For evenly sampled data with white Gaussian noise, the power at a single frequency of a normalised periodogram is exponentially distributed. The probability of exceeding a level zz at one frequency is eze^{-z}; the probability of not exceeding it at any of MM independent frequencies is (1ez)M(1-e^{-z})^M; and the false-alarm probability for the highest peak is one minus that.

Two features of the result matter more than the algebra.

The threshold rises logarithmically in MM. A hundredfold wider search costs ln100=4.6\ln 100 = 4.6 in power, which for the levels typically involved is about a fifty per cent increase in the threshold. That is the good news: widening a search is affordable.

MM is not the number of grid points. It is the number of independent frequencies, which is set by the length of the time series — frequencies separated by less than 1/T1/T are not independent, because their sinusoids are not orthogonal over the data. Oversampling the grid, which is necessary to find the peak accurately, does not increase MM; extending the time baseline does.

Estimating MM for unevenly sampled data is genuinely hard, and the standard approximations differ by factors of a few. The asymmetry to remember is that overestimating MM is conservative and underestimating it is not.

A peak worth 10.8 in a narrow search is worth nothing in a wide one. The probability that noise alone produces a peak at least as tall as a given power, for searches over four different numbers of independent frequencies. A single frequency examined in isolation gives a one-per-cent chance at a power of 4.6; searching fifty thousand frequencies for the same one-per-cent chance requires 15.4. The threshold rises as the logarithm of the width of the search, which is why the penalty is survivable — but it is a penalty, it is often not applied, and the number of independent frequencies in an unevenly sampled time series is not the number of frequencies on the grid. Overestimating that count is conservative and underestimating it is not, which is the one asymmetry worth remembering.
Fig. 2 The same curves at a narrower range of search widths and with the conventional levels marked, including the three-sigma equivalent. What the plot makes visible is how little the threshold moves for a large change in the search: three decades of extra frequencies cost about seven in power, from six to thirteen. That is the reason a period search over a wide range is a sensible thing to do at all, and it is also the reason the correction is easy to neglect — the number it changes is not dramatic, and the peaks it rejects are the interesting ones.

There is a further complication that is specific to unevenly sampled data and that has no clean answer. For evenly sampled data the periodogram’s frequencies really are orthogonal, and MM is exactly the number of Fourier frequencies. For uneven sampling they are not orthogonal, neighbouring frequencies are correlated by an amount that depends on the sampling pattern, and the effective MM has to be estimated — usually by counting the number of local maxima in the periodogram of a noise realisation, which is a Monte Carlo estimate dressed as an analytic one. That is not a criticism: it is the honest situation, and it is a hint about where the argument ends up.

Why the formula is usually wrong

The derivation assumes the noise is white: independent from measurement to measurement, with no correlation in time. Almost no astronomical noise is.

A star’s surface convects, rotates and spots; the atmosphere’s transparency drifts; an instrument’s temperature follows the dome. All of these produce noise that is correlated over hours to months, which in the frequency domain means power concentrated at low frequencies. The periodogram of such noise is not flat: it rises towards long periods, often as an inverse power law.

A threshold that is not a horizontal line. A periodogram of noise that is not white, with a signal injected at 0.031 cycles a day. The lower dashed line is the one-per-cent false-alarm threshold computed on the assumption of white noise; the upper curve is the same threshold computed against the noise the data actually have, whose power rises towards low frequencies as an inverse power law. The injected peak clears the first and does not clear the second. Almost every astronomical time series has noise of this kind — stellar surfaces, atmospheric transparency, instrument temperature — and almost every long-period detection sits in the part of the spectrum where the two thresholds differ most. The remedy is not a cleverer statistic but an empirical one: estimating the threshold by permuting or bootstrapping the data themselves, which reproduces whatever correlation structure they have without anybody having to name it.
Fig. 3 A periodogram of noise that is not white, with a signal injected at a long period. The lower dashed line is the one-per-cent threshold computed assuming white noise; the upper curve is the same threshold computed against the noise the data actually have. The injected peak clears the first and does not clear the second. Almost every long-period detection sits in the part of the spectrum where the two differ most.

The consequence is that the false-alarm probability computed from the white-noise formula is not merely approximate — it is wrong in a direction that manufactures detections, and it is most wrong at exactly the long periods where the interesting signals are. A planet in a habitable orbit, a stellar cycle, a long-term trend: all of them sit where the noise is reddest.

A threshold that is not a horizontal line. A periodogram of noise that is not white, with a signal injected at 0.045 cycles a day. The lower dashed line is the one-per-cent false-alarm threshold computed on the assumption of white noise; the upper curve is the same threshold computed against the noise the data actually have, whose power rises towards low frequencies as an inverse power law. The injected peak clears the first and does not clear the second. Almost every astronomical time series has noise of this kind — stellar surfaces, atmospheric transparency, instrument temperature — and almost every long-period detection sits in the part of the spectrum where the two thresholds differ most. The remedy is not a cleverer statistic but an empirical one: estimating the threshold by permuting or bootstrapping the data themselves, which reproduces whatever correlation structure they have without anybody having to name it.
Fig. 4 Redder noise with a higher knee, which is what an active star or a ground-based photometric series looks like. The gap between the two thresholds has widened and the region in which a white-noise analysis is badly misleading now covers most of the useful frequency range. The shape of the noise spectrum is not a nuisance parameter here; it is the dominant term in whether anything has been detected, and it has to be measured from the data rather than assumed.

It is worth naming the reason the noise is red rather than treating it as an empirical fact. Almost every physical process that contributes noise to an astronomical time series has a timescale: granulation lives for minutes, active regions for weeks, a stellar cycle for years, a dome’s thermal response for hours. A process with a timescale has a correlation time, and a correlation time in the time domain is a knee in the frequency domain, below which the power rises. So red noise is not a nuisance peculiar to bad data — it is what the data look like when the noise is made of things rather than of counting. Photon noise is white precisely because photons arrive independently, and it is the only component of the budget that is.

The size of the discrepancy is worth quantifying because it decides whether this is a refinement or a catastrophe. If the noise power at the frequency of interest is a factor RR above the white level, then the threshold in power is a factor RR higher, and since the false-alarm probability falls exponentially in the power, the probability attached to a fixed peak height is wrong by a factor of order ez(R1)e^{z(R-1)}. For a peak at the nominal one-per-cent level and R=3R = 3, that is a factor of several thousand. A “one in a hundred” detection becomes a “one in a few” one, which is not a detection at all. The exponential is what turns a modest error in the noise model into a complete reversal of the conclusion.

What to do instead

Three approaches, and the third is the one that works.

Model the noise. Fit a parameterised covariance — a Gaussian process with a kernel chosen to represent rotation, say — simultaneously with the signal, and compare models by their evidence rather than by peak heights. This is principled, it is now standard in radial-velocity work, and its answer depends on the kernel chosen.

Whiten the data. Filter out the low frequencies before searching. That works and it removes the long-period signals along with the long-period noise, which is a solution only when the signals of interest are short-period.

Bootstrap. Estimate the threshold from the data themselves: scramble the measurements in a way that destroys any real periodicity while preserving whatever correlation structure the noise has, recompute the periodogram, and repeat many times. The distribution of the highest peak across those trials is the null distribution, with no assumption about the noise at all.

The bootstrap’s difficulty is preserving the right structure. Shuffling the measurements at random destroys the correlation as well as the signal, and gives back the white-noise answer. Shuffling in blocks preserves correlation on scales shorter than the block, and choosing the block length is choosing an answer. There is no free lunch, and the honest version of the procedure states the block length alongside the result.

The sampling has a spectrum of its own, and it spikes at one cycle per day. The spectral window |Σ exp(−2πiνt)|²/N² of the same 162 observation times, with no data in it at all — this is a picture of when the telescope was pointed, not of what it saw. The spike at 0.9999 cycles per day reaches 0.82 of the zero-frequency value, because observations are taken at night and nights are one day apart. Around it sit sidelobes about a cycle per year away, at 0.9970, from the months each year in which the target is not up. A periodogram of real data is this function convolved with the true spectrum, which is the whole reason a single sinusoid produces more than one peak: every feature here is copied to every real frequency. The remedy is not a better statistic but a better window — a second telescope at a different longitude fills the nightly gap, and its window has no spike at one.
Fig. 5 A completely different way for a periodogram to lie, and one this collection has met before: the sampling itself has a spectrum, and it is convolved with the true one. A single sinusoid observed on a schedule with structure produces several peaks, and the tallest is not always the true one. Window-function artefacts and false alarms are different failures — one produces a peak at the wrong frequency, the other a peak where there is nothing — and a detection has to survive both tests.

And there is a fourth response that is not statistical at all and is often decisive: get a second, independent observable. A radial-velocity signal that is a planet moves every spectral line by the same fraction; one that is an active region does not, and the line-by-line velocities separate them. A photometric period that is rotation is accompanied by a chromospheric emission signal at the same period; an orbit is not. Neither test involves a periodogram, and both are far more discriminating than any threshold, because they test a different prediction rather than a stronger version of the same one.

It is worth adding what the field learned about presenting these results, because the convention has changed within a working lifetime. A detection used to be reported as a period and a false-alarm probability. It is now reported, in careful work, as a posterior on the period alongside an explicit noise model, with the analysis repeated under two or three plausible noise models to show how much the conclusion depends on the choice. That is more honest and it is also less quotable, and the tension between those two is why the older convention persists in abstracts.

What was actually measured

The clearest evidence that this matters is the list of published detections that were later withdrawn.

Planets that were rotation. Several radial-velocity planets around active stars have been retracted after the signal was identified with the star’s rotation period or one of its harmonics. In each case the original detection had a formal false-alarm probability well below one per cent, computed on the white-noise assumption, and the signal was coherent over years — which is exactly what a long-lived active region produces and what a white-noise analysis cannot distinguish from an orbit.

Periods that were the window. Signals at one year, one sidereal day, one lunar month and their aliases appear in almost every ground-based time series, and their appearance in a detection is a well-known warning sign rather than a coincidence — the sampling has a spectrum of its own and it is convolved with everything.

The empirical check. The most convincing single demonstration is to run a period search on data known to contain no signal — a comparison star, a scrambled series, a synthetic series with realistic noise — and count how often a “detection” appears. Where this has been done systematically, the rate of spurious detections above a nominal one-per-cent threshold has come out at several per cent to tens of per cent, depending on how red the noise is.

A true peak at 0.3103 and a false one at 0.6897, from the same data. The Lomb–Scargle periodogram of 162 simulated observations taken over 419 nights from one site, of a star carrying a 4.2-unit sinusoid at 0.3103 cycles per day under 3 units of Gaussian noise per point. The injected signal is recovered at 0.3103 cycles per day, within the 0.0024 resolution element the baseline allows. The second peak, at 0.6899, is 87 per cent as tall and corresponds to nothing: it is 1 − f, the signal reflected in the one-cycle-per-day spike of the sampling. Neither peak is more real than the other in this picture — deciding between them needs a second site at a different longitude, or a run long enough for the seasonal window to separate them. The dashed line is the power a pure-noise series would exceed once in 1000 trials, computed from 586 independent frequencies rather than from the 4691 grid points searched; using the grid count instead would put the line 2.08 higher and reject a real detection.
Fig. 6 A periodogram with a genuine signal in it, for contrast. The peak is tall, it is narrow, and it sits well above the surrounding power. What no single such plot can show is whether the surrounding power is white — and that, rather than the height of the peak, is what decides whether the detection stands. Two periodograms that look identical can have false-alarm probabilities differing by orders of magnitude.
Two periods, 3.22 days and 1.45, through the same observations. Twelve days of the same simulated run, with the best-fitting sinusoid at the injected 0.3103 cycles per day and at its one-day alias 0.6897 drawn together. The two curves differ by a root-mean-square of 4.49 units between the samples and 1.42 at them — a factor of 3.2 — and that is the definition of an alias rather than a coincidence: sampling once a night cannot distinguish a wave that advances 0.31 of a cycle between observations from one that advances the same fraction backwards, and 1 − f is exactly that wave. The discrimination that exists comes from the scatter in the observation times, a few hours either way as the target rises and sets, and from the length of the run; both are why the periodogram prefers one peak slightly over the other rather than not at all. A dataset does not contain a period. It contains a set of periods it cannot rule out, and the width of that set is a property of the schedule.
Fig. 7 The third way a period search misleads, and the one with the longest history: aliasing. A signal sampled less often than twice per cycle appears at a different frequency, and on a nightly observing schedule the aliases of a true period sit at predictable places — one over the period plus or minus one per day, plus or minus one per year. A detection at one of those places is not evidence of anything until the alternative has been excluded, and excluding it requires either observations from a second longitude or a genuinely irregular schedule.

One further practical point about how the threshold should be used, because it is routinely inverted. The number of independent frequencies depends on the length of the time series, so a longer series is penalised — it searches more frequencies at the same resolution. That sounds like a reason not to observe longer and it is not: the penalty is logarithmic and the gain in the signal’s own power is linear in the number of observations, so extending a series always wins. The same arithmetic says something about survey design. Observing a thousand stars ten times each and one star ten thousand times give very different sensitivities to a periodic signal, and the difference is not captured by the total number of measurements. Every survey draws a different sky, and the shape of a period search’s sensitivity is one of the things a survey’s cadence decides before any data are taken.

Where the picture stops

The picture stops in three places, and the second is the one that trips careful people.

The threshold is not a decision. A false-alarm probability answers a narrow question: how often noise alone would produce a peak this tall. It says nothing about how often a real signal would, which is what actually decides whether to believe a detection. A one-per-cent false-alarm rate on a search where genuine signals are rare produces mostly false positives, and the correct framing is a comparison of models rather than a rejection of a null.

Repeating a search on more data is not independent. Adding a season of observations and re-running the search uses the old data again, so the two answers are correlated and the effective number of trials is not doubled. Programmes that observe until a signal crosses a threshold and then stop have a false-alarm rate substantially higher than the nominal one, for the same reason that stopping an experiment when the result looks good is a known way of manufacturing significance.

And the highest peak is not always the right one. The formula describes the maximum over the search. A real signal at a different frequency, a harmonic, or an alias can be present alongside it, and evaluating the false-alarm probability of the second-highest peak requires conditioning on the first — which nobody does, and which matters for multi-planet systems where the second signal is the interesting one.

One more belongs on that list, and it concerns what “the noise” means when the signal is present. All three approaches above estimate the null distribution from the data, and the data contain the signal — so the estimated noise is inflated by whatever real periodicity is there, and the threshold is too high. That is conservative, which is the right way to be wrong, and it means that the significance of a strong detection is systematically understated by a bootstrap. For a marginal detection the effect is small; for a strong one in a series with few other signals it can be substantial, and the standard remedy is to subtract the best-fit signal before estimating the null and then iterate.

Why this belongs with periodograms rather than with statistics

The reason to file the argument here is that the failure is specific to the way period searches are done, and it recurs in every subfield that does them.

A period search is a maximisation over a continuum of models, and the maximum of a fitted quantity is not distributed like the quantity. That is a completely general statement — it is the same effect that gives an experiment searching a mass range a look-elsewhere penalty, and the same effect that makes the best-fitting eccentricity of a circular orbit come out positive. What is specific to astronomy is that the noise is red, the sampling is uneven, and the number of independent trials is therefore not calculable in closed form. Every one of those three makes the standard correction inapplicable, and all three are properties of observing from a rotating planet with a schedule.

The practical distillation is short. Estimate the null from the data rather than from a formula. A bootstrap costs an afternoon of computing and makes no assumption that can be wrong; a formula costs nothing and assumes something that is false. The window function has to be examined for the same reason — the sampling is a property of the observing programme, and the only reliable description of it is the programme’s own record of when it observed.

A final observation, and it is about a asymmetry that decides how the field behaves. A false detection is published, cited, and then quietly retracted years later; a missed detection is never published at all and nobody knows it happened. The costs of the two errors are therefore wildly asymmetric in visibility and roughly symmetric in scientific damage, which pushes the whole enterprise towards thresholds that are too permissive. The corrective is not a stricter number — it is the practice of demanding a second, physically different observable before a signal is called a discovery, which is what a transit that runs late provides for a planet found in velocities and what a chromospheric index provides against one that is not.

It is worth stating the one habit that protects against nearly all of this, because it is cheap and it is not universal. Before believing any peak, run the identical search on the identical data with the signal destroyed — shuffle the measurements against their timestamps, or shift each night’s data by a random amount, or search a set of stars known to be constant. Whatever the noise really is, the shuffled version has the same sampling, the same gaps, the same number of points and the same aliasing structure, and the distribution of its tallest peaks is the null distribution the threshold should have come from. It costs a few seconds of computation and it makes no assumption about the noise at all. Nearly every retracted period in the literature would have failed that test, and the reason it is skipped is that it is nobody’s method — it produces no formula and it appears in no textbook chapter.

Where the ladder goes next

The immediate next rung is the model comparison that ought to replace the threshold: how the evidence for a periodic model is weighed against the evidence for a correlated-noise one, and why the answer depends on priors that are rarely stated. The rung beyond it is the case where the noise and the signal are the same physical object — a star whose rotation is both the contaminant and the quantity of interest.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

BootstrapDetection thresholdFalse-alarm probabilityIndependent frequenciesLook elsewhere effectMultiple testingNull hypothesisPeriodogramRed noiseStellar jitter