The tallest peak in nothing at all
Assumes Periodograms and Photon noise.
A periodogram turns a time series into a spectrum of power against frequency, and the eye is drawn to its highest peak. The question that has to be answered before that peak means anything is how tall a peak noise alone would have produced.
The answer is not “small”. Searching many frequencies means taking the maximum of many random numbers, and the maximum of many draws sits far out in the tail of the distribution any one of them came from. A search over fifty thousand frequencies in pure noise produces a peak that would be a five-sigma detection if it had been the only frequency examined.
The effect has names in several fields — the look-elsewhere effect in particle physics, multiple testing in statistics, the false-alarm probability in time-series astronomy — and the fact that it was named separately three times says something about how easy it is to meet without recognising.
Where the formula comes from
For evenly sampled data with white Gaussian noise, the power at a single frequency of a normalised periodogram is exponentially distributed. The probability of exceeding a level at one frequency is ; the probability of not exceeding it at any of independent frequencies is ; and the false-alarm probability for the highest peak is one minus that.
Two features of the result matter more than the algebra.
The threshold rises logarithmically in . A hundredfold wider search costs in power, which for the levels typically involved is about a fifty per cent increase in the threshold. That is the good news: widening a search is affordable.
is not the number of grid points. It is the number of independent frequencies, which is set by the length of the time series — frequencies separated by less than are not independent, because their sinusoids are not orthogonal over the data. Oversampling the grid, which is necessary to find the peak accurately, does not increase ; extending the time baseline does.
Estimating for unevenly sampled data is genuinely hard, and the standard approximations differ by factors of a few. The asymmetry to remember is that overestimating is conservative and underestimating it is not.
There is a further complication that is specific to unevenly sampled data and that has no clean answer. For evenly sampled data the periodogram’s frequencies really are orthogonal, and is exactly the number of Fourier frequencies. For uneven sampling they are not orthogonal, neighbouring frequencies are correlated by an amount that depends on the sampling pattern, and the effective has to be estimated — usually by counting the number of local maxima in the periodogram of a noise realisation, which is a Monte Carlo estimate dressed as an analytic one. That is not a criticism: it is the honest situation, and it is a hint about where the argument ends up.
Why the formula is usually wrong
The derivation assumes the noise is white: independent from measurement to measurement, with no correlation in time. Almost no astronomical noise is.
A star’s surface convects, rotates and spots; the atmosphere’s transparency drifts; an instrument’s temperature follows the dome. All of these produce noise that is correlated over hours to months, which in the frequency domain means power concentrated at low frequencies. The periodogram of such noise is not flat: it rises towards long periods, often as an inverse power law.
The consequence is that the false-alarm probability computed from the white-noise formula is not merely approximate — it is wrong in a direction that manufactures detections, and it is most wrong at exactly the long periods where the interesting signals are. A planet in a habitable orbit, a stellar cycle, a long-term trend: all of them sit where the noise is reddest.
It is worth naming the reason the noise is red rather than treating it as an empirical fact. Almost every physical process that contributes noise to an astronomical time series has a timescale: granulation lives for minutes, active regions for weeks, a stellar cycle for years, a dome’s thermal response for hours. A process with a timescale has a correlation time, and a correlation time in the time domain is a knee in the frequency domain, below which the power rises. So red noise is not a nuisance peculiar to bad data — it is what the data look like when the noise is made of things rather than of counting. Photon noise is white precisely because photons arrive independently, and it is the only component of the budget that is.
The size of the discrepancy is worth quantifying because it decides whether this is a refinement or a catastrophe. If the noise power at the frequency of interest is a factor above the white level, then the threshold in power is a factor higher, and since the false-alarm probability falls exponentially in the power, the probability attached to a fixed peak height is wrong by a factor of order . For a peak at the nominal one-per-cent level and , that is a factor of several thousand. A “one in a hundred” detection becomes a “one in a few” one, which is not a detection at all. The exponential is what turns a modest error in the noise model into a complete reversal of the conclusion.
What to do instead
Three approaches, and the third is the one that works.
Model the noise. Fit a parameterised covariance — a Gaussian process with a kernel chosen to represent rotation, say — simultaneously with the signal, and compare models by their evidence rather than by peak heights. This is principled, it is now standard in radial-velocity work, and its answer depends on the kernel chosen.
Whiten the data. Filter out the low frequencies before searching. That works and it removes the long-period signals along with the long-period noise, which is a solution only when the signals of interest are short-period.
Bootstrap. Estimate the threshold from the data themselves: scramble the measurements in a way that destroys any real periodicity while preserving whatever correlation structure the noise has, recompute the periodogram, and repeat many times. The distribution of the highest peak across those trials is the null distribution, with no assumption about the noise at all.
The bootstrap’s difficulty is preserving the right structure. Shuffling the measurements at random destroys the correlation as well as the signal, and gives back the white-noise answer. Shuffling in blocks preserves correlation on scales shorter than the block, and choosing the block length is choosing an answer. There is no free lunch, and the honest version of the procedure states the block length alongside the result.
And there is a fourth response that is not statistical at all and is often decisive: get a second, independent observable. A radial-velocity signal that is a planet moves every spectral line by the same fraction; one that is an active region does not, and the line-by-line velocities separate them. A photometric period that is rotation is accompanied by a chromospheric emission signal at the same period; an orbit is not. Neither test involves a periodogram, and both are far more discriminating than any threshold, because they test a different prediction rather than a stronger version of the same one.
It is worth adding what the field learned about presenting these results, because the convention has changed within a working lifetime. A detection used to be reported as a period and a false-alarm probability. It is now reported, in careful work, as a posterior on the period alongside an explicit noise model, with the analysis repeated under two or three plausible noise models to show how much the conclusion depends on the choice. That is more honest and it is also less quotable, and the tension between those two is why the older convention persists in abstracts.
What was actually measured
The clearest evidence that this matters is the list of published detections that were later withdrawn.
Planets that were rotation. Several radial-velocity planets around active stars have been retracted after the signal was identified with the star’s rotation period or one of its harmonics. In each case the original detection had a formal false-alarm probability well below one per cent, computed on the white-noise assumption, and the signal was coherent over years — which is exactly what a long-lived active region produces and what a white-noise analysis cannot distinguish from an orbit.
Periods that were the window. Signals at one year, one sidereal day, one lunar month and their aliases appear in almost every ground-based time series, and their appearance in a detection is a well-known warning sign rather than a coincidence — the sampling has a spectrum of its own and it is convolved with everything.
The empirical check. The most convincing single demonstration is to run a period search on data known to contain no signal — a comparison star, a scrambled series, a synthetic series with realistic noise — and count how often a “detection” appears. Where this has been done systematically, the rate of spurious detections above a nominal one-per-cent threshold has come out at several per cent to tens of per cent, depending on how red the noise is.
One further practical point about how the threshold should be used, because it is routinely inverted. The number of independent frequencies depends on the length of the time series, so a longer series is penalised — it searches more frequencies at the same resolution. That sounds like a reason not to observe longer and it is not: the penalty is logarithmic and the gain in the signal’s own power is linear in the number of observations, so extending a series always wins. The same arithmetic says something about survey design. Observing a thousand stars ten times each and one star ten thousand times give very different sensitivities to a periodic signal, and the difference is not captured by the total number of measurements. Every survey draws a different sky, and the shape of a period search’s sensitivity is one of the things a survey’s cadence decides before any data are taken.
Where the picture stops
The picture stops in three places, and the second is the one that trips careful people.
The threshold is not a decision. A false-alarm probability answers a narrow question: how often noise alone would produce a peak this tall. It says nothing about how often a real signal would, which is what actually decides whether to believe a detection. A one-per-cent false-alarm rate on a search where genuine signals are rare produces mostly false positives, and the correct framing is a comparison of models rather than a rejection of a null.
Repeating a search on more data is not independent. Adding a season of observations and re-running the search uses the old data again, so the two answers are correlated and the effective number of trials is not doubled. Programmes that observe until a signal crosses a threshold and then stop have a false-alarm rate substantially higher than the nominal one, for the same reason that stopping an experiment when the result looks good is a known way of manufacturing significance.
And the highest peak is not always the right one. The formula describes the maximum over the search. A real signal at a different frequency, a harmonic, or an alias can be present alongside it, and evaluating the false-alarm probability of the second-highest peak requires conditioning on the first — which nobody does, and which matters for multi-planet systems where the second signal is the interesting one.
One more belongs on that list, and it concerns what “the noise” means when the signal is present. All three approaches above estimate the null distribution from the data, and the data contain the signal — so the estimated noise is inflated by whatever real periodicity is there, and the threshold is too high. That is conservative, which is the right way to be wrong, and it means that the significance of a strong detection is systematically understated by a bootstrap. For a marginal detection the effect is small; for a strong one in a series with few other signals it can be substantial, and the standard remedy is to subtract the best-fit signal before estimating the null and then iterate.
Why this belongs with periodograms rather than with statistics
The reason to file the argument here is that the failure is specific to the way period searches are done, and it recurs in every subfield that does them.
A period search is a maximisation over a continuum of models, and the maximum of a fitted quantity is not distributed like the quantity. That is a completely general statement — it is the same effect that gives an experiment searching a mass range a look-elsewhere penalty, and the same effect that makes the best-fitting eccentricity of a circular orbit come out positive. What is specific to astronomy is that the noise is red, the sampling is uneven, and the number of independent trials is therefore not calculable in closed form. Every one of those three makes the standard correction inapplicable, and all three are properties of observing from a rotating planet with a schedule.
The practical distillation is short. Estimate the null from the data rather than from a formula. A bootstrap costs an afternoon of computing and makes no assumption that can be wrong; a formula costs nothing and assumes something that is false. The window function has to be examined for the same reason — the sampling is a property of the observing programme, and the only reliable description of it is the programme’s own record of when it observed.
A final observation, and it is about a asymmetry that decides how the field behaves. A false detection is published, cited, and then quietly retracted years later; a missed detection is never published at all and nobody knows it happened. The costs of the two errors are therefore wildly asymmetric in visibility and roughly symmetric in scientific damage, which pushes the whole enterprise towards thresholds that are too permissive. The corrective is not a stricter number — it is the practice of demanding a second, physically different observable before a signal is called a discovery, which is what a transit that runs late provides for a planet found in velocities and what a chromospheric index provides against one that is not.
It is worth stating the one habit that protects against nearly all of this, because it is cheap and it is not universal. Before believing any peak, run the identical search on the identical data with the signal destroyed — shuffle the measurements against their timestamps, or shift each night’s data by a random amount, or search a set of stars known to be constant. Whatever the noise really is, the shuffled version has the same sampling, the same gaps, the same number of points and the same aliasing structure, and the distribution of its tallest peaks is the null distribution the threshold should have come from. It costs a few seconds of computation and it makes no assumption about the noise at all. Nearly every retracted period in the literature would have failed that test, and the reason it is skipped is that it is nobody’s method — it produces no formula and it appears in no textbook chapter.
Where the ladder goes next
The immediate next rung is the model comparison that ought to replace the threshold: how the evidence for a periodic model is weighed against the evidence for a correlated-noise one, and why the answer depends on priors that are rarely stated. The rung beyond it is the case where the noise and the signal are the same physical object — a star whose rotation is both the contaminant and the quantity of interest.
What links here
Essays that link to this one from their own argument.
The objects this essay names
Each one links to every other essay that touches it.
BootstrapDetection thresholdFalse-alarm probabilityIndependent frequenciesLook elsewhere effectMultiple testingNull hypothesisPeriodogramRed noiseStellar jitter