Starlight

Odds for a planet that depend on a guess about its size

A false-alarm probability asks how often noise alone would produce a peak this tall. A Bayes factor asks something better — how much more probable the data are with a planet than without one — and pays for it with two numbers nobody measured. The width allowed for the period barely matters. The width allowed for the amplitude can turn strong evidence into none, from the same velocities.

Assumes Periodograms and The two-body problem.

The standard test for a periodic signal asks a question about the noise: if there were no planet, how often would a peak this tall appear somewhere in the periodogram? That is a false-alarm probability, and its value depends on how many frequencies were searched and on whether the noise is what the formula assumes. It says nothing directly about the planet. A small false-alarm probability means the noise model is a poor explanation of the data; it does not say how good the planet is as an alternative.

The Bayesian test asks the question that was actually wanted. Given the measured velocities, how much more probable are they if the star has a planet than if it does not? The ratio is the Bayes factor, and multiplied by whatever prior odds one holds for a planet, it gives the posterior odds. It compares two explanations rather than rejecting one, and it can favour the null hypothesis as well as the signal. That makes it the natural replacement for a threshold, and radial-velocity surveys have adopted it widely.

The price is that each explanation has to be stated completely enough to predict the data before they are seen. A model “with a planet” has to say which periods and which amplitudes it thinks likely, and the answer changes the Bayes factor. Some of those choices matter very little. One matters without limit.

The evidence for a period, one trial period at a time. The natural logarithm of the Bayes factor for a sinusoid at each trial period against no signal, for 120 radial velocities spread over 359 days with 2.0 m/s white noise per point and a planet of semi-amplitude 1.0 m/s at 11.3 days. The amplitude prior is Gaussian with scale 3.0 m/s in each component. At the planet's period the evidence is 14.4 — odds of about 1.8·10⁶ to one — and almost everywhere else it is near zero or negative, because a sinusoid that fits nothing is penalised for the amplitude range it wasted. The quantity that answers the question "is there a planet at some period" is not the peak but the average of the Bayes factor over the prior on the period, here uniform in the logarithm of the period from 1.5 to 1,000 days, and that average is ln B = 7.7: 6.7 lower than the peak, the Bayesian form of the look-elsewhere penalty, paid because the data had to find the period as well as the amplitude.
Fig. 1 The Bayes factor for a sinusoid at each trial period against pure noise, for 120 velocities over a year with 2 m/s errors and a planet of 1 m/s at 11.3 days. At the planet’s period the logarithm of the Bayes factor reaches 14.4. Averaged over a prior uniform in the logarithm of the period from 1.5 to 1,000 days it is 7.7 — the price of not knowing the period in advance.

A Bayes factor at one period

Suppose the period were known. Then a circular orbit adds asin⁡ωt+bcos⁡ωta\sin\omega t + b\cos\omega t to the velocities, with two unknown amplitudes, and the model is linear in them. Give each a Gaussian prior with scale σa\sigma_a — so the semi-amplitude K=a2+b2K = \sqrt{a^2 + b^2} has a Rayleigh prior peaking at σa\sigma_a — and give the noise a Gaussian covariance CC. Everything is Gaussian, the integrals over the amplitudes can be done exactly, and the logarithm of the Bayes factor has a closed form:

ln⁡B(ω)=12 bTM−1b  −  12ln⁡det⁡ ⁣(I+σa2A),\ln B(\omega) = \tfrac12\, \mathbf b^{\mathsf T} M^{-1} \mathbf b \;-\; \tfrac12 \ln\det\!\left(I + \sigma_a^2 A\right),

with A=XTC−1XA = X^{\mathsf T} C^{-1} X and b=XTC−1y\mathbf b = X^{\mathsf T} C^{-1} \mathbf y built from the two columns XX of sines and cosines, and M=I/σa2+AM = I/\sigma_a^2 + A.

The two terms are the whole of Bayesian model comparison in miniature. The first is the reward: how much better the data are fitted with the sinusoid than without it, which for white noise is essentially the Lomb–Scargle power, and grows as the square of the signal-to-noise ratio. The second is the penalty, the Occam factor. It is the logarithm of the ratio between how widely the prior spread its belief over amplitudes and how narrowly the data confine them. A model that allowed a wide range of amplitudes and needed only a small part of that range is charged for the rest.

At 120 points with 2 m/s errors, each amplitude component is measured to about 0.26 m/s. A prior scale of 3 m/s is about twelve times wider, twice over, and the penalty is roughly ln⁡122≈5\ln 12^2 \approx 5. That is why the scan in the first figure sits below zero at almost every period: a sinusoid that fits nothing still pays for the amplitudes it was prepared to have.

The penalty for not knowing the period

The planet’s period is not known in advance, and the question “is there a planet” means “is there a planet at some period”. The Bayes factor for that is the average of the one-period Bayes factor over the prior on the period. It is an average of BB, not of ln⁡B\ln B, so a single tall peak dominates it, but the average is taken over the whole prior range and most of the range contributes almost nothing.

In the first figure the peak reaches ln⁡B=14.4\ln B = 14.4 and the average is 7.7. The difference, 6.7, is the Bayesian look-elsewhere penalty. It is roughly the logarithm of the number of distinct periods the prior allowed, each weighted by how much prior belief it received, and it plays the same part as the trials factor that turns a single-frequency false-alarm probability into a global one. The difference is that it is not bolted on afterwards. It falls out of the definition of the question, and it is paid in the same currency as the evidence.

The count of distinct periods is set by the observations, not by the grid. With a baseline of 359 days two frequencies closer together than about one part in 359 cycles per day cannot be told apart. Between 1.5 and 1,000 days there are some 240 resolvable frequencies, and the gaps in the sampling spread the evidence among aliases as well. A search over the same range with a baseline ten times longer would have ten times as many distinct periods and would pay about ln⁡10≈2.3\ln 10 \approx 2.3 more.

Which period prior

What the width of the period prior costs. The global ln Bayes factor for the same 1 m/s planet at 11.3 days, recomputed for different period priors. Blue: a prior uniform in the logarithm of the period, from 1.5 days up to the longest period on the horizontal axis; widening it from 20 to 3,000 days lowers ln B from 8.7 to 7.6, because the prior mass near the true period falls only as the logarithm of the width. Red: the prior's short end moved instead, from 8 days down to 0.5 with the long end at 1,000 days, for a prior uniform in frequency — ln B goes from 9.3 to 6.7, because a frequency-uniform prior puts almost all its weight on short periods and each halving of the shortest period halves the weight left near the planet. The dashed curve is the log-uniform prior over the same short ends: 8.1 to 7.7. Two defensible priors disagree by 0.9 in ln B at the widest setting — a factor of 3 in the odds.
Fig. 2 The global evidence for the same planet under different period priors. Widening a log-uniform prior from 20 to 3,000 days lowers ln⁡B\ln B from 8.7 to 7.6. A prior uniform in frequency is far more sensitive to its short end: moving that end from 8 days to half a day lowers ln⁡B\ln B from 9.3 to 6.7, because nearly all of such a prior’s weight lies at short periods.

There are two common choices for the period prior, and both are defensible. A prior uniform in the logarithm of the period says that no scale is preferred: a planet is as likely between 1 and 10 days as between 10 and 100. The known planet population is roughly like that over the range where surveys are sensitive. A prior uniform in frequency says that the data have equal resolution everywhere in frequency, which is true of the periodogram but not of planets, and it puts most of its belief at short periods, because most of the frequency range lies there.

The figure measures how much the choice matters for this data set. Under a log-uniform prior the evidence falls by only 1.1 as the upper end of the range goes from 20 days to 3,000, because the fraction of the prior near 11.3 days falls only as the logarithm of the range. A frequency-uniform prior is much more sensitive to where its short end is placed: moving it from 8 days to half a day costs 2.6 in ln⁡B\ln B, because each halving of the shortest period roughly doubles the frequency range and halves the weight near the planet. At the widest setting the two priors disagree by 0.9, a factor of about two and a half in the odds.

That is a real dependence but a bounded one. It is logarithmic in the choices, and no reasonable period prior changes a Bayes factor of 7.7 into one favouring no planet. The period prior is the part of the model that the data constrain most tightly, and the penalty for being vague about it is small.

Which amplitude prior

The amplitude prior is different in kind, and it helps to translate amplitudes into planets first. A star’s reflex velocity has a semi-amplitude of about 28.4 m/s for a planet of one Jupiter mass on a one-year orbit around a star like the Sun, scaling linearly with the planet’s mass and as the inverse cube root of its period. At 11.3 days the factor from the period is about 3.2, so the planet in these figures, at 1 m/s, has a minimum mass of about 3.5 Earth masses — a super-Earth, of the kind that surveys find around a large fraction of Sun-like stars. A prior scale of 3 m/s says planets up to about ten Earth masses are expected. A scale of 200 m/s says planets of two Jupiter masses are as plausible as small ones.

The same data, and odds that depend on a guess about the amplitude. The global ln Bayes factor for the 1 m/s planet at 11.3 days against the scale of the Gaussian prior on the sinusoid's two amplitude components, from 0.3 to 1,000 m/s, the data held fixed. The evidence is largest, ln B = 8.8, near a scale of 1 m/s, comparable to the true amplitude. Narrower, the prior forbids the amplitude the data demand. Broader, it spreads its belief over amplitudes the data rule out, and every factor of ten in the scale costs 4.61 in ln B — two powers of ten, one for each amplitude component, so 2 ln 10 = 4.61 — with no limit. Past a scale of about 200 m/s the odds favour no planet at all, from the same measurements. This is Lindley's paradox: with a proper prior broad enough, any signal can be made to lose, and an improper one makes the Bayes factor undefined.
Fig. 3 The same data, with the global evidence recomputed for amplitude priors from 0.3 to 1,000 m/s. The evidence peaks at ln⁡B=8.8\ln B = 8.8 for a prior scale of about 1 m/s, comparable to the true amplitude. Beyond that, each factor of ten in the prior scale costs 4.61 in ln⁡B\ln B, without limit; above about 200 m/s the same velocities favour no planet.

A narrow prior forbids the amplitude the data need, and the evidence falls. That side is unremarkable. The other side is the problem. Once the prior scale is much larger than the amplitude the data allow, the reward term stops changing — the fit is the same — and the Occam penalty keeps growing, by ln⁡σa2\ln \sigma_a^2, because the prior spreads more and more of its belief over amplitudes that the data rule out. Two amplitude components each cost ln⁡10\ln 10 per decade of prior width, and the figure’s slope is exactly 2ln⁡10=4.612\ln 10 = 4.61 per decade.

There is no floor. A prior of 200 m/s — hardly absurd for a survey that includes stars with hot Jupiters, whose semi-amplitudes run to hundreds of metres a second — turns a Bayes factor of about two thousand to one into even odds, and a prior of 1,000 m/s makes the data favour no planet by about forty to one. A prior that is uniform over all amplitudes, the usual expression of ignorance, is improper: it cannot be normalised, the Occam factor is infinite, and the Bayes factor is zero whatever the data.

This is Lindley’s paradox, described in 1957: when a precise hypothesis is compared with a vague one, the Bayes factor can favour the precise hypothesis however strongly a significance test rejects it, provided the vague one is vague enough. In a planet search the precise hypothesis is “no planet”, which predicts zero amplitude exactly, and the vague one is “a planet of some size”. A frequentist test at the same data would report the 11.3-day peak as highly significant, whatever anyone thought about amplitudes in advance.

The paradox is not a flaw in the arithmetic. It says that the question “how much more probable are the data with a planet” has no answer until “a planet” is specified, and the specification has to include how big it is likely to be. The honest ways to proceed are to derive the amplitude prior from something measured — the occurrence rates of planets by mass and period, converted to semi-amplitudes for the star in hand — or to report the Bayes factor as a function of the prior scale, as this figure does, so that a reader with a different prior can read off their own answer.

How much evidence a planet earns

The first figure used one data set. Other data sets of identical quality — the same times, the same errors, the same planet, different noise — give different answers, and the spread is part of the result.

How much evidence a planet of each size earns. The global ln Bayes factor for a planet at 11.3 days against its true semi-amplitude, for 8 independent data sets of about 120 points with 2.0 m/s noise at each amplitude: the bars span the smallest to the largest, the line joins the medians. With no planet the median is −3.0, in favour of no signal, as it should be. The evidence rises slowly while the planet is smaller than the noise per point and then very fast, because ln B grows as the square of the signal-to-noise ratio: 0 m/s, −3.0; 0.25 m/s, −2.9; 0.5 m/s, −2.8; 0.75 m/s, −2.7; 1 m/s, −2.0; 1.25 m/s, 1.1; 1.5 m/s, 6.3; 1.75 m/s, 12.5; 2 m/s, 19.7. The median first passes ln B = 5 — odds of about 150 to one, a conventional threshold for strong evidence — at 1.5 m/s. Across data sets of identical quality the spread is several units of ln B, which is larger than most of the disagreements between priors.
Fig. 4 The global evidence against the true semi-amplitude, for eight independent noise realisations at each amplitude. With no planet the median is ln⁡B=−3.0\ln B = -3.0: the evidence correctly favours no signal. It stays near that until the planet approaches the per-point noise, then climbs as the square of the amplitude, passing ln⁡B=5\ln B = 5 at about 1.5 m/s. The bars span several units at every amplitude above 1 m/s.

The first thing the figure shows is a property significance tests lack. With no planet, and with planets well below the noise, the Bayes factor is negative: the data actively favour no signal, at odds of about twenty to one. A false-alarm probability can only fail to reject; it can never report evidence of absence. A negative Bayes factor can, and for survey statistics — deciding how many stars have no detectable planet, rather than how many have one — that matters.

The second is that the first figure’s data set was a fortunate one. At 1 m/s its ln⁡B\ln B of 7.7 lies near the top of its bar, and the median at that amplitude is −2.0-2.0. Half of all data sets with the same planet and the same measurements would favour no planet. The planet crosses the conventional threshold of ln⁡B=5\ln B = 5, odds of about 150 to one, only in the median at 1.5 m/s. Between those amplitudes the outcome is decided by the particular noise, and the width of the bars — several units of ln⁡B\ln B — is larger than the whole disagreement between the two period priors.

The third is the shape. Below the noise level, ln⁡B\ln B barely moves, because the reward term is smaller than the Occam penalty; above it, ln⁡B\ln B climbs as the square of the amplitude, because the reward grows as the squared signal-to-noise ratio and the penalty only as its logarithm. A planet twice as large does not give twice the evidence; it gives four times the log-evidence, and odds larger by a power.

When the noise is a star

So far the noise has been white. A real star’s velocities are not. Spots and bright regions on a rotating star suppress the light from one limb and then the other, shift the centroids of the thousands of lines that make up a velocity measurement, and produce a signal that repeats with the rotation period while the spots last and wanders as they evolve. It is periodic enough to look like a planet and irregular enough not to be one.

A spotted star, and a planet that is its rotation. ln Bayes factor against trial period for a star with no planet, whose velocities carry 2.0 m/s white noise and a quasi-periodic signal from rotating spots — a Gaussian process with a 27-day rotation period, 3.0 m/s amplitude and spots that live about 40 days. Blue: analysed with white noise only. The evidence reaches 33 at 417.9 days, 23 near the rotation period and 1 near half of it, and the global ln B is 30.3 — overwhelming odds for a planet that does not exist. Red: the same data with the activity included in the noise model, the kernel's parameters held at their true values. The peaks are gone and the global ln B is −1.8. Nothing in the data changed; the alternative the planet was compared with did.
Fig. 5 A star with no planet, whose velocities carry 2 m/s of white noise and a quasi-periodic activity signal with a 27-day rotation period and spots lasting about 40 days. With a white-noise model the evidence peaks near the rotation period and even higher at long periods, where the slow evolution of the spots looks like a sinusoid; the global ln⁡B\ln B is 30.3. With the activity included in the noise model it is −1.8-1.8.

Analysed with a white-noise model, the spotted star produces overwhelming evidence for a planet that does not exist. There is a forest of peaks around 27 days, and the tallest peak of all lies at a few hundred days, where the spot pattern’s slow drift over the observing season resembles a long-period orbit. The global Bayes factor is e30e^{30}. Every part of the calculation is correct except the null hypothesis, which assumed noise the star does not have. A Bayes factor compares two models, and if both are wrong it reports which is less wrong.

The standard remedy is to describe the activity as a Gaussian process: a noise model in which the covariance between two measurements depends on their separation in time through a kernel. The quasi-periodic kernel used here is

k(τ)=h2exp⁡ ⁣(−τ22λ2−Γsin⁡2πτProt),k(\tau) = h^2 \exp\!\left(-\frac{\tau^2}{2\lambda^2} - \Gamma \sin^2\frac{\pi\tau}{P_{\mathrm{rot}}}\right),

which says that velocities one rotation apart are correlated, strongly if the spots have not changed and weakly if they have, with λ\lambda the spot lifetime. With the kernel in the null model, and its parameters held at their true values, the Bayes factor drops to −1.8-1.8. The data are the same. What changed is the explanation the planet had to beat.

The planet the remedy hides

A noise model flexible enough to absorb a quasi-periodic signal at the rotation period is also flexible enough to absorb a real planet there.

A real planet near the rotation period, absorbed by the model that removed the false one. The global ln Bayes factor for a real planet of 3.0 m/s semi-amplitude on the same spotted star, analysed with the activity kernel in the noise model, against the planet's orbital period; 3 data sets at each period, the bar spanning them and the line through the middle one. Away from the rotation period and its half the planet is found easily, with a median ln B near 42. At 27 days, the rotation period, the median falls to −1.2, and at 13.5 days, half of it, to 13.7: a quasi-periodic kernel with that period can reproduce a sinusoid at that period as well as the planet can, so the data cannot say which one they contain, and a model comparison that is honest about the activity is also blind to a planet whose orbit matches the star's spin.
Fig. 6 A real 3 m/s planet on the same spotted star, analysed with the activity model, against its orbital period. Away from the rotation period the planet earns ln⁡B\ln B of about 40. At 27 days the median falls to −1.2-1.2, and at half the rotation period to 13.7: the activity kernel reproduces a sinusoid at those periods as well as the planet does, and the comparison cannot tell them apart.

A 3 m/s planet, larger than the per-point noise, is found with odds of e40e^{40} or so at most periods. When its period matches the star’s rotation, the evidence falls to nothing. The kernel can generate a sinusoid at exactly that period, with the right amplitude, as a plausible draw of spot noise, and the Bayes factor has no information with which to prefer an orbit. Near half the rotation period, where the kernel has its first harmonic, the evidence is reduced but not removed.

This is not a defect of the method. It is an honest statement that velocities alone cannot separate a planet from spots at the same period. The separation needs different information: activity indicators measured from the same spectra, which follow the spots and not the planet; the dependence of the velocity signal on wavelength, which spots produce and orbits do not; or photometry, which sees the spots directly. Modern analyses fit the velocities and the indicators jointly, with shared Gaussian processes, and the evidence for a planet near the rotation period then comes from what the indicators fail to show.

What the number can and cannot be used for

A Bayes factor is a statement about two fully specified models, and it shares that property with every inference that has to choose among answers the data cannot separate, from an image reconstructed from an interferometer’s incomplete measurements to a cosmological fit whose conclusion depends on the functional form it was given. Three of the choices this one depends on have appeared here. The period prior moves it by a unit or two. The noise model can move it by thirty. The amplitude prior can move it by any amount at all. None of that is hidden in the formalism; it is written in the models.

Those dependences are why independent analyses of the same velocities have reported evidences that differ by several units of ln⁡B\ln B for the same planet — through different amplitude priors, different noise kernels, and different numerical methods for averaging over parameters where, unlike here, no closed form exists. They are also why the analysis in this essay holds the activity kernel’s parameters fixed. Letting them float is standard practice, and the evidence then includes a further Occam factor for the kernel’s own flexibility, which is itself sensitive to the priors on the kernel’s parameters.

The Bayes factor is at its most useful in comparisons where the contested choices cancel: two planet models with the same amplitude prior, one with an extra planet or an eccentric orbit; the same model on two data sets; one survey’s detections against its non-detections, computed the same way. It is at its least useful as an absolute threshold for announcing a discovery, which is how it is often used, because the threshold’s meaning depends on a prior that a reader cannot check unless it is stated.

Still open: the prior as a measurement

The amplitude prior does not have to be a guess. The distribution of planet masses and periods is measured by the same surveys that use the priors, and the prior for a star can be built from that distribution, converted to semi-amplitudes for the star’s mass and the orbit’s inclination. The difficulty is circular: the measured distribution below a few metres a second comes from exactly the marginal detections whose evidence depends on the prior. Whether a population model and the individual detections that define it can be fitted together consistently, at the amplitudes where the next generation of instruments will find Earth-mass planets around Sun-like stars, is not settled. Until it is, the most useful thing a planet claim near the noise can report is not one Bayes factor but the curve of the fourth figure: the evidence against the width of the amplitude prior, so that the reader’s own prior can be read off it.

About the same objects

Not linked from either essay — found by the objects both name.

The objects this essay names

Each one links to every other essay that touches it.

Bayes factorFalse-alarm probabilityGaussian processLindley's paradoxLook elsewhere effectPeriodogramPriorRadial velocityRed noiseStarspotsStellar activityStellar rotation