Odds for a planet that depend on a guess about its size
Assumes Periodograms and The two-body problem.
The standard test for a periodic signal asks a question about the noise: if there were no planet, how often would a peak this tall appear somewhere in the periodogram? That is a false-alarm probability, and its value depends on how many frequencies were searched and on whether the noise is what the formula assumes. It says nothing directly about the planet. A small false-alarm probability means the noise model is a poor explanation of the data; it does not say how good the planet is as an alternative.
The Bayesian test asks the question that was actually wanted. Given the measured velocities, how much more probable are they if the star has a planet than if it does not? The ratio is the Bayes factor, and multiplied by whatever prior odds one holds for a planet, it gives the posterior odds. It compares two explanations rather than rejecting one, and it can favour the null hypothesis as well as the signal. That makes it the natural replacement for a threshold, and radial-velocity surveys have adopted it widely.
The price is that each explanation has to be stated completely enough to predict the data before they are seen. A model “with a planet” has to say which periods and which amplitudes it thinks likely, and the answer changes the Bayes factor. Some of those choices matter very little. One matters without limit.
A Bayes factor at one period
Suppose the period were known. Then a circular orbit adds to the velocities, with two unknown amplitudes, and the model is linear in them. Give each a Gaussian prior with scale — so the semi-amplitude has a Rayleigh prior peaking at — and give the noise a Gaussian covariance . Everything is Gaussian, the integrals over the amplitudes can be done exactly, and the logarithm of the Bayes factor has a closed form:
with and built from the two columns of sines and cosines, and .
The two terms are the whole of Bayesian model comparison in miniature. The first is the reward: how much better the data are fitted with the sinusoid than without it, which for white noise is essentially the Lomb–Scargle power, and grows as the square of the signal-to-noise ratio. The second is the penalty, the Occam factor. It is the logarithm of the ratio between how widely the prior spread its belief over amplitudes and how narrowly the data confine them. A model that allowed a wide range of amplitudes and needed only a small part of that range is charged for the rest.
At 120 points with 2 m/s errors, each amplitude component is measured to about 0.26 m/s. A prior scale of 3 m/s is about twelve times wider, twice over, and the penalty is roughly . That is why the scan in the first figure sits below zero at almost every period: a sinusoid that fits nothing still pays for the amplitudes it was prepared to have.
The penalty for not knowing the period
The planet’s period is not known in advance, and the question “is there a planet” means “is there a planet at some period”. The Bayes factor for that is the average of the one-period Bayes factor over the prior on the period. It is an average of , not of , so a single tall peak dominates it, but the average is taken over the whole prior range and most of the range contributes almost nothing.
In the first figure the peak reaches and the average is 7.7. The difference, 6.7, is the Bayesian look-elsewhere penalty. It is roughly the logarithm of the number of distinct periods the prior allowed, each weighted by how much prior belief it received, and it plays the same part as the trials factor that turns a single-frequency false-alarm probability into a global one. The difference is that it is not bolted on afterwards. It falls out of the definition of the question, and it is paid in the same currency as the evidence.
The count of distinct periods is set by the observations, not by the grid. With a baseline of 359 days two frequencies closer together than about one part in 359 cycles per day cannot be told apart. Between 1.5 and 1,000 days there are some 240 resolvable frequencies, and the gaps in the sampling spread the evidence among aliases as well. A search over the same range with a baseline ten times longer would have ten times as many distinct periods and would pay about more.
Which period prior
There are two common choices for the period prior, and both are defensible. A prior uniform in the logarithm of the period says that no scale is preferred: a planet is as likely between 1 and 10 days as between 10 and 100. The known planet population is roughly like that over the range where surveys are sensitive. A prior uniform in frequency says that the data have equal resolution everywhere in frequency, which is true of the periodogram but not of planets, and it puts most of its belief at short periods, because most of the frequency range lies there.
The figure measures how much the choice matters for this data set. Under a log-uniform prior the evidence falls by only 1.1 as the upper end of the range goes from 20 days to 3,000, because the fraction of the prior near 11.3 days falls only as the logarithm of the range. A frequency-uniform prior is much more sensitive to where its short end is placed: moving it from 8 days to half a day costs 2.6 in , because each halving of the shortest period roughly doubles the frequency range and halves the weight near the planet. At the widest setting the two priors disagree by 0.9, a factor of about two and a half in the odds.
That is a real dependence but a bounded one. It is logarithmic in the choices, and no reasonable period prior changes a Bayes factor of 7.7 into one favouring no planet. The period prior is the part of the model that the data constrain most tightly, and the penalty for being vague about it is small.
Which amplitude prior
The amplitude prior is different in kind, and it helps to translate amplitudes into planets first. A star’s reflex velocity has a semi-amplitude of about 28.4 m/s for a planet of one Jupiter mass on a one-year orbit around a star like the Sun, scaling linearly with the planet’s mass and as the inverse cube root of its period. At 11.3 days the factor from the period is about 3.2, so the planet in these figures, at 1 m/s, has a minimum mass of about 3.5 Earth masses — a super-Earth, of the kind that surveys find around a large fraction of Sun-like stars. A prior scale of 3 m/s says planets up to about ten Earth masses are expected. A scale of 200 m/s says planets of two Jupiter masses are as plausible as small ones.
A narrow prior forbids the amplitude the data need, and the evidence falls. That side is unremarkable. The other side is the problem. Once the prior scale is much larger than the amplitude the data allow, the reward term stops changing — the fit is the same — and the Occam penalty keeps growing, by , because the prior spreads more and more of its belief over amplitudes that the data rule out. Two amplitude components each cost per decade of prior width, and the figure’s slope is exactly per decade.
There is no floor. A prior of 200 m/s — hardly absurd for a survey that includes stars with hot Jupiters, whose semi-amplitudes run to hundreds of metres a second — turns a Bayes factor of about two thousand to one into even odds, and a prior of 1,000 m/s makes the data favour no planet by about forty to one. A prior that is uniform over all amplitudes, the usual expression of ignorance, is improper: it cannot be normalised, the Occam factor is infinite, and the Bayes factor is zero whatever the data.
This is Lindley’s paradox, described in 1957: when a precise hypothesis is compared with a vague one, the Bayes factor can favour the precise hypothesis however strongly a significance test rejects it, provided the vague one is vague enough. In a planet search the precise hypothesis is “no planet”, which predicts zero amplitude exactly, and the vague one is “a planet of some size”. A frequentist test at the same data would report the 11.3-day peak as highly significant, whatever anyone thought about amplitudes in advance.
The paradox is not a flaw in the arithmetic. It says that the question “how much more probable are the data with a planet” has no answer until “a planet” is specified, and the specification has to include how big it is likely to be. The honest ways to proceed are to derive the amplitude prior from something measured — the occurrence rates of planets by mass and period, converted to semi-amplitudes for the star in hand — or to report the Bayes factor as a function of the prior scale, as this figure does, so that a reader with a different prior can read off their own answer.
How much evidence a planet earns
The first figure used one data set. Other data sets of identical quality — the same times, the same errors, the same planet, different noise — give different answers, and the spread is part of the result.
The first thing the figure shows is a property significance tests lack. With no planet, and with planets well below the noise, the Bayes factor is negative: the data actively favour no signal, at odds of about twenty to one. A false-alarm probability can only fail to reject; it can never report evidence of absence. A negative Bayes factor can, and for survey statistics — deciding how many stars have no detectable planet, rather than how many have one — that matters.
The second is that the first figure’s data set was a fortunate one. At 1 m/s its of 7.7 lies near the top of its bar, and the median at that amplitude is . Half of all data sets with the same planet and the same measurements would favour no planet. The planet crosses the conventional threshold of , odds of about 150 to one, only in the median at 1.5 m/s. Between those amplitudes the outcome is decided by the particular noise, and the width of the bars — several units of — is larger than the whole disagreement between the two period priors.
The third is the shape. Below the noise level, barely moves, because the reward term is smaller than the Occam penalty; above it, climbs as the square of the amplitude, because the reward grows as the squared signal-to-noise ratio and the penalty only as its logarithm. A planet twice as large does not give twice the evidence; it gives four times the log-evidence, and odds larger by a power.
When the noise is a star
So far the noise has been white. A real star’s velocities are not. Spots and bright regions on a rotating star suppress the light from one limb and then the other, shift the centroids of the thousands of lines that make up a velocity measurement, and produce a signal that repeats with the rotation period while the spots last and wanders as they evolve. It is periodic enough to look like a planet and irregular enough not to be one.
Analysed with a white-noise model, the spotted star produces overwhelming evidence for a planet that does not exist. There is a forest of peaks around 27 days, and the tallest peak of all lies at a few hundred days, where the spot pattern’s slow drift over the observing season resembles a long-period orbit. The global Bayes factor is . Every part of the calculation is correct except the null hypothesis, which assumed noise the star does not have. A Bayes factor compares two models, and if both are wrong it reports which is less wrong.
The standard remedy is to describe the activity as a Gaussian process: a noise model in which the covariance between two measurements depends on their separation in time through a kernel. The quasi-periodic kernel used here is
which says that velocities one rotation apart are correlated, strongly if the spots have not changed and weakly if they have, with the spot lifetime. With the kernel in the null model, and its parameters held at their true values, the Bayes factor drops to . The data are the same. What changed is the explanation the planet had to beat.
The planet the remedy hides
A noise model flexible enough to absorb a quasi-periodic signal at the rotation period is also flexible enough to absorb a real planet there.
A 3 m/s planet, larger than the per-point noise, is found with odds of or so at most periods. When its period matches the star’s rotation, the evidence falls to nothing. The kernel can generate a sinusoid at exactly that period, with the right amplitude, as a plausible draw of spot noise, and the Bayes factor has no information with which to prefer an orbit. Near half the rotation period, where the kernel has its first harmonic, the evidence is reduced but not removed.
This is not a defect of the method. It is an honest statement that velocities alone cannot separate a planet from spots at the same period. The separation needs different information: activity indicators measured from the same spectra, which follow the spots and not the planet; the dependence of the velocity signal on wavelength, which spots produce and orbits do not; or photometry, which sees the spots directly. Modern analyses fit the velocities and the indicators jointly, with shared Gaussian processes, and the evidence for a planet near the rotation period then comes from what the indicators fail to show.
What the number can and cannot be used for
A Bayes factor is a statement about two fully specified models, and it shares that property with every inference that has to choose among answers the data cannot separate, from an image reconstructed from an interferometer’s incomplete measurements to a cosmological fit whose conclusion depends on the functional form it was given. Three of the choices this one depends on have appeared here. The period prior moves it by a unit or two. The noise model can move it by thirty. The amplitude prior can move it by any amount at all. None of that is hidden in the formalism; it is written in the models.
Those dependences are why independent analyses of the same velocities have reported evidences that differ by several units of for the same planet — through different amplitude priors, different noise kernels, and different numerical methods for averaging over parameters where, unlike here, no closed form exists. They are also why the analysis in this essay holds the activity kernel’s parameters fixed. Letting them float is standard practice, and the evidence then includes a further Occam factor for the kernel’s own flexibility, which is itself sensitive to the priors on the kernel’s parameters.
The Bayes factor is at its most useful in comparisons where the contested choices cancel: two planet models with the same amplitude prior, one with an extra planet or an eccentric orbit; the same model on two data sets; one survey’s detections against its non-detections, computed the same way. It is at its least useful as an absolute threshold for announcing a discovery, which is how it is often used, because the threshold’s meaning depends on a prior that a reader cannot check unless it is stated.
Still open: the prior as a measurement
The amplitude prior does not have to be a guess. The distribution of planet masses and periods is measured by the same surveys that use the priors, and the prior for a star can be built from that distribution, converted to semi-amplitudes for the star’s mass and the orbit’s inclination. The difficulty is circular: the measured distribution below a few metres a second comes from exactly the marginal detections whose evidence depends on the prior. Whether a population model and the individual detections that define it can be fitted together consistently, at the amplitudes where the next generation of instruments will find Earth-mass planets around Sun-like stars, is not settled. Until it is, the most useful thing a planet claim near the noise can report is not one Bayes factor but the curve of the fourth figure: the evidence against the width of the amplitude prior, so that the reader’s own prior can be read off it.
About the same objects
Not linked from either essay — found by the objects both name.
- A velocity measured from a shape radial velocity · stellar rotation
- One eccentric planet, or two circular ones periodogram · radial velocity
The objects this essay names
Each one links to every other essay that touches it.
Bayes factorFalse-alarm probabilityGaussian processLindley's paradoxLook elsewhere effectPeriodogramPriorRadial velocityRed noiseStarspotsStellar activityStellar rotation