Concept

Detection threshold — where it appears

The smallest signal a survey can claim. For a rare-event search it is set by the false-positive rate across all the samples examined rather than by the noise on any one. Lowering it raises the real rate and the spurious one together.

Named by 6 essays across 3 fields — each of them below, with the objects they name alongside it.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 1.4 Earth masses, cut off at 1.21 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.

Every survey draws a different sky

The first exoplanets found were enormous and impossibly close to their stars. That was not a discovery about planets. It was a measurement of what a 10 m/s spectrograph watching for three years is able to see.

exoplanets · Detection bias
Four defensible boxes, and a factor of 2.9 between the answers. Above: the occurrence surface in starlight received against planet radius, with four published definitions of "an Earth-size planet in the habitable zone" drawn on it as rectangles. The surface is a stated parameterisation — two lognormal populations at 1.3 and 2.4 Earth radii, with the radius valley at 1.9 between them, normalised so that the whole of it comes to 0.5 planets per star between one and four Earth radii inside a hundred days. Integrating it over the four boxes gives 17%, 8%, 6%, 10% — a factor of 2.9 between the conservative zone and a broad definition, before any error bar is attached to any of them. The dashed curve is what a transit survey can actually see: at one year around a Sun-like star the smallest detectable planet is 1.6 Earth radii, so the lower-left corner of every box contains no detections at all and the rate there is an extrapolation of a fitted surface rather than a count of anything. Below: the same integral with only one corner moved. Holding the flux range fixed and sliding the radius bound from 1.5 to 1.75 Earth radii — a quarter of an Earth radius, well inside the uncertainty of a measured planetary radius — changes the answer by 43 per cent, which is larger than every error bar quoted with any of these numbers. The published values of η⊕ span two per cent to sixty; roughly a factor of 2.9 of that is definition, and the rest is how far each author was willing to extrapolate past the dashed line.

The part of a rate that is a definition

Published values for the frequency of Earth-size planets in habitable zones span two per cent to sixty. The spread is not measurement error — it is where the box was drawn, on a surface that is steepest exactly at the corner every author has to choose, and outside the last detection.

exoplanets · Occurrence rates
A peak worth 10.8 in a narrow search is worth nothing in a wide one. The probability that noise alone produces a peak at least as tall as a given power, for searches over four different numbers of independent frequencies. A single frequency examined in isolation gives a one-per-cent chance at a power of 4.6; searching fifty thousand frequencies for the same one-per-cent chance requires 15.4. The threshold rises as the logarithm of the width of the search, which is why the penalty is survivable — but it is a penalty, it is often not applied, and the number of independent frequencies in an unevenly sampled time series is not the number of frequencies on the grid. Overestimating that count is conservative and underestimating it is not, which is the one asymmetry worth remembering.

The tallest peak in nothing at all

A periodogram of pure noise has peaks in it, and the tallest is not small. How tall it has to be before it means something depends on how many frequencies were searched and on what the noise actually is — and astronomical noise is almost never the white noise the standard formula assumes.

starlight · Periodograms
Below 1.3 km a shadow stops getting smaller. The relative rate at which a star is occulted, against the smallest body a survey can detect, for size distributions with slopes 3.5, 4, 4.5. The cross-section of a body is not its own diameter: diffraction gives every shadow a minimum width of about the Fresnel scale, √(λD/2), which at 40 AU and 550 nm is 1.28 km. Above that the cross-section grows with the body, so lowering the limit gains events as the limit to the power -1.5 for the shallowest distribution drawn; below it the cross-section stops shrinking and only the number of bodies keeps rising, which is a shallower gain by one power. Pushing the limit from 40 km to 0.2 multiplies the rate by 3106196 at the steepest slope and 14131 at the shallowest — so the rate a survey measures is a measurement of the size distribution, which is the quantity a collisional history predicts and nothing else can reach at these sizes.

A population counted by shadows that never repeat

A body a kilometre across at forty astronomical units is a hundred million times too faint to image and casts a shadow just as dark as a large one. Monitoring enough stars fast enough catches those shadows — each one a single unrepeatable event of a fraction of a second, and the measurement is not any event but the rate.

sky · Occultations
What comes back is a ramp, not a threshold. Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The measured curve is a gamma cumulative distribution of shape 4.65 and scale 0.98 beginning at 4.1, which is the form a survey's own injection tests are fitted with; the dashed line is the step at 7.1 that a threshold calculation assumes instead. Half the injections are recovered at 8.33, 1.2 units above the nominal threshold — the ramp is a property of the search and the cut is a separate decision, so the two need not meet anywhere in particular. The rest of the disagreement is the area between the curves. The pipeline does not reach 99 per cent efficiency until 14.9, four units above the threshold, and it recovers 45 per cent one unit above it. Over a population whose signal-to-noise falls as s^-2 — which is what a planet population looks like, because there are far more small planets than large ones — the step function counts 1.21 times as many detections as the ramp does. That factor is not an error bar. It multiplies every occurrence rate computed without it, and it is larger for the small planets than for the large ones, because the small ones live where the ramp is.

The threshold that is not a threshold

A survey's detection limit is quoted as a number — seven point one — and a pipeline does not behave that way. Half the injected signals come back at the threshold, and full efficiency arrives four units above it.

exoplanets · Detection bias
Three biases against eccentricity, and they do not agree. Four quantities against orbital eccentricity, each relative to a circular orbit of the same semi-major axis, averaged over the argument of periastron. The transit probability rises as (1 − e²)⁻¹, because an eccentric planet spends part of its orbit inside its own semi-major axis: at e = 0.5 a transit is 1.33 times as likely. The transit duration falls as √(1 − e²), so the event carries less signal-to-noise, and the two together — probability times the square root of the time in transit — come to 1.24 at the same eccentricity. They very nearly cancel, and that is the surprise: a transit survey has almost no eccentricity bias at all. The radial-velocity curve is the one that does. A Keplerian of eccentricity e puts less of its variance in the fundamental and more into harmonics no sinusoidal search is looking at — 68 per cent remains at e = 0.6 and 47 per cent at e = 0.8 — so a velocity survey loses amplitude exactly where a transit survey does not. What no figure here can show is which of these the measured eccentricity distribution is made of, because the correction depends on a detection pipeline rather than on geometry, and the two surveys have to be corrected separately before their answers can be compared.

Every method prefers a circle, and not for the same reason

A transit is more likely on an eccentric orbit and shorter when it happens, and the two very nearly cancel. A velocity curve loses amplitude to harmonics no sinusoidal search is looking at, and that one does not cancel at all.

exoplanets · Detection bias

Named alongside it

The objects these essays reach for when they reach for this one.

Selection effectSignal-to-noiseSurvey completenessDetection limitFalse-positiveOccurrence rateArgument of periastronBootstrapCollisional cascadeCompletenessDefinitional uncertaintyDiffraction

All concepts