Exoplanets

The planets that were not seen

An occurrence rate is a count divided by a probability, and the probability can be a five-hundredth. Everything difficult about saying how common planets are lives in that denominator.

Assumes Detection bias and Transits.

Kepler found about 2,700 confirmed planets, each of them a dip of a few hundred parts per million repeating on a fixed period. There are approximately 200 billion stars in the Galaxy. Nothing about the first number implies anything about the second until the question “how many did the mission miss” has been answered — and that question is nearly the whole of the subject.

The answer is a division, and the denominator is a probability that can be as small as a five-hundredth.

The valley is a slope, and the sign of it decides the mechanism. Planet radius against orbital period, both logarithmic, with the centre of the radius valley drawn as R ∝ P^-0.09. The exponent is the measurement that separates two explanations: photoevaporation, in which the star strips the atmospheres of close-in planets, predicts a valley that moves to smaller radii at longer periods — a negative exponent — while a gas-poor formation origin predicts a positive one. The line drawn here has slope -0.09, so it is a picture of photoevaporation; the measured value is negative, at about −0.09. The bands above and below the line are where the two populations sit.
Fig. 1 The valley’s slope, which is the measurement that separates two explanations. The centre of the radius gap moves with orbital period as RP0.09R \propto P^{-0.09}, and the sign of that exponent is the discriminant: photoevaporation, in which the star strips the atmospheres of close-in planets, predicts a valley that moves to smaller radii at longer periods, while a gas-poor formation origin predicts the opposite. The measured value is negative. That is a statement about which planets are missing and why, read off the boundary between two populations rather than from either.

The shape of the calculation

For a given kind of planet — a cell in the plane of period against radius — the occurrence rate per star is

f=Ndetectedjpgeom,jpdetect,j,f = \frac{N_{\text{detected}}}{\sum_j p_{\text{geom},j}\, p_{\text{detect},j}},

summed over the survey’s stars. The two probabilities are different in kind and it is worth keeping them apart.

The geometric probability is R/aR_\star/a, computable from the star’s radius and the period through the harmonic law. It involves no instrument and no judgement. For a survey of a hundred thousand stars, it is a hundred thousand exact numbers.

The detection probability is not analytic at all. It is the fraction of the time that a signal of those parameters, injected into that star’s actual light curve, is recovered by the actual pipeline. It depends on the star’s noise, the number of transits in the window, the shape of the gaps, and every heuristic in the search. The only honest way to obtain it is to inject millions of synthetic transits and count.

For an Earth analogue the product of the two is around 0.0047×0.30.00140.0047 \times 0.3 \approx 0.0014. Each detected planet in that cell stands for about seven hundred.

What that does to the error bar

A rate estimated from a handful of detections carries an uncertainty that has nothing to do with measurement precision. If four planets are found where the effective search covered 2,800 star-equivalents, the rate is 0.14 per cent per star with a Poisson uncertainty of ±50\pm 50 per cent from the count alone.

A gap where planets should be. The number of planets per star per interval of log radius, for orbital periods under a hundred days, corrected for detection efficiency. There are two peaks — super-Earths near 1.3 R⊕ and sub-Neptunes near 2.4 R⊕ — and a deficit between them at 1.89 R⊕, where the occurrence falls to 33% of the peak. The gap is not a gap in what can be detected: detection efficiency rises smoothly through it, so a smooth underlying distribution could not produce a dip there.
Fig. 2 And the corrected histogram the whole argument rests on. Every planet in it is a detection divided by a detection probability, and the probability is a geometric factor times a completeness that has to be measured by injecting fake signals into the real photometry and counting how many come back. The valley is a feature of the quotient and not of the numerator, which is why it took a decade and a purpose-built completeness analysis to see something that is, in the corrected distribution, unmissable.

And the count is only the beginning. Three further corrections each carry their own uncertainty, and each has at some point been the dominant one:

  • False positives. Eclipsing binaries blended into the aperture mimic small planets, and the rate at which they do so has to be modelled from the stellar population and the pixel-level data. For candidates near the detection limit the false-positive fraction is not small.
  • Reliability. The pipeline itself produces spurious detections from instrumental artefacts and correlated noise. Measuring that rate requires running the search on inverted or scrambled light curves, where no real transit can exist, and counting what comes out.
  • Stellar radii. The planet’s radius is kRk R_\star, so an error in RR_\star moves a planet between cells. Before Gaia, radii were uncertain by tens of per cent and the migration between cells was substantial.

The number the whole thing exists to produce

η\eta_\oplus — the frequency of roughly Earth-sized planets in roughly Earth-like orbits around roughly Sun-like stars — is where every one of these difficulties is at its worst simultaneously. The geometric probability is at its smallest, the number of transits in a four-year window is at its fewest, the signal is at its faintest, and the stellar noise at the relevant timescale is at its most troublesome.

Published values from the same Kepler data have ranged from about 2 per cent to about 60 per cent. That is not a disagreement about arithmetic. It is a disagreement about definitions and about extrapolation:

  • What counts as Earth-sized? 0.5–1.5 RR_\oplus or 0.8–1.25?
  • What counts as the habitable zone? The conservative limits, or the optimistic ones, which differ by a factor of nearly two in flux?
  • What counts as Sun-like? G dwarfs only, or FGK?
  • And how far is the extrapolation? Kepler detected very few planets in the actual cell of interest. Most estimates fit a smooth function to the well-measured cells at shorter periods and extend it, and the answer depends on the functional form assumed.

The current consensus sits somewhere near 10–20 per cent for conservative definitions, with an uncertainty of a factor of two. Its most defensible reading is that Earth-sized planets in temperate orbits are common enough that the nearest one is probably within about ten parsecs — which is the number that decides whether the next generation of telescopes is worth building.

The same planets, before the correction. The radius distribution of the same sample without correcting for detection efficiency. A small planet is harder to find, and the efficiency rises steeply with radius, so the uncorrected histogram climbs monotonically and the valley is not visible at all. The correction is not a cosmetic step: it is the whole difference between a distribution with structure and one without.
Fig. 3 The uncorrected histogram computed with the detection efficiency rising more steeply than the geometric R2R^2 — an index of 2.4, which is what a real pipeline delivers once the number of transits and the noise properties are folded in. The valley is buried more completely still. Every departure of the efficiency from the geometric expectation makes the correction larger, and the corrections that matter most are the ones that are hardest to measure.

Why the correction is largest where the interest is

There is a structural cruelty in this arithmetic that is worth naming, because it is not an accident of the Kepler mission and will not be fixed by a better one.

The completeness correction is a dividing factor, and it is smallest — that is, the division is largest — exactly where the planets are hardest to find. Hardest to find means small, cool and far from the star, which is the description of the planets the question was asked about. So the parameter space divides into a region where the correction is a few per cent and the answer is uninteresting, and a region where the correction is a factor of several hundred and the answer is what everyone wants.

This is the same trap as the one that makes the last rung of the distance ladder carry every earlier error, and it has the same partial escape: measure the well-measured region very well, establish the shape of the distribution there, and extrapolate a function rather than a number. The extrapolation is then explicit and can be argued about, which is better than a division whose largeness is hidden in a denominator.

The rates that are solid

Away from that corner, the results are much firmer, because the completeness is high and the counts are large.

A gap where planets should be. The number of planets per star per interval of log radius, for orbital periods under a hundred days, corrected for detection efficiency. There are two peaks — super-Earths near 1.3 R⊕ and sub-Neptunes near 2.4 R⊕ — and a deficit between them at 1.89 R⊕, where the occurrence falls to 33% of the peak. The gap is not a gap in what can be detected: detection efficiency rises smoothly through it, so a smooth underlying distribution could not produce a dip there.
Fig. 4 The completeness-corrected occurrence of planets against radius, for periods under a hundred days. The area under this curve is the number of such planets per star, and it comes out at roughly one — so the typical Sun-like star has a planet inside Mercury’s orbit, which the solar system does not. This is a corrected distribution, and the correction is the difference between a measurement and a catalogue.

The distribution is a measurement, not a catalogue. Every number below is a count divided by a completeness, and the division is what turns a list of what was noticed into a statement about what is there. That is the same operation a corrected radius distribution depends on, and it is the reason a feature in a corrected histogram can be believed.

Small planets are common. Roughly half of Sun-like stars have a planet between 1 and 4 Earth radii with a period under 100 days, and the average number of such planets per star is of order one. The solar system, with nothing at all inside Mercury’s orbit, is in this respect unusual.

The same planets, before the correction. The radius distribution of the same sample without correcting for detection efficiency. A small planet is harder to find, and the efficiency rises steeply with radius, so the uncorrected histogram climbs monotonically and the valley is not visible at all. The correction is not a cosmetic step: it is the whole difference between a distribution with structure and one without.
Fig. 5 The same sample without the correction. Detection efficiency rises steeply with radius, so the uncorrected histogram climbs monotonically and says only that big planets are easier to find. Every statement in this essay is a statement about the difference between this figure and the one above it.

Around M dwarfs they are more common still. Rates of 2–3 planets per star for periods under 200 days, which is why the nearest transiting temperate planets are all around small stars.

Giants are rare close in and commoner further out. Hot Jupiters occur around about 1 per cent of Sun-like stars; giant planets between 1 and 10 AU around perhaps 10–20 per cent.

And the geometry cuts the other way for multi-planet systems. A system of several coplanar planets transits as a unit or not at all, so the statistics of how many transiting planets a star shows carries information about mutual inclinations. The observed distribution has more single-transiting systems than a population of flat systems would give — the “Kepler dichotomy” — which is evidence either for a population of genuinely single planets or for a spread of mutual inclinations. That inference is made entirely from counts of how often a geometric coincidence recurs.

The same planets, before the correction. The radius distribution of the same sample without correcting for detection efficiency. A small planet is harder to find, and the efficiency rises steeply with radius, so the uncorrected histogram climbs monotonically and the valley is not visible at all. The correction is not a cosmetic step: it is the whole difference between a distribution with structure and one without.
Fig. 6 And a survey half as sensitive at every radius. The efficiency saturates later, so the tilt it imposes on the histogram extends further up in radius and the distribution rises monotonically over the whole drawn range. The shallower the survey, the more of its own histogram is a picture of itself, and that is the sense in which an occurrence rate is a statement about a pipeline before it is one about planets.

One number, and what it would take to improve it

It is worth being concrete about what would move η\eta_\oplus from a factor-of-two number to a ten-per-cent one, because the answer is not “more analysis”.

A transit survey’s yield in the Earth-analogue cell scales with the number of stars observed, the duration of the observation, and the photometric precision — and the first of those is the only one that can be bought cheaply. Doubling the mission duration adds transits linearly and helps as the square root in signal-to-noise, but it also doubles the periods that can be established, which is the more valuable half. Quadrupling the number of target stars quadruples the yield only if the added stars are as quiet and as well characterised as the originals, and they are not: the brightest stars are used first.

There is also a floor that no mission design removes. Sun-like stars vary intrinsically at the level of tens of parts per million over hours to days, and an Earth transit is 84 ppm lasting thirteen hours. The signal and the noise are the same size and have similar timescales, so beyond a certain point the measurement is limited by the target rather than by the instrument — which is why the estimates converge on a factor of two rather than on a number.

What was actually measured

The injection–recovery campaigns. For the final Kepler catalogue, several million synthetic transits were injected into the pixel-level data and put through the full pipeline. The result is a completeness map: a detection probability for every cell of period and radius, for every target star. It is the single most laborious product of the mission and the one everything quantitative rests on.

Inverted light curves. To measure the pipeline’s own false-alarm rate, the search was run on light curves with time reversed and on scrambled data, where no astrophysical transit can be present. Anything found is by construction a false alarm, and the rate at which it happens sets the reliability correction. In the longest-period, smallest-radius cells — exactly the η\eta_\oplus cell — reliability turned out to be a substantial correction rather than a footnote.

The revision by Gaia. When Gaia’s parallaxes fixed the distances of the Kepler field’s stars, many were found to be larger than catalogued — subgiants misclassified as dwarfs — and their planets grew in proportion. Occurrence rates for small planets shifted, and the fraction of stars with Earth-sized planets moved. The planets were the same planets; the denominators changed.

And the number that turned out to be a coincidence of definition. Early statements that “one in five Sun-like stars has an Earth-sized planet in its habitable zone” and later statements of one in fifty were computed from overlapping data. The difference was mostly the definition of the habitable zone and the treatment of reliability. The lesson the field took is that a rate quoted without its definitions is not a measurement.

Where the picture stops

Extrapolation is doing the work in the most-quoted number. Where a survey has few or no detections, the rate is inferred from a fitted trend. That is legitimate, and it is not a measurement of the cell in question.

Occurrence is not habitability. The habitable zone is defined by a flux, not by a climate, and the definition itself is a model with assumptions in it. A rate quoted for the habitable zone inherits every one of them.

Stellar samples are not stellar populations. Kepler observed a particular field, at a particular Galactic latitude, with a particular selection of targets. Whether its rates apply to the thin disc as a whole, or to metal-poor halo stars, is assumed rather than shown.

And the rarest configurations are the least well measured, permanently. Long-period planets need long missions; there is no analysis that gets around the requirement of watching for a decade to establish a decade-long period. That constraint is not statistical and no amount of cleverness removes it.

And multiplicity confuses the count. A rate quoted “per star” and a rate quoted “per planet” differ by the average multiplicity, which is itself a corrected quantity with its own completeness problem. Systems with several transiting planets are found preferentially, since a flat system either transits as a unit or not at all — the mutual inclinations do the same work here that a single inclination does for one planet, and they are measured only through their statistical effect on how many planets are seen at once.

The valley is a slope, and the sign of it decides the mechanism. Planet radius against orbital period, both logarithmic, with the centre of the radius valley drawn as R ∝ P^-0.15. The exponent is the measurement that separates two explanations: photoevaporation, in which the star strips the atmospheres of close-in planets, predicts a valley that moves to smaller radii at longer periods — a negative exponent — while a gas-poor formation origin predicts a positive one. The line drawn here has slope -0.15, so it is a picture of photoevaporation; the measured value is negative, at about −0.09. The bands above and below the line are where the two populations sit.
Fig. 7 The valley’s slope drawn at 0.15-0.15 rather than the measured 0.09-0.09. Both are negative, so both are pictures of photoevaporation; what differs is how quickly the valley moves with period, and that is a quantitative statement about how the stripping depends on the incident flux. The sign settles the mechanism and the magnitude constrains it, and the second needs many more planets than the first — which is exactly the completeness problem this essay is about, met a second time.

The gaps in the record

Completeness as described above is a function of period and radius, and there is a third variable hiding inside it: when the survey was looking.

A transit search normally requires three events before it will claim a period, and whether three events are available depends on the fraction of the mission during which data were being taken. No survey observes continuously. A spacecraft rolls to keep its solar panels pointed, downlinks its data, enters safe mode after a cosmic-ray hit, and loses days at a time; a ground-based survey loses the daytime and the weather.

The consequence is a window function, and its effect is not smooth in period. A planet whose period is close to a multiple of the observing cadence has its transits repeatedly falling in the same phase of the gap structure, so all three can be missed while a planet with a period a few per cent different is caught. The completeness surface therefore has narrow spikes of low sensitivity at particular periods, superimposed on the broad decline with period that the signal-to-noise argument produces.

For a space mission with a duty cycle above ninety per cent the effect is small and is folded into the injection–recovery result automatically, since the injections are made into the real time series with its real gaps. For a ground-based survey it is severe: the one-day alias is the dominant feature of the sensitivity, and periods near integer numbers of days are systematically under-recovered.

That is why the injections have to go into the actual data rather than into a model of it. An analytic completeness estimate would be a smooth function, and the real one is not.

The generalisation

Estimating a population from a biased sample with a computable selection function is one of the standard structures of quantitative science, and the exoplanet version is unusually clean.

The equivalent in cosmology is the luminosity function, recovered from a flux-limited galaxy catalogue by weighting each object by the inverse of the volume in which it could have been seen — the 1/Vmax1/V_{\max} method, which is exactly the same division by a detection probability. In particle physics it is the acceptance correction. In ecology it is capture–recapture. In every case the difficult part is the same: the correction is largest exactly where the data are thinnest, so the least reliable part of the answer is the part the answer was wanted for.

The comparison worth making is with the distance ladder, where a chain of calibrations propagates fractional errors multiplicatively and the last rung inherits all of them. An occurrence rate has the same shape: a count, divided by a product of four corrections, extrapolated one step beyond where it was measured.

The valley is a slope, and the sign of it decides the mechanism. Planet radius against orbital period, both logarithmic, with the centre of the radius valley drawn as R ∝ P^-0.09. The exponent is the measurement that separates two explanations: photoevaporation, in which the star strips the atmospheres of close-in planets, predicts a valley that moves to smaller radii at longer periods — a negative exponent — while a gas-poor formation origin predicts a positive one. The line drawn here has slope -0.09, so it is a picture of photoevaporation; the measured value is negative, at about −0.09. The bands above and below the line are where the two populations sit.
Fig. 8 The same slope with the valley’s centre at 2.3 Earth radii rather than 2.0. The two populations move with it and the tilt is unchanged. Where the valley sits is set by the core masses the planets happen to have and is therefore a statement about formation, while the tilt is a statement about what happened afterwards — and a survey that measures one well and the other badly has answered only half of the question it set out to.

A rate is not a probability about any one star

One further distinction is worth drawing, because the numbers above are routinely misread.

An occurrence rate of 0.2 planets per star in some cell does not mean that a given star has a one-in-five chance of holding such a planet. It means that the average number per star is 0.2, and the distribution behind that average is unknown. If planets come in systems — and they do — then a population in which a fifth of stars have one planet each and a population in which a twentieth have four each give the same rate and describe very different galaxies.

Separating those requires multiplicity statistics, which are harder still, because the number of detected planets in a system depends on the mutual inclinations as well as on how many there are. The best current statement is that the planets are clustered: systems tend to have several or none, and the ones with several tend to have them evenly spaced with similar sizes — a regularity strong enough to have earned the name “peas in a pod”, and one that no correction for selection removes.

One more reading covers the case the two mechanisms are hardest to tell apart in.

The valley is a slope, and the sign of it decides the mechanism. Planet radius against orbital period, both logarithmic, with the centre of the radius valley drawn as R ∝ P^-0.05. The exponent is the measurement that separates two explanations: photoevaporation, in which the star strips the atmospheres of close-in planets, predicts a valley that moves to smaller radii at longer periods — a negative exponent — while a gas-poor formation origin predicts a positive one. The line drawn here has slope -0.05, so it is a picture of photoevaporation; the measured value is negative, at about −0.09. The bands above and below the line are where the two populations sit.
Fig. 9 The valley drawn nearly flat in period. A shallow slope is what a core-powered mechanism produces and a steeper negative one is what photoevaporation produces, and the measured slope sits between them with error bars that overlap both — which is why the question is still open after a decade of larger samples.

The valley is a statement about a population and its depth is a statement about a survey, and separating the two is the whole of what a completeness correction does — which is why the same catalogue has supported two incompatible mechanisms for a decade.

Where this goes next

The corrected distributions have structure in them that the raw catalogues did not, and the sharpest piece of that structure is a deficit — a radius at which planets are markedly less common than on either side of it, which appeared only once the stellar radii were good enough to see it.

Later rungs on this anchor: injection and recovery in detail. Reliability from inverted light curves. The Kepler dichotomy and mutual inclinations. Occurrence around M dwarfs. Metallicity dependence of occurrence. η\eta_\oplus and its definitions. Occurrence rates from microlensing for cold planets. The frequency of free-floating planets. Multiplicity statistics. And how many stars a direct-imaging mission must observe to expect one detection, which is the same arithmetic run in reverse.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 19 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

CompletenessError propagationEta earthExtrapolationFalse-positiveHabitable zoneOccurrence rateSmall number statisticsSurvey completenessTransit probability