The planets that were not seen
Assumes Detection bias and Transits.
Kepler found about 2,700 confirmed planets, each of them a dip of a few hundred parts per million repeating on a fixed period. There are approximately 200 billion stars in the Galaxy. Nothing about the first number implies anything about the second until the question “how many did the mission miss” has been answered — and that question is nearly the whole of the subject.
The answer is a division, and the denominator is a probability that can be as small as a five-hundredth.
The shape of the calculation
For a given kind of planet — a cell in the plane of period against radius — the occurrence rate per star is
summed over the survey’s stars. The two probabilities are different in kind and it is worth keeping them apart.
The geometric probability is , computable from the star’s radius and the period through the harmonic law. It involves no instrument and no judgement. For a survey of a hundred thousand stars, it is a hundred thousand exact numbers.
The detection probability is not analytic at all. It is the fraction of the time that a signal of those parameters, injected into that star’s actual light curve, is recovered by the actual pipeline. It depends on the star’s noise, the number of transits in the window, the shape of the gaps, and every heuristic in the search. The only honest way to obtain it is to inject millions of synthetic transits and count.
For an Earth analogue the product of the two is around . Each detected planet in that cell stands for about seven hundred.
What that does to the error bar
A rate estimated from a handful of detections carries an uncertainty that has nothing to do with measurement precision. If four planets are found where the effective search covered 2,800 star-equivalents, the rate is 0.14 per cent per star with a Poisson uncertainty of per cent from the count alone.
And the count is only the beginning. Three further corrections each carry their own uncertainty, and each has at some point been the dominant one:
- False positives. Eclipsing binaries blended into the aperture mimic small planets, and the rate at which they do so has to be modelled from the stellar population and the pixel-level data. For candidates near the detection limit the false-positive fraction is not small.
- Reliability. The pipeline itself produces spurious detections from instrumental artefacts and correlated noise. Measuring that rate requires running the search on inverted or scrambled light curves, where no real transit can exist, and counting what comes out.
- Stellar radii. The planet’s radius is , so an error in moves a planet between cells. Before Gaia, radii were uncertain by tens of per cent and the migration between cells was substantial.
The number the whole thing exists to produce
— the frequency of roughly Earth-sized planets in roughly Earth-like orbits around roughly Sun-like stars — is where every one of these difficulties is at its worst simultaneously. The geometric probability is at its smallest, the number of transits in a four-year window is at its fewest, the signal is at its faintest, and the stellar noise at the relevant timescale is at its most troublesome.
Published values from the same Kepler data have ranged from about 2 per cent to about 60 per cent. That is not a disagreement about arithmetic. It is a disagreement about definitions and about extrapolation:
- What counts as Earth-sized? 0.5–1.5 or 0.8–1.25?
- What counts as the habitable zone? The conservative limits, or the optimistic ones, which differ by a factor of nearly two in flux?
- What counts as Sun-like? G dwarfs only, or FGK?
- And how far is the extrapolation? Kepler detected very few planets in the actual cell of interest. Most estimates fit a smooth function to the well-measured cells at shorter periods and extend it, and the answer depends on the functional form assumed.
The current consensus sits somewhere near 10–20 per cent for conservative definitions, with an uncertainty of a factor of two. Its most defensible reading is that Earth-sized planets in temperate orbits are common enough that the nearest one is probably within about ten parsecs — which is the number that decides whether the next generation of telescopes is worth building.
Why the correction is largest where the interest is
There is a structural cruelty in this arithmetic that is worth naming, because it is not an accident of the Kepler mission and will not be fixed by a better one.
The completeness correction is a dividing factor, and it is smallest — that is, the division is largest — exactly where the planets are hardest to find. Hardest to find means small, cool and far from the star, which is the description of the planets the question was asked about. So the parameter space divides into a region where the correction is a few per cent and the answer is uninteresting, and a region where the correction is a factor of several hundred and the answer is what everyone wants.
This is the same trap as the one that makes the last rung of the distance ladder carry every earlier error, and it has the same partial escape: measure the well-measured region very well, establish the shape of the distribution there, and extrapolate a function rather than a number. The extrapolation is then explicit and can be argued about, which is better than a division whose largeness is hidden in a denominator.
The rates that are solid
Away from that corner, the results are much firmer, because the completeness is high and the counts are large.
The distribution is a measurement, not a catalogue. Every number below is a count divided by a completeness, and the division is what turns a list of what was noticed into a statement about what is there. That is the same operation a corrected radius distribution depends on, and it is the reason a feature in a corrected histogram can be believed.
Small planets are common. Roughly half of Sun-like stars have a planet between 1 and 4 Earth radii with a period under 100 days, and the average number of such planets per star is of order one. The solar system, with nothing at all inside Mercury’s orbit, is in this respect unusual.
Around M dwarfs they are more common still. Rates of 2–3 planets per star for periods under 200 days, which is why the nearest transiting temperate planets are all around small stars.
Giants are rare close in and commoner further out. Hot Jupiters occur around about 1 per cent of Sun-like stars; giant planets between 1 and 10 AU around perhaps 10–20 per cent.
And the geometry cuts the other way for multi-planet systems. A system of several coplanar planets transits as a unit or not at all, so the statistics of how many transiting planets a star shows carries information about mutual inclinations. The observed distribution has more single-transiting systems than a population of flat systems would give — the “Kepler dichotomy” — which is evidence either for a population of genuinely single planets or for a spread of mutual inclinations. That inference is made entirely from counts of how often a geometric coincidence recurs.
One number, and what it would take to improve it
It is worth being concrete about what would move from a factor-of-two number to a ten-per-cent one, because the answer is not “more analysis”.
A transit survey’s yield in the Earth-analogue cell scales with the number of stars observed, the duration of the observation, and the photometric precision — and the first of those is the only one that can be bought cheaply. Doubling the mission duration adds transits linearly and helps as the square root in signal-to-noise, but it also doubles the periods that can be established, which is the more valuable half. Quadrupling the number of target stars quadruples the yield only if the added stars are as quiet and as well characterised as the originals, and they are not: the brightest stars are used first.
There is also a floor that no mission design removes. Sun-like stars vary intrinsically at the level of tens of parts per million over hours to days, and an Earth transit is 84 ppm lasting thirteen hours. The signal and the noise are the same size and have similar timescales, so beyond a certain point the measurement is limited by the target rather than by the instrument — which is why the estimates converge on a factor of two rather than on a number.
What was actually measured
The injection–recovery campaigns. For the final Kepler catalogue, several million synthetic transits were injected into the pixel-level data and put through the full pipeline. The result is a completeness map: a detection probability for every cell of period and radius, for every target star. It is the single most laborious product of the mission and the one everything quantitative rests on.
Inverted light curves. To measure the pipeline’s own false-alarm rate, the search was run on light curves with time reversed and on scrambled data, where no astrophysical transit can be present. Anything found is by construction a false alarm, and the rate at which it happens sets the reliability correction. In the longest-period, smallest-radius cells — exactly the cell — reliability turned out to be a substantial correction rather than a footnote.
The revision by Gaia. When Gaia’s parallaxes fixed the distances of the Kepler field’s stars, many were found to be larger than catalogued — subgiants misclassified as dwarfs — and their planets grew in proportion. Occurrence rates for small planets shifted, and the fraction of stars with Earth-sized planets moved. The planets were the same planets; the denominators changed.
And the number that turned out to be a coincidence of definition. Early statements that “one in five Sun-like stars has an Earth-sized planet in its habitable zone” and later statements of one in fifty were computed from overlapping data. The difference was mostly the definition of the habitable zone and the treatment of reliability. The lesson the field took is that a rate quoted without its definitions is not a measurement.
Where the picture stops
Extrapolation is doing the work in the most-quoted number. Where a survey has few or no detections, the rate is inferred from a fitted trend. That is legitimate, and it is not a measurement of the cell in question.
Occurrence is not habitability. The habitable zone is defined by a flux, not by a climate, and the definition itself is a model with assumptions in it. A rate quoted for the habitable zone inherits every one of them.
Stellar samples are not stellar populations. Kepler observed a particular field, at a particular Galactic latitude, with a particular selection of targets. Whether its rates apply to the thin disc as a whole, or to metal-poor halo stars, is assumed rather than shown.
And the rarest configurations are the least well measured, permanently. Long-period planets need long missions; there is no analysis that gets around the requirement of watching for a decade to establish a decade-long period. That constraint is not statistical and no amount of cleverness removes it.
And multiplicity confuses the count. A rate quoted “per star” and a rate quoted “per planet” differ by the average multiplicity, which is itself a corrected quantity with its own completeness problem. Systems with several transiting planets are found preferentially, since a flat system either transits as a unit or not at all — the mutual inclinations do the same work here that a single inclination does for one planet, and they are measured only through their statistical effect on how many planets are seen at once.
The gaps in the record
Completeness as described above is a function of period and radius, and there is a third variable hiding inside it: when the survey was looking.
A transit search normally requires three events before it will claim a period, and whether three events are available depends on the fraction of the mission during which data were being taken. No survey observes continuously. A spacecraft rolls to keep its solar panels pointed, downlinks its data, enters safe mode after a cosmic-ray hit, and loses days at a time; a ground-based survey loses the daytime and the weather.
The consequence is a window function, and its effect is not smooth in period. A planet whose period is close to a multiple of the observing cadence has its transits repeatedly falling in the same phase of the gap structure, so all three can be missed while a planet with a period a few per cent different is caught. The completeness surface therefore has narrow spikes of low sensitivity at particular periods, superimposed on the broad decline with period that the signal-to-noise argument produces.
For a space mission with a duty cycle above ninety per cent the effect is small and is folded into the injection–recovery result automatically, since the injections are made into the real time series with its real gaps. For a ground-based survey it is severe: the one-day alias is the dominant feature of the sensitivity, and periods near integer numbers of days are systematically under-recovered.
That is why the injections have to go into the actual data rather than into a model of it. An analytic completeness estimate would be a smooth function, and the real one is not.
The generalisation
Estimating a population from a biased sample with a computable selection function is one of the standard structures of quantitative science, and the exoplanet version is unusually clean.
The equivalent in cosmology is the luminosity function, recovered from a flux-limited galaxy catalogue by weighting each object by the inverse of the volume in which it could have been seen — the method, which is exactly the same division by a detection probability. In particle physics it is the acceptance correction. In ecology it is capture–recapture. In every case the difficult part is the same: the correction is largest exactly where the data are thinnest, so the least reliable part of the answer is the part the answer was wanted for.
The comparison worth making is with the distance ladder, where a chain of calibrations propagates fractional errors multiplicatively and the last rung inherits all of them. An occurrence rate has the same shape: a count, divided by a product of four corrections, extrapolated one step beyond where it was measured.
A rate is not a probability about any one star
One further distinction is worth drawing, because the numbers above are routinely misread.
An occurrence rate of 0.2 planets per star in some cell does not mean that a given star has a one-in-five chance of holding such a planet. It means that the average number per star is 0.2, and the distribution behind that average is unknown. If planets come in systems — and they do — then a population in which a fifth of stars have one planet each and a population in which a twentieth have four each give the same rate and describe very different galaxies.
Separating those requires multiplicity statistics, which are harder still, because the number of detected planets in a system depends on the mutual inclinations as well as on how many there are. The best current statement is that the planets are clustered: systems tend to have several or none, and the ones with several tend to have them evenly spaced with similar sizes — a regularity strong enough to have earned the name “peas in a pod”, and one that no correction for selection removes.
One more reading covers the case the two mechanisms are hardest to tell apart in.
The valley is a statement about a population and its depth is a statement about a survey, and separating the two is the whole of what a completeness correction does — which is why the same catalogue has supported two incompatible mechanisms for a decade.
Where this goes next
The corrected distributions have structure in them that the raw catalogues did not, and the sharpest piece of that structure is a deficit — a radius at which planets are markedly less common than on either side of it, which appeared only once the stellar radii were good enough to see it.
Later rungs on this anchor: injection and recovery in detail. Reliability from inverted light curves. The Kepler dichotomy and mutual inclinations. Occurrence around M dwarfs. Metallicity dependence of occurrence. and its definitions. Occurrence rates from microlensing for cold planets. The frequency of free-floating planets. Multiplicity statistics. And how many stars a direct-imaging mission must observe to expect one detection, which is the same arithmetic run in reverse.
What this makes readable
Essays that name this one as a prerequisite.
- A gap in a histogram that says how planets are built exoplanets
- The part of a rate that is a definition exoplanets
About the same objects
Not linked from either essay — found by the objects both name.
- How many planets a star has is not a measurement occurrence rate · survey completeness · transit probability
- Every method prefers a circle, and not for the same reason survey completeness · transit probability
What links here
The 8 of 19 essays linking to this one that name the most of the same objects.
- The part of a rate that is a definition exoplanets
- The threshold that is not a threshold exoplanets
- A count with a knee in it galaxies
- A gap in a histogram that says how planets are built exoplanets
- A planet measured by the light it removes exoplanets
- A planet radius is a stellar radius exoplanets
- One number sets the zone exoplanets
- The band moves and the orbit does not exoplanets
The objects this essay names
Each one links to every other essay that touches it.
CompletenessError propagationEta earthExtrapolationFalse-positiveHabitable zoneOccurrence rateSmall number statisticsSurvey completenessTransit probability