The part of a rate that is a definition
Assumes Occurrence rates, Detection bias and Habitable zone.
The rung below corrected a survey for the planets it could not have seen: the transit probability, the completeness of the pipeline, and the geometry of a survey are all computable, and dividing the detections by them turns a catalogue into a rate.
That machinery works. It is not what is wrong with , the frequency of Earth-size planets in the habitable zones of Sun-like stars — the single number the whole field is asked for, and one whose published values span a factor of thirty.
Four boxes, all defensible
The quantity has to be defined by a range of radius and a range of insolation, and nothing in nature marks either.
The radius. Where does “Earth-size” stop? At 1.0? At 1.25, allowing for measurement error? At 1.5, the usual boundary of “super-Earth”? At 1.75, below the radius valley and therefore plausibly rocky? Every one of those has been used, and the last two differ by an amount smaller than the uncertainty on a typical measured planetary radius.
The insolation. The conservative habitable zone runs from about 0.32 to 1.78 times the Earth’s insolation, from the maximum-greenhouse limit to the runaway-greenhouse limit. The optimistic zone, based on the evidence that Venus and Mars once had liquid water, runs from about 0.2 to 2.2. The choice between them is a choice about climate modelling and not about planets at all.
And the stars. “Sun-like” has meant G dwarfs only, GK dwarfs, and FGK dwarfs, and the occurrence rate per star differs between those samples by a factor of about two on its own — because smaller stars have more small planets, which is itself a statement about what a survey can see rather than about the sky.
Why the corner is the worst possible place to put a boundary
If the occurrence surface were flat, none of this would matter much: moving a boundary by ten per cent would move the answer by ten per cent.
It is not flat. It has a radius valley at about 1.8 Earth radii — the gap between rocky planets and those with substantial hydrogen envelopes — and the number of planets per unit radius changes by a large factor across it. Every definition of “Earth-size” puts its upper bound near that feature, because that is what the feature means, and near a steep feature a small change in a boundary is a large change in an integral.
The insolation direction is gentler but longer. Occurrence rises with orbital period out to about ten days and is roughly flat beyond, so the habitable zone of a Sun-like star sits on a plateau — which is helpful, and which is also the region where the surface is least constrained by data.
The corner is also outside the data
The second problem is worse than the first, and it is not a matter of definition.
A transit survey detects a planet when the accumulated signal-to-noise crosses a threshold, conventionally about 7. The signal is the transit depth, which goes as the square of the planet-to-star radius ratio; the noise is the photometric scatter divided by the square root of the number of transits observed; and the number of transits is the mission duration divided by the period.
Put a Sun-like star of typical photometric quality into that. At one year’s period, the smallest detectable planet is between one and a half and two Earth radii — larger than an Earth.
So the corner of every box in the first figure — small radius, long period — contains no detections whatever. The occurrence rate quoted there is an extrapolation of a surface fitted to the region where detections exist, and its value depends on the functional form assumed for the extrapolation as much as on anything measured.
What a completeness correction actually is
It is worth being concrete about the machinery, because the criticism above is not a criticism of it.
A survey’s completeness at a given radius and period is the probability that a planet with those parameters, orbiting one of the survey’s stars, would have been detected. It is measured rather than computed: synthetic transits are injected into the real photometry of the real target stars, the real pipeline is run over the result, and the fraction recovered is counted. Millions of injections give a completeness surface.
That is an honest and expensive procedure, and it captures things no analytic estimate would — the detrending algorithm’s tendency to absorb a shallow long-period signal, the effect of data gaps, the human vetting step.
What it cannot do is produce information where there is none. Completeness at one Earth radius and one year is a few per cent for the best stars and zero for most, and dividing a detection count of zero by a completeness of zero is not a rate.
The other correction, which points the other way
Completeness corrects for planets that were missed. There is a second correction, discovered later and pointing the opposite way, for planets that were not there.
A transit survey’s candidate list contains false positives: eclipsing binaries blended with the target, instrumental artefacts that repeat, and statistical fluctuations that cross the threshold. Near the detection limit — which is exactly where the small long-period planets are — the false-positive rate becomes comparable to the true detection rate, because the number of noise events crossing a threshold rises steeply as the threshold is approached from above and the number of real planets does not.
The measure of this is reliability, and it is estimated by running the same pipeline over deliberately scrambled data, where every detection is by construction spurious. For the longest-period, smallest candidates in the main transit survey, reliability estimates run below 50 per cent.
Correcting for completeness alone therefore overestimates the rate, by dividing a contaminated numerator by a small number. That correction was applied in later analyses and moved published values of down by factors of two to three — which accounts for a good part of the spread between early and late estimates, and which is a real change in the answer rather than a redefinition.
The honest form of the number
Given all of that, what should be quoted?
A surface, with the boxes drawn on it. That is the first figure of this essay, and it says everything a single number cannot: what the occurrence rate is as a function of radius and insolation, where the detections stop, and what any particular definition integrates to.
A number with its definition attached, always. “ for planets of 1 to 1.5 Earth radii receiving 0.32 to 1.78 times the Earth’s insolation, around G dwarfs” is a statement. “” is not.
And an explicit statement of how far the extrapolation reaches. A rate quoted at a point where the survey’s completeness is below ten per cent is dominated by the fitted model, and saying so is not a caveat but part of the result.
How large the definitional term is, measured
The first figure puts a number on the definitional spread, and it is worth separating from the rest.
Holding everything else fixed — the same surface, the same completeness, the same reliability — and integrating over four published definitions gives a factor of about three. Moving one boundary alone, the radius ceiling, from 1.5 to 1.75 Earth radii moves the answer by roughly half again.
The full published spread is a factor of thirty. So the accounting is approximately: a factor of three from definition, a factor of two to three from reliability corrections that were absent from the early estimates and present in the later ones, and the remainder from how far each analysis was willing to extrapolate its fitted surface into the empty corner.
There is a fourth term, and it is the one hardest to put an interval on: the four definitions are not independent draws. Each was chosen by somebody who had already seen a version of the occurrence surface, and a boundary chosen after the fact lands where the author judged the data would support it. So the factor of three between the published definitions is not the width of the space of defensible choices; it is the width of the choices that were actually made, by a small number of groups reading overlapping data. The honest space is wider, and there is no way to measure it — which is the strongest argument for quoting the surface rather than any integral over it. The same objection applies with more force to the agreement between published values than to their spread: two groups adopting nearly the same boundary and getting nearly the same rate have not confirmed each other, because the boundary is most of what the answer is.
None of those three is a statistical error, and none of them shrinks with more data of the same kind. That is the point of the rung: the error bar quoted with a value of describes the smallest of the terms that make it uncertain.
Why the number is asked for anyway
It is worth saying what is actually used for, because that determines how much of this matters.
It is used to size a mission. A direct-imaging observatory looking for terrestrial planets in habitable zones needs to know how many stars to survey, and that number scales as one over . A factor of three in the rate is a factor of three in the required telescope time, or roughly a factor of one and a half in the aperture — which is a difference of billions of currency units and years of schedule.
For that purpose the definitional spread is not a philosophical difficulty; it is the dominant term in the requirement. And the sensible response, which the mission studies have adopted, is not to argue about the number but to carry the whole range through the design and see where it breaks.
The same problem, in a subject that solved it
The difficulty is not peculiar to planets, and one neighbouring field has an instructive convention.
Galaxy luminosity functions are quoted as functions, not as counts. Nobody publishes “the number of bright galaxies per cubic megaparsec”; they publish a Schechter function with three parameters and its covariance, and anyone who wants a count over a particular range integrates it themselves. The convention exists because the same argument applies — the bright end is steep, the faint end is incomplete, and any single number is a statement about where somebody drew a line.
The stellar initial mass function is quoted the same way, as a broken power law with its break points, and for the same reason.
Occurrence rates are moving in that direction and are not there yet, because a single number is what a mission proposal needs and a function is not. The compromise that has emerged is to publish the fitted surface and its posterior and to quote the integral over one named box, with the box in the sentence.
The boundary is made of somebody else’s distance
The radius ceiling is the definitional choice the answer is most sensitive to, and the feature it is drawn against — the radius valley at about 1.8 Earth radii — is itself a derived quantity resting on a chain the essay has not yet followed.
A transit measures a ratio: the planet’s radius divided by the star’s. Nothing about the depth of a transit says how big anything is in kilometres. To get a planet radius, the stellar radius must be supplied from elsewhere, and it is normally obtained from the star’s luminosity and temperature — which needs a distance.
That makes the whole radius axis of the occurrence surface a function of the parallaxes available when the analysis was done. And they changed. Before Gaia, stellar radii for the main transit survey’s targets came from photometric estimates with uncertainties of 20 to 40 per cent; afterwards they came from parallaxes good to a per cent, and a substantial fraction of the targets turned out to be subgiants — larger than assumed, with correspondingly larger planets.
The consequence is the one that matters here. The radius valley was not clearly visible until the stellar radii improved, because a 30 per cent scatter in the horizontal axis smears out a feature a few tenths of an Earth radius wide. It appeared in 2017 in a sample with spectroscopically determined stellar parameters and was confirmed when Gaia’s parallaxes were applied to the whole catalogue.
So the sequence runs: a parallax fixes a stellar radius, a stellar radius fixes a planet radius, a distribution of planet radii reveals a valley, and the valley is where every author now puts the ceiling of the box that defines . The definitional boundary of the number is downstream of an astrometric mission, which is not a criticism of anybody’s choice — it is what it means for a boundary to be drawn against a real physical feature rather than at a round number.
Per star, and which stars
The other half of the definition is the denominator, and it is chosen at least as freely as the numerator.
An occurrence rate is quoted per star, so the answer depends entirely on which stars are in the sample — and the dependence is not a nuisance term. Small stars have more small planets: the occurrence of planets between one and two Earth radii per M dwarf is several times the rate per G dwarf, in the period ranges where both are measurable, and that is a measured difference rather than a completeness artefact, since small planets are easier to detect around small stars.
Which means the answer to “how common are Earth-size planets in habitable zones” depends on whether the question is asked per star, per Sun-like star, or per star of any kind. Per star of any kind is the largest number by a wide margin, because three quarters of the stars in the Galaxy are M dwarfs.
The habitable zone of an M dwarf is also a different place. It sits at a hundredth of an astronomical unit or so, inside the tidal locking radius, so a planet there keeps one face to its star permanently; and the star is a flare star for much of its life, delivering ultraviolet and particle fluxes orders of magnitude above what the Earth receives. Whether such a planet is habitable is a question about atmospheres that nobody can currently answer.
So the denominator carries an assumption about habitability every bit as large as the insolation boundary carries, and it is usually made silently — by choosing a stellar sample and reporting a rate per star of that sample, with the reason for the choice left in the section heading.
What would settle it
A longer baseline. The detection threshold improves as the fourth root of the mission duration, since the number of transits grows linearly and the signal-to-noise as its square root. Doubling a four-year mission to eight buys about twenty per cent in the smallest detectable radius — which is not nothing and is not enough.
A quieter sample. The floor at long period is set by stellar variability rather than by photons, so the useful improvement is to observe stars that are quieter, which means older and slower-rotating, which means fewer of them.
Or a different observable entirely. Direct imaging measures a planet’s reflected light rather than its silhouette, and its detection boundary runs the other way — easier at wide separations, harder close in. A survey with that boundary would place its detections in exactly the region the transit surveys extrapolate into, which is the argument for building one.
There is one more knob in the calculation that is not a definition of the planet at all, and moving it changes the answer by as much as the box does.
A threshold of 7.1 was chosen because it gives about one false positive across the whole Kepler catalogue under Gaussian noise, and the noise is not Gaussian. Analyses that adopt a higher floor for that reason are not being conservative about the planets; they are being conservative about the statistics, and the two are not the same decision. The occurrence rate that comes out is therefore a function of somebody’s tolerance for false positives, which is not a property of the planetary population.
That is the third definitional term in a number often quoted with two significant figures, and it compounds with the other two rather than averaging against them. The box says which planets count, the completeness correction says how many were missed, and the threshold says which detections were real — and a plausible choice at each of the three moves the answer by a factor approaching two.
The honest presentation is therefore a surface with the boxes drawn on it, which is what the figures in this essay are, rather than a number with an error bar. An error bar communicates that the answer is uncertain; a family of boxes communicates that the question has not been asked precisely enough for an answer to exist. Those are different admissions and only the second is true here. The data are good enough to say that Earth-size planets in temperate orbits are common; they are not good enough, and no reasonable amount of further data would be, to attach a number to a phrase that four groups define four ways.
Where this ladder goes next
This rung has taken a rate that looked like a measurement and separated it into three parts: a genuinely measured surface, a boundary that somebody chose, and an extrapolation past the last detection.
The rung above is the same accounting applied to the thing the rate is a proxy for. An Earth-size planet at the right distance is not an inhabited one, and the chain from occurrence rate to anything about life passes through several more factors of exactly this kind — each of them a definition with a number attached.
Beside it lies the survey design question run forwards rather than backwards: given a detection boundary and an occurrence surface, what does a mission actually yield, and how does that yield depend on the parts of the surface nobody has measured.
And below it, the habit: a quantity defined by a boundary is partly a measurement of the boundary. Nothing about the sky changes when the radius bound moves from 1.5 to 1.75; the answer changes by more than every error bar quoted with it, and the only defence is to publish the surface rather than the integral.
About the same objects
Not linked from either essay — found by the objects both name.
- A planet that was the star's own rotation false-positive · window function
- A population counted by shadows that never repeat detection threshold · false-positive
- The threshold that is not a threshold detection threshold · occurrence rate
What links here
Essays that link to this one from their own argument.
- How many planets a star has is not a measurement exoplanets
- The band moves and the orbit does not exoplanets
- The line beyond which ice counts as rock exoplanets
- Two explosions told apart by a missing line stars
The objects this essay names
Each one links to every other essay that touches it.
CompletenessDefinitional uncertaintyDetection thresholdEta earthExtrapolationFalse-positiveHabitable zoneOccurrence rateReliabilitySurvey yieldVettingWindow function