Exoplanets

Every survey draws a different sky

The first exoplanets found were enormous and impossibly close to their stars. That was not a discovery about planets. It was a measurement of what a 10 m/s spectrograph watching for three years is able to see.

Assumes Reflex velocity and Transits.

In 1995 the known planets outside the solar system numbered one, and it was a Jupiter at 0.05 AU. By 2000 there were around thirty, and most were giants inside 1 AU on eccentric orbits. The immediate reading — that the solar system’s architecture is unusual, that giant planets normally sit close in — turned out to be almost exactly backwards.

The census was not measuring planets. It was measuring what a spectrograph of a particular precision, operated for a particular number of years, is capable of noticing.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 1.4 Earth masses, cut off at 1.21 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 1 Each method’s detection boundary in the plane of planet mass against orbital distance, computed from the threshold on its own observable. The radial-velocity line rises as a\sqrt{a}; the astrometric line falls as 1/a1/a; the transit line is horizontal because a depth is a statement about radius, cut off on the right by the need for repeats. The solar system is drawn on top, and for two decades the only one of its planets inside any of these regions was Jupiter.

Each boundary is an equation, not an impression

The useful thing about a selection effect in this field is that it can be written down.

Radial velocity. The semi-amplitude is K28.4 m/s(mp/MJ)(M/M)2/3(P/yr)1/3K \approx 28.4\ \text{m/s}\,(m_p/M_J)(M_\star/M_\odot)^{-2/3}(P/\text{yr})^{-1/3}. Setting KK to the instrument’s threshold and solving for mass gives a boundary rising as P1/3P^{1/3}, or a1/2a^{1/2}. At 1 m/s around a solar-mass star, the least detectable planet is about 1.4 Earth masses at 0.1 AU and about 11 Earth masses at 5 AU.

Transits. The depth is (Rp/R)2(R_p/R_\star)^2, so a threshold on depth is a threshold on radius and is independent of orbital distance. The distance enters twice elsewhere: through the geometric probability R/aR_\star/a, and through the requirement that the planet transit several times within the observing window. The second is the harder wall — a mission lasting four years cannot establish a period longer than about sixteen months, whatever its photometric precision.

Astrometry. The star’s angular displacement is (mp/M)(a/d)(m_p/M_\star)(a/d), which grows with aa. Astrometry is the one method that becomes more sensitive further out, and it is therefore the natural complement to everything else on the diagram.

Direct imaging. A wall at the diffraction limit and a floor set by contrast, both discussed in their own place.

Four methods, four boundaries with four different slopes. None of them is a statement about planets.

The hot Jupiters were not a discovery about planet formation

Consider what the 1995 instruments could see. A precision of about 13 m/s, three years of data, and a few hundred stars observed. The detectable region is everything above K=13K = 13 m/s with a period short enough to have completed several cycles.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 13 m/s needs mass rising as √a; astrometry at 100 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 1.4 Earth masses, cut off at 1.00 AU by the need for three transits in 3 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 2 The same plane at the precision available when the first exoplanet was found. Only the top-left corner is accessible: massive planets, close in. A hot Jupiter sits deep inside the region; Jupiter itself sits just outside it, at 12.5 m/s and a twelve-year period, which is why the solar system’s own architecture would have been undetectable by the survey that found the first exoplanet.

A 0.5 Jupiter-mass planet at 0.05 AU produces K=57K = 57 m/s and a full cycle every four days. It is not merely detectable; it is the most detectable thing in the entire plane. Every one of the biases points the same way: large mass, short period, and — because a short period gives dozens of cycles per season — an unambiguous phase coverage that makes the fit trivially convincing.

So the early census found hot Jupiters because hot Jupiters are the easiest possible object to find, and it found few of anything else because everything else was below threshold. It later turned out that hot Jupiters orbit only about 1 per cent of Sun-like stars, which makes them one of the rarest kinds of planet known. A class of object that is rare in nature and dominant in a catalogue is the signature of a selection effect operating at full strength.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 0.1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 1.4 Earth masses, cut off at 1.21 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 3 The same boundaries with the radial-velocity precision improved tenfold, to a tenth of a metre a second. The whole curve moves down by a factor of a hundred in mass at fixed separation — the threshold mass goes as the precision — and it moves down parallel to itself, keeping the same a\sqrt{a} slope. An instrumental improvement translates the boundary and does not rotate it, so it enlarges the region a method sees without changing the shape of what that region excludes. The Earth on this diagram is still outside it.

The same effect, one field over

None of this is peculiar to exoplanets. It is the same structure as the bias that inflates every brightness-limited sample of stars: a survey that can see to a fixed flux limit sees intrinsically bright objects to a greater distance, so its volume is larger for bright objects and its sample is over-representative of them.

The exoplanet version is more tractable than the stellar one for a specific reason: the detectability of a planet is a deterministic function of a small number of parameters, all of which are also the parameters being measured. There is no equivalent of a scatter in absolute magnitude to fold in. That is why exoplanet occurrence rates are quoted with error bars that mean something, and why the inversion below is possible at all.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 0.3 m/s needs mass rising as √a; astrometry at 5 µas needs it falling as 1/a, which is the only method that gets easier further out; a 20 ppm transit is a threshold on radius and so a horizontal line at about 0.1 Earth masses, cut off at 2.23 AU by the need for three transits in 10 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 4 What a decade of a much better instrument reaches. Every boundary moves and none of them moves the same way: the radial-velocity floor falls linearly with the precision, the transit floor as the square root of the depth, and the mission limit as the two-thirds power of the duration. So improving an instrument does not enlarge the surveyed region uniformly — it changes its shape, and the population that appears is a different population rather than more of the same one.
What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 0.0 Earth masses, cut off at 0.81 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 5 The same methods around a star of three tenths of a solar mass. Every boundary moves, and they do not move together: the radial-velocity threshold falls, because a lighter star is pulled harder by the same planet; the transit depth rises for the same planet radius, because the stellar disc is smaller. The selection function is a function of the star as well as of the instrument, which is why the exoplanet census around M dwarfs is a different census rather than a deeper one, and why comparing occurrence rates across host masses is the delicate part.

The bias inside the bias

Two effects sit underneath the boundaries and are easy to miss, because both operate on the shape of what is found rather than on whether anything is found at all.

Eccentric orbits are easier to detect at fixed mass. The radial-velocity semi-amplitude carries a factor 1/1e21/\sqrt{1-e^2}, so a planet on an e=0.6e = 0.6 orbit produces a signal 25 per cent larger than the same planet on a circle. Near the threshold that is the difference between a detection and a non-detection, and the effect biases the measured eccentricity distribution upward — which matters, because the eccentricity distribution is the main evidence about how giant planets got where they are. Multiplanet systems are harder. A velocity curve containing three planets is not three times easier to detect than one; the signals interfere, and a periodogram searching for a single sinusoid can miss all three. Early surveys therefore under-counted multiplicity systematically, and the correction requires re-searching archival data with multi-signal models — which has repeatedly turned single-planet systems into multiple ones.

Inverting a selection function

The recipe is simple to state and laborious to carry out.

For each star in the survey, and for each cell of the parameter space, compute the probability that a planet of those parameters around that star would have been detected. That means injecting synthetic signals into the actual data — the actual noise, the actual observing cadence, the actual gaps for weather and daylight — and running the actual detection pipeline on them. The fraction recovered is the completeness of that cell for that star.

Then the occurrence rate in a cell is

f=number of planets found in the cellstars(completeness of the cell for that star),f = \frac{\text{number of planets found in the cell}}{\sum_{\text{stars}} (\text{completeness of the cell for that star})},

with the geometric transit probability included in the completeness for transit surveys.

The denominator does the work, and it can be brutal. For an Earth-analogue cell in a transit survey, the completeness is the product of a geometric probability of 0.5 per cent and a detection efficiency that may itself be 30 per cent — so each detected planet stands for about six hundred. A rate derived from three detections and a denominator like that carries an uncertainty dominated by the fact that three is a small number.

What was actually measured

The mass distribution, once corrected. Giant planets are rare and small planets are common. The corrected distribution rises steeply toward lower masses: roughly dN/dlogmm0.5dN/d\log m \propto m^{-0.5} or steeper below Neptune, so there are several times more super-Earths than Neptunes and several times more Neptunes than Jupiters. In the raw catalogues of 2000 the ordering was the reverse.

The period distribution. Giant-planet occurrence rises with distance out to a few AU, with a pile-up at 3–10 days that is the hot Jupiters and a broad increase beyond 1 AU that was invisible until surveys had run for a decade. Jupiter analogues — giants between 3 and 7 AU on low-eccentricity orbits — occur around roughly 5–10 per cent of Sun-like stars.

The metallicity correlation, which survived the correction. Giant planets are far more common around stars rich in iron: the occurrence rises by roughly a factor of ten from a third of the solar iron abundance to twice it. This was checked hard against selection, because metal-rich stars have more spectral lines and therefore better radial-velocity precision, which would produce exactly this correlation artificially. Correcting for it leaves the effect intact, and it is now one of the strongest pieces of evidence for core accretion as the formation route for giants.

And the correlation that did not survive. Several early claims — an excess of planets around stars of particular ages, a deficit at particular periods — turned out to track the observing programme rather than the sky. The distinguishing test in each case was whether the feature moved when the sample or the threshold changed.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 1.4 Earth masses, cut off at 2.23 AU by the need for three transits in 10 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 6 The same diagram for a ten-year mission rather than four. Only one boundary moves — the one set by how many orbits have to be seen — and it extends to the right, to longer periods. Nothing about sensitivity to mass has changed. Duration and precision are different axes of a survey’s reach, and the planets a longer mission adds are not fainter ones but slower ones, which is why the outer solar system’s analogues are a scheduling problem rather than an instrumental one.

Where the picture stops

A boundary is a threshold, not a cliff. Detection is probabilistic near the limit: a signal at exactly the threshold is found perhaps half the time, depending on the phase, the sampling and the noise realisation. Drawing the boundary as a line is a simplification that the injection–recovery method exists to replace.

Thresholds are not stationary. An instrument is upgraded, a pipeline is improved, a star is observed for another decade. A survey’s completeness is a function of when it is evaluated, and comparing rates from two epochs of the same programme means recomputing both.

Some biases are not in the instrument. Targets are chosen: bright stars, quiet stars, stars without close companions, stars of particular spectral types. The radial-velocity samples systematically exclude active young stars — precisely the ones whose planets would say most about formation — because their jitter makes them unpromising. That is a selection effect in the target list rather than in the detector, and it does not appear anywhere in the completeness calculation.

Threshold effects sharpen distributions artificially. A population that falls smoothly toward small masses, cut by a threshold, produces a catalogue with a sharp edge — and an edge in a histogram is exactly what a physical process would produce too. Distinguishing the two requires the completeness map; without one, every threshold looks like a discovery. The one real edge in the current census is believed because the completeness rises smoothly through it, not because it looks sharp.

A boundary computed for one star is not a boundary for the sample. Every expression above carries a stellar mass and a stellar radius, so each star in a survey has its own set of curves. The diagram drawn here is for a solar-mass star; around a 0.3 M☉ dwarf the transit line sits four times lower in radius and the velocity line about twice lower in mass, which is most of why the small planets in the census orbit small stars. Summing over a heterogeneous sample means summing over thousands of different diagrams, not scaling one.

And a non-detection is not always a null. A planet can be missed because it was there and too faint, or because the pipeline mistook it for a systematic, or because a human vetted it away. The last is the hardest to model, and it is why the modern surveys automate their vetting: an automated mistake can be characterised by injecting signals, and a human one cannot.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 0.0 Earth masses, cut off at 0.81 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.3 AU at 5 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 7 The same instruments pointed at an M dwarf five parsecs away. Everything moves inward and up: a lighter star wobbles more for the same planet and its habitable zone is at a tenth of an astronomical unit, so the same radial-velocity precision reaches Earth masses there and Jupiter masses around a solar-type star. That is not a discovery about where planets are — it is the reason the first temperate planets found were around the smallest stars.

The solar system, observed from outside

The most compact way to feel the size of these effects is to ask what an alien survey of the Sun would find, using the instruments this one has.

A radial-velocity programme at 1 m/s running for thirty years would find Jupiter easily at 12.5 m/s, Saturn marginally at 2.8 m/s and a 29-year period, and nothing else: Neptune is 0.28 m/s, the Earth 0.09 m/s, and Venus and Mars below that. A transit survey would need to be pointed at the Sun from within about half a degree of the ecliptic to see anything at all, and having got there would find the Earth once a year at 84 ppm and Jupiter once a decade at 1 per cent. An astrometric programme at 20 µas from ten parsecs would find Jupiter comfortably and Saturn with effort.

The catalogue entry would read: one giant planet at 5 AU, possibly a second, both on nearly circular orbits. Four terrestrial planets, four giant planets, and everything about the architecture that took centuries to work out would be reduced to that.

It is a useful discipline, because the census of other systems is exactly that thin. What is known about a typical exoplanetary system is one or two of its members, and the rest is inferred from statistics over systems rather than observed in any of them.

It is also the reason the field’s response has been to build instruments whose completeness in that cell is not near zero, rather than to argue harder about the extrapolation. A denominator estimated by simulation is a denominator nobody can check.

One more property of these boundaries is worth stating, because it is what makes the inversion in the last section possible at all. Every one of them is a sharp line in the diagram and a soft one in reality: a planet a little below a threshold is detected some of the time, with a probability that depends on the noise, the number of transits, the stellar activity and the observer’s tolerance for false positives. The completeness function that replaces these lines is measured by injecting synthetic planets into real data and counting how many are recovered — which is the only honest way to get it, and which means a survey’s selection function is itself a measurement with its own error bars.

The generalisation

Every observational science that samples a population through an instrument has this problem, and the ones that solved it did so the same way.

Palaeontology has the same structure: the fossil record over-represents hard-shelled marine organisms in subsiding basins, and the correction is a model of preservation probability rather than an apology. Epidemiology has it as ascertainment bias. Cosmology has it as the Malmquist bias and as the Eddington bias, and the correction of supernova samples for them is a step of the analysis that changes the inferred expansion history.

What is unusual about the exoplanet case is the completeness of the accounting. The detectability of a signal is a computable function of the parameters being inferred, the instrument’s noise is measurable, and the pipeline can be run on synthetic data indefinitely. The result is that the inversion can be done honestly, and the honest answer is that most of what a raw catalogue shows is a picture of the telescope.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 20 ppm transit is a threshold on radius and so a horizontal line at about 0.1 Earth masses, cut off at 1.21 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 8 And the transit boundary at twenty parts per million rather than a hundred. This is the regime a space photometer reaches and a ground-based one cannot, and it is the change that brought Earth-sized planets inside the line for the first time. The boundary is horizontal, because a transit depth is a ratio of radii and does not care how far out the planet is — the one method whose reach in mass is independent of separation, and therefore the one whose selection function is easiest to invert.

The number that moved by a factor of five

The clearest demonstration that a selection function is the whole measurement is a quantity everybody wants and nobody agrees on: the fraction of Sun-like stars with an Earth-sized planet in the habitable zone.

Published values from the same survey data have ranged from a couple of per cent to about half. That is not a disagreement about the data, because the data are one catalogue; it is a disagreement about three things, and they are worth separating because each is a different kind of problem.

The first is the extrapolation. The cell in question — Earth radius, one-year period, Sun-like star — is where the completeness is at its lowest: a transit probability under half a per cent, and a detection efficiency that falls off a cliff for a signal needing three transits in four years of data. So the rate is not measured in that cell; it is measured in adjacent cells where the completeness is decent, and extrapolated in. Different assumed functional forms for how the rate varies with radius and period give different answers, and the data cannot choose between them.

The second is the definition. The habitable zone’s boundaries are computed from a climate model, and the conservative and optimistic boundaries differ by nearly a factor of two in the width of the annulus. A rate quoted per star depends on how wide a target was drawn, and papers quoting different widths are not measuring the same quantity.

The third is reliability. Near the threshold the candidates are marginal, and some fraction of them are instrumental artefacts rather than planets. Estimating that fraction requires injecting false signals as well as real ones, and a rate computed without a reliability correction is systematically high in exactly the cell where the numerator is smallest.

Three detections, a denominator estimated by simulation, an extrapolation, and a definition — that is the whole of the most quoted number in the field, and its factor of five is an honest reflection of where the uncertainty lives. It is not in the photometry.

One more reading of the same domain plot makes the essay’s point about the star rather than about the instrument.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 3.2 Earth masses, cut off at 1.39 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 9 What each method can see around a star half again the Sun’s mass. Every boundary moves and they do not move together: the radial-velocity limit worsens with the star’s mass, the transit depth worsens with its radius, and the imaging limit improves with its brightness. The sky each method draws is a function of the star it is pointed at.

Where this goes next

Once the selection function is written down, the interesting quantity is not the catalogue but the rate — how many planets per star, of what kind, at what distance. That calculation is what turns a list into a statement about the galaxy, and its most famous output is a number that has moved by a factor of five between successive analyses of the same data.

Later rungs on this anchor: injection and recovery in practice. Completeness maps for Kepler and TESS. The metallicity correlation and its tests. Target-selection bias in radial-velocity samples. Eccentricity bias, and why circular orbits are easier to find. The bias against multiplanet systems in velocity data. Detection thresholds as a function of stellar activity. Vetting, automated and human. And the biases of the microlensing surveys, which are the only ones whose sensitivity peaks in the middle of the parameter space rather than at an edge.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 31 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Detection limitDetection thresholdHot jupiterMalmquist biasMass distributionPeriod distributionSelection effectSignal-to-noiseSurvey completenessSurvey duration