Exoplanets

How many planets a star has is not a measurement

Draw five thousand identical five-planet systems, scatter their orbital planes by half a degree, and a third of the detections show all five. Scatter them by ten degrees and two thirds show exactly one. Every system has five.

Assumes Detection bias and Transits.

A catalogue reports how many planets each star has, and the number is read as a property of the star. It is not. It is a property of the star, the survey, and one angle nobody has ever measured — and the third of those moves the answer further than the first two together.

How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 5 planets, at four mutual inclination dispersions. 40,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 12, 16, 21, 27, 34 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.5° of dispersion 33 per cent of the detected systems show all 5 planets and the mean apparent multiplicity is 3.13; at 10° it is 1.39, with 68 per cent of them showing exactly one. Every one of those systems has 5 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 8%, 9%, 12%, 19% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 2.2. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 1 The multiplicity distribution a transit catalogue would contain, for systems that all truly hold five planets, at four mutual inclination dispersions. Forty thousand systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 12 to 34 stellar radii; the bars are conditioned on at least one planet transiting, which is what puts a system in a catalogue at all. At half a degree of dispersion the mean apparent multiplicity is 3.13 and a third of the detections show all five. At ten degrees it is 1.39, with 68 per cent showing exactly one. Every one of those systems has five planets.

Two reasons a planet is missed, and only one of them is about the system

A planet transits if the line of sight lies within an angle R/aR_\star/a of its orbital plane. That is a narrow band, and it is narrower for the outer planets: at twelve stellar radii the band is five degrees wide, and at thirty-four it is under two.

So even a perfectly flat system is usually seen incompletely, and the first version of the calculation above was built on the assumption that it would not be. A coplanar system viewed from a random direction shows its inner planets and loses its outer ones, and the loss is geometric rather than dynamical — a consequence of the orbits being different sizes rather than of their being tilted relative to each other.

That is the part of the shortfall no better observation can fix, because it is not a sensitivity limit. The outer planet is not faint; it does not cross the star at all.

The second reason is the mutual inclination. If the planes are tilted with respect to one another by a degree or two, then a viewing direction that satisfies one planet’s condition may not satisfy another’s, and the effect compounds along the system.

How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 5 planets, at four mutual inclination dispersions. 40,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 12, 16, 21, 27, 34 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.2° of dispersion 35 per cent of the detected systems show all 5 planets and the mean apparent multiplicity is 3.15; at 6° it is 1.64, with 53 per cent of them showing exactly one. Every one of those systems has 5 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 8%, 9%, 10%, 16% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 1.9. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 2 The same systems at a finer range of dispersions, from a fifth of a degree up to six. Even at the tightest, where the planes are flatter than the Solar System’s, only a minority of detections show the full set — because the geometric term is still there and does not care about the tilt. The dispersion at which the mutual inclination starts to dominate is around two degrees, which is where the tilt becomes comparable to the transit band of the outermost planet.

The compounding, written out

The reason the two effects combine so badly is that they act on the same quantity and multiply rather than add.

A planet transits when cosi<R/a|\cos i| < R_\star/a, where ii is measured from the line of sight. In a flat system all the planets share one ii, so the condition is a single inequality evaluated at several values of aa — and the outermost planet’s inequality is the binding one. The probability of seeing all of them is the probability of seeing the outermost, which is R/amaxR_\star/a_{\text{max}} and nothing else.

Tilt the planes and the conditions stop being nested. Each planet now has its own effective inclination, and the events are nearly independent once the tilt exceeds the narrowest band. The probability of seeing all five becomes a product of five small numbers rather than the smallest of them.

That transition — from “the outermost decides” to “all five must succeed” — is the whole of what the dispersion does, and it happens over a range of about one degree. Below it the multiplicity distribution is set by the range of semi-major axes; above it by the tilt. The drawn figures straddle the transition, which is why the leftmost and rightmost dispersions look like different physics.

The Solar System sits on the flat side. Its planets’ mutual inclinations are of order two degrees, and its transit bands are much narrower than that — the Earth’s is a third of a degree — so a distant observer of the Solar System would see at most one or two planets transit, almost never more. The census that first opened this field would have found nothing in it at all, twice over.

A quantity that moves the other way

There is a second number the survey reports, and the two are usually quoted in the same sentence: the fraction of stars found to have any planet at all.

That fraction goes up with the dispersion — 8 per cent, 9, 12 and 19 across the four cases drawn — because scattering the orbital planes gives each of them an independent chance of crossing the line of sight. A flat system is all or nothing from any given direction; a scattered one is a lottery with five tickets.

So a survey observing a population of scattered systems finds more stars with planets and fewer planets per star. Both numbers move, in opposite directions, and neither on its own says which has happened.

How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 3 planets, at four mutual inclination dispersions. 40,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 12, 16, 21, 27, 34 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.5° of dispersion 58 per cent of the detected systems show all 3 planets and the mean apparent multiplicity is 2.34; at 10° it is 1.23, with 79 per cent of them showing exactly one. Every one of those systems has 3 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 8%, 9%, 11%, 16% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 1.9. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 3 Three-planet systems rather than five, at the same dispersions. The mean apparent multiplicity falls to 2.15 at the tightest and 1.24 at the loosest, and the fraction with any detection falls too, because there are fewer tickets. The two observables — the fraction of stars with a planet, and the mean number seen per detected star — depend on the true count and the dispersion together, and there are two of each, so a pair of measurements can in principle determine a pair of unknowns. In practice the determination is weak, because the two observables are correlated along nearly the same direction.

The dichotomy that may not exist

The pattern this produces in a real catalogue is a well-known one. Transit surveys find many systems with a single transiting planet and many with three or more, and comparatively few with exactly two — and the number of singles is larger than a single population of flat multi-planet systems can produce.

The standard reading is that there are two populations: compact multi-planet systems that are nearly coplanar, and a separate population of genuinely single planets, or of systems dynamically hot enough that only one is ever seen.

The alternative reading is that there is one population with a broader distribution of mutual inclinations than the flat case assumes, and that the excess singles are multi-planet systems seen badly.

How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 7 planets, at four mutual inclination dispersions. 40,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 12, 16, 21, 27, 34 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.5° of dispersion 0 per cent of the detected systems show all 7 planets and the mean apparent multiplicity is 3.08; at 12° it is 1.32, with 73 per cent of them showing exactly one. Every one of those systems has 7 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 8%, 9%, 14%, 20% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 2.3. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 4 Seven-planet systems across a wider range of dispersions. At twelve degrees the catalogue would contain overwhelmingly singles from a population in which every star has seven planets — 73 per cent of detections showing one, against a mean apparent multiplicity of 3.08 at half a degree — while the fraction of stars with any detection is at its highest of the four. That is what the excess of singles looks like if it is produced by inclination rather than by a second population, and the distribution of the others is what distinguishes the two: a dynamically hot seven-planet system still produces a tail of doubles and triples, and a genuinely single population produces none.

The evidence that has accumulated favours something between the two readings. The mutual inclination dispersion of the compact systems is measured, indirectly, from the ratio of doubles to triples and from transit duration ratios, and it comes out between one and three degrees — flatter than the Solar System’s giant planets and not perfectly flat.

The ratios that carry the information

If the multiplicity distribution is a poor instrument for measuring the true count, it is a surprisingly good one for measuring the dispersion — provided the question is asked as a ratio rather than as a count.

The number of doubles divided by the number of triples does not depend on how many stars were observed, on the detection efficiency at a fixed planet size, or on the true multiplicity’s normalisation. It depends almost entirely on the tilt, because a triple requires three successes where a double requires two, and the odds ratio between them is set by how independent the successes are.

That is the measurement that has actually been made. Fitting the observed ratios of singles to doubles to triples, against a forward model with the true multiplicity distribution and the dispersion as free parameters, gives a dispersion of one to three degrees for the compact systems — and the answer is stable against most of what is uncertain in the forward model, because the ratios are.

A ratio of two badly measured things is often a well measured thing, and it is worth recognising the shape when it recurs. The number of contact points in a transit fixes a set of ratios that survive an unknown stellar radius; a duration ratio between two planets of one system gives an eccentricity difference without the stellar density. The habit is the same: find the combination the nuisance parameter cancels out of.

What flatness is worth knowing

The dispersion is not merely a nuisance parameter. It is one of the few dynamical quantities a transit survey can constrain at all, and it separates formation histories that agree about everything else.

A system assembled quietly in a disc and left alone stays flat, because the disc was flat and nothing has since tilted it. A system that has scattered — or that has been perturbed by a distant companion — carries a mutual inclination proportional to the violence.

So the quantity that most damages the multiplicity measurement is also the most interesting thing the multiplicity measurement could deliver. A resonant chain of planets is the extreme case: assembling one requires smooth migration and nothing else, and such chains are found flat, which is a consistency the geometry could have refused to give.

How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 5 planets, at four mutual inclination dispersions. 40,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 8, 11, 15, 20, 27 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.5° of dispersion 28 per cent of the detected systems show all 5 planets and the mean apparent multiplicity is 2.95; at 10° it is 1.53, with 59 per cent of them showing exactly one. Every one of those systems has 5 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 13%, 13%, 15%, 24% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 1.9. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 5 The same five-planet systems drawn more compactly, from eight to twenty-seven stellar radii. Every transit band is wider and the geometric shortfall is milder: the mean apparent multiplicity rises at every dispersion, and at half a degree nearly half the detections now show all five. The scale of the system and the flatness of it are degenerate in the multiplicity distribution, which is why the measurement is made on systems whose periods are known rather than on the catalogue as a whole.
How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 5 planets, at four mutual inclination dispersions. 80,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 12, 16, 21, 27, 34 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.5° of dispersion 33 per cent of the detected systems show all 5 planets and the mean apparent multiplicity is 3.13; at 10° it is 1.38, with 69 per cent of them showing exactly one. Every one of those systems has 5 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 8%, 9%, 12%, 19% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 2.3. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 6 The same construction at twice the sample, which changes nothing measurable and is worth showing for that reason: the bars move by less than their own width, so the structure in the drawn distributions is the geometry rather than the sampling. What the extra trials buy is confidence in the tails, which is where the distinguishing information lives — the ratio of triples to quadruples at the tightest dispersion is what a real fit is most sensitive to, and it is built from the rarest bins.

Why a velocity survey is worse rather than better

It is natural to expect that radial velocity escapes the geometry, since it detects a planet at any inclination. It escapes that geometry and acquires a different problem.

A velocity search fits one Keplerian at a time, and the second planet is found in what is left over after the first has been removed. That residual, and a residual is not a light curve — it has had a model subtracted from it, and the subtraction absorbs part of whatever else was there. Each successive planet is found at a lower effective amplitude than it would have had in isolation, and the detection of the $n$th requires the first n1n-1 to have been modelled correctly.

The failure mode is specific and has produced retractions: two planets near a period ratio of 2:1 can be fitted by one planet on an eccentric orbit, because an eccentric velocity curve’s second harmonic looks like a companion at half the period. That ambiguity is the same harmonic decomposition the eccentricity bias turns on, read in the opposite direction.

So a velocity survey’s multiplicity is limited by model degeneracy where a transit survey’s is limited by geometry, and the two limits are unrelated. Comparing their multiplicity distributions is therefore informative in a way comparing their occurrence rates is not — the biases have nothing in common, so agreement means something.

What the number is used for, and why it survives being wrong

Multiplicity is quoted in two very different arguments, and only one of them is damaged by everything above.

The first is a statement about architecture: how many planets a typical star has, and therefore how common systems like the Solar System are. That argument is destroyed by the biases — the answer depends on a dispersion measured from the same data, and the error bars are dominated by the model rather than by the counting.

The second is a statement about relative architecture, and it is robust. Whatever the biases, they act the same way on two samples observed the same way, so comparing the multiplicity of planets around metal-rich and metal-poor stars, or around single stars and binaries, or at different periods, gives a difference that the common bias cancels out of.

Almost every result that has survived from multiplicity statistics is of the second kind. Compact multi-planet systems are less common around stars with a close stellar companion; the fraction of stars with a detected system rises with metallicity for giant planets and not for small ones; multi-planet systems have more circular orbits than singles. Each is a comparison, and each is a difference large enough that the shared correction cannot have produced it.

That is the general escape from an unmeasurable systematic, and it is worth naming because it is available far more often than it is taken: if the systematic is the same in two samples, measure the difference and stop trying to correct it. An occurrence rate defined as a ratio to another occurrence rate is the same move applied to a different quantity.

An exact geometry and an assumed distribution

The transit condition is exact geometry: the impact parameter is the projection of the orbit’s offset from the line of sight, scaled by the semi-major axis in stellar radii, and it is compared against one.

The viewing direction is drawn isotropically, which is not an assumption but a fact about where the Earth is relative to a randomly chosen star.

The inclination distribution is Rayleigh, which is an assumption, and it is the standard one because it is what a random walk of small deflections about a common plane produces. A system whose planes were set by a single large event would have a different distribution with the same dispersion, and would give a different multiplicity distribution — the tails matter more than the width, since a system is lost from the full-multiplicity bin by its worst-tilted planet rather than by its typical one.

How many planets a star has is the hardest thing a catalogue measures. The multiplicity distribution a transit catalogue would contain, for systems that all truly hold 5 planets, at four mutual inclination dispersions. 40,000 systems are drawn per dispersion with an isotropic viewing direction and Rayleigh-distributed inclinations about a common plane, at semi-major axes of 20, 32, 50, 80, 125 stellar radii; the bars are conditioned on at least one planet transiting, which is what makes a system appear in a catalogue at all. At 0.5° of dispersion 9 per cent of the detected systems show all 5 planets and the mean apparent multiplicity is 2.43; at 10° it is 1.16, with 85 per cent of them showing exactly one. Every one of those systems has 5 planets. The entire difference between a catalogue of singles and a catalogue of compact multiples is one number that nothing in the light curve measures. And the two effects run in opposite directions: the fraction of stars showing any planet RISES with the dispersion — 5%, 6%, 8%, 11% across the four — because scattering the orbits gives more of them a chance to cross the line of sight, while the number seen per detected star falls by a factor of 2.1. A survey that scatters its systems finds more stars with planets and fewer planets per star, and neither number on its own says which has happened. What no figure here can show is the true dispersion, because the observable is the ratio of those two and a system with fewer planets and a tighter plane reproduces it exactly.
Fig. 7 The same five planets placed further out, from twenty to a hundred and twenty-five stellar radii — periods of weeks to a year rather than days. Every transit band narrows, the fraction of stars with any detection collapses, and the multiplicity distribution is dominated by singles at every dispersion drawn. Nothing about the systems has changed except their size. This is the regime the surveys are now pushed into, and the multiplicity measured there is very nearly a measurement of nothing: at these separations a catalogue reports the geometry rather than the systems.

Every planet assumed detectable if it transits

Every planet is assumed detectable if it transits. A real survey loses small planets to the detection efficiency ramp, and the loss is worse for the outer planets because they transit less often. So the real multiplicity shortfall is the geometric one compounded with a sensitivity one, and the two are not independent — both worsen outward.

The planets are given one semi-major axis list and no size distribution. In a real system the outer planets are not the same size as the inner ones, and whether they are systematically larger or smaller changes the compounding above. The drawn figures hold that fixed to isolate the geometry.

And no figure here shows a real catalogue. Every distribution drawn is a forward model of an assumed population. Running the comparison backwards — from a catalogue to a dispersion — requires a population model for the true multiplicity as well, and the answer is a joint constraint on two distributions rather than a measurement of one.

The shape of the difficulty

What makes multiplicity harder than every other quantity a survey reports is that it is a property of a system rather than of an object, and a survey detects objects.

An occurrence rate is a count of planets divided by a count of stars, and both are extensive: a planet missed lowers the numerator by one and does nothing else. A multiplicity distribution is not extensive. A planet missed moves a whole system from one bin to another, and the system it moves to is populated by systems that genuinely belong there.

So the errors do not average out. In a rate, missing ten per cent of planets costs ten per cent. In a multiplicity distribution, missing ten per cent of planets moves a much larger fraction of systems, because a five-planet system is lost from the top bin if any one of its five is missed.

That is the general lesson and it is not confined to planets. Any measurement of a multiplicity — binary star fractions, satellite counts of galaxies, cluster richness — has the same structure, and in each the quantity is easier to state than to measure for exactly this reason.

The galaxy version is the closest analogue and has the same two-sided failure. A cluster’s richness is a count of member galaxies, the count is incomplete at the faint end, and the faint end is where almost all the galaxies are — so a richness is a number whose meaning depends entirely on the limit it was counted to. The parallel extends to the fix: richnesses are quoted within a stated luminosity range for the same reason multiplicities are quoted within a stated period and radius range, and comparing two of either across different ranges is meaningless.

What planets have that galaxies do not is a hidden angle. A cluster’s members are all there to be counted, however hard; a system’s planets are not there at all unless the geometry cooperates, and no depth of observation changes that. It is the one place in the subject where the sample is limited by orientation rather than by sensitivity, and it makes multiplicity the only quantity here whose measurement cannot be improved by building a larger telescope.

Still open: whether the singles are a population

The question all of this has been circling is whether the excess of single transiting planets is a second population or a viewing angle, and the honest answer is that both contribute and the split is not settled.

What would settle it is a measurement of mutual inclination that does not come from counting. There is one, and it is beginning to be available: a planet’s transit duration relative to the duration its period and the stellar density predict depends on its impact parameter, so a system with several transiting planets gives several impact parameters and therefore the relative tilts directly.

The limitation is that it can only be applied to systems where several planets transit — which is to say, to the flattest systems, the ones least in dispute. The systems whose dispersion matters most are the ones showing a single planet, and they supply exactly one impact parameter and no tilt at all.

That circularity is not unusual and it is not fatal. It means the measurement constrains the flat end of the distribution well and the tail badly, so the inferred dispersion is a lower limit on the population’s spread. Whether there is a separate hot population hiding behind it is a question the method is structurally unable to answer, and saying so is more useful than a number would be.

From here: the correction nobody has yet made jointly

Each argument so far has isolated one bias and computed it exactly. What none of them has done is combine them, and the combination is not a product — the completeness ramp, the eccentricity bias and the multiplicity geometry are correlated through the period, and correcting for them one at a time double-counts the overlap.

What is needed is that joint correction: a forward model of the whole selection function, with the population as its parameters, fitted to the catalogue rather than applied to it. It is how every occurrence rate published since about 2015 has actually been derived, and it inverts the logic of every correction so far — instead of correcting the data towards the truth, it corrupts a candidate truth until it matches the data.

About the same objects

Not linked from either essay — found by the objects both name.

The objects this essay names

Each one links to every other essay that touches it.

CoplanarityDetection limitMutual inclinationOccurrence rateOrbital inclinationPlanetary systemRadial velocitySelection effectSurvey completenessTransit probability