Exoplanets

A planet that is never confirmed, only validated

A background eclipsing binary diluted by the target's light reproduces a planetary transit depth exactly, and no amount of better photometry separates the two. Most known planets are therefore the output of a probability calculation rather than a detection, and the honest statement about them is a statement about a false-positive rate.

Assumes Transits, Binary stars and Detection bias.

The first two rungs of this ladder read a light curve as though it were certainly a planet. A depth gives a radius ratio; four contact points give an impact parameter and a stellar density. Both are correct arithmetic, and both begin from an assumption that a transit survey cannot check.

The assumption is that the dip is a planet. It very often is not, and the alternative that matters is disarmingly simple: an eclipsing binary somewhere else in the aperture, whose deep eclipse is diluted by the target star’s light down to a planetary depth.

Two objects, the same 1.055 per cent, and only one of them a planet. A planet of 0.1027 stellar radii transiting at impact parameter 0.3, and a background eclipsing binary of radius ratio 0.62 whose 38 per cent eclipse is diluted by the target's light to 1.055 per cent — the same depth to nine decimal places, because the dilution was chosen to make it so. A blend contributing 2.7 per cent of the light in the aperture can manufacture any planetary depth at all, so the depth is not evidence about what produced it. Two things in the same photometry are. The ingress occupies 20.4 per cent of the planet's transit and 80.3 per cent of the blend's, a factor of 3.9: the shape of the shoulders is set by the radius ratio of whatever is actually eclipsing, and dilution scales a curve without changing its shape. And the duration with the period gives the mean density of the star being crossed — 1.41 g/cm³ here against 0.15, a factor of 9 — so a blend usually implies a host of a completely different kind from the one the spectrum shows. Neither test needs an observation the survey did not already make, and neither of them proves a planet: they reject specific alternatives, and what is left is a probability.
Fig. 1 The problem, drawn. A planet of a tenth of a stellar radius and a background eclipsing binary whose 38 per cent eclipse is diluted to the same depth to nine decimal places, because the dilution was chosen to make it so. A blend contributing under three per cent of the light in the aperture can manufacture any planetary depth at all, so the depth is not evidence about what produced it. What differs is the shape of the shoulders and the stellar density the duration implies.

Why the aperture is the problem

A transit survey measures the total light in a region of sky — a few pixels of a space telescope, several arcseconds across — and reports its variation. It does not resolve what is inside.

Within that region there may be the target star, physically bound companions of it, and entirely unrelated stars at any distance along the line of sight. All of them contribute light; any of them may be eclipsing.

The observed fractional depth of an eclipse of a contaminating star is

δobs=δtrueFblendFtotal,\delta_{\text{obs}} = \delta_{\text{true}}\,\frac{F_{\text{blend}}}{F_{\text{total}}},

and since the second factor can take any value between zero and one, any true depth maps to any observed depth by a suitable choice of contaminant brightness. The observable is a product of two unknowns.

The rate at which this happens is not small. Roughly one main-sequence star in two hundred is an eclipsing binary at the depths and periods a survey is sensitive to, which is comparable with the rate of transiting planets — and the background stars in a crowded field are far more numerous than the targets. For Kepler’s giant-planet candidates the astrophysical false-positive rate is of order a third to a half; for small planets in multiple systems it is a few per cent, and the difference between those two numbers is itself a piece of physics.

The discriminants that work

Four tests do useful work, and the useful ones are geometric rather than statistical.

The centroid shifts. If the eclipsing source is not the target, the flux-weighted centre of the image moves during the event, towards the target and away from the contaminant. Differencing images in and out of transit locates the source of the variability directly, and for Kepler this was the single most effective test — it rejected a large fraction of candidates outright and, where it did not reject, it bounded how far away the true source could be.

Odd and even eclipses differ. A binary of two unequal stars produces two eclipses per orbit of unequal depth. If the true period is twice the assumed one, alternate “transits” are the primary and secondary eclipses of a binary, and their depths differ. Comparing the mean depth of odd-numbered events with even-numbered ones is free and catches a specific and common error.

A secondary eclipse appears. A planet’s secondary eclipse is at the tens-of-parts-per-million level; a stellar companion’s is at per cent level. Detecting any secondary deeper than a planet could plausibly produce rejects the candidate.

The secondary eclipse. The planet passing behind its star. The step down is the planet's own light being removed — 1800 parts per million of the system — and it is the only measurement in which the planet is subtracted rather than the star.
Fig. 2 The scale the test works at. A planet is seen when it disappears, and the depth of that disappearance is the planet-to-star flux ratio — parts per million for a temperate planet, a couple of thousand for the hottest known. Anything at the per-cent level is a star, and the test needs no modelling at all: it is a depth compared against a ceiling.

The duration implies a stellar density. The four contacts fix the geometry, and the transit duration with the period gives a/Ra/R_\star and hence the mean density of whatever is being crossed:

ρ=3πGP2(aR)3.\rho_\star = \frac{3\pi}{GP^2}\left(\frac{a}{R_\star}\right)^3.

If that density disagrees with the spectroscopic classification of the target — a duration implying a giant when the spectrum says dwarf — the light curve is of something else. It is the most powerful of the four for a well-characterised target, and it is why a survey’s stellar catalogue matters as much as its photometry.

The test has a well-known weakness which is worth naming, because it is the same degeneracy in another guise. An eccentric orbit changes the transit duration — faster at periapsis, slower at apoapsis — so a density discrepancy can be attributed to eccentricity rather than to a blend. With one transit and no other information the two cannot be separated, and the standard practice is to treat the density mismatch as a probability rather than a veto.

A transit of a planet 0.05 of its star's radius. The star's brightness through one transit, computed by integrating the uniform stellar disc over the region the planet covers. The depth is 0.25%, which is exactly (Rp/R⋆)² = 0.0025. The four contact points are where the two discs are externally and internally tangent, at separations 1 ± 0.05 stellar radii.
Fig. 3 The same geometry with the planet half the size — a quarter of a per cent deep rather than one per cent. Everything about the shape is unchanged: four contacts, a flat bottom, the same duration. A transit depth is one number and the shape carries the rest, and the smaller the depth the less of that shape survives the noise. This is why validation becomes necessary rather than optional below about a per cent: the discriminants of the previous section are all statements about the parts of the curve that have gone.

The blends that are planets

Not every blend is a false positive. Some are real planets around the wrong star, and those are the most insidious of all, because every test above passes.

If the transiting object is a planet orbiting a companion to the target rather than the target itself, the light curve is a genuine planetary transit diluted by the target’s light. The centroid barely shifts, because the companion is bound and therefore at the same position. There is no odd–even difference and no detectable secondary. The implied stellar density is that of the companion, which for a bound pair is often similar.

What is wrong is the radius. A depth diluted by the target’s light is shallower than the true one, so the derived radius ratio is too small — and it is being applied to the wrong star’s radius as well. A planet transiting a fainter bound companion can have its radius underestimated by a factor of two or three, which moves it across every category boundary the field cares about.

Kepler’s catalogue was combed for this after the fact, and the estimates are that a substantial minority of its host stars have a bound companion contributing enough light to matter. The corrections are individually modest and collectively significant: they shift the radius distribution, and the radius distribution is where most of the science is.

A false positive that is a real planet is still a wrong answer, and it is the case a test designed to detect “not a planet” cannot catch.

A transit of a planet 0.1 of its star's radius. The star's brightness through one transit, computed by integrating the uniform stellar disc over the region the planet covers. The depth is 1%, which is exactly (Rp/R⋆)² = 0.01. The four contact points are where the two discs are externally and internally tangent, at separations 1 ± 0.1 stellar radii.
Fig. 4 The same planet on a grazing chord, at an impact parameter of 0.85 rather than 0.3. The depth is still one per cent by the algebra and the drawn curve is V-shaped rather than flat-bottomed, because the planet is never fully inside the stellar disc. A V-shaped transit is the classic signature of an eclipsing binary blend — and it is also what a real planet on a grazing orbit produces, which is exactly why the shape is a discriminant that fails in a knowable direction.

Where the tests run out

For a large fraction of candidates none of the four fires. The centroid shift is below the noise, the odd–even difference is consistent with zero, no secondary is detected, and the implied density is compatible with the target.

That is not a confirmation. It is a failure to reject, and the set of blends that survive all four tests is not empty: a background binary close enough to the target on the sky, faint enough that the centroid barely moves, with two nearly identical components so that odd and even eclipses match, and diluted enough that the secondary is buried.

The real test — the one that would settle it — is a mass. A radial-velocity measurement of the target star at the candidate’s period, or a transit-timing signal from a companion, establishes that the object is planetary and not stellar.

Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 1.03 times that floor and after 10000 it is 1.00, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 5 And the reason that is usually impossible. Where integrating longer stops helping: a candidate around a fifteenth-magnitude star needs a radial-velocity precision of metres per second to weigh a small planet, and the photon budget for such a measurement on a four-metre telescope is nights per point. The Kepler field was chosen for photometry, not for spectroscopy, and most of its targets are too faint for the follow-up that would confirm them.

Validation

The response was to change the question. Instead of asking whether the object is a planet, ask how much more probable the planet hypothesis is than every blend hypothesis, given everything measured — and report the answer as a probability.

The calculation has three parts.

A likelihood for each hypothesis. Fit the light curve with a planet model and with each class of blend, and record how well each does. Shape carries most of the discrimination here.

A prior for each hypothesis. How many eclipsing binaries of the required kind are expected within the aperture, at that galactic latitude and depth? This comes from a model of the stellar population plus the measured binary statistics, and it is where the calculation gets its teeth: a very faint background star is a priori unlikely to be within a fraction of an arcsecond of the target.

Constraints from imaging. High-resolution imaging of the target — adaptive optics or speckle — rules out companions outside a fraction of an arcsecond, which removes most of the parameter space the prior would otherwise allow.

Multiply and normalise, and the result is a false-positive probability. By convention a candidate with an FPP below about one per cent is called a validated planet.

That threshold deserves a moment’s scepticism, because it is doing something a probability cut usually does not. Applied to a catalogue of four thousand candidates, a one-per-cent threshold admits, by construction, of order forty false positives — and nobody knows which forty. For a statement about the population that is entirely acceptable and is in fact the point: the contamination is quantified. For a statement about an individual object it is much less comfortable, and a paper about one interesting validated planet is resting on a number computed from a Galactic star-count model.

What validation costs

Three things follow, and they are not small.

A validated planet is a statement about a population model. If the assumed density of background eclipsing binaries is wrong, every FPP computed with it is wrong in the same direction. The errors do not average out across a catalogue; they are coherent, and a systematic in the prior shifts thousands of objects together. That is a different kind of exposure from the usual one: a photometric zero point that is wrong by a per cent makes every radius wrong by half a per cent and nothing else, whereas a stellar-density model that is wrong by a factor of two moves a whole catalogue across the threshold at once. The quantity most worth publishing alongside a validated sample is therefore not the individual probabilities but the sensitivity of the sample to the model that produced them, and it is published far less often than it is computed.

The multiplicity boost is doing a lot of work. A star with several transit-like signals at unrelated periods is very unlikely to have several blends, and that argument reduces the FPP of every candidate in a multiple system by a large factor. It is a good argument. It is also the reason the validated sample is dominated by multiple systems, and therefore why the validated sample is not a random sample of planets.

And the radius depends on the dilution. Even a genuine planet’s radius is wrong if there is unaccounted light in the aperture: an undetected companion contributing 20 per cent of the flux makes the measured depth too shallow and the derived radius about 10 per cent too small. Gaia has since found companions to a large fraction of transit hosts, and the resulting corrections have moved a good many planets across category boundaries.

A transit of a planet 0.1 of its star's radius. The star's brightness through one transit, computed by integrating the uniform stellar disc over the region the planet covers. The depth is 1%, which is exactly (Rp/R⋆)² = 0.01. The four contact points are where the two discs are externally and internally tangent, at separations 1 ± 0.1 stellar radii.
Fig. 6 The same planet on a wider orbit, at forty stellar radii instead of fifteen. The transit is deeper in time — longer ingress relative to nothing, a longer flat bottom — because the planet crosses more slowly, and the depth has not moved at all. The duration is the second number a transit gives, and with the period it returns the stellar density; a candidate whose duration is inconsistent with its host’s density is a blend, and that test needs no extra observation.

What a brighter survey changed

The difficulty described above is partly a property of the transit method and partly a property of one particular survey’s design, and separating the two is worth doing because the second half has since changed.

The field observed by Kepler was chosen for photometric stability: a single patch of sky, observed continuously for four years, at a galactic latitude where the star counts are high enough to fill the detector. That maximised the number of targets and it also maximised the number of contaminating stars per aperture, and it put the median target at a magnitude too faint for radial-velocity follow-up on any telescope.

Its successor made the opposite trade. Surveying the whole sky in short segments, it observes targets several magnitudes brighter, and a brighter target is one whose spectrum can be taken. The cost is a much shorter observing window per field and a much coarser pixel scale, which makes each aperture larger on the sky and therefore more likely to contain a contaminant.

So the false-positive problem did not go away; it changed shape. There are more blends per candidate and there is a real prospect of weighing the survivors. A candidate around a ninth-magnitude star can be observed spectroscopically in an hour, so the question “is it a planet” becomes answerable rather than merely estimable, and a substantial fraction of that survey’s candidates now have measured masses.

The consequence for the census is subtle. The two surveys’ catalogues are not the same kind of object: one is a large, deep, mostly validated sample with a modelled reliability, and the other is a shallower, brighter, more often confirmed one. Combining them into a single occurrence rate means combining a measured contamination with a modelled one, and the combination is only as good as the modelled half.

A survey’s false-positive rate is a design parameter as much as a fact about the sky, and it is decided by the pixel scale and the target brightness before any data are taken.

The candidate that transits once

There is a category of candidate for which almost none of the tests in this essay can be applied, and it is the category that reaches the planets everybody most wants.

A survey that observes a field for a month detects a planet on a hundred-day orbit at most once. A single event gives a depth and a duration and nothing else: there is no period, so there is no odd–even comparison, no phase at which to look for a secondary, and no way to fold the data to improve the signal.

The stellar-density test fails too, because it needs the period. What can be done instead is to invert it — assume the star’s density from its spectrum, and derive the period the duration implies — but that gives a period only under the assumption of a circular orbit, and a long-period planet has no reason to be on one.

So a single-transit candidate is a shape, a depth and a date. What remains is the shape: the ratio of ingress duration to total duration still constrains the radius ratio independently of the depth, and a blend that would produce that depth has a different shoulder profile. That test survives, and it is nearly the only one that does.

Everything else has to come from outside the photometry. High-resolution imaging bounds the companions. A radial-velocity campaign, if the star is bright enough, detects the long-period reflex motion and gives the period directly. And a second event, years later, converts the candidate into an ordinary one.

The planets on the widest orbits a survey can reach are the ones its own machinery is least able to vet, which is the same shape of selection this collection keeps finding: the interesting end of a distribution is the end where the method is weakest.

One whole orbit. The system's total brightness through one orbit: the transit at phase 0, the slow rise and fall of the planet's illuminated hemisphere between, and the secondary eclipse at phase 0.5 where the planet's own light is removed. The transit is 1%; the secondary eclipse is 1800 ppm, about 6 times shallower.
Fig. 7 What the whole orbit looks like when the photometry is good enough to see it. The transit at phase zero, the planet’s illuminated hemisphere rising and falling through the orbit, and the secondary eclipse at phase 0.5 where the planet’s own light is removed. Every one of these features is an independent check that the object is a planet, and a system with a measurable secondary eclipse is not validated but confirmed — which is why the phrase in this essay’s title applies to the faint majority and not to the bright few.

What the honest statement is

A validated planet is not a detection with an unusually careful write-up. It is a different kind of claim.

A confirmed planet has a mass and a radius, and the object is known to be planetary because it has been weighed. A validated planet has a light curve, a set of failed rejections and a probability, and it is a planet in the sense that the alternative is improbable given a model of the sky.

Both are useful. Occurrence rates computed from validated samples are the basis of essentially everything known about how common small planets are, and the practice is defensible because the false-positive rate is estimated rather than ignored. The rate at which planets were missed is estimated the same way — by injecting synthetic signals and counting recoveries — so an occurrence rate is a measured count divided by a modelled completeness and multiplied by a modelled reliability, with two of the three terms coming from simulations.

What is not defensible is quoting the two as though they were the same. A catalogue entry that says “confirmed” and a catalogue entry that says “validated” differ by whether anything was measured, and the number of objects in the second category exceeds the number in the first by a wide margin.

Three transit depths, on the only axis that holds them. The flux deficit of three planets crossing the same star at the same impact parameter, computed by integrating the stellar disc under each one. The depths are Jupiter 1.05%, Neptune 0.125%, Earth 84 ppm — a range of a factor of 125, which is why the axis is logarithmic. The shape is identical because the geometry is; only the height differs, and it differs as the square of the radius ratio.
Fig. 8 And the scale that makes the whole difficulty inevitable. Three transit depths on the only axis that holds them: an Earth around a Sun removes 84 parts per million. A blend needs to dilute a stellar eclipse by a factor of five thousand to imitate that — which sounds like a lot until one notices that a star five thousand times fainter than the target is entirely ordinary, and there are many of them within a few arcseconds of anything.

One more blend makes the limiting case explicit.

Two objects, the same 1.055 per cent, and only one of them a planet. A planet of 0.1027 stellar radii transiting at impact parameter 0.3, and a background eclipsing binary of radius ratio 0.62 whose 38 per cent eclipse is diluted by the target's light to 1.055 per cent — the same depth to nine decimal places, because the dilution was chosen to make it so. A blend contributing 2.7 per cent of the light in the aperture can manufacture any planetary depth at all, so the depth is not evidence about what produced it. Two things in the same photometry are. The ingress occupies 20.4 per cent of the planet's transit and 80.3 per cent of the blend's, a factor of 3.9: the shape of the shoulders is set by the radius ratio of whatever is actually eclipsing, and dilution scales a curve without changing its shape. And the duration with the period gives the mean density of the star being crossed — 1.41 g/cm³ here against 0.15, a factor of 9 — so a blend usually implies a host of a completely different kind from the one the spectrum shows. Neither test needs an observation the survey did not already make, and neither of them proves a planet: they reject specific alternatives, and what is left is a probability.
Fig. 9 An eclipsing binary with a companion 0.62 of the primary’s radius, blended into the same aperture. Its eclipse is deep and its diluted depth is planetary, and the light curve alone cannot tell it from a small planet around a bright star — which is the ambiguity validation is a statistical answer to.

A validated planet is a statistical statement about a population of false positives, and a confirmed one is a measurement of a mass. The catalogue holds many more of the first than the second, and the distinction is rarely carried into the papers that use it.

Where this ladder goes next

This rung has established what a transit survey can and cannot conclude on its own, and what the word validated is doing.

The rung above is the statistical treatment of the whole catalogue: propagating a per-object false-positive probability through an occurrence-rate calculation, rather than applying a cut and treating what survives as clean. The difference matters most exactly where the rates are most interesting.

Beside it lies the use of the false positives themselves. The eclipsing binaries a planet survey rejects are a large and uniformly selected sample, and they have produced some of the best stellar radii and masses available — the discarded half of the data turning out to be worth as much as the half that was wanted.

And below it, the habit: a measurement that cannot distinguish two hypotheses has not measured which one is true, however precise it is. The depth of a transit is measured to parts per million and says nothing whatever about what produced it, and recognising that is the difference between a candidate and a planet.