Exoplanets

The threshold that is not a threshold

A survey's detection limit is quoted as a number — seven point one — and a pipeline does not behave that way. Half the injected signals come back at the threshold, and full efficiency arrives four units above it.

Assumes Detection bias and Transits.

Every survey draws a different sky, and the boundaries of the sky it draws were drawn there as lines: a velocity precision gives a mass floor, a photometric scatter gives a radius floor, a mission length gives a period ceiling. Each line has planets on one side and none on the other.

No pipeline has ever behaved like that, and the difference is not a refinement. It is the difference between a selection function that can be written down and one that has to be measured, and the only way to measure it is to put signals into the data that are known not to be there and count how many come back.

What comes back is a ramp, not a threshold. Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The measured curve is a gamma cumulative distribution of shape 4.65 and scale 0.98 beginning at 4.1, which is the form a survey's own injection tests are fitted with; the dashed line is the step at 7.1 that a threshold calculation assumes instead. Half the injections are recovered at 8.33, 1.2 units above the nominal threshold — the ramp is a property of the search and the cut is a separate decision, so the two need not meet anywhere in particular. The rest of the disagreement is the area between the curves. The pipeline does not reach 99 per cent efficiency until 14.9, four units above the threshold, and it recovers 45 per cent one unit above it. Over a population whose signal-to-noise falls as s^-2 — which is what a planet population looks like, because there are far more small planets than large ones — the step function counts 1.21 times as many detections as the ramp does. That factor is not an error bar. It multiplies every occurrence rate computed without it, and it is larger for the small planets than for the large ones, because the small ones live where the ramp is.
Fig. 1 Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The dashed step at 7.1 is what a threshold calculation assumes. The measured curve is a gamma cumulative distribution — the form a survey’s own injection tests are fitted with — and the two cross where they should, with half the injections recovered at 8.3. The rest of the disagreement is the area between them. The pipeline does not reach 99 per cent efficiency until 14.9, and over a population whose signal-to-noise falls as s2s^{-2} the step counts 1.21 times as many detections as the ramp.

Why a threshold is a fiction even when it is enforced

The threshold is real in one sense: a search does apply a cut, and a candidate below it is discarded. The fiction is in the number the cut is applied to.

A transit search computes a detection statistic by correlating the light curve against a template — a box of a given depth, duration and period — and taking the largest value over the grid of templates. That statistic is the signal-to-noise of the best-matching template, not of the planet. The two differ for several reasons at once, and every one of them costs efficiency.

The template grid is finite, so a real planet almost never lands on a grid point. The mismatch in period alone smears the folded transit and costs a few per cent of depth. The light curve has been detrended to remove stellar variability and instrumental drift, and the detrending removes some of the transit with it — more for a long transit than a short one, because a long dip resembles the trends being taken out. And the noise is not white: a transit sitting on top of correlated noise has a real signal-to-noise different from the one a white-noise formula computes.

Each of those is a small effect and they compound in the same direction. The result is not a threshold shifted to a different number; it is a threshold smeared into a ramp, because the size of the loss varies from one light curve to the next.

What comes back is a ramp, not a threshold. Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The measured curve is a gamma cumulative distribution of shape 2.2 and scale 1.6 beginning at 3.4, which is the form a survey's own injection tests are fitted with; the dashed line is the step at 7.1 that a threshold calculation assumes instead. Half the injections are recovered at 6.40, within a unit of the nominal threshold, which is the reassuring half of the result. The rest of the disagreement is the area between the curves. The pipeline does not reach 99 per cent efficiency until 14.6, four units above the threshold, and it recovers 75 per cent one unit above it. Over a population whose signal-to-noise falls as s^-2 — which is what a planet population looks like, because there are far more small planets than large ones — the step function counts 0.85 times as many detections as the ramp does — fewer, because this ramp begins well below the cut and recovers signals the threshold discards. That factor is not an error bar. It multiplies every occurrence rate computed without it, and it is larger for the small planets than for the large ones, because the small ones live where the ramp is.
Fig. 2 A less crisp pipeline against the same nominal cut: a recovery curve that begins at 3.4 and is half complete by 6.40, well below the threshold, and still climbing at 14. The sign of the error reverses. This search recovers three quarters of the injections one unit above the cut and a useful fraction below it, so the step — which counts nothing under 7.1 and everything over — takes only 0.85 of the ramp’s yield. What distinguishes this from the previous figure is not the cut, which is 7.1 in both, but where the ramp sits relative to it.

The measurement, which is simply an experiment

The procedure is the one thing in this essay with no theory in it.

Take the real light curve of a real star, with its real noise and its real systematics. Add a transit signal with known depth, period and epoch. Run the complete search pipeline — detrending, folding, template matching, the vetting that rejects eclipsing binaries and instrumental artefacts, all of it — and record whether the injected signal is recovered and whether its parameters come back right. Repeat a few hundred thousand times over the survey’s whole parameter space.

What comes out is the detection efficiency as a function of every parameter that matters: the signal strength, the period, the number of transits, the stellar variability, the data gaps. It is a lookup table rather than a formula, and it is the only honest form the selection function takes.

The two properties that make it worth the expense are that it is end to end and that it uses real noise. A simulation of the pipeline would miss whatever the pipeline does that nobody documented; a synthetic light curve would miss whatever the star does that nobody modelled. Injection into real data tests both at once, and the cost is that the answer belongs to that dataset and cannot be transferred to another.

What comes back is a ramp, not a threshold. Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The measured curve is a gamma cumulative distribution of shape 4.65 and scale 0.98 beginning at 4.1, which is the form a survey's own injection tests are fitted with; the dashed line is the step at 10 that a threshold calculation assumes instead. Half the injections are recovered at 8.33, 1.7 units below the nominal threshold — the ramp is a property of the search and the cut is a separate decision, so the two need not meet anywhere in particular. The rest of the disagreement is the area between the curves. The pipeline does not reach 99 per cent efficiency until 14.9, four units above the threshold, and it recovers 87 per cent one unit above it. Over a population whose signal-to-noise falls as s^-2 — which is what a planet population looks like, because there are far more small planets than large ones — the step function counts 0.75 times as many detections as the ramp does — fewer, because this ramp begins well below the cut and recovers signals the threshold discards. That factor is not an error bar. It multiplies every occurrence rate computed without it, and it is larger for the small planets than for the large ones, because the small ones live where the ramp is.
Fig. 3 The first pipeline again against a stricter cut of 10. The recovery curve does not move, because it is a property of the search and not of the threshold — the same injections come back at the same rate, and only the line drawn through them has shifted. The step now sits well right of the ramp’s midpoint and discards a band of signals the search did in fact find, so it counts 0.75 of the ramp’s yield. The same pipeline, the same data, the same planets, and a correction that has changed both size and sign because one number was moved.

The same experiment in velocities

Transit surveys are where the technique was developed and it is not confined to them, though the version used on radial velocities is structurally different in a way that is worth noticing.

A velocity search fits a Keplerian and asks whether the fit is significant, against a model in which there is no planet. Injecting a synthetic Keplerian into a real velocity time series and re-running the fit gives the same kind of efficiency map: the fraction of injected planets recovered, as a function of mass and period.

Where it differs is in what limits it. A transit survey’s efficiency is limited by noise; a velocity survey’s is limited by sampling. A planet whose period happens to be close to a year, or to a month, or to a day, is aliased with the observing cadence and can be invisible at any amplitude — so the recovery map of a velocity survey has holes in it at particular periods, and the holes are properties of the observing schedule rather than of the instrument.

That makes the velocity version harder to summarise and easier to interpret. A hole in a recovery map is an unambiguous statement: no planet of that period would have been found, whatever its mass. A ramp is a statement about probability, and probabilities compound with the population.

The two surveys’ selection functions therefore have to be measured separately and cannot be combined into one, which is why an occurrence rate derived from transits and one derived from velocities are compared rather than merged even when they concern the same planets.

What the ramp costs a rate

The reason this matters is that nobody wants a detection efficiency for its own sake. What is wanted is an occurrence rate — planets per star, of a given size, at a given period — and a rate is a count divided by a completeness.

The count is what the survey found. The completeness is the fraction of the planets that exist which the survey would have found, and that fraction is the integral of the detection efficiency over the population. Using a step function for the efficiency puts the completeness too high, and a completeness too high puts the rate too low.

The size of the error depends on where the population sits on the ramp, and for the planets anybody cares about it sits badly. Small planets are more numerous than large ones, so the distribution of signal-to-noise rises steeply towards the threshold; the population is concentrated exactly in the region where the ramp and the step disagree most.

What comes back is a ramp, not a threshold. Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The measured curve is a gamma cumulative distribution of shape 4.65 and scale 0.98 beginning at 4.1, which is the form a survey's own injection tests are fitted with; the dashed line is the step at 7.1 that a threshold calculation assumes instead. Half the injections are recovered at 8.33, 1.2 units above the nominal threshold — the ramp is a property of the search and the cut is a separate decision, so the two need not meet anywhere in particular. The rest of the disagreement is the area between the curves. The pipeline does not reach 99 per cent efficiency until 14.9, four units above the threshold, and it recovers 45 per cent one unit above it. Over a population whose signal-to-noise falls as s^-3.2 — which is what a planet population looks like, because there are far more small planets than large ones — the step function counts 1.29 times as many detections as the ramp does. That factor is not an error bar. It multiplies every occurrence rate computed without it, and it is larger for the small planets than for the large ones, because the small ones live where the ramp is.
Fig. 4 The same efficiency against a steeper population — signal-to-noise falling as s3.2s^{-3.2} rather than s2s^{-2}, which is what a size distribution rising more sharply towards small planets produces. The curves are identical and the overstatement rises to 1.29, because more of the population now sits in the ramp. The correction depends on the thing being measured, so an occurrence rate computed with a step function is wrong by a factor that itself depends on the occurrence rate’s own shape, and the calculation has to be done self-consistently or not at all.

That dependence has a second consequence that is easy to miss. Because the factor is larger for small planets than for large ones, ignoring the ramp does not merely lower every rate — it tilts the size distribution, making small planets look rarer relative to large ones than they are. A histogram of planet radii corrected with a step function has the wrong shape, not merely the wrong normalisation.

Where the ramp does most of its damage

It is worth naming the place in the parameter space where all of this is at its worst, because it is also the place the surveys were built for.

A planet’s detection statistic grows as the square root of the number of transits observed. A short-period planet delivers dozens; a planet at one astronomical unit around a solar-type star delivers three in four years, and three is the minimum a search will accept as periodic at all. So the longest periods a survey can reach are reached by planets sitting at the very bottom of the ramp, where the efficiency is a few tens of per cent and its uncertainty is comparable to its value.

The completeness correction at long period is therefore both the largest and the least well determined, and it multiplies precisely the quantity the whole enterprise is pointed at. The occurrence of Earth-size planets in the habitable zone is a number obtained by dividing a handful of detections by a completeness of a few per cent, and a few per cent is a region where the injection experiment has the fewest recoveries to average over.

That is a different situation from a merely noisy measurement. A correction of a factor of thirty is not a correction; it is an extrapolation, and its error bar is dominated by whether the efficiency curve has the right shape rather than by how many planets were found. Where the boundaries of the habitable zone are drawn contributes a comparable uncertainty from an entirely unrelated direction, and the two are usually quoted as one range.

A bias that runs the other way

There is a second effect of the same measurement, and it pushes in the opposite direction, which is why the sign of the total correction is not obvious in advance.

A detection near the threshold is detected partly because it was lucky. Noise that happened to deepen the transit pushes a marginal signal over the cut; noise that happened to fill it in pushes one under. Since there are more planets just below the threshold than just above it, more are promoted than demoted, and the detected sample near the cut contains planets whose measured depths are systematically too large.

That is the Eddington bias, and it is the same arithmetic that inflates a luminosity function’s bright end when a scatter is applied to a falling distribution. The magnitude depends on the product of the noise and the steepness of the population, and near a detection threshold both are at their largest.

So a marginal detection is doubly untrustworthy: less likely to have been found than a threshold calculation assumes, and — if found — measured too big. The first depresses the count and the second inflates the individual radii, and disentangling them requires the injection experiment to record not just whether a signal was recovered but with what parameters.

What comes back is a ramp, not a threshold. Detection efficiency against signal-to-noise: the fraction of synthetic transits injected into real photometry that the pipeline afterwards finds. The measured curve is a gamma cumulative distribution of shape 9 and scale 0.55 beginning at 5.6, which is the form a survey's own injection tests are fitted with; the dashed line is the step at 7.1 that a threshold calculation assumes instead. Half the injections are recovered at 10.37, 3.3 units above the nominal threshold — the ramp is a property of the search and the cut is a separate decision, so the two need not meet anywhere in particular. The rest of the disagreement is the area between the curves. The pipeline does not reach 99 per cent efficiency until 15.2, four units above the threshold, and it recovers 4 per cent one unit above it. Over a population whose signal-to-noise falls as s^-2 — which is what a planet population looks like, because there are far more small planets than large ones — the step function counts 1.69 times as many detections as the ramp does. That factor is not an error bar. It multiplies every occurrence rate computed without it, and it is larger for the small planets than for the large ones, because the small ones live where the ramp is.
Fig. 5 The other failure: a pipeline whose losses are uniform, so the ramp is steep, but large, so it starts late. Half the injections are back only at 10.37, and the step at 7.1 counts 1.69 times the ramp’s yield — worse than the broad case, from a curve that looks far more like the idealisation. A sharp recovery curve is not a safe one. What the step needs in order to be right is not that the ramp be narrow but that it be narrow and centred on the cut, and no published search has ever reported both.

What the analytic boundaries are still for

None of this makes the drawn domains useless, and it is worth being clear about what they do.

What each method can see. Planet mass against orbital distance, both logarithmic, with the detection threshold of each method drawn as the boundary it actually is. Radial velocity at 1 m/s needs mass rising as √a; astrometry at 20 µas needs it falling as 1/a, which is the only method that gets easier further out; a 100 ppm transit is a threshold on radius and so a horizontal line at about 1.4 Earth masses, cut off at 1.21 AU by the need for three transits in 4 years; direct imaging begins outside the diffraction limit, 0.6 AU at 10 parsecs for a 39 m aperture at 10 µm. The solar system is drawn on top: for two decades every one of its planets except Jupiter lay outside every region, which is the whole of what the early census was measuring.
Fig. 6 The detection thresholds of four methods drawn as boundaries in the mass–distance plane: radial velocity needing mass to rise as the square root of separation, astrometry needing it to fall as the inverse, a transit threshold horizontal in radius and cut off by the mission length, and direct imaging beginning outside the diffraction limit. These lines are exactly right about slopes and exactly wrong about edges. Each is a locus of constant signal-to-noise, and constant signal-to-noise is a real and useful thing — it says how a method’s sensitivity scales, which is what decides whether a new instrument opens a region or merely deepens one. What none of them is, is the edge of a detected population.

The distinction is between scaling and completeness. A boundary’s slope is a statement about physics — a transit depth is a ratio of areas, a velocity amplitude falls as the cube root of the period — and physics is not smeared by a pipeline. Its position is a statement about a threshold, and that is the part injection and recovery replaces.

Which is why both pictures are needed and neither substitutes for the other. The boundary says which planets are in reach; the recovery curve says what fraction of those in reach were actually found.

What the ramp’s shape says about the pipeline

The two fitted parameters are not merely a convenience. They separate two different failures, and reading them apart is how a survey learns what to fix.

The offset — where the curve leaves zero — is where the search first becomes capable of finding anything at all, and it can sit on either side of the nominal cut. When it sits below, a signal between the offset and the cut is sometimes found by the search and then thrown away by the threshold, which is a deliberate choice rather than a loss: the cut is set where the false-alarm rate becomes intolerable, and the survey is buying purity with completeness. When it sits above, the cut is doing nothing at all and the search’s own losses are the whole of the selection.

The shape parameter is how sharply the curve rises, and it measures the scatter of the losses rather than their size. A search that lost a fixed twenty per cent of every signal’s depth would still have a step, at a threshold twenty per cent higher. A step becomes a ramp only when the loss varies from one target to another — so the shape parameter is a direct readout of how heterogeneous the data are, and a pipeline improvement that makes the detrending more uniform steepens the curve even if it does not make it more sensitive on average.

Those are different problems with different fixes, and a survey that quoted only its threshold could not tell them apart. The second is the one that has improved most over the life of the transit surveys, which is why the completeness of later data releases is higher at a given signal-to-noise without the noise itself having fallen.

A functional form fitted to an experiment, not to a theory

The gamma-function form used here is not a theory of anything. It is the two-parameter family the Kepler and TESS injection tests were fitted with, chosen because it rises from zero, saturates at one, and has a shape parameter that interpolates between a gentle ramp and a step.

The fitted parameters are what the experiment produced, and they differ between surveys, between data releases of the same survey, and between stellar types within one release. A quiet solar-type star and an active M dwarf have different recovery curves from the same pipeline, because the detrending removes more from the second.

And the numbers move between reprocessings. Kepler’s final data release changed its measured completeness enough to move published occurrence rates by tens of per cent, with no new observations and no new planets — the pipeline had been improved, the injection experiment rerun, and the selection function was a different function. That is a peculiar kind of measurement to depend on, and there is no alternative to it.

Four defensible boxes, and a factor of 2.9 between the answers. Above: the occurrence surface in starlight received against planet radius, with four published definitions of "an Earth-size planet in the habitable zone" drawn on it as rectangles. The surface is a stated parameterisation — two lognormal populations at 1.3 and 2.4 Earth radii, with the radius valley at 1.9 between them, normalised so that the whole of it comes to 0.5 planets per star between one and four Earth radii inside a hundred days. Integrating it over the four boxes gives 17%, 8%, 6%, 10% — a factor of 2.9 between the conservative zone and a broad definition, before any error bar is attached to any of them. The dashed curve is what a transit survey can actually see: at one year around a Sun-like star the smallest detectable planet is 1.6 Earth radii, so the lower-left corner of every box contains no detections at all and the rate there is an extrapolation of a fitted surface rather than a count of anything. Below: the same integral with only one corner moved. Holding the flux range fixed and sliding the radius bound from 1.5 to 1.75 Earth radii — a quarter of an Earth radius, well inside the uncertainty of a measured planetary radius — changes the answer by 43 per cent, which is larger than every error bar quoted with any of these numbers. The published values of η⊕ span two per cent to sixty; roughly a factor of 2.9 of that is definition, and the rest is how far each author was willing to extrapolate past the dashed line.
Fig. 7 What the correction is ultimately for: the occurrence of Earth-size planets in the habitable zone, under four published definitions of what that phrase means. The vertical spread between the definitions is a disagreement about boundaries; the horizontal uncertainty within any one of them is where the completeness correction lives, and it is the larger of the two for the definitions that reach to the longest periods. A planet at one astronomical unit around a solar-type star transits three times in four years at a signal-to-noise near the threshold, which is to say in the middle of the ramp.

One axis where a real efficiency has five

The efficiency is drawn against one variable and depends on several. A real recovery curve is a function of period as well as signal strength, because a long period means few transits and few transits mean the detrending has more opportunity to remove the signal. Collapsing that onto one axis hides a dependence that is the main reason long-period completeness is so poor.

Vetting is not in the drawn curve at all. An injected signal that the search finds may still be rejected by the automated checks that remove eclipsing binaries and systematics, and the fraction lost to vetting is a second efficiency that has to be measured by injecting signals into the vetting stage. It behaves nothing like a ramp; it is close to flat in signal strength and depends on properties of the light curve instead.

And the population’s signal-to-noise distribution is assumed to be a power law, which is a stand-in for a real planet distribution convolved with a real distribution of stellar noise. The overstatement factors quoted are therefore illustrative of a size, not a published correction for any survey.

Nor is the noise the drawn axis assumes the noise a real search fights. Signal-to-noise here is the white-noise quantity that counting statistics gives. A real light curve carries correlated noise on the timescale of a transit — from stellar granulation, from spacecraft pointing drift, from thermal cycling — and correlated noise does not average down as the square root of the number of points. Two stars with identical white-noise levels can have detection efficiencies that differ by a factor, and the injection experiment measures that difference without ever naming it.

The habit this belongs to

The general shape is that an instrument’s response has to be measured with the instrument rather than derived from its specification, and astronomy arrives at that conclusion repeatedly and reluctantly.

A photometric system is calibrated by observing standard stars rather than by integrating the filter’s transmission curve, because the response of a real system is not the product of its documented parts. A spectrograph’s line-spread function is measured from arc lamps rather than computed from the grating. A detector’s flat field is taken, not modelled.

A pipeline’s detection efficiency is the same kind of quantity, and it is the last one to have been treated that way. For the first decade of transit surveys, completeness was computed analytically, and the occurrence rates of that period are now known to have been wrong in the direction this essay describes.

The reason it took a decade is that the experiment is expensive and its result is unglamorous. Injecting half a million signals into a survey’s data and re-running the whole search costs as much computing as the original search did, and produces a table.

There is also a reason of temperament. A threshold is a number a survey can publish about itself, and a recovery curve is an admission that the number is approximate. The transition from one to the other happened when the quantity being derived stopped being a list of planets and became a rate — because a list is not damaged by a fuzzy selection function and a rate is destroyed by one.

That transition is visible in the language. Early transit papers describe a detection limit; later ones describe a detection efficiency; the most recent describe a completeness map and publish it as a data product alongside the catalogue. A candidate that is validated statistically rather than confirmed individually belongs to the same shift — both are consequences of a survey whose output is a population rather than a set of objects.

Still open: what the pipeline is efficient at

Every number here concerns whether a signal is found. Nothing concerns whether what is found is what was injected.

A recovery is usually scored as a match in period and epoch, which means a signal recovered at twice or half the true period counts as a recovery in some analyses and a failure in others. The two conventions give completeness curves that differ by several per cent at long period, and the difference propagates directly into any rate quoted there.

And the deeper version of the same question is about the parameters rather than the period. A planet recovered with a radius fifteen per cent too large has been detected, and has been mismeasured, and the rate computed from it is right while the size distribution computed from it is not. A bias of exactly that kind follows — one that does not change whether a planet is found, only what it appears to be, and which every method has in a different direction.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Detection limitDetection thresholdEddington biasInjection recoveryOccurrence ratePhoton noiseSelection effectSignal-to-noiseSurvey completenessTransit depth