A star subtracted using the star
Assumes Direct imaging and Seeing.
The difficulty of imaging a planet is not the angular separation. A young Jupiter half an arcsecond from its star is comfortably resolved by an eight-metre telescope. The difficulty is the contrast: the planet is between a millionth and a billionth of the star’s brightness, and what stands between them is not the star’s own image but the halo of scattered light around it.
That halo is not smooth. It is a field of speckles — the residual wavefront errors after adaptive-optics correction, imaged as a swarm of diffraction-limited spots at the same brightness a planet would have. They come and go on timescales from milliseconds to hours, and the slowly varying ones are indistinguishable from a companion.
Nine orders of magnitude, half an arcsecond apart is the statement of the problem. This essay is about the only technique that has solved any part of it, and about the price that technique charges.
Why the reference has to come from the data
The natural approach is to observe a reference star, one with no companion, and subtract its image. That fails for a specific and instructive reason.
The speckle pattern is a snapshot of the instrument’s wavefront error, which depends on the telescope’s flexure with elevation, on the temperature of the optics, on the adaptive-optics system’s state and on the atmosphere. Two stars observed an hour apart at different elevations have different speckle patterns, and the difference is larger than the planet.
So the reference has to be built from images taken under conditions as nearly identical as possible, which means from the target’s own sequence. The only thing that must differ between the images being combined is the position of the sky — which is what pupil-stabilised observing provides.
That is a clean trick and it has an unavoidable consequence: any real companion is in every frame used to build the reference.
There is a second variant of the same idea that exploits a different axis, and it is worth naming because the two are often combined. Spectral differential imaging uses the fact that a speckle’s position scales with wavelength while a planet’s does not: observing at several wavelengths simultaneously and rescaling each image moves the speckles and holds the planet fixed, so a median over wavelength is again a speckle reference. It has the same structure and the same defect — a planet whose spectrum is featureless is present in every rescaled frame and subtracts itself — and the two techniques are used together because their self-subtraction is worst in different places.
What the self-subtraction costs
How much of the planet survives depends on how far it moves between the frames that are combined. The movement is an arc: the field rotation angle times the separation. Far from the star the arc is many resolution elements and the planet is effectively absent from the reference; close in the arc is a fraction of one and the planet is fully present.
The result is a throughput that rises with separation, and the transition is at the separation where the arc equals about one resolution element.
The consequence is that a contrast curve — the standard product of a high-contrast observation, giving the faintest companion detectable against separation — is not a property of the telescope. It is a property of the telescope, the algorithm, the sequence’s field rotation, and how the throughput was estimated. Two groups analysing the same data with different algorithms produce different curves.
The throughput has to be injected
Since no calculation predicts the surviving fraction for a real companion, it is measured: fake planets of known brightness are injected into the raw data at a grid of separations and position angles, the whole reduction is run, and the recovered brightness is compared against the injected one.
That is expensive — the reduction is run hundreds of times — and it is the only defensible way to produce a contrast curve. It also has to be done with the fake planets far enough from any real one that the two do not interfere, and at enough position angles that the azimuthal variation is sampled.
The injection-recovery procedure has a second use that is at least as important. It measures not only the throughput but the noise after the reduction, which is not the noise before it: an algorithm that subtracts aggressively reduces the speckles and the planet together, and the ratio of the two is what a detection limit is made of.
There is a subtlety in the injection that took the field some time to settle. A fake planet injected at the same separation as a real one biases the reduction, because the reference is built from frames containing both; and a fake planet injected into data that already contain a real one measures the throughput in the presence of that planet, which is not the same as in its absence. The standard practice is to inject at several position angles well away from any candidate, to run the whole reduction each time, and to quote the mean and scatter of the recovered fractions — with the scatter itself being informative, since a large azimuthal variation means the speckle field is structured rather than random.
One more property of the injection is worth stating because it decides how the results are used. The recovered throughput depends on the injected brightness: a bright fake planet dominates the frames it appears in, biases the reference towards itself, and is subtracted harder than a faint one. So the throughput is not a single number but a function of contrast, and a curve computed by injecting bright planets overestimates the sensitivity to faint ones. The standard is to inject at a brightness near the expected detection limit and to iterate, which converges in a couple of passes and which is a further reason the reduction is run hundreds of times. The same non-linearity appears in any pipeline that estimates its own background, and the same remedy applies.
The algorithm is a projection, and that is exploitable
The reduction described so far is a median: build a reference frame from the sequence, subtract it, derotate, combine. Almost nobody does that any more, and the algorithms that replaced it make the self-subtraction both worse and calculable.
The modern reduction expresses the reference as a fitted combination of the other frames rather than as their median. One family finds, for each small region of each frame, the linear combination of other frames that minimises the residual there; another decomposes the sequence into principal components and projects each frame onto the first several of them. Both remove far more of the speckle field than a median does, because both are allowed to fit it.
Both also fit the planet, and for the same reason. A linear combination chosen to minimise the residual in a region containing a planet will use whatever freedom it has to reduce the planet’s own flux, and the more components it is given the more of the planet it removes. That is why the number of components retained is the field’s central tuning parameter: it trades speckle suppression against throughput along a curve with an optimum that depends on the dataset, the separation and the rotation, and there is no theory that locates it.
What the fitting buys back is that the operation is linear. A projection onto a fixed set of components is a matrix, so the effect of the reduction on any known signal can be computed rather than measured: take the planet’s point-spread function, put it through the same projection with the same components, and see what emerges. That is forward modelling, and it gives the throughput analytically for the price of one matrix multiplication instead of hundreds of reductions.
The catch is that the components themselves depend on the data, and the data contain the planet — so the projection is only linear once the components are fixed, and the components were computed from frames the planet was in. In practice the effect is second order for a faint companion and is not for a bright one, which is why forward modelling is used to measure a candidate’s photometry and astrometry, and injection-recovery is still used to establish the detection limit. The two methods answer different questions: one asks what happened to a source that is there, the other asks how faint a source could have been and gone unnoticed.
There is a practical consequence for anybody comparing published results. A photometric measurement of an imaged planet is algorithm-dependent at the tens-of-per-cent level unless a forward model or an injection was used to correct it, and the older literature frequently reports the raw recovered flux. Several early estimates of imaged planets’ luminosities — and therefore of their inferred masses, which depend steeply on luminosity — were low for exactly this reason.
Where the detection limit really sits
Two further effects push the achievable contrast away from what the photon statistics suggest, and both are worst where the planets are.
The speckle noise is correlated. Residual speckles are not independent from pixel to pixel; they have the size of a resolution element and they persist over minutes. So the noise in an aperture is not the square root of the counts, and estimating it requires measuring the scatter of the reduced image at the same separation rather than assuming Poisson statistics.
The number of independent samples is small. At a separation of two resolution elements there are only about a dozen independent resolution elements around the annulus. Estimating a noise level from twelve samples, and then applying a five-sigma threshold, is statistically quite different from doing it with thousands — the threshold has to be raised substantially to keep the false-alarm rate at the intended level, by an amount that grows sharply as the separation shrinks.
That second effect is the same look-elsewhere arithmetic that inflates the threshold in a period search, arriving in a spatial rather than a frequency domain, and it was not applied in the field’s first decade.
There is a third effect and it is the one that limits the very best observations. After all the speckles that field rotation can remove have been removed, what is left is the static part of the wavefront error — the difference between what the wavefront sensor measures and what the science camera sees, which does not rotate with the sky because it is in the instrument. Field rotation cannot touch it and no amount of observing time averages it down. Reducing it is a matter of measuring the science camera’s own wavefront, which requires an instrument that can do that, and it is the single largest term in the error budget of any high-contrast system.
What was actually measured
Three results establish both the technique’s power and its pitfalls.
The directly imaged planetary systems. A handful of young, wide, massive planets have been imaged and followed over orbital arcs of years. Every one of them was found beyond the separation at which self-subtraction is severe, which is a statement about the technique as much as about the planets.
The retracted candidates. Several claimed companions at small separations have not been confirmed, and the leading explanation in each case is a residual speckle combined with an underestimated threshold. The field’s response was to adopt the small-sample correction and to require detections in more than one epoch and more than one band.
And the injection-recovery standard. Contrast curves published before about 2014 and after are not comparable, because the earlier ones generally did not correct for throughput. The difference is a factor of two to five at small separations, always in the direction of the earlier curves being optimistic.
And a fourth, and it concerns what “detection” means in this field. A five-sigma point source in a reduced image is a candidate, not a planet: it has to be recovered at a second epoch, at a position consistent with orbital motion rather than with the background star it might be, and ideally in a second band. The common-proper-motion test — showing that the candidate moves with the star rather than staying fixed as the star’s own proper motion carries it across a static background — is what actually establishes companionship, and it takes a year or more. Several published candidates have failed it. That the confirmation is an astrometric measurement rather than a photometric one is worth noticing: the detection is a contrast problem and the confirmation is a position problem, and they have entirely different systematics.
Why an algorithm belongs in the error bar
The underlying point deserves stating on its own because it is unusual and becoming less so.
When the reference for a measurement is estimated from the data, the analysis is part of the instrument. A contrast curve depends on the number of principal components retained, on how the frames were selected, on the rotation available and on the injection procedure — and none of those is a property of the telescope. Reporting a detection limit without reporting the algorithm is like reporting a magnitude without reporting the filter.
The field has converged on the right practice, which is to publish the reduction alongside the result and to characterise the pipeline by injection rather than by calculation. That is now standard in high-contrast imaging and it is not standard everywhere else, though the same argument applies wherever a background is estimated from the data — a detrended transit light curve has exactly the same structure, and a signal removed by the detrending is exactly this self-subtraction under another name.
The deeper reason to take the parallel seriously is that both cases have the same remedy. Whatever the algorithm does to the data, do it to a known signal injected into the same data, and measure what comes out. That converts an unknowable systematic into a measured throughput, and it costs computer time rather than telescope time.
A final observation about what the technique has and has not delivered. Direct imaging has found of order a dozen planets in twenty years, against thousands from transits and radial velocities, and it is nevertheless the only method that produces a spectrum of a planet’s atmosphere with no star in the beam. So its yield is small and its information per object is enormous, which is the opposite trade from the survey methods. What a survey could have seen is the right way to compare the two: the imaged planets are the young, wide, massive tail of the population, and the contrast curves in this essay are exactly the statement of which part of the parameter space that tail occupies.
It is also worth saying what makes the field’s practice unusually good, since this essay has spent most of its length on difficulties. High-contrast imaging adopted injection-recovery, small-sample statistics and published pipelines faster and more thoroughly than most observational subfields adopt anything, and it did so because its systematics were unmissable: a false detection of a planet is embarrassing in a way that a slightly wrong flux is not, and the retractions came quickly enough to force the correction. The discipline of characterising a pipeline by feeding it a known signal is the same one a well-run survey applies to its own completeness, and here it arrived under pressure rather than by design.
It is worth ending on the scale of the numbers. The signal being extracted is a source between a millionth and a billionth of the star it sits beside, recovered from a residual after a reference built from the same data has been subtracted. That any of it survives is remarkable, and the fact that the surviving fraction has to be measured rather than calculated is the honest price of the arrangement.
Where the ladder goes next
Later rungs on this ladder start with the alternative that does not remove the planet: reference-star differential imaging with a library of thousands of archival exposures, from which the best-matching reference is constructed without ever using the target’s own frames. The one above that is the interferometric route — nulling the star’s light rather than subtracting it, which changes the inner working angle by an order of magnitude and brings a different set of systematics.
About the same objects
Not linked from either essay — found by the objects both name.
- The threshold that is not a threshold detection limit · injection recovery
What links here
Essays that link to this one from their own argument.
The objects this essay names
Each one links to every other essay that touches it.
Angular differential imagingContrast curveDetection limitInjection recoveryThe point-spread functionPrincipal component analysisSelf subtractionSmall sample statisticsSpeckle noiseThroughput