Exoplanets

A star subtracted using the star

Imaging a planet means removing a halo of scattered starlight a hundred million times brighter than the planet, and no model of that halo is good enough to subtract. So it is built from the star's own exposures — and since the planet is in those exposures too, it subtracts part of itself.

Assumes Direct imaging and Seeing.

The difficulty of imaging a planet is not the angular separation. A young Jupiter half an arcsecond from its star is comfortably resolved by an eight-metre telescope. The difficulty is the contrast: the planet is between a millionth and a billionth of the star’s brightness, and what stands between them is not the star’s own image but the halo of scattered light around it.

That halo is not smooth. It is a field of speckles — the residual wavefront errors after adaptive-optics correction, imaged as a swarm of diffraction-limited spots at the same brightness a planet would have. They come and go on timescales from milliseconds to hours, and the slowly varying ones are indistinguishable from a companion.

A planet that subtracts 73 per cent of itself at the inner working angle. The fraction of a planet's flux that survives an angular differential imaging subtraction, against its separation from the star in resolution elements, for a sequence covering 25 degrees of field rotation. The reference image is built from the target's own frames, so a planet that has not moved far between them is present in the reference and is removed along with the speckles. How far it moves is the arc length, which is proportional to the separation — so the self-subtraction is severe close in and negligible far out, and the half-throughput point is at 2.1 resolution elements here. The consequence for any contrast curve is that it is a statement about an algorithm as well as about an instrument: the depth reached has to be measured by injecting fake planets into the data and recovering them, because no calculation predicts what fraction of a real one survives.
Fig. 1 The technique that removes them, and its cost. With the telescope’s pupil held fixed the speckles stay put while the sky rotates, so a median over a long sequence is an image of the speckles with any companion averaged away. Subtracting it leaves the planet — partly. A planet close to the star moves only a fraction of a resolution element during the sequence, so it is present in the reference and removes part of itself.

Nine orders of magnitude, half an arcsecond apart is the statement of the problem. This essay is about the only technique that has solved any part of it, and about the price that technique charges.

Why the reference has to come from the data

The natural approach is to observe a reference star, one with no companion, and subtract its image. That fails for a specific and instructive reason.

The speckle pattern is a snapshot of the instrument’s wavefront error, which depends on the telescope’s flexure with elevation, on the temperature of the optics, on the adaptive-optics system’s state and on the atmosphere. Two stars observed an hour apart at different elevations have different speckle patterns, and the difference is larger than the planet.

So the reference has to be built from images taken under conditions as nearly identical as possible, which means from the target’s own sequence. The only thing that must differ between the images being combined is the position of the sky — which is what pupil-stabilised observing provides.

That is a clean trick and it has an unavoidable consequence: any real companion is in every frame used to build the reference.

There is a second variant of the same idea that exploits a different axis, and it is worth naming because the two are often combined. Spectral differential imaging uses the fact that a speckle’s position scales with wavelength while a planet’s does not: observing at several wavelengths simultaneously and rescaling each image moves the speckles and holds the planet fixed, so a median over wavelength is again a speckle reference. It has the same structure and the same defect — a planet whose spectrum is featureless is present in every rescaled frame and subtracts itself — and the two techniques are used together because their self-subtraction is worst in different places.

What the self-subtraction costs

How much of the planet survives depends on how far it moves between the frames that are combined. The movement is an arc: the field rotation angle times the separation. Far from the star the arc is many resolution elements and the planet is effectively absent from the reference; close in the arc is a fraction of one and the planet is fully present.

The result is a throughput that rises with separation, and the transition is at the separation where the arc equals about one resolution element.

A planet that subtracts 29 per cent of itself at the inner working angle. The fraction of a planet's flux that survives an angular differential imaging subtraction, against its separation from the star in resolution elements, for a sequence covering 60 degrees of field rotation. The reference image is built from the target's own frames, so a planet that has not moved far between them is present in the reference and is removed along with the speckles. How far it moves is the arc length, which is proportional to the separation — so the self-subtraction is severe close in and negligible far out, and the half-throughput point is at 1.2 resolution elements here. The consequence for any contrast curve is that it is a statement about an algorithm as well as about an instrument: the depth reached has to be measured by injecting fake planets into the data and recovering them, because no calculation predicts what fraction of a real one survives.
Fig. 2 The same relation for a longer sequence covering sixty degrees of rotation rather than twenty-five. The half-throughput point moves inward in exact proportion, because the arc is the product of the rotation and the separation, so doubling the rotation halves the separation at which a given fraction survives. That is the practical argument for observing through transit, when the parallactic angle rotates fastest — the observing strategy is chosen to buy inner working angle rather than signal-to-noise.

The consequence is that a contrast curve — the standard product of a high-contrast observation, giving the faintest companion detectable against separation — is not a property of the telescope. It is a property of the telescope, the algorithm, the sequence’s field rotation, and how the throughput was estimated. Two groups analysing the same data with different algorithms produce different curves.

A planet that subtracts 84 per cent of itself at the inner working angle. The fraction of a planet's flux that survives an angular differential imaging subtraction, against its separation from the star in resolution elements, for a sequence covering 12 degrees of field rotation. The reference image is built from the target's own frames, so a planet that has not moved far between them is present in the reference and is removed along with the speckles. How far it moves is the arc length, which is proportional to the separation — so the self-subtraction is severe close in and negligible far out, and the half-throughput point is at 4.3 resolution elements here. The consequence for any contrast curve is that it is a statement about an algorithm as well as about an instrument: the depth reached has to be measured by injecting fake planets into the data and recovering them, because no calculation predicts what fraction of a real one survives.
Fig. 3 A short sequence with only twelve degrees of rotation, which is what a target observed away from transit or for a limited window gives. The half-throughput separation has moved outward to several resolution elements, and inside that the planet is largely gone. Since the interesting separations for a young giant planet are a few resolution elements, an observation with insufficient rotation is not a shallower version of a good one — it is blind in exactly the region it was aimed at.

The throughput has to be injected

Since no calculation predicts the surviving fraction for a real companion, it is measured: fake planets of known brightness are injected into the raw data at a grid of separations and position angles, the whole reduction is run, and the recovered brightness is compared against the injected one.

That is expensive — the reduction is run hundreds of times — and it is the only defensible way to produce a contrast curve. It also has to be done with the fake planets far enough from any real one that the two do not interfere, and at enough position angles that the azimuthal variation is sampled.

The injection-recovery procedure has a second use that is at least as important. It measures not only the throughput but the noise after the reduction, which is not the noise before it: an algorithm that subtracts aggressively reduces the speckles and the planet together, and the ratio of the two is what a detection limit is made of.

A planet that subtracts 86 per cent of itself at the inner working angle. The fraction of a planet's flux that survives an angular differential imaging subtraction, against its separation from the star in resolution elements, for a sequence covering 25 degrees of field rotation. The reference image is built from the target's own frames, so a planet that has not moved far between them is present in the reference and is removed along with the speckles. How far it moves is the arc length, which is proportional to the separation — so the self-subtraction is severe close in and negligible far out, and the half-throughput point is at 2.3 resolution elements here. The consequence for any contrast curve is that it is a statement about an algorithm as well as about an instrument: the depth reached has to be measured by injecting fake planets into the data and recovering them, because no calculation predicts what fraction of a real one survives.
Fig. 4 A more aggressive algorithm: one that builds its reference from a larger number of principal components and therefore removes more of the speckle field and more of the planet. The floor at small separations is now two per cent rather than twelve, which is to say that at the inner working angle the planet is essentially gone. Whether that is worth it depends on how much of the speckle noise went with it, and the trade is optimised empirically for every dataset — there is no theory that says how many components to use.

There is a subtlety in the injection that took the field some time to settle. A fake planet injected at the same separation as a real one biases the reduction, because the reference is built from frames containing both; and a fake planet injected into data that already contain a real one measures the throughput in the presence of that planet, which is not the same as in its absence. The standard practice is to inject at several position angles well away from any candidate, to run the whole reduction each time, and to quote the mean and scatter of the recovered fractions — with the scatter itself being informative, since a large azimuthal variation means the speckle field is structured rather than random.

One more property of the injection is worth stating because it decides how the results are used. The recovered throughput depends on the injected brightness: a bright fake planet dominates the frames it appears in, biases the reference towards itself, and is subtracted harder than a faint one. So the throughput is not a single number but a function of contrast, and a curve computed by injecting bright planets overestimates the sensitivity to faint ones. The standard is to inject at a brightness near the expected detection limit and to iterate, which converges in a couple of passes and which is a further reason the reduction is run hundreds of times. The same non-linearity appears in any pipeline that estimates its own background, and the same remedy applies.

The algorithm is a projection, and that is exploitable

The reduction described so far is a median: build a reference frame from the sequence, subtract it, derotate, combine. Almost nobody does that any more, and the algorithms that replaced it make the self-subtraction both worse and calculable.

The modern reduction expresses the reference as a fitted combination of the other frames rather than as their median. One family finds, for each small region of each frame, the linear combination of other frames that minimises the residual there; another decomposes the sequence into principal components and projects each frame onto the first several of them. Both remove far more of the speckle field than a median does, because both are allowed to fit it.

Both also fit the planet, and for the same reason. A linear combination chosen to minimise the residual in a region containing a planet will use whatever freedom it has to reduce the planet’s own flux, and the more components it is given the more of the planet it removes. That is why the number of components retained is the field’s central tuning parameter: it trades speckle suppression against throughput along a curve with an optimum that depends on the dataset, the separation and the rotation, and there is no theory that locates it.

What the fitting buys back is that the operation is linear. A projection onto a fixed set of components is a matrix, so the effect of the reduction on any known signal can be computed rather than measured: take the planet’s point-spread function, put it through the same projection with the same components, and see what emerges. That is forward modelling, and it gives the throughput analytically for the price of one matrix multiplication instead of hundreds of reductions.

The catch is that the components themselves depend on the data, and the data contain the planet — so the projection is only linear once the components are fixed, and the components were computed from frames the planet was in. In practice the effect is second order for a faint companion and is not for a bright one, which is why forward modelling is used to measure a candidate’s photometry and astrometry, and injection-recovery is still used to establish the detection limit. The two methods answer different questions: one asks what happened to a source that is there, the other asks how faint a source could have been and gone unnoticed.

There is a practical consequence for anybody comparing published results. A photometric measurement of an imaged planet is algorithm-dependent at the tens-of-per-cent level unless a forward model or an injection was used to correct it, and the older literature frequently reports the raw recovered flux. Several early estimates of imaged planets’ luminosities — and therefore of their inferred masses, which depend steeply on luminosity — were low for exactly this reason.

Where the detection limit really sits

Two further effects push the achievable contrast away from what the photon statistics suggest, and both are worst where the planets are.

The speckle noise is correlated. Residual speckles are not independent from pixel to pixel; they have the size of a resolution element and they persist over minutes. So the noise in an aperture is not the square root of the counts, and estimating it requires measuring the scatter of the reduced image at the same separation rather than assuming Poisson statistics.

The number of independent samples is small. At a separation of two resolution elements there are only about a dozen independent resolution elements around the annulus. Estimating a noise level from twelve samples, and then applying a five-sigma threshold, is statistically quite different from doing it with thousands — the threshold has to be raised substantially to keep the false-alarm rate at the intended level, by an amount that grows sharply as the separation shrinks.

That second effect is the same look-elsewhere arithmetic that inflates the threshold in a period search, arriving in a spatial rather than a frequency domain, and it was not applied in the field’s first decade.

Contrast against separation, which is where direct imaging lives. Planet-to-star brightness ratio against apparent separation, both logarithmic, for a system 10 parsecs away. The reflected-light curves are A_g(R_p/a)² and fall as the inverse square of the orbit; the thermal curve is the ratio of two Planck functions at 10 µm and does not, which is why every imaged planet so far is young and hot rather than merely large. The vertical lines are diffraction limits λ/D — nothing inside a telescope's own line is reachable by it at any contrast at all. An Earth at ten parsecs sits at 6.3e-10, which is 4 orders of magnitude below the faintest planet yet imaged. The four imaged planets are plotted at their measured near-infrared contrasts rather than at 10 µm, because the near infrared is the band they were found in — which is itself part of the argument, since a young planet is hot enough to be bright where its star is not.
Fig. 5 The curve the whole apparatus produces: achievable contrast against separation, with the sensitivity of a given telescope to planets of a given mass and age. Everything in this essay lives in the shape of that curve at small separations — the rise towards the star is the self-subtraction and the small-sample penalty together, and the flat outer part is where the photon noise takes over and the algorithm has stopped mattering.

There is a third effect and it is the one that limits the very best observations. After all the speckles that field rotation can remove have been removed, what is left is the static part of the wavefront error — the difference between what the wavefront sensor measures and what the science camera sees, which does not rotate with the sky because it is in the instrument. Field rotation cannot touch it and no amount of observing time averages it down. Reducing it is a matter of measuring the science camera’s own wavefront, which requires an instrument that can do that, and it is the single largest term in the error budget of any high-contrast system.

What was actually measured

Three results establish both the technique’s power and its pitfalls.

The directly imaged planetary systems. A handful of young, wide, massive planets have been imaged and followed over orbital arcs of years. Every one of them was found beyond the separation at which self-subtraction is severe, which is a statement about the technique as much as about the planets.

The retracted candidates. Several claimed companions at small separations have not been confirmed, and the leading explanation in each case is a residual speckle combined with an underestimated threshold. The field’s response was to adopt the small-sample correction and to require detections in more than one epoch and more than one band.

And the injection-recovery standard. Contrast curves published before about 2014 and after are not comparable, because the earlier ones generally did not correct for throughput. The difference is a factor of two to five at small separations, always in the direction of the earlier curves being optimistic.

196 nanometres of wavefront, and a Strehl that depends entirely on the colour. Above, the error budget of an adaptive-optics system, term by term, in nanometres of residual wavefront. The terms are independent and add in quadrature, so the total is 196 nanometres and is dominated by the two largest; removing the smallest term entirely would improve it by under six per cent, which is why an optimisation programme that does not know which term is largest achieves nothing. Below, what that residual delivers: the Strehl ratio, the fraction of the light in the diffraction core, against wavelength. It is the exponential of minus the square of the residual measured in radians, and a radian is a wavelength, so the same physical error is four times smaller in phase at two microns than at half a micron. The system drawn here delivers 1 per cent of its light into the core at 550 nanometres and 73 per cent at 2200, and nothing about it changed between the two.
Fig. 6 Where the speckles come from in the first place: the residual wavefront after adaptive-optics correction, whose error budget decides how bright the halo is. The term that matters most for high contrast is the static one — the non-common-path error between the wavefront sensor and the science camera — because a static wavefront error produces a static speckle, and a static speckle is exactly what field rotation cannot remove and what looks most like a planet.

And a fourth, and it concerns what “detection” means in this field. A five-sigma point source in a reduced image is a candidate, not a planet: it has to be recovered at a second epoch, at a position consistent with orbital motion rather than with the background star it might be, and ideally in a second band. The common-proper-motion test — showing that the candidate moves with the star rather than staying fixed as the star’s own proper motion carries it across a static background — is what actually establishes companionship, and it takes a year or more. Several published candidates have failed it. That the confirmation is an astrometric measurement rather than a photometric one is worth noticing: the detection is a contrast problem and the confirmation is a position problem, and they have entirely different systematics.

Why an algorithm belongs in the error bar

The underlying point deserves stating on its own because it is unusual and becoming less so.

When the reference for a measurement is estimated from the data, the analysis is part of the instrument. A contrast curve depends on the number of principal components retained, on how the frames were selected, on the rotation available and on the injection procedure — and none of those is a property of the telescope. Reporting a detection limit without reporting the algorithm is like reporting a magnitude without reporting the filter.

The field has converged on the right practice, which is to publish the reduction alongside the result and to characterise the pipeline by injection rather than by calculation. That is now standard in high-contrast imaging and it is not standard everywhere else, though the same argument applies wherever a background is estimated from the data — a detrended transit light curve has exactly the same structure, and a signal removed by the detrending is exactly this self-subtraction under another name.

The deeper reason to take the parallel seriously is that both cases have the same remedy. Whatever the algorithm does to the data, do it to a known signal injected into the same data, and measure what comes out. That converts an unknowable systematic into a measured throughput, and it costs computer time rather than telescope time.

A final observation about what the technique has and has not delivered. Direct imaging has found of order a dozen planets in twenty years, against thousands from transits and radial velocities, and it is nevertheless the only method that produces a spectrum of a planet’s atmosphere with no star in the beam. So its yield is small and its information per object is enormous, which is the opposite trade from the survey methods. What a survey could have seen is the right way to compare the two: the imaged planets are the young, wide, massive tail of the population, and the contrast curves in this essay are exactly the statement of which part of the parameter space that tail occupies.

It is also worth saying what makes the field’s practice unusually good, since this essay has spent most of its length on difficulties. High-contrast imaging adopted injection-recovery, small-sample statistics and published pipelines faster and more thoroughly than most observational subfields adopt anything, and it did so because its systematics were unmissable: a false detection of a planet is embarrassing in a way that a slightly wrong flux is not, and the retractions came quickly enough to force the correction. The discipline of characterising a pipeline by feeding it a known signal is the same one a well-run survey applies to its own completeness, and here it arrived under pressure rather than by design.

It is worth ending on the scale of the numbers. The signal being extracted is a source between a millionth and a billionth of the star it sits beside, recovered from a residual after a reference built from the same data has been subtracted. That any of it survives is remarkable, and the fact that the surviving fraction has to be measured rather than calculated is the honest price of the arrangement.

Where the ladder goes next

Later rungs on this ladder start with the alternative that does not remove the planet: reference-star differential imaging with a library of thousands of archival exposures, from which the best-matching reference is constructed without ever using the target’s own frames. The one above that is the interferometric route — nulling the star’s light rather than subtracting it, which changes the inner working angle by an order of magnitude and brings a different set of systematics.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Angular differential imagingContrast curveDetection limitInjection recoveryThe point-spread functionPrincipal component analysisSelf subtractionSmall sample statisticsSpeckle noiseThroughput