Starlight

The comparison stars are part of the measurement

Measuring a star against others on the same frame cancels everything the atmosphere and the instrument do to all of them at once — a two per cent change in transparency vanishes completely. What does not vanish is the comparison stars' own noise, which the target inherits, and the variability of any comparison that is not constant, which the target reports as its own.

Assumes Photon noise, Extinction and Variable stars.

A star’s brightness measured on its own is at the mercy of everything between it and the detector. Thin cloud passes. The airmass changes as the star rises or sets, and with it the fraction of light the atmosphere absorbs. The detector’s sensitivity drifts with temperature. Each of those multiplies the star’s measured flux by a factor that can easily be a per cent or more, and none of them has anything to do with the star.

All of them multiply every other star on the same frame by the same factor. So instead of measuring the target’s flux, measure the ratio of its flux to that of other stars observed simultaneously through the same air on the same detector. Every common factor divides out. This is differential photometry, and it is the reason the ground-based photometry of variable stars and transiting planets reaches precisions far below the stability of any atmosphere.

The counting error of a single measurement and the best way to extract it from the pixels set how precisely one star’s flux can be measured on one frame. The ratio of two such measurements carries both errors, and that is the whole cost of the method: what cancels is everything common, and what remains is everything individual — including the comparison stars’ own individual noise.

Ten comparison stars as bright as the target cost 5 per cent in precision; ten 2 magnitudes fainter cost 23. The precision of a V = 12 target measured relative to an ensemble of comparison stars on the same 60-second frames, against the number of comparison stars, on a logarithmic precision axis, for comparisons 1 mag brighter, as bright as the target, 1 mag fainter, 2 mag fainter. A change in the atmosphere's transparency of 2.0 per cent — which would put the target's raw brightness out by 20.0 mmag — multiplies every star by the same factor and vanishes from the ratio. What is left is the target's own noise, 0.89 mmag, plus the ensemble's, which falls as the inverse square root of the number of stars in it. With comparisons as bright as the target the result is σ√(1 + 1/N): one comparison costs 41 per cent, ten cost 5. Fainter comparisons are noisier and need many more to reach the same point; brighter ones help, but the target's own noise is a floor the ensemble can only approach. Scintillation is treated as independent from star to star, which is right for stars more than a few arcseconds apart on a large telescope and makes it part of the noise that does not cancel. The figure also cannot show the defining weakness: every comparison star is assumed constant, and a variable among them injects its variability into every measurement made against the ensemble.
Fig. 1 The precision of a V = 12 target measured against an ensemble of comparison stars on the same 60-second frames of a 1 m telescope, against the number of comparisons, for comparisons 1 magnitude brighter, as bright, 1 fainter and 2 fainter. A transparency change of 2 per cent — 20.0 mmag in the target’s raw brightness — cancels completely. What remains is the target’s own noise, 0.89 mmag, plus the ensemble’s: with 30 comparisons, 0.90, 0.91, 0.93 and 0.97 mmag for the four cases. Ten comparisons as bright as the target cost 5 per cent in precision; ten 2 magnitudes fainter cost 28.

The arithmetic of a ratio

Let the target’s measured flux be TT and each comparison’s CiC_i. Any factor kk that multiplies all of them — transparency, extinction, gain — cancels exactly from T/CˉT/\bar C, where Cˉ\bar C is a weighted combination of the comparisons. The fractional errors that do not cancel add in quadrature:

σdiff2=σT2+σens2,1σens2=i1σi2\sigma_{\rm diff}^2 = \sigma_T^2 + \sigma_{\rm ens}^2, \qquad \frac{1}{\sigma_{\rm ens}^2} = \sum_i \frac{1}{\sigma_i^2}

with the comparisons combined using inverse-variance weights. For N comparisons each as noisy as the target, σens2=σT2/N\sigma_{\rm ens}^2 = \sigma_T^2/N, and the differential precision is

σdiff=σT1+1/N\sigma_{\rm diff} = \sigma_T\sqrt{1 + 1/N}

One comparison star makes the measurement 41 per cent noisier than the target alone. Four make it 12 per cent noisier, ten 5 per cent, a hundred half a per cent. The target’s own noise is a floor the ensemble approaches and cannot pass. A differential measurement is never better than the target’s own photon statistics, only better than its atmosphere.

The quadratic addition also says which comparisons are worth having. A comparison star twice as noisy as the target contributes a quarter as much to 1/σens21/\sigma_{\rm ens}^2 as one equally noisy, so it takes four of them to do the work of one. For a source-limited target, a comparison two magnitudes fainter is two and a half times noisier and needs six times as many copies; brighter comparisons help more, but only up to the point where the target’s own noise dominates, which a handful of bright comparisons reaches.

When the target is faint

Ten comparison stars as bright as the target cost 5 per cent in precision; ten 2 magnitudes fainter cost 60. The precision of a V = 17 target measured relative to an ensemble of comparison stars on the same 60-second frames, against the number of comparison stars, on a logarithmic precision axis, for comparisons as bright as the target, 1 mag fainter, 2 mag fainter. A change in the atmosphere's transparency of 2.0 per cent — which would put the target's raw brightness out by 22.0 mmag — multiplies every star by the same factor and vanishes from the ratio. What is left is the target's own noise, 9.11 mmag, plus the ensemble's, which falls as the inverse square root of the number of stars in it. With comparisons as bright as the target the result is σ√(1 + 1/N): one comparison costs 41 per cent, ten cost 5. Fainter comparisons are noisier and need many more to reach the same point; brighter ones help, but the target's own noise is a floor the ensemble can only approach. Scintillation is treated as independent from star to star, which is right for stars more than a few arcseconds apart on a large telescope and makes it part of the noise that does not cancel. The figure also cannot show the defining weakness: every comparison star is assumed constant, and a variable among them injects its variability into every measurement made against the ensemble.
Fig. 2 The same calculation for a V = 17 target, whose own noise on these frames is 9.11 mmag. A transparency change of 2 per cent, 22.0 mmag, still cancels. Ten comparisons as bright cost 5 per cent, ten a magnitude fainter cost 16 per cent, and fainter comparisons cost more than they did for the bright target, because at V = 17 and fainter the noise rises steeply with magnitude as the sky takes over.

For a faint target the comparison stars’ quality matters more, not less. The noise of a sky-limited star rises as 100.4m10^{0.4m}, twice as steeply as a source-limited star’s, so a comparison one magnitude fainter than a V = 17 target is much noisier, relative to it, than a comparison one magnitude fainter than a V = 12 target. The curves separate further. Picking comparisons for a faint target is a search for stars at least as bright as the target in a field where bright stars are rare, and in a small field of view the best available ensemble can be a handful of stars fainter than the target — the regime where the 1+1/N\sqrt{1 + 1/N} penalty is large.

This is why wide-field cameras suit differential photometry of faint stars so well: a larger field holds more bright stars, and the ensemble’s variance falls with every one. It is also why a target that is the brightest star in its field is the hardest to measure precisely from the ground, however many photons it delivers: every available comparison is noisier than it, and the ensemble’s contribution can exceed the target’s own.

The scintillation question

One kind of atmospheric noise is not always common to every star, and whether it cancels decides the precision for bright targets.

Scintillation — the twinkling produced by turbulence high in the atmosphere, focusing and defocusing starlight on scales of centimetres — changes a star’s measured flux by a fraction that is independent of how bright the star is. For a bright star in a short exposure it is often the largest term in the error budget. Two stars close enough together on the sky that their light passes through the same turbulent cells scintillate together, and in the ratio their scintillation cancels. Two stars further apart see different turbulence and scintillate independently, and the ratio carries both.

The angular scale that separates the two cases is set by the height of the turbulence and the size of the telescope: light from two stars at an angle θ passes through a turbulent layer at height hh displaced by hθh\theta, and the scintillation decorrelates once that displacement exceeds the telescope’s aperture. For a metre-class telescope and turbulence ten kilometres up, that is about twenty arcseconds. Comparison stars are almost always further away than that.

Ten comparison stars as bright as the target cost 5 per cent in precision; ten 2 magnitudes fainter cost 28. The precision of a V = 12 target measured relative to an ensemble of comparison stars on the same 60-second frames, against the number of comparison stars, on a logarithmic precision axis, for comparisons 1 mag brighter, as bright as the target, 1 mag fainter, 2 mag fainter. A change in the atmosphere's transparency of 2.0 per cent — which would put the target's raw brightness out by 20.0 mmag — multiplies every star by the same factor and vanishes from the ratio. What is left is the target's own noise, 0.78 mmag, plus the ensemble's, which falls as the inverse square root of the number of stars in it. With comparisons as bright as the target the result is σ√(1 + 1/N): one comparison costs 41 per cent, ten cost 5. Fainter comparisons are noisier and need many more to reach the same point; brighter ones help, but the target's own noise is a floor the ensemble can only approach. Scintillation is left out, which is right for stars close enough to share the same turbulence and wrong otherwise. The figure also cannot show the defining weakness: every comparison star is assumed constant, and a variable among them injects its variability into every measurement made against the ensemble.
Fig. 3 The V = 12 target with scintillation left out, as it would be if every comparison shared the target’s turbulence. The target’s own noise falls from 0.89 to 0.78 mmag, and the gains from more comparisons are unchanged in proportion: ten as bright still cost 5 per cent, ten 2 magnitudes fainter still cost 28. The eleven-hundredths of a millimagnitude between this figure and the first is the scintillation that a real ensemble, spread over arcminutes, cannot cancel.

The difference looks small because the target is faint enough for photons to matter. Scintillation’s share of the budget is a fixed fraction of the flux, while photon noise falls as the square root of the flux, so the brighter the target the larger the share the atmosphere takes. Four magnitudes brighter, a star delivers forty times as many photons, its own shot noise shrinks by a factor of six, and the term that does not shrink is the one the ensemble cannot remove.

Ten comparison stars as bright as the target cost 5 per cent in precision; ten 2 magnitudes fainter cost 28. The precision of a V = 8 target measured relative to an ensemble of comparison stars on the same 60-second frames, against the number of comparison stars, on a logarithmic precision axis, for comparisons as bright as the target, 2 mag fainter. A change in the atmosphere's transparency of 2.0 per cent — which would put the target's raw brightness out by 20.0 mmag — multiplies every star by the same factor and vanishes from the ratio. What is left is the target's own noise, 0.12 mmag, plus the ensemble's, which falls as the inverse square root of the number of stars in it. With comparisons as bright as the target the result is σ√(1 + 1/N): one comparison costs 41 per cent, ten cost 5. Fainter comparisons are noisier and need many more to reach the same point; brighter ones help, but the target's own noise is a floor the ensemble can only approach. Scintillation is left out, which is right for stars close enough to share the same turbulence and wrong otherwise. The figure also cannot show the defining weakness: every comparison star is assumed constant, and a variable among them injects its variability into every measurement made against the ensemble.
Fig. 4 A bright V = 8 target, with scintillation assumed common. Its own photon noise on a 60-second frame is only 0.12 mmag, and the transparency change of 20.0 mmag cancels. Ten comparisons as bright cost 5 per cent and ten 2 magnitudes fainter cost 28. For a star this bright, independent scintillation would dominate everything else in the budget, and the figure shows the precision only a very small field, or a telescope large enough to average over many turbulent cells, could deliver.

For bright targets, then, differential photometry from the ground removes transparency and leaves scintillation. The remedies are to use a larger telescope — scintillation falls as the aperture to the minus two-thirds power, because a larger mirror averages over more turbulent cells — to take longer exposures, or to leave the atmosphere entirely. The last is why the precision of space photometry for bright stars is not limited by the same term at all.

One comparison and one check

The ensemble’s arithmetic explains a habit that visual and photoelectric observers of variable stars kept for a century before anybody wrote the formula down. An observer measured the variable against one comparison star and, separately, measured a second star — the check — against the same comparison. The variable-minus-comparison differences were the light curve. The check-minus-comparison differences were supposed to be constant, and their scatter was the measurement’s precision.

That is the smallest ensemble there is, with its own verification built in. One comparison costs a factor of 2\sqrt 2 in noise when it is as bright as the target, which observers accepted because photomultipliers measured one star at a time and a second comparison meant a second pointing. The check star answered the question the arithmetic cannot: whether the comparison was constant. A comparison that varied showed up in the check’s differences as surely as in the variable’s, and a light curve whose check star was flat could be trusted at least against that failure.

The light curves of Cepheids that first calibrated the distance scale were measured this way, as were most eclipsing binaries’ and most pulsating stars’. Charge-coupled detectors put hundreds of stars on every frame and made the check star free: every comparison can check every other, which is the procedure described below, and the old practice of choosing one comparison carefully became a practice of choosing many and testing them all.

A comparison that is not constant

The whole method rests on an assumption the arithmetic does not check: every comparison star is constant. A comparison that varies moves the ensemble, and the target’s ratio to the ensemble moves the opposite way.

One comparison varying by 10 mmag puts 0.3 mmag into the target. The false signal that appears in a target's differential light curve when one of 30 equally weighted comparison stars is itself variable, with an amplitude of 10 millimagnitudes and a period of 3 hours, over an eight-hour night. The ensemble mean moves by the variable's amplitude divided by the number of stars, ±0.33 mmag, and the target appears to move by the same amount in the opposite sense. The shaded band is the target's precision after binning its 60-second measurements into quarter-periods, ±0.14 mmag: the injected signal stands above it and would be detected as a real variation of the target. A large ensemble dilutes one bad comparison without removing it, and the dilution is exactly what makes it hard to find. Surveys that reach part-per-million precision therefore check every comparison against every other — each star measured against the ensemble of the rest — and discard the ones that fail, which is a test the ensemble can pass only if most of its members are constant.
Fig. 5 One of thirty equally weighted comparison stars varies by 10 mmag with a 3-hour period. The ensemble mean moves by the amplitude divided by the number of stars, ±0.33 mmag, and the target appears to vary by the same amount in the opposite sense. The shaded band is the target’s precision after binning its 60-second points into quarter-periods, ±0.14 mmag: the injected signal stands above it and would be read as a real variation of the target.

That is a small amplitude, and it is a dangerous one. Variability at the level of a hundredth of a magnitude is common among field stars — pulsations, spots, eclipsing companions, rotation — and a random choice of thirty comparisons has a good chance of including one. Diluted by thirty, it produces a third of a millimagnitude in the target, with a period that is the comparison’s, which is exactly the size of signal a careful ground-based observation of a transiting planet or a low-amplitude variable is trying to measure.

One comparison varying by 10 mmag puts 2.0 mmag into the target. The false signal that appears in a target's differential light curve when one of five equally weighted comparison stars is itself variable, with an amplitude of 10 millimagnitudes and a period of 3 hours, over an eight-hour night. The ensemble mean moves by the variable's amplitude divided by the number of stars, ±2.00 mmag, and the target appears to move by the same amount in the opposite sense. The shaded band is the target's precision after binning its 60-second measurements into quarter-periods, ±0.15 mmag: the injected signal stands above it and would be detected as a real variation of the target. A large ensemble dilutes one bad comparison without removing it, and the dilution is exactly what makes it hard to find. Surveys that reach part-per-million precision therefore check every comparison against every other — each star measured against the ensemble of the rest — and discard the ones that fail, which is a test the ensemble can pass only if most of its members are constant.
Fig. 6 The same variable comparison in an ensemble of five. The false signal in the target grows to ±2.00 mmag, six times larger, while the target’s binned precision is ±0.15 mmag. A small ensemble does not hide a bad comparison; it announces it. The dilution that makes a large ensemble statistically better is the same dilution that makes a bad member hard to see.

The dilution cuts both ways, and that is the instructive part. A small ensemble with one variable member produces a large false signal, obvious enough to notice. A large ensemble with the same member produces a small one, below any threshold for suspicion and above the precision the large ensemble achieves. Increasing the ensemble reduces the random error as 1/N1/\sqrt{N} and reduces a single bad member’s effect only as 1/N1/N — but it also lowers the noise against which that effect would be noticed, so the ratio of false signal to precision falls only as 1/N1/\sqrt{N} too. The bad comparison is never diluted into irrelevance; it is diluted into invisibility.

How ensembles are checked

The remedy is to treat every comparison as a target. Measure each comparison against the ensemble of all the others, look at the scatter of each resulting light curve, and discard any star whose scatter is larger than its photon noise predicts or which shows a period. Repeat until every remaining star is consistent with being constant. The procedure converges only if most of the stars are constant to begin with, which is usually true and is an assumption rather than a result.

Large surveys go further. With thousands of stars on each frame, the common variations — transparency, focus, the telescope’s pointing drifts — can be estimated from the whole field as a set of trends, and each star’s light curve is corrected by the combination of trends that best describes it. That is differential photometry generalised: instead of dividing by an average of comparison stars, each star is fitted with a model built from all of them. It works because the systematic effects are shared and the stars’ own variations are not — and it has a failure mode of exactly the kind drawn above. A real signal that happens to resemble one of the trends is partly fitted away, and the depth of a transit dimmed by the removal of its own trend comes out shallower than it is.

Noise that does not average down

Every figure here gives a precision for a single frame. A light curve has hundreds of frames, and binning them should improve the precision as the square root of their number — for white noise. Differential photometry from the ground rarely delivers white noise for long.

The residuals after an ensemble correction contain what does not cancel perfectly: slow drifts from colour-dependent extinction, small changes in the image’s position on pixels of slightly different sensitivity, focus changes that alter how much light a fixed aperture catches, and scintillation’s own low-frequency tail. These are correlated from one frame to the next. Binned into groups of a few minutes, the scatter falls as 1/n1/\sqrt n at first and then flattens, and the bin size at which it flattens is set by the timescale of the correlated drifts rather than by any photon count.

That flattening is the same systematic floor that stops a single long exposure, appearing in a time series instead of in an integration. It matters most for signals whose own timescale is comparable with the drifts: a transit lasting a few hours sits exactly in the range where correlated noise of a few tenths of a millimagnitude can mimic or mask it. Analyses of ground-based transits therefore measure the noise at the transit’s own timescale, by comparing the scatter of binned residuals with the scatter white noise would predict, and inflate their uncertainties by the ratio. An uncertainty computed from single-frame scatter alone can be too small by a factor of two or three, and a transit depth quoted with it looks more precise than it is.

From space the atmosphere is gone and the correlated noise is not. The telescope’s pointing jitters, its focus breathes with the spacecraft’s temperature, and every star on the detector shares those effects in slightly different proportions. The correction used by space photometry missions is the ensemble idea at scale: a set of common trends extracted from thousands of quiet stars, fitted to each star’s light curve, and subtracted. Its failure mode is also the ensemble’s — a real astrophysical variation that resembles a trend is partly fitted away — and it is one reason the planets that were not seen include some whose transits were removed along with the instrument’s drifts.

Colour and the atmosphere’s selective absorption

Transparency cancels exactly only if the atmosphere absorbs every star’s light by the same fraction, and it does not quite. Extinction depends on wavelength — the atmosphere dims blue light more than red through scattering — so a blue comparison and a red target are dimmed by slightly different factors as the airmass changes through a night. The ratio between them drifts with airmass, a second-order extinction term proportional to the difference in their colours.

For a broad filter and a colour difference of half a magnitude between target and comparisons, the drift over a night’s range of airmass can be a millimagnitude or more — larger than the random precision of a bright target. The remedies are comparisons of similar colour to the target, narrower filters, and a colour term fitted as part of the ensemble correction. None of them is perfect, and a smooth drift with airmass is the characteristic residual of differential photometry done carefully everywhere else: it looks like a trend in the target’s light curve over the night, and a slow variation of the target can hide inside it.

What the figures leave out

The noise model inherits everything the single-star budget assumed: a Gaussian image, a flat and known background, a fixed read noise. It treats each comparison’s errors as independent, which is true of photon and sky noise and false of anything that affects neighbouring pixels together, such as a detector artefact. It assumes the comparison stars are weighted optimally by their variances, which requires knowing those variances, and a comparison whose noise is underestimated gets too much weight. And it has no flat-field error, the systematic floor that stops every ground-based photometric measurement somewhere near a few tenths of a millimagnitude however many comparisons are used, because the target and the comparisons sit on different pixels whose sensitivities are known only to a precision of their own.

A reference cancels what it shares and passes on what it does not

A measurement relative to a reference is exactly as good as the reference, in both senses. It is freed from every error the reference shares, which is its power; and it inherits every error the reference has that the target does not share, which is its cost. An ensemble lowers the random part of that cost as fast as its members are added, and lowers a single systematic part — one bad member — only as fast as that member is diluted, so that the ensemble’s growing precision is what makes the systematic hard to see.

The structure recurs wherever a signal is measured against a comparison: a radial velocity against a reference spectrum, a timing against a clock ensemble, an interferometric phase against a calibrator. A phase that survives what corrupts it is the cleanest version, where the combination cancels the corruption exactly and inherits nothing; differential photometry is the ordinary version, where the cancellation is exact for the common factor and the price is paid in the comparison’s own noise.

Still open: when the floor is the star itself

With enough comparisons and a large enough telescope, or above the atmosphere, the target’s own photon noise is the last random term. For stars observed from space to a few parts per million, even that is not the floor: the star’s surface is covered with convective granulation whose brightness fluctuations, averaged over the disc, amount to tens of parts per million on timescales of hours. That noise is not instrumental and does not cancel against anything. How large it is, how it scales with the star’s surface gravity and temperature, and whether it can be modelled out of a light curve rather than merely averaged over, is what decides whether an Earth-sized planet crossing a Sun-like star can be detected in a single transit.

About the same objects

Not linked from either essay — found by the objects both name.

The objects this essay names

Each one links to every other essay that touches it.

Atmospheric extinctionCommon mode noiseComparison starDifferential photometryInverse variance weightingLight curvePhoton noiseScintillationSystematic errorTransit photometryVariable stars