Starlight

The error bar that comes from counting

A brightness is a number of photons, so the precision of the measurement is fixed before any instrument is chosen. What follows is a slope of exactly 0.2 magnitudes of error per magnitude of star, a slope of 0.4 once the sky wins, and a floor that neither of them explains.

Assumes Magnitudes and Photometric systems.

Every quantity in astronomy that is not an angle or a time is, at bottom, a count of photons. That is a stronger statement than it sounds. It means that the precision of a brightness measurement is decided before the telescope is designed, before the detector is chosen and before the night begins: a source that delivers NN photons can be measured to N\sqrt N and no better, however clever the passband it is measured through, and everything an instrument builder can do is arrange to lose as few of them as possible.

The consequences are a set of straight lines with exact slopes, a crossing point that moves with the phase of the Moon, and a floor that has nothing to do with photons at all.

What the error bar is made of, on a 1 m in 60 s. The four contributions to a photometric error, against the brightness of the star, for a 1-metre aperture, a 60-second exposure and a sky of 21 magnitudes per square arcsecond. The star's own photons give a line of slope exactly 0.2 — σ ∝ N^−1/2 and N ∝ 10^−0.4m, so a magnitude of extra faintness costs a fifth of a magnitude of precision, and no instrument changes that. The sky and the read noise are fixed counts, so their lines have slope 0.4, twice as steep, and they overtake the star at V = 18.25 — that crossing is the faint limit of the night, and it moves when the Moon rises rather than when the telescope changes. Scintillation is flat, because the atmosphere modulates a bright star and a faint one by the same fraction: at 4.09e-4 relative it is 0.44 millimagnitudes here and it is what caps the bright end, up to about V = 10.8. Below all of them is the systematic floor at 0.3 millimagnitudes, which is flat-fielding and colour terms and does not integrate down at all.
Fig. 1 The four contributions to a photometric error, against the brightness of the star, for a one-metre aperture, a sixty-second exposure and a sky of 21 magnitudes per square arcsecond. The star’s own photons give a line of slope exactly 0.2σN1/2\sigma \propto N^{-1/2} and N100.4mN \propto 10^{-0.4m} — so a magnitude of extra faintness costs a fifth of a magnitude of precision, and no instrument changes that. The sky and the read noise are fixed counts, so their lines have slope 0.4, twice as steep, and they overtake the star at V=18.3V = 18.3. Scintillation is flat, because the atmosphere modulates a bright star and a faint one by the same fraction.

Where the two slopes come from

The variance in an aperture photometry measurement is a sum, and the sum is short:

σN2=N+Nsky+npix(R2+Dt),\sigma_N^2 = N_\star + N_{\rm sky} + n_{\rm pix}\left(R^2 + D t\right),

with NN_\star the star’s detected photons, NskyN_{\rm sky} the sky’s over the same aperture, RR the read noise per pixel and DD the dark current. The first two terms are Poisson — a count’s variance is the count — and the third is the detector’s own contribution, independent of anything on the sky.

A magnitude error is that divided by the signal and scaled: σm=1.0857σN/N\sigma_m = 1.0857\,\sigma_N/N_\star, where 1.0857=2.5/ln101.0857 = 2.5/\ln 10 is the number of magnitudes in an ee-fold. Then:

  • while the star dominates, σN=N\sigma_N = \sqrt{N_\star}, so σmN1/2100.2m\sigma_m \propto N_\star^{-1/2} \propto 10^{0.2m}: slope 0.2;
  • once a fixed background dominates, σN\sigma_N is constant, so σmN1100.4m\sigma_m \propto N_\star^{-1} \propto 10^{0.4m}: slope 0.4.

Neither exponent is fitted or approximate. They come from Pogson’s definition of the magnitude scale meeting the fact that a count’s variance is the count, and any measured photometric error curve that does not show them is showing something else.

The faint limit belongs to the night

The crossing point in the opening figure — where the sky’s noise overtakes the star’s — is what an observer means by the limiting magnitude, and it is worth being clear about what it depends on.

It moves with the sky brightness, which varies by three magnitudes per square arcsecond between a dark site at new moon and the same site at full moon. Three magnitudes of sky is a factor of four in the noise from it, which moves the crossing by about 1.5 magnitudes.

It moves with the aperture used for the photometry, since the sky contribution scales with the area of the aperture and the star’s does not. Halving the aperture radius quarters the sky counts and loses perhaps 20% of the star’s, which is why optimal photometry weights the pixels by the point-spread function rather than summing them equally.

And it moves with the seeing, for the same reason — a night with 1″ images allows a smaller aperture than one with 3″ images, and therefore a fainter limit, on the same telescope with the same sky.

What it does not do is move much with the telescope. A larger aperture collects more from the star and more from the sky in the same proportion, so in the sky-limited regime the gain is only D\sqrt{D} rather than DD. A telescope four times the area reaches 0.75 magnitudes fainter, not 1.5.

That last point is the one most often got wrong, and it has a corollary worth stating. Because the sky-limited gain goes as D\sqrt{D} while the cost of a telescope goes as something between D2D^2 and D2.7D^{2.7}, the marginal magnitude bought by aperture is the most expensive thing in observational astronomy — and the marginal magnitude bought by a darker sky, a better detector or a sharper night is nearly free by comparison. The history of the subject since 1980 is largely a history of taking the cheap magnitudes first.

Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 1.49 times that floor and after 10000 it is 1.06, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 2 The same accounting where the systematic floor has been pushed two orders of magnitude down. Photon noise falls as the square root of the count, so a large aperture and a long exposure can in principle reach any precision — and in practice every instrument has a floor below which correlated errors, not counting statistics, set the answer. Where that floor sits decides what science is possible: at three parts in ten thousand a hot Jupiter’s atmosphere is measurable, and at three parts in a hundred thousand an Earth’s is.

The floor that is not photons

At the bright end all the counting terms become negligible, and the honest expectation is that a bright star can be measured arbitrarily well by exposing longer. It cannot, and the reason is the atmosphere.

Scintillation is the twinkling: refractive-index variations high in the atmosphere focus and defocus the beam, so the flux arriving at the aperture varies by a fraction that has nothing to do with how bright the source is. Young’s approximation gives it as

σscint=0.09D2/3X1.75eh/8000(2t)1/2\sigma_{\rm scint} = 0.09\,D^{-2/3}X^{1.75}e^{-h/8000}(2t)^{-1/2}

with DD the aperture in centimetres, XX the airmass and hh the site altitude, and for a one-metre telescope at airmass 1.2 in a sixty-second exposure that is about 0.45 millimagnitudes.

Three things in that expression are worth reading. The D2/3D^{-2/3} says a bigger telescope helps, but weakly — a hundred times the area buys a factor of 4.6. The X1.75X^{1.75} says airmass matters enormously, which is why photometric campaigns are scheduled around the meridian, and why refraction and airmass are not merely a pointing correction. And the t1/2t^{-1/2} says it does average down, so scintillation is not a floor in the strict sense.

Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 1.09 times that floor and after 10000 it is 1.01, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 3 Which raises the question of what is. The same budget against exposure time for a star at V=10V = 10: both random terms fall as t1/2t^{-1/2} — straight lines of slope 12-\tfrac12 on these axes — and one term does not fall at all. Flat-fielding error, differential colour terms and the imperfect match between a target’s spectrum and its comparison’s are systematic: they repeat, so averaging leaves them exactly where they were. After a thousand seconds the total is 1.29 times that floor and after ten thousand it is 1.03, which is the practical statement that a ground-based night is over long before the photons run out.

What was actually measured

The best evidence that the floor is real is the gap between what is achieved from the ground and what is achieved above it.

The best ground-based differential photometry, on bright stars with well-matched comparisons, reaches about 0.3 millimagnitudes per point and is limited by second-order extinction and flat-fielding rather than by counting. Kepler, in a 30-minute integration on a 12th-magnitude star, achieved 20 parts per million — 0.02 millimagnitudes — which is a factor of fifteen better, on a telescope of 0.95 m against ground-based apertures of similar size or larger.

The difference is not photons and it is not aperture. It is that a space telescope has no atmosphere to scintillate, no airmass term, no variable extinction, and a thermal environment stable enough that the flat field does not move. What a transit survey buys by leaving the ground is the elimination of the flat part of the curve above, and that is what makes an Earth-sized planet detectable at all — which is why the occurrence rate of small planets is a number that could only be measured from space.

Three planets against one night's error bar. The same total error, with three transit depths laid across it as horizontal lines: a Jupiter at 10000 parts per million, a Neptune at 1100 parts per million, an Earth at 84 parts per million. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. The marks are where a dip is seven times the error on one 60-second measurement, which is a Jupiter down to V = 13.4, and a Neptune and an Earth nowhere at all. That last is the figure's real content: those depths are below the bright-end floor, so no single exposure of any star on any night contains them, however bright the star and however large the telescope. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is not one measurement but a few hundred through the dip and dozens of repeats, and folding them is exactly a √N.
Fig. 4 The consequence, laid out. Three transit depths as horizontal lines across the same total error: a Jupiter at 10,000 parts per million, a Neptune at 1,100 and an Earth at 84. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. At a signal-to-noise of seven per point a Jupiter is detectable down to V=13.4V = 13.4 and the other two nowhere at all, because their depths are below the bright-end floor. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is a few hundred points through the dip and dozens of repeats, and folding them is exactly a N\sqrt N.

The other place counting sets the limit

Photometry is not the only measurement built on a count. A radial velocity is too, and the arithmetic there is the same with an extra step.

A spectrograph spreads the same photons over thousands of resolution elements, so each has few; the velocity precision is set by how well the position of a line can be measured, and that goes as the line’s width divided by the signal-to-noise per element, summed over all the lines. The result is again N1/2\propto N^{-1/2}, and again there is a floor — set here by the stability of the instrument and by the star’s own surface, whose convective motions shift line centroids by metres per second in ways no calibration removes.

How many photons is that, exactly

It helps to put a number on the abstraction, and the number is startling in both directions.

A star of magnitude zero delivers about 8.8×1098.8\times10^{9} photons per second per square metre through the V band — a figure obtainable two ways, from the AB zero point of 3,631 janskys and from the Johnson flux density in ergs, which agree to a few per cent and are checked against each other in the figures above. Through a one-metre telescope with a realistic 30% end-to-end throughput, that is 2×1092\times10^{9} detected photons a second.

Scale it down. At V=15V = 15 the same telescope collects 2,100 photons a second, so a sixty-second exposure has 126,000 of them and a photon-limited precision of 3 millimagnitudes. At V=25V = 25 — the depth of a deep survey image — it collects 0.2 photons a second, and an hour’s exposure yields 760 photons, which is a 4% measurement.

Seven hundred and sixty photons is the entire observational basis for a galaxy at that magnitude, and that is what the phrase “faint object astronomy” means arithmetically. A single detected photon at V=25V=25 through a one-metre telescope represents an event that happened, on average, five seconds ago at the aperture and some billions of years ago at the source.

Where the model stops

The budget above assumes several things that a real night violates.

It assumes the noise is Gaussian and independent between exposures, which it is not: cosmic rays, satellite trails and cosmic-ray-like detector events give a tail that no N\sqrt N describes, and correlated red noise from slowly varying transparency is the reason transit surveys quote a “red noise” term separately.

It assumes the sky is uniform across the aperture. Near a bright star, near the Moon, or in a crowded field, it is not, and the sky estimate becomes the dominant error rather than the sky counts.

It assumes the zero point is known. It is not measured on the same exposure as the star, and a catalogue magnitude is an extrapolation off the end of a graph — the airmass extrapolation to zero, which has its own error and its own colour dependence.

Three planets against one night's error bar. The same total error, with three transit depths laid across it as horizontal lines: a Jupiter at 10000 parts per million, a Neptune at 1100 parts per million, an Earth at 84 parts per million. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. The marks are where a dip is seven times the error on one 3600-second measurement, which is a Jupiter down to V = 20.0, and a Neptune and an Earth nowhere at all. That last is the figure's real content: those depths are below the bright-end floor, so no single exposure of any star on any night contains them, however bright the star and however large the telescope. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is not one measurement but a few hundred through the dip and dozens of repeats, and folding them is exactly a √N.
Fig. 5 And the same three-term budget for a space telescope, where the sky is a hundred times darker. The limiting magnitude is where the source count stops beating the noise the sky contributes into the same aperture; above the atmosphere that background is zodiacal light rather than airglow and it is five magnitudes fainter per square arcsecond. The depth gained has nothing to do with resolution and everything to do with which of the three terms is dominant.

Three regimes lie outside the budget above, and in each of them a different term takes over.

The same budget read for two very different instruments shows which of the regimes each of them lives in, and the answer is not the one aperture alone would suggest.

What the error bar is made of, on a 8 m in 600 s. The four contributions to a photometric error, against the brightness of the star, for a 8-metre aperture, a 600-second exposure and a sky of 21 magnitudes per square arcsecond. The star's own photons give a line of slope exactly 0.2 — σ ∝ N^−1/2 and N ∝ 10^−0.4m, so a magnitude of extra faintness costs a fifth of a magnitude of precision, and no instrument changes that. The sky and the read noise are fixed counts, so their lines have slope 0.4, twice as steep, and they overtake the star at V = 18.25 — that crossing is the faint limit of the night, and it moves when the Moon rises rather than when the telescope changes. Scintillation is flat, because the atmosphere modulates a bright star and a faint one by the same fraction: at 3.23e-5 relative it is 0.04 millimagnitudes here and it is what caps the bright end, up to about V = 12.3. Below all of them is the systematic floor at 0.3 millimagnitudes, which is flat-fielding and colour terms and does not integrate down at all.
Fig. 6 The error budget for an eight-metre telescope in a ten-minute exposure. The source-limited regime extends to much fainter magnitudes than for a one-metre, and the sky-limited slope beyond it is identical — because the sky’s contribution scales with the collecting area exactly as the source’s does.
Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 2.49 times that floor and after 10000 it is 1.23, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 7 And the precision reachable from above the atmosphere with a two-and-a-half-metre aperture. The photon-limited curve falls as the square root of the exposure until it meets the systematic floor, and where that happens is the only number that matters for a transit measurement.

The sky that is made of sources

The budget treats the sky as a smooth background with Poisson noise, and at faint limits that is wrong in a way that sets a floor no exposure removes.

The sky is not smooth. A large fraction of it is the summed light of galaxies too faint to detect individually, and those galaxies are distributed in a pattern that does not average away. Within any resolution element the number of faint sources fluctuates, so the background varies from place to place by more than its own Poisson noise.

That is confusion noise, and its magnitude depends on the beam size and on the counts of sources fainter than the detection limit. It is negligible in an optical image with sub-arcsecond resolution, because a resolution element contains a small fraction of a faint galaxy on average. It is dominant in the far infrared and in the radio, where the beam is tens of arcseconds and each one contains many.

The consequence is a hard limit. Integrating longer reduces the photon noise and does not touch the confusion, so a survey reaches the confusion limit and stops improving — and the only escape is a smaller beam, which means a larger telescope or an interferometer.

Several of the major far-infrared surveys were confusion-limited within hours of starting, and the majority of their observing time bought sky coverage rather than depth, because depth was unavailable at any price.

There is a statistical technique that extracts information below the limit, and it is worth naming because it inverts the usual approach. Instead of detecting sources, it measures the distribution of pixel values — the fluctuations themselves — and fits the source counts that would produce that distribution. What comes out is a number of sources per unit flux, below the detection limit, without any individual source having been detected.

A noise term that cannot be removed can still be measured, and where the noise is made of the objects being surveyed, measuring it is a survey.

The other end, where there are too many photons

Everything above concerns the faint limit. The bright end has its own difficulties and they are not the same ones reversed.

A detector holds a finite charge per pixel, and beyond it the pixel saturates: additional photons produce no additional signal, and on many devices the excess charge spills into neighbouring pixels along a column. A saturated star is not merely badly measured, it is corrupting its neighbours.

Below saturation there is non-linearity. A well that is nearly full responds less than proportionally to further light, by a per cent or more in the top part of the range, so bright stars are systematically under-measured unless the response is calibrated and inverted.

And the shutter has a finite travel time. A very short exposure, needed to avoid saturating a bright star, gives a different effective exposure time at the centre of the field than at the edge — the shutter is still opening at one end while it has finished at the other — which puts a spatial gradient into the photometry that varies inversely with the requested exposure.

The consequence is a dynamic range problem. An image cannot simultaneously measure a fourth-magnitude star and a twentieth-magnitude one, twelve magnitudes apart, because no exposure serves both. Surveys therefore take a short exposure and a long one of every field, and the two are joined by stars in the middle of the range that appear unsaturated in one and well measured in the other.

The brightest stars in the sky are among the worst-measured, and the catalogues of naked-eye stars are still based, for the very brightest, on visual and photographic work rather than on modern detectors.

When the counts are too few to be Gaussian

The budget assumes a Poisson distribution can be treated as a Gaussian of the same variance, which is excellent above a few tens of counts and false below it. In several branches of astronomy the counts are single digits.

An X-ray observation of a faint source may collect five photons in a hundred kiloseconds. A gamma-ray observation may collect none at all and report a limit. In that regime a chi-squared statistic — which is derived from a Gaussian likelihood — is not merely imprecise; it is biased, because it weights the bins by their observed counts and a bin that happened to be low is given more weight than one that happened to be high.

The remedy is to use the Poisson likelihood directly, which for fitting purposes goes by the name of the Cash statistic. It has the same asymptotic behaviour as chi-squared at high counts and the right behaviour at low ones, and it handles empty bins, which chi-squared cannot because their variance is zero.

The difference is not academic. Fitting a spectrum with chi-squared in the low-count regime returns systematically low fluxes, and the bias grows as the counts fall — which is exactly where the faintest and most interesting sources are.

There is a further consequence for what a detection means. With a handful of counts, the significance of a source is not a number of standard deviations, because the distribution is not symmetric and has no meaningful standard deviation. It is a probability of the background alone producing that many counts, and the two ways of quoting it can differ by a great deal.

The arithmetic in this essay is a large-number approximation, and the branches of astronomy that work with the fewest photons are the ones that had to abandon it.

And what those photons buy in the one currency a survey cares about, which is how faint a planet it can see.

Three planets against one night's error bar. The same total error, with three transit depths laid across it as horizontal lines: a Jupiter at 10000 parts per million, a Neptune at 1100 parts per million, an Earth at 84 parts per million. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. The marks are where a dip is seven times the error on one 600-second measurement, which is a Jupiter down to V = 17.6, and a Neptune and an Earth nowhere at all. That last is the figure's real content: those depths are below the bright-end floor, so no single exposure of any star on any night contains them, however bright the star and however large the telescope. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is not one measurement but a few hundred through the dip and dozens of repeats, and folding them is exactly a √N.
Fig. 8 Three planet sizes against host magnitude for the same instrument on a darker sky. Each curve is where the transit depth equals the noise, and the vertical spacing between them is the square of the radius ratio — so a survey that reaches one magnitude fainter finds planets a few per cent smaller, and no more than that.

Where this ladder goes next

This rung establishes the budget and its exponents. The rungs above it are about what is done when the floor is reached.

The first is differential photometry as a technique rather than a remedy: measuring a target against comparison stars in the same field, on the same pixels, through the same air, so that everything common cancels. It is why a transit survey stares at one field rather than scanning, and it converts an absolute measurement problem into a relative one — with its own new failure mode, which is that the comparison stars vary too.

Beyond it lies the question of what the floor is made of, which is now the active problem in the field: the granulation and oscillations of the star itself set a limit on radial-velocity precision at about half a metre per second, and on photometric precision at tens of parts per million, and neither is instrumental. The noise that stops the next generation of measurements is the target’s own surface, and separating a planet’s signal from a star’s own restlessness is the work that decides whether an Earth around a Sun is detectable at all.

What this makes readable

Essays that name this one as a prerequisite.

What links here

The 8 of 21 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Aperture photometryDifferential photometryFaint limitPhoton noiseRead noiseScintillationSignal-to-noiseSky backgroundSystematic floorZero point