Starlight

The faint star is measured against a brighter sky

For anything at the edge of detection the dominant source of noise is not the object. It is the sky in the same aperture, which is brighter than the star and is subtracted rather than measured — and once that is true, every rule of thumb about apertures, exposure times and image quality changes.

Assumes Photon noise, Seeing and Magnitudes.

A star at the limit of a telescope’s reach is not measured against darkness. It is measured against a sky that, within the patch of the detector the star’s light falls on, is very often brighter than the star.

That single fact restructures the whole business of faint photometry. It changes what an exposure time buys, what a larger mirror buys, and — the least intuitive of the three — it makes the sharpness of the image worth as much as the size of the collecting area. None of it follows from the star; all of it follows from the background.

What the error bar is made of, on a 1 m in 60 s. The four contributions to a photometric error, against the brightness of the star, for a 1-metre aperture, a 60-second exposure and a sky of 21 magnitudes per square arcsecond. The star's own photons give a line of slope exactly 0.2 — σ ∝ N^−1/2 and N ∝ 10^−0.4m, so a magnitude of extra faintness costs a fifth of a magnitude of precision, and no instrument changes that. The sky and the read noise are fixed counts, so their lines have slope 0.4, twice as steep, and they overtake the star at V = 18.25 — that crossing is the faint limit of the night, and it moves when the Moon rises rather than when the telescope changes. Scintillation is flat, because the atmosphere modulates a bright star and a faint one by the same fraction: at 4.09e-4 relative it is 0.44 millimagnitudes here and it is what caps the bright end, up to about V = 10.8. Below all of them is the systematic floor at 0.3 millimagnitudes, which is flat-fielding and colour terms and does not integrate down at all.
Fig. 1 What an error bar is actually made of, for a short exposure on a small telescope. The star contributes its own Poisson noise, and so does the sky in the aperture, and so does the detector’s read noise and its dark current. Which term dominates depends entirely on how bright the source is: at the bright end the star’s own counting statistics are everything, and at the faint end they are a minority of the variance.

The two regimes

Let the star deliver SS photons into the measuring aperture during the exposure and the sky deliver BB into the same aperture. The measurement is the total minus an estimate of the background, so the signal is SS, and the variance is S+BS + Bbecause counting is Poisson and independent variances add. The signal-to-noise ratio is

SS+B.\frac{S}{\sqrt{S+B}}.

Two limits fall out, and everything else in this essay is a consequence of which one applies.

Source-limited, SBS \gg B: the ratio is S\sqrt{S}. Doubling the exposure improves it by 2\sqrt2; quadrupling the collecting area doubles it.

Background-limited, BSB \gg S: the ratio is S/BS/\sqrt{B}. Both SS and BB are proportional to the exposure time and to the collecting area, so the ratio still improves as the square root of each — but BB also depends on the size of the aperture on the sky, and SS does not.

That last clause is the whole of it. In the bright regime the shape of the image is irrelevant: every photon from the star is counted whether it lands in one pixel or a thousand. In the faint regime, spreading the same starlight over four times the solid angle admits four times the sky, and halves the signal-to-noise.

Three planets against one night's error bar. The same total error, with three transit depths laid across it as horizontal lines: a Jupiter at 10000 parts per million, a Neptune at 1100 parts per million, an Earth at 84 parts per million. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. The marks are where a dip is seven times the error on one 60-second measurement, which is a Jupiter down to V = 13.4, and a Neptune and an Earth nowhere at all. That last is the figure's real content: those depths are below the bright-end floor, so no single exposure of any star on any night contains them, however bright the star and however large the telescope. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is not one measurement but a few hundred through the dip and dozens of repeats, and folding them is exactly a √N.
Fig. 2 The same accounting on a metre-class telescope in a minute, which is what a survey actually has. The limiting magnitude is where the source count stops beating the noise from the sky in the same aperture, and the sky is the dominant term for everything faint — so depth improves as the square root of exposure and of collecting area, and directly with how small the point-spread function is, because a tighter image collects the same source photons against less sky. That last dependence is why seeing is worth as much as aperture.

Why a sharper image is worth as much as a bigger mirror

Put the two dependences together for a background-limited source. The star’s counts go as the collecting area, so as D2D^2. The sky’s counts go as the collecting area and as the solid angle of the aperture, which goes as the square of the image width θ\theta. So

SB    D2D2θ2  =  Dθ.\frac{S}{\sqrt{B}} \;\propto\; \frac{D^2}{\sqrt{D^2\theta^2}} \;=\; \frac{D}{\theta}.

Signal-to-noise on a faint source goes as the diameter over the seeing, not as the collecting area. Doubling the mirror diameter and halving the image width are worth exactly the same amount, and the second is very much cheaper.

That equivalence is the economic argument behind adaptive optics, behind putting telescopes on high dry mountains, and behind space telescopes of otherwise unremarkable aperture outperforming much larger ground-based ones on faint point sources. A 2.4-metre telescope above the atmosphere delivering 0.05-arcsecond images has the same faint-source figure of merit as a 40-metre telescope delivering 0.8-arcsecond ones.

Resolution stops improving at 10 cm of aperture. Angular resolution against aperture at 500 nm, both logarithmic. The falling line is diffraction alone, 1.03 λ/D, which is what a telescope in vacuum delivers and has no floor. The curve is the same telescope under an atmosphere of Fried parameter r₀ = 10 cm, combining diffraction and seeing in quadrature: it follows the diffraction line while D < r₀ and then bends onto a plateau at 0.98 λ/r₀ = 1.01″. An amateur's 100 mm at 0.1 m would resolve 1.062″ above the air and delivers 1.47″ through it; a metre at 1 m would resolve 0.106″ above the air and delivers 1.02″ through it; the VLT at 8.2 m would resolve 0.013″ above the air and delivers 1.01″ through it; the ELT at 39 m would resolve 0.003″ above the air and delivers 1.01″ through it. At 39 m the atmosphere is costing a factor of 371: the aperture is 390 coherence lengths across and every one of the 152,100 patches it collects arrives with a phase of its own. What the extra aperture still buys is photons and speckles, and those two are what adaptive optics and speckle interferometry respectively spend to get the falling line back.
Fig. 3 The resolution a ground-based aperture actually achieves, against its diameter. Past about ten centimetres the atmosphere rather than the optics sets the image width, so the resolution curve flattens completely — and with it the 1/θ1/\theta half of the figure of merit. Everything gained above that aperture is gained through DD alone, which is why the returns from raw size are so much worse than the collecting area suggests.

It also explains a fact that surprises people the first time they meet it: a larger telescope does not see a fainter sky. Sky brightness is a surface brightness, and surface brightness is what distance and aperture cannot touch. A bigger mirror collects more sky photons and more star photons in the same proportion; what it changes is the counting statistics, not the contrast.

Two numbers worth carrying

The crossover between the regimes is where S=BS = B, and for a point source it is close to the sky brightness per square arcsecond minus about a magnitude — so at a dark site in V band, anything fainter than about magnitude 21 is background limited and anything brighter is not. Almost everything a large telescope is built to observe is on the faint side of that line, and almost everything an amateur instrument observes is on the bright side, which is why the two communities have such different instincts about what helps.

The second number is how the depth scales. In the background-limited regime the faintest detectable flux goes as θB/(Dt)\theta\sqrt{B}/(D\sqrt{t}), so reaching one magnitude deeper costs a factor of 6.3 in time, or 2.5 in diameter, or 2.5 in image width. Reaching three magnitudes deeper — which is roughly what each generation of survey has managed — costs a factor of 250 in time. That is why depth is bought with aperture and sharpness rather than with patience.

What the sky is made of

The background is not one thing, and its composition decides how it behaves. From a dark site the optical night sky is, roughly in order: airglow from the upper atmosphere — hydroxyl radicals recombining at ninety kilometres, in a dense forest of near-infrared bands; zodiacal light, sunlight scattered from interplanetary dust; scattered starlight and unresolved faint galaxies; and, increasingly, artificial light. At V band a genuinely dark site sits near 21.9 magnitudes per square arcsecond, which sounds faint and is not: a one-arcsecond aperture at that brightness contains as many photons as a star of magnitude 21.9, and a great deal of interesting astronomy happens fainter than that.

Two of those components are not fixed. Airglow varies by tens of per cent over a night and structurally across the sky, so a background measured in one part of a wide field is not the background in another; and the zodiacal light depends on where the telescope is pointing relative to the ecliptic, by a factor of several. Neither averages down with time, because neither is noise — both are signal belonging to something else, and both have to be modelled rather than beaten.

In the near infrared the situation is far worse. The hydroxyl bands make the K-band sky roughly 13 magnitudes per square arcsecond — some two thousand times brighter, per unit area, than the optical sky — which is why infrared instruments read out fast, subtract aggressively, and were transformed by getting above the atmosphere.

Where integrating longer stops helping

The square-root law is unforgiving but it is at least a law. What ends it is not the law but the assumptions underneath it.

Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 1.09 times that floor and after 10000 it is 1.01, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 4 Precision against exposure time. The curve falls as the inverse square root until it hits a floor and then does not fall at all. The floor is set by whatever does not average down — flat-fielding errors, imperfect subtraction of the background, variations in atmospheric transparency between the source and its comparison stars, small changes in where the star sits on the detector — and once it is reached, more time is a waste.

The systematic floor is where nearly every hard measurement in this collection actually lives. The metre per second that is not the star is one instance of it; a transit depth measured to parts in ten thousand is another. Beating a floor requires a different measurement, not a longer one — differential photometry against comparison stars on the same frame, chopping between source and sky, or observing the same object through two apertures at once.

The common structure in all three is that they turn an absolute measurement into a ratio. An absolute flux depends on the atmosphere’s transparency, the telescope’s throughput and the detector’s gain, each of which drifts; a ratio between two stars imaged through the same optics in the same second depends on none of them. Almost every high-precision result in this collection is a ratio of that kind, which is also why so many of them measure changes exquisitely and absolute values poorly. A transit depth is a ratio and a stellar radius is not, and the second is a great deal harder.

Three planets against one night's error bar. The same total error, with three transit depths laid across it as horizontal lines: a Jupiter at 10000 parts per million, a Neptune at 1100 parts per million, an Earth at 84 parts per million. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. The marks are where a dip is seven times the error on one 60-second measurement, which is a Jupiter down to V = 13.4, and a Neptune and an Earth nowhere at all. That last is the figure's real content: those depths are below the bright-end floor, so no single exposure of any star on any night contains them, however bright the star and however large the telescope. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is not one measurement but a few hundred through the dip and dozens of repeats, and folding them is exactly a √N.
Fig. 5 Three planetary signals against one night’s error bar. This is the version of the argument that matters for exoplanet work: the question is never whether the star is bright enough to see, but whether the fractional precision achievable in one transit duration is smaller than the depth being looked for. That is a comparison between two small numbers, and the noise term in it is dominated by the background for every faint host.

The exposure that is too long

There is a limit from the other direction, and it is a practical one rather than a statistical one.

A long exposure accumulates cosmic-ray hits, which is a nuisance, and saturates its brightest stars, which is worse — a saturated comparison star is useless, and comparison stars are how the systematic floor is beaten. It also integrates over whatever the object was doing, which for anything variable destroys the measurement it was meant to make.

So the practice is to take many short exposures and add them, at which point the read noise enters NN times instead of once. For a background-limited optical exposure that is a negligible penalty, and the ability to reject cosmic rays, discard frames taken through cloud and track the point-spread function from frame to frame is worth far more. In the read-noise-limited regimes it is not negligible at all, and the choice of exposure time becomes a genuine optimisation with the two noise terms pulling opposite ways.

What a limiting magnitude actually means

A telescope’s quoted limiting magnitude is a statement with four hidden arguments in it: the exposure time, the seeing, the sky brightness, and the signal-to-noise ratio deemed to constitute a detection.

The last is the one most often left out, and it matters more than it appears to. A five-sigma detection threshold on a survey covering millions of independent resolution elements will produce thousands of false positives from noise alone, which is why survey thresholds are set nearer seven or ten sigma and why the “limiting magnitude” of a survey is a full magnitude brighter than the naive calculation gives.

What it does to a survey’s catalogue

The scaling has a consequence for what a deep image contains, and it is not the one a limiting magnitude suggests.

A survey does not detect everything brighter than its limit and nothing fainter. It detects a fraction of the objects at each magnitude, rising from zero well below the limit to one well above it, and the width of that transition is set by the scatter in the measurement. Near the limit the fraction is a half by construction, and the objects that make it in are preferentially those that happened to be measured high — so the catalogue’s faintest entries are systematically brighter than their true fluxes, by an amount that grows as the completeness falls.

That is one of the reasons a magnitude has to say which light before it can be compared to anything, and it is why every serious survey publishes a completeness function rather than a limiting magnitude. It is also why the faint end of any counted distribution — a luminosity function’s faint slope most of all — is the part that has to be argued for rather than read off.

For extended sources the story is different again and worse. A galaxy of low surface brightness can be arbitrarily large and still undetectable, because its light per unit area never rises far enough above the sky for any aperture to help; the limiting quantity is a surface brightness rather than a flux, and a survey’s surface-brightness limit and its point-source limit are independent numbers.

The imaging version of the same argument

Direct imaging of a faint companion beside a bright star is the same competition with the background replaced by the star itself.

Three planets against one night's error bar. The same total error, with three transit depths laid across it as horizontal lines: a Jupiter at 10000 parts per million, a Neptune at 1100 parts per million, an Earth at 84 parts per million. The lines are flat because a transit depth is a ratio — the same fraction of the light whether the star is bright or faint — while the error bar is not. The marks are where a dip is seven times the error on one 3600-second measurement, which is a Jupiter down to V = 20.0, and a Neptune and an Earth nowhere at all. That last is the figure's real content: those depths are below the bright-end floor, so no single exposure of any star on any night contains them, however bright the star and however large the telescope. What the figure cannot show is the recovery, and it is the whole of how the planets were found: a transit is not one measurement but a few hundred through the dip and dozens of repeats, and folding them is exactly a √N.
Fig. 6 And the same accounting for a six-metre space telescope in an hour, where the sky is a hundred times darker. Above the atmosphere the background is zodiacal light rather than airglow, and it falls by five magnitudes an arcsecond squared — so the same integration reaches four or five magnitudes deeper for reasons that have nothing to do with resolution. Every number on this curve is the same three-term budget: what the source delivers, what the sky delivers into the same aperture, and what the detector adds.

The lesson transfers exactly. What limits the measurement is not how many photons the companion delivers but how well the thing underneath it can be subtracted, and the way forward is not more time but a better estimate of the pedestal — from a reference star, from the same star at a different roll angle, or from the wavelength dependence that distinguishes a speckle from a planet.

Where the picture stops

Read noise has not gone away, it has moved. Modern detectors have read noise of a few electrons, which is negligible against a bright sky in a long optical exposure and dominant in a short one or in a narrow filter. Any regime where the exposure has to be short — fast photometry, speckle imaging, spectroscopy at high dispersion where each pixel gets few photons — is read-noise limited, and there the old t\sqrt{t} intuition fails in the opposite direction.

The background is not smooth. Subtracting a sky level assumes the sky is flat across the aperture, and near a bright galaxy, inside a nebula, or in a crowded field it is not. Then the limiting quantity is how well the structure underneath the source can be modelled, and no amount of integration improves it.

And the aperture is not the only estimator. Fitting the known point-spread function to the data — profile-fitting photometry — weights each pixel by how much signal it is expected to contain, and recovers a factor approaching the square root of two over a plain aperture in the background-limited case. On a crowded field it recovers a great deal more, because it can fit overlapping sources simultaneously.

The background that is other stars

The budget treats the background as sky, and in a crowded field most of it is not. It is the overlapping wings of neighbouring stars, and it behaves quite differently.

Sky is smooth and its fluctuations are Poisson, so a larger aperture admits more of it and the noise grows as the square root of the area. The wings of neighbouring stars are structured: they have a definite value at every point that depends on where the neighbours are, and they do not average away at all.

The consequence is that in a crowded field the limiting quantity is not photon statistics but how well the neighbours can be modelled and removed. That is crowding noise, and it sets a floor that no exposure time touches.

The remedy is to fit rather than to sum. Profile-fitting photometry takes the known shape of the point-spread function, fits all the overlapping sources in a region simultaneously, and reports each one’s amplitude — so the neighbours are subtracted by construction rather than treated as background.

That works, and it degrades gracefully rather than failing sharply. As the field gets denser the fits become more strongly correlated, the uncertainty on each star grows, and eventually a star is indistinguishable from a chance superposition of fainter ones.

Where that limit falls depends on the image sharpness, and here the sharpness enters twice rather than once. A narrower point-spread function reduces the sky in the aperture, as this essay has already argued, and it also reduces the overlap between neighbours — so in a crowded field the return on image quality is steeper than the linear one the sky-limited case gives.

The two most demanding photometric environments are opposite in character: an empty field at the detection limit is sky-limited and beaten with aperture and time, and a globular cluster core is crowding-limited and beaten with resolution alone.

The floor that is left when the atmosphere is gone

Going above the atmosphere removes the airglow, the scattered moonlight and the scintillation, and it does not leave a black sky.

In the optical the residual is zodiacal light — sunlight scattered from interplanetary dust, brightest near the ecliptic and near the Sun and never absent. From a spacecraft in the outer solar system it would be lower, and from anywhere in the inner one it is the floor. That is why the deepest optical images are taken at high ecliptic latitude, and why the sky brightness of a space telescope varies by a factor of several with pointing.

In the infrared the residual is the telescope itself. Everything warm radiates, and at ten microns a mirror at room temperature is an enormously bright source compared with anything being observed — so an infrared telescope’s background is set by its own optics unless the optics are cold.

That is the whole reason a large infrared observatory is a cryogenic instrument. Cooling the mirror to a few tens of kelvin moves its own emission out of the band being observed, and the sunshield that maintains it is the largest structure on the spacecraft.

The floor that remains is again zodiacal, this time in emission rather than in scattering: the same dust, warmed by sunlight, radiating at a couple of hundred kelvin. It is what an infrared space telescope’s sensitivity is quoted against.

Every observation is against something, and going to a darker place changes which thing rather than removing it.

The two regimes and the floor beneath them are worth reading for an instrument at the other end of the range from the one used above.

What the error bar is made of, on a 4 m in 300 s. The four contributions to a photometric error, against the brightness of the star, for a 4-metre aperture, a 300-second exposure and a sky of 19 magnitudes per square arcsecond. The star's own photons give a line of slope exactly 0.2 — σ ∝ N^−1/2 and N ∝ 10^−0.4m, so a magnitude of extra faintness costs a fifth of a magnitude of precision, and no instrument changes that. The sky and the read noise are fixed counts, so their lines have slope 0.4, twice as steep, and they overtake the star at V = 16.25 — that crossing is the faint limit of the night, and it moves when the Moon rises rather than when the telescope changes. Scintillation is flat, because the atmosphere modulates a bright star and a faint one by the same fraction: at 7.25e-5 relative it is 0.08 millimagnitudes here and it is what caps the bright end, up to about V = 11.8. Below all of them is the systematic floor at 0.3 millimagnitudes, which is flat-fielding and colour terms and does not integrate down at all.
Fig. 7 The error budget for a four-metre telescope under a bright sky. The sky-limited regime begins much brighter, because a bright sky puts more photons into every aperture — so a large telescope at a poor site can be worse than a small one at a good site for faint work.
Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 1.66 times that floor and after 10000 it is 1.08, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 8 And the precision reachable with a one-metre aperture against a systematic floor a hundred parts per million. Integrating longer stops helping at a well-defined exposure, and beyond it the only remaining improvements are in the instrument rather than in the observing.

Where this ladder goes next

Later rungs on this anchor: the optimal extraction weighting and how much it is worth; the sky subtraction problem in the infrared, where the background varies on a timescale shorter than the exposure; differential photometry and the ensemble of comparison stars that beats a transparency systematic; the statistics of thresholds in a survey with many trials; and the point at which precision stops being a noise question and becomes a calibration question, which is where every part-per-million measurement in this collection ends up.

What links here

The 8 of 10 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

AirglowAperture photometryBackground limitedExposure timeLimiting magnitudePhoton noiseThe point-spread functionRead noiseSeeing discSignal-to-noise ratioSky backgroundSystematic floor