Starlight

The best aperture throws away a tenth

Aperture photometry counts every pixel inside a circle equally and every pixel outside it not at all. For a faint star against its sky the best circle is two-thirds of the seeing wide, catches 71.5 per cent of the light, and reaches 90.2 per cent of the signal-to-noise that weighting each pixel by what it is worth achieves — a loss of 23 per cent in exposure time that no algorithm can beat by more.

Assumes Photon noise, Seeing and Magnitudes.

A star’s image on a detector is not a point. The atmosphere spreads it into a blob a second of arc or so across, brightest in the middle and fading into wings, and the sky adds a smooth glow under all of it. Measuring the star’s brightness means deciding which pixels to add up. The simplest decision — draw a circle, add everything inside, subtract the sky’s contribution, ignore everything outside — is aperture photometry, and it is how most stellar brightnesses in history have been measured.

The circle has a radius, and the radius is a trade. Too small, and starlight in the wings is thrown away. Too large, and the extra pixels contribute more sky noise than starlight. When the sky is brighter than the star, which for every faint object is the case, the trade has a sharp optimum, and the optimum is worth knowing exactly because it answers a second question as well: how much better could any method do?

A bright star wants a wide aperture and a faint one wants 0.68 of the seeing. Signal-to-noise of simple aperture photometry against the aperture radius, in units of the seeing's full width at half maximum (1″), each divided by what optimal pixel weighting achieves for the same star, for stars of V = 12, 17, 20, 23 observed for 60 s through a 1 m telescope under a sky of 21 mag/arcsec². A small aperture loses starlight; a large one admits sky, and the balance depends on which dominates. For a bright star its own photons are most of the noise, so a wider aperture keeps gaining light almost for free and the best radius is large — 1.63 FWHM at V = 12, reaching 100.0 per cent of the optimum. For a star fainter than its sky the best radius shrinks to 0.680 FWHM and the best aperture reaches only 90.5 per cent of what weighting each pixel by its share of starlight divided by its variance achieves. That residual is exact in the background-limited limit: the best aperture captures 71.5 per cent of the light and 0.902 of the optimal signal-to-noise, so optimal weighting is worth 11 per cent in signal-to-noise, or 23 per cent in exposure time, and no more. The image is taken to be Gaussian; a real point-spread function has broader wings, which makes a fixed aperture a little worse and the optimal weights harder to know.
Fig. 1 Signal-to-noise of aperture photometry against aperture radius, in units of the seeing FWHM (1″), each divided by what optimal pixel weighting achieves, for V = 12, 17, 20 and 23 on a 1 m telescope in 60 s under a 21 mag/arcsec² sky. The bright star’s best radius is 1.63 FWHM and reaches 100.0 per cent of the optimum; V = 17 reaches 98.2 per cent at 1.06 FWHM; V = 20, 93.0 per cent at 0.76; and V = 23, far below its sky, 90.5 per cent at 0.68 FWHM. In the background-limited limit the best aperture captures 71.5 per cent of the light and 0.902 of the optimal signal-to-noise.

Two regimes, two answers

The signal inside a circle of radius rr is Nf(r)N f(r), where NN is the star’s total count and ff is the fraction of a Gaussian image of width σ enclosed, 1er2/2σ21 - e^{-r^2/2\sigma^2}. The noise has two parts: the star’s own Poisson noise, Nf\sqrt{N f}, and the background’s, bπr2\sqrt{b\,\pi r^2}, where bb is the variance per square arcsecond from the sky and the detector’s read noise together.

When the star is bright its own noise dominates. Then the signal-to-noise is about Nf\sqrt{N f}, which only ever rises with rr: every extra ring of pixels adds starlight and adds noise proportional to the square root of that starlight, a trade that is always favourable. The optimum is as wide as the background allows, and for the V = 12 star in the figure that is one and a half seeing-widths, capturing essentially all the light.

When the star is faint its sky dominates. Then the signal-to-noise is about Nf(r)/bπr2N f(r)/\sqrt{b\pi r^2}, and ff saturates while the area keeps growing. Differentiating gives a maximum at

ropt=1.585σ=0.673 FWHMr_{\rm opt} = 1.585\,\sigma = 0.673\ {\rm FWHM}

where the circle contains 71.5 per cent of the starlight. The optimum does not depend on how faint the star is, how bright the sky is or how big the telescope is — only on the shape of the image. A background-limited star is best measured through a circle two-thirds of the seeing wide, and that circle deliberately throws away more than a quarter of the star’s light.

What weighting the pixels is worth

The circle treats pixels as all-or-nothing. A better estimate weights each pixel by how much it tells about the star: its expected share of the star’s light divided by its variance. A pixel near the centre, where the star is bright against the sky, gets a large weight; one in the wings, mostly sky, gets a small weight but not zero. This is optimal extraction — introduced for spectra in the 1980s and applied to images in the same form — and for Gaussian noise it is the best linear estimate there is.

Its signal-to-noise can be written down:

(SN)opt2=(NP)2NP+bdA\left(\frac{S}{N}\right)^2_{\rm opt} = \int \frac{(N P)^2}{N P + b}\,dA

where PP is the image’s normalised profile. In the background-limited limit this is N2/(4πσ2b)N^2/(4\pi\sigma^2 b): the image behaves as if it occupied an effective area of 4πσ24\pi\sigma^2, four times the area of the one-σ circle. The best aperture, at 1.585σ, achieves 0.902 of that signal-to-noise. So optimal weighting improves on the best possible aperture by eleven per cent in signal-to-noise.

In exposure time, which is what an observer actually spends, that is 1/0.9022=1.231/0.902^2 = 1.23: optimal weighting buys 23 per cent.

Optimal weighting buys 21 per cent of exposure time for a faint star and nothing for a bright one. The exposure time saved by weighting each pixel optimally rather than using the best simple aperture, as a fraction, against the star's magnitude, for 60-second exposures through a 1 m telescope with 1″ seeing under a 21 mag/arcsec² sky. The saving is zero for bright stars, whose own photons dominate the noise so that every pixel's weight is the same and an aperture wide enough to include them all is already optimal. It rises through the magnitudes where the star and the sky under its image are comparable — the star's photons equal the background inside a patch four times the image's area near V ≈ 19.2 — and reaches 21.1 per cent by V = 22, approaching the background-limited ceiling of 1/0.902² − 1 = 23 per cent for stars far below the sky. It passes half of that saving near V = 19.1. The saving is a real and permanent fifth to a quarter of a survey's time at its faint limit, and it is also a ceiling: no weighting, fitting or algorithm extracts more from the same pixels, because the optimal weights are already the best linear estimate a Gaussian noise permits.
Fig. 2 The exposure time saved by optimal weighting over the best simple aperture, against magnitude, for 60 s on a 1 m telescope with 1″ seeing under a 21 mag/arcsec² sky. Nothing is saved for bright stars. The saving rises through the magnitudes where star and sky under the image are comparable — near V ≈ 19.2 on this set-up, where it passes half its ceiling at V = 19.1 — and reaches 21.1 per cent by V = 22, approaching the background-limited ceiling of 23 per cent.

That number is both a gain and a ceiling. It is a gain because it is free: the same exposure, the same telescope, processed better, gives the same signal-to-noise as an exposure 23 per cent longer processed simply. For a survey whose faint limit is set by time, that is a fifth of the survey. And it is a ceiling because no linear method does better than the optimal weights, and for Gaussian noise no method at all does better than the Cramér–Rao bound the optimal weights reach. A claim that some new algorithm extracts much more than a quarter extra from the same pixels, on a point source in a flat background, is a claim to beat a theorem.

What the seeing does and does not change

A bright star wants a wide aperture and a faint one wants 0.68 of the seeing. Signal-to-noise of simple aperture photometry against the aperture radius, in units of the seeing's full width at half maximum (2″), each divided by what optimal pixel weighting achieves for the same star, for stars of V = 12, 17, 20, 23 observed for 60 s through a 1 m telescope under a sky of 21 mag/arcsec². A small aperture loses starlight; a large one admits sky, and the balance depends on which dominates. For a bright star its own photons are most of the noise, so a wider aperture keeps gaining light almost for free and the best radius is large — 1.47 FWHM at V = 12, reaching 99.9 per cent of the optimum. For a star fainter than its sky the best radius shrinks to 0.677 FWHM and the best aperture reaches only 90.3 per cent of what weighting each pixel by its share of starlight divided by its variance achieves. That residual is exact in the background-limited limit: the best aperture captures 71.5 per cent of the light and 0.902 of the optimal signal-to-noise, so optimal weighting is worth 11 per cent in signal-to-noise, or 23 per cent in exposure time, and no more. The image is taken to be Gaussian; a real point-spread function has broader wings, which makes a fixed aperture a little worse and the optimal weights harder to know.
Fig. 3 The same stars under 2″ seeing instead of 1″. Measured in units of the seeing, almost nothing moves: V = 12 is best at 1.47 FWHM and reaches 99.9 per cent of its optimum, and V = 23 is best at 0.677 FWHM and reaches 90.3 per cent. The optimum aperture scales with the image, and the ratio to optimal weighting is a property of the image’s shape. What worse seeing changes is the optimum itself — twice the image width admits four times the sky — which does not appear on axes normalised by it.

The normalised figure hides the thing that matters most about seeing, and it is worth being explicit about what it hides. In the background-limited limit the optimal signal-to-noise is N/4πσ2bN/\sqrt{4\pi\sigma^2 b}, so it goes as 1/σ1/\sigma: doubling the image width halves the signal-to-noise and quadruples the exposure time needed. A sharper image is worth as much as a bigger mirror, and every telescope’s faint limit is set by its site as much as by its size. What the normalised figure shows instead is that the relative merit of apertures and weights does not care about the seeing — an observer who scales the aperture with the night’s image width is always within the same fraction of the best.

That scaling is the practical lesson. A fixed aperture of, say, three arcseconds in radius is too large on a night of 0.8-arcsecond seeing — it admits about thirty times the sky area a faint star needs — and too small on a night of five arcseconds. Photometric pipelines that set the aperture by the measured width of each frame’s stars are recovering most of the difference between a fixed circle and the optimum without doing any weighting at all.

Under a bright sky

A bright star wants a wide aperture and a faint one wants 0.68 of the seeing. Signal-to-noise of simple aperture photometry against the aperture radius, in units of the seeing's full width at half maximum (1″), each divided by what optimal pixel weighting achieves for the same star, for stars of V = 14, 18, 21 observed for 60 s through a 1 m telescope under a sky of 18 mag/arcsec². A small aperture loses starlight; a large one admits sky, and the balance depends on which dominates. For a bright star its own photons are most of the noise, so a wider aperture keeps gaining light almost for free and the best radius is large — 1.15 FWHM at V = 14, reaching 99.0 per cent of the optimum. For a star fainter than its sky the best radius shrinks to 0.680 FWHM and the best aperture reaches only 90.5 per cent of what weighting each pixel by its share of starlight divided by its variance achieves. That residual is exact in the background-limited limit: the best aperture captures 71.5 per cent of the light and 0.902 of the optimal signal-to-noise, so optimal weighting is worth 11 per cent in signal-to-noise, or 23 per cent in exposure time, and no more. The image is taken to be Gaussian; a real point-spread function has broader wings, which makes a fixed aperture a little worse and the optimal weights harder to know.
Fig. 4 Stars of V = 14, 18 and 21 under a sky of 18 mag/arcsec², as bright as moonlight makes it. The V = 14 star, which was comfortably source-limited under a dark sky, now reaches only 99.0 per cent of its optimum at 1.15 FWHM; V = 21 is fully background-limited, best at 0.680 FWHM and 90.5 per cent. A brighter sky moves every star towards the background-limited end of the curve, and the ceiling of 23 per cent applies to more of the catalogue.

Moonlight makes the sky three magnitudes brighter in the visible, and a bright sky shifts the boundary between the two regimes to brighter stars. Under a dark sky of 21 magnitudes per square arcsecond on a metre-class telescope the crossing is near V = 19; under an 18-magnitude sky it moves three magnitudes brighter. More of the objects a survey cares about fall into the regime where the aperture has to be small and the weighting helps most.

What the error bar is made of, on a 1 m in 60 s. The four contributions to a photometric error, against the brightness of the star, for a 1-metre aperture, a 60-second exposure and a sky of 18 magnitudes per square arcsecond. The star's own photons give a line of slope exactly 0.2 — σ ∝ N^−1/2 and N ∝ 10^−0.4m, so a magnitude of extra faintness costs a fifth of a magnitude of precision, and no instrument changes that. The sky and the read noise are fixed counts, so their lines have slope 0.4, twice as steep, and they overtake the star at V = 15.25 — that crossing is the faint limit of the night, and it moves when the Moon rises rather than when the telescope changes. Scintillation is flat, because the atmosphere modulates a bright star and a faint one by the same fraction: at 4.09e-4 relative it is 0.44 millimagnitudes here and it is what caps the bright end, up to about V = 10.8. Below all of them is the systematic floor at 0.3 millimagnitudes, which is flat-fielding and colour terms and does not integrate down at all.
Fig. 5 The four contributions to the error under the same bright sky, on a 1 m telescope in 60 s. The star’s own photons give a line of slope 0.2 magnitudes of error per magnitude; the sky and read noise give slope 0.4 and overtake the star at V = 15.25. Scintillation caps the bright end at 0.44 millimagnitudes up to about V = 10.8, and a systematic floor sits at 0.3 millimagnitudes. The crossing at 15.25 is where the aperture’s trade changes from one regime to the other.

The error budget shows why the transition is not sharp in practice. The sky noise overtakes the star’s own noise at V = 15.25 under this sky, but the aperture’s optimum depends on the ratio of starlight to sky within the image, which for a one-arcsecond image is a much smaller patch than the square arcsecond the sky magnitude is quoted in. The regimes overlap for a couple of magnitudes, which is exactly the range over which the extraction gain rises from zero to its ceiling.

The larger telescope

Optimal weighting buys 23 per cent of exposure time for a faint star and nothing for a bright one. The exposure time saved by weighting each pixel optimally rather than using the best simple aperture, as a fraction, against the star's magnitude, for 300-second exposures through a 4 m telescope with 1″ seeing under a 21 mag/arcsec² sky. The saving is zero for bright stars, whose own photons dominate the noise so that every pixel's weight is the same and an aperture wide enough to include them all is already optimal. It rises through the magnitudes where the star and the sky under its image are comparable — the star's photons equal the background inside a patch four times the image's area near V ≈ 20.1 — and reaches 22.7 per cent by V = 26, approaching the background-limited ceiling of 1/0.902² − 1 = 23 per cent for stars far below the sky. It passes half of that saving near V = 20.0. The saving is a real and permanent fifth to a quarter of a survey's time at its faint limit, and it is also a ceiling: no weighting, fitting or algorithm extracts more from the same pixels, because the optimal weights are already the best linear estimate a Gaussian noise permits.
Fig. 6 The saving for a 4 m telescope with 300 s exposures under the dark sky. The curve has the same shape, displaced to fainter stars: it passes half its ceiling near V = 20.0 and reaches 22.7 per cent by V = 26. A larger telescope collects more star and more sky in the same proportion, so the magnitude at which the two balance moves only because the longer exposure beats down the read noise.

The transition moves less than the change of telescope suggests, and the reason is in the ratio. A 4-metre mirror collects sixteen times the starlight of a 1-metre mirror and sixteen times the sky, so the ratio of star to sky within the image is the same for the same star. What shifts the curve is the read noise, a fixed number of electrons per pixel regardless of aperture, which matters less as the sky counts grow with longer exposures and larger mirrors. The regime a star is in is set by the sky and the seeing, not by the telescope.

Where the method began: a spectrum is a stack of profiles

Optimal weighting was first worked out not for images of stars but for spectra, and the setting makes its logic especially plain. A long-slit spectrum on a detector is a strip: along one direction, wavelength; across it, the star’s image smeared by the slit and the seeing. Every column of the strip is a one-dimensional profile of the star against the sky, and adding up a column to get the flux at that wavelength is aperture photometry in one dimension.

The traditional extraction summed a fixed number of pixels across the strip. The optimal extraction weights each by the profile divided by its variance, exactly as above, and it gains the same kind of factor — larger when the spectrum is faint against the sky, as it always is in the wings of absorption lines that carry the composition, where the star has fewest photons. It also gives something aperture photometry cannot: a cosmic-ray hit shows up as a pixel wildly inconsistent with the profile, and the profile tells the extraction to ignore it. A method built to squeeze out signal-to-noise turned out to be the method that cleans the data too, because both uses rest on the same thing — knowing what shape the signal should have.

The bright stars where none of this matters

At the other end of the magnitude range the arithmetic reverses completely, and practice follows it.

A transit survey observes stars bright enough that the light a planet removes — a part in ten thousand for an Earth round a Sun — can be measured at all. Such stars are deep in the source-limited regime, where the best aperture is large, the weighting gains nothing, and the sky is irrelevant. What limits them is not photons against sky but the detector: the star’s image sits on a handful of pixels whose sensitivities differ by a per cent, and as pointing jitter moves the image by a fraction of a pixel, the measured brightness changes by far more than any transit.

The remedy is to make the image worse. Ground-based observers of bright transits deliberately defocus their telescopes, spreading each star over hundreds of pixels so that pixel-to-pixel sensitivity differences average away and the image’s position within any one pixel stops mattering. On a faint, sky-limited star that would be ruinous — a four-times-wider image admits sixteen times the sky. On a bright, source-limited one it costs almost nothing, because the sky it admits is negligible against the star, and it buys a large reduction in the systematic that actually limits the measurement. A response measured pixel by pixel is never measured perfectly, and averaging over many pixels is the cheapest correction for what the flat field missed.

The same star, under the same sky, is therefore measured through a tight weighted aperture on a sharp image if it is faint and through a wide unweighted one on a deliberately blurred image if it is bright. Both are optimal. The regime decides.

Objects that are not points

Galaxies have their own profiles, and the same trade applies with the galaxy’s light profile in place of the seeing. A galaxy much larger than the seeing is best measured through an aperture scaled to its own size, and because a galaxy’s outer light falls off gently, the best aperture for a faint galaxy excludes a large fraction of its light — often most of it. A galaxy’s surface brightness does not depend on its distance, so a low-surface-brightness galaxy is background-limited at every distance, and the fraction of its light a sky-limited aperture can profitably include is small whatever telescope is used.

That is why galaxy magnitudes come in so many kinds — isophotal, Petrosian, model-fitted, Kron — each a different answer to where to stop adding pixels, each capturing a different fraction of the light, and each consistent only with itself. A total magnitude for a faint galaxy is an extrapolation from the part of the profile that could be measured, and the extrapolation is the same kind of profile model that optimal weighting needs.

Where the optimal weights come from

Optimal weighting needs the image’s profile, and that is its practical weakness. The weights are the expected share of the star in each pixel, which requires knowing the point-spread function — its width, its shape, its wings — in every frame and at every position on the detector. Behind adaptive optics that function has a sharp diffraction-limited core sitting on a broad seeing halo, and the fraction of light in each changes minute by minute with the atmosphere; the optimal weights change with it. A wrong profile gives wrong weights, and the resulting estimate is still unbiased for a constant profile error but no longer optimal; a profile that varies across the field and is modelled as constant gives a position-dependent bias.

In crowded fields the weighting becomes point-spread-function fitting: each star is modelled as a scaled copy of the profile at its position, overlapping stars are fitted simultaneously, and the fitted amplitude is the brightness. That is the method that makes photometry possible in globular clusters and the centres of galaxies, where no aperture can be drawn round one star without including another. Its statistical efficiency is the optimal weighting’s, and its practical accuracy is limited by how well the profile is known — which is why the brightest isolated stars in each frame are measured first, to learn the profile, and the faint crowded ones second, with it.

What the model leaves out

The image here is a Gaussian. Real seeing-limited images have broader wings than a Gaussian — closer to a Moffat profile — and those wings hold a few per cent of the light out to several times the FWHM. That makes a small fixed aperture lose more light than the Gaussian predicts, and it makes an aperture correction necessary: a measurement of how much light a small aperture misses, determined on bright stars and applied to faint ones. The ratio between aperture and optimal extraction is slightly worse for real profiles than the 0.902 drawn.

The detector is also taken to sample the image finely. The pixels here are a fifth of an arcsecond, five to a seeing-width, and the integrals are smooth. A camera whose pixels are as large as the image — common in wide-field survey instruments and in space telescopes with small optics — cannot draw a circle of 0.673 FWHM at all: an aperture is a set of whole pixels, and where a star falls within its central pixel changes both how much light the aperture catches and what the weights should be. Undersampled photometry carries an intra-pixel sensitivity term that no choice of aperture removes, and the ceiling of 23 per cent is a statement about well-sampled data that such cameras do not reach.

The background is taken to be flat and perfectly known. In practice it is estimated from an annulus around the star, and the estimate has its own noise, which adds to the sky term; in crowded or nebulous fields the background varies under the star and its estimate is biased. And the noise is taken to be Gaussian, which is right for sky-dominated counts and wrong for the very faintest sources, where counting in small numbers makes the Poisson distribution’s asymmetry matter.

A cut is a weighting with only two weights

An estimate built by including some data and excluding the rest is a weighted estimate with weights of one and zero, and the best such binary choice is always worse than the best continuous one. For a signal spread over a profile against a uniform noise, the loss from the best binary choice is a fixed fraction set by the profile’s shape — eleven per cent in signal-to-noise for a Gaussian — and that fraction is also the most any continuous weighting can recover.

The same arithmetic appears wherever a measurement integrates a known shape against noise: a matched filter in radio astronomy, the weighting of radial-velocity lines by their depth and sharpness, the choice of a window around a transit. In each case the matched weights are the ceiling, a well-chosen box is within a known fraction of it, and the size of that fraction says how much processing is worth before a single extra photon is collected.

Still open: what cancels when stars are measured against each other

Everything above measures one star against its own sky. The atmosphere does not stay constant from one exposure to the next: thin cloud, changing extinction and the slow drift of a detector’s response all multiply a star’s counts by a factor that has nothing to do with the star. Measuring the target against other stars on the same frame cancels every such factor at once, and the precision that results is set by what does not cancel — the comparison stars’ own noise, which the ensemble inherits, and the variability of any comparison star that turns out not to be constant.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Aperture photometryBackground limitedCramer rao boundInverse variance weightingOptimal extractionPhoton noiseThe point-spread functionRead noiseSeeingSignal-to-noiseSky background