The best aperture throws away a tenth
Assumes Photon noise, Seeing and Magnitudes.
A star’s image on a detector is not a point. The atmosphere spreads it into a blob a second of arc or so across, brightest in the middle and fading into wings, and the sky adds a smooth glow under all of it. Measuring the star’s brightness means deciding which pixels to add up. The simplest decision — draw a circle, add everything inside, subtract the sky’s contribution, ignore everything outside — is aperture photometry, and it is how most stellar brightnesses in history have been measured.
The circle has a radius, and the radius is a trade. Too small, and starlight in the wings is thrown away. Too large, and the extra pixels contribute more sky noise than starlight. When the sky is brighter than the star, which for every faint object is the case, the trade has a sharp optimum, and the optimum is worth knowing exactly because it answers a second question as well: how much better could any method do?
Two regimes, two answers
The signal inside a circle of radius is , where is the star’s total count and is the fraction of a Gaussian image of width σ enclosed, . The noise has two parts: the star’s own Poisson noise, , and the background’s, , where is the variance per square arcsecond from the sky and the detector’s read noise together.
When the star is bright its own noise dominates. Then the signal-to-noise is about , which only ever rises with : every extra ring of pixels adds starlight and adds noise proportional to the square root of that starlight, a trade that is always favourable. The optimum is as wide as the background allows, and for the V = 12 star in the figure that is one and a half seeing-widths, capturing essentially all the light.
When the star is faint its sky dominates. Then the signal-to-noise is about , and saturates while the area keeps growing. Differentiating gives a maximum at
where the circle contains 71.5 per cent of the starlight. The optimum does not depend on how faint the star is, how bright the sky is or how big the telescope is — only on the shape of the image. A background-limited star is best measured through a circle two-thirds of the seeing wide, and that circle deliberately throws away more than a quarter of the star’s light.
What weighting the pixels is worth
The circle treats pixels as all-or-nothing. A better estimate weights each pixel by how much it tells about the star: its expected share of the star’s light divided by its variance. A pixel near the centre, where the star is bright against the sky, gets a large weight; one in the wings, mostly sky, gets a small weight but not zero. This is optimal extraction — introduced for spectra in the 1980s and applied to images in the same form — and for Gaussian noise it is the best linear estimate there is.
Its signal-to-noise can be written down:
where is the image’s normalised profile. In the background-limited limit this is : the image behaves as if it occupied an effective area of , four times the area of the one-σ circle. The best aperture, at 1.585σ, achieves 0.902 of that signal-to-noise. So optimal weighting improves on the best possible aperture by eleven per cent in signal-to-noise.
In exposure time, which is what an observer actually spends, that is : optimal weighting buys 23 per cent.
That number is both a gain and a ceiling. It is a gain because it is free: the same exposure, the same telescope, processed better, gives the same signal-to-noise as an exposure 23 per cent longer processed simply. For a survey whose faint limit is set by time, that is a fifth of the survey. And it is a ceiling because no linear method does better than the optimal weights, and for Gaussian noise no method at all does better than the Cramér–Rao bound the optimal weights reach. A claim that some new algorithm extracts much more than a quarter extra from the same pixels, on a point source in a flat background, is a claim to beat a theorem.
What the seeing does and does not change
The normalised figure hides the thing that matters most about seeing, and it is worth being explicit about what it hides. In the background-limited limit the optimal signal-to-noise is , so it goes as : doubling the image width halves the signal-to-noise and quadruples the exposure time needed. A sharper image is worth as much as a bigger mirror, and every telescope’s faint limit is set by its site as much as by its size. What the normalised figure shows instead is that the relative merit of apertures and weights does not care about the seeing — an observer who scales the aperture with the night’s image width is always within the same fraction of the best.
That scaling is the practical lesson. A fixed aperture of, say, three arcseconds in radius is too large on a night of 0.8-arcsecond seeing — it admits about thirty times the sky area a faint star needs — and too small on a night of five arcseconds. Photometric pipelines that set the aperture by the measured width of each frame’s stars are recovering most of the difference between a fixed circle and the optimum without doing any weighting at all.
Under a bright sky
Moonlight makes the sky three magnitudes brighter in the visible, and a bright sky shifts the boundary between the two regimes to brighter stars. Under a dark sky of 21 magnitudes per square arcsecond on a metre-class telescope the crossing is near V = 19; under an 18-magnitude sky it moves three magnitudes brighter. More of the objects a survey cares about fall into the regime where the aperture has to be small and the weighting helps most.
The error budget shows why the transition is not sharp in practice. The sky noise overtakes the star’s own noise at V = 15.25 under this sky, but the aperture’s optimum depends on the ratio of starlight to sky within the image, which for a one-arcsecond image is a much smaller patch than the square arcsecond the sky magnitude is quoted in. The regimes overlap for a couple of magnitudes, which is exactly the range over which the extraction gain rises from zero to its ceiling.
The larger telescope
The transition moves less than the change of telescope suggests, and the reason is in the ratio. A 4-metre mirror collects sixteen times the starlight of a 1-metre mirror and sixteen times the sky, so the ratio of star to sky within the image is the same for the same star. What shifts the curve is the read noise, a fixed number of electrons per pixel regardless of aperture, which matters less as the sky counts grow with longer exposures and larger mirrors. The regime a star is in is set by the sky and the seeing, not by the telescope.
Where the method began: a spectrum is a stack of profiles
Optimal weighting was first worked out not for images of stars but for spectra, and the setting makes its logic especially plain. A long-slit spectrum on a detector is a strip: along one direction, wavelength; across it, the star’s image smeared by the slit and the seeing. Every column of the strip is a one-dimensional profile of the star against the sky, and adding up a column to get the flux at that wavelength is aperture photometry in one dimension.
The traditional extraction summed a fixed number of pixels across the strip. The optimal extraction weights each by the profile divided by its variance, exactly as above, and it gains the same kind of factor — larger when the spectrum is faint against the sky, as it always is in the wings of absorption lines that carry the composition, where the star has fewest photons. It also gives something aperture photometry cannot: a cosmic-ray hit shows up as a pixel wildly inconsistent with the profile, and the profile tells the extraction to ignore it. A method built to squeeze out signal-to-noise turned out to be the method that cleans the data too, because both uses rest on the same thing — knowing what shape the signal should have.
The bright stars where none of this matters
At the other end of the magnitude range the arithmetic reverses completely, and practice follows it.
A transit survey observes stars bright enough that the light a planet removes — a part in ten thousand for an Earth round a Sun — can be measured at all. Such stars are deep in the source-limited regime, where the best aperture is large, the weighting gains nothing, and the sky is irrelevant. What limits them is not photons against sky but the detector: the star’s image sits on a handful of pixels whose sensitivities differ by a per cent, and as pointing jitter moves the image by a fraction of a pixel, the measured brightness changes by far more than any transit.
The remedy is to make the image worse. Ground-based observers of bright transits deliberately defocus their telescopes, spreading each star over hundreds of pixels so that pixel-to-pixel sensitivity differences average away and the image’s position within any one pixel stops mattering. On a faint, sky-limited star that would be ruinous — a four-times-wider image admits sixteen times the sky. On a bright, source-limited one it costs almost nothing, because the sky it admits is negligible against the star, and it buys a large reduction in the systematic that actually limits the measurement. A response measured pixel by pixel is never measured perfectly, and averaging over many pixels is the cheapest correction for what the flat field missed.
The same star, under the same sky, is therefore measured through a tight weighted aperture on a sharp image if it is faint and through a wide unweighted one on a deliberately blurred image if it is bright. Both are optimal. The regime decides.
Objects that are not points
Galaxies have their own profiles, and the same trade applies with the galaxy’s light profile in place of the seeing. A galaxy much larger than the seeing is best measured through an aperture scaled to its own size, and because a galaxy’s outer light falls off gently, the best aperture for a faint galaxy excludes a large fraction of its light — often most of it. A galaxy’s surface brightness does not depend on its distance, so a low-surface-brightness galaxy is background-limited at every distance, and the fraction of its light a sky-limited aperture can profitably include is small whatever telescope is used.
That is why galaxy magnitudes come in so many kinds — isophotal, Petrosian, model-fitted, Kron — each a different answer to where to stop adding pixels, each capturing a different fraction of the light, and each consistent only with itself. A total magnitude for a faint galaxy is an extrapolation from the part of the profile that could be measured, and the extrapolation is the same kind of profile model that optimal weighting needs.
Where the optimal weights come from
Optimal weighting needs the image’s profile, and that is its practical weakness. The weights are the expected share of the star in each pixel, which requires knowing the point-spread function — its width, its shape, its wings — in every frame and at every position on the detector. Behind adaptive optics that function has a sharp diffraction-limited core sitting on a broad seeing halo, and the fraction of light in each changes minute by minute with the atmosphere; the optimal weights change with it. A wrong profile gives wrong weights, and the resulting estimate is still unbiased for a constant profile error but no longer optimal; a profile that varies across the field and is modelled as constant gives a position-dependent bias.
In crowded fields the weighting becomes point-spread-function fitting: each star is modelled as a scaled copy of the profile at its position, overlapping stars are fitted simultaneously, and the fitted amplitude is the brightness. That is the method that makes photometry possible in globular clusters and the centres of galaxies, where no aperture can be drawn round one star without including another. Its statistical efficiency is the optimal weighting’s, and its practical accuracy is limited by how well the profile is known — which is why the brightest isolated stars in each frame are measured first, to learn the profile, and the faint crowded ones second, with it.
What the model leaves out
The image here is a Gaussian. Real seeing-limited images have broader wings than a Gaussian — closer to a Moffat profile — and those wings hold a few per cent of the light out to several times the FWHM. That makes a small fixed aperture lose more light than the Gaussian predicts, and it makes an aperture correction necessary: a measurement of how much light a small aperture misses, determined on bright stars and applied to faint ones. The ratio between aperture and optimal extraction is slightly worse for real profiles than the 0.902 drawn.
The detector is also taken to sample the image finely. The pixels here are a fifth of an arcsecond, five to a seeing-width, and the integrals are smooth. A camera whose pixels are as large as the image — common in wide-field survey instruments and in space telescopes with small optics — cannot draw a circle of 0.673 FWHM at all: an aperture is a set of whole pixels, and where a star falls within its central pixel changes both how much light the aperture catches and what the weights should be. Undersampled photometry carries an intra-pixel sensitivity term that no choice of aperture removes, and the ceiling of 23 per cent is a statement about well-sampled data that such cameras do not reach.
The background is taken to be flat and perfectly known. In practice it is estimated from an annulus around the star, and the estimate has its own noise, which adds to the sky term; in crowded or nebulous fields the background varies under the star and its estimate is biased. And the noise is taken to be Gaussian, which is right for sky-dominated counts and wrong for the very faintest sources, where counting in small numbers makes the Poisson distribution’s asymmetry matter.
A cut is a weighting with only two weights
An estimate built by including some data and excluding the rest is a weighted estimate with weights of one and zero, and the best such binary choice is always worse than the best continuous one. For a signal spread over a profile against a uniform noise, the loss from the best binary choice is a fixed fraction set by the profile’s shape — eleven per cent in signal-to-noise for a Gaussian — and that fraction is also the most any continuous weighting can recover.
The same arithmetic appears wherever a measurement integrates a known shape against noise: a matched filter in radio astronomy, the weighting of radial-velocity lines by their depth and sharpness, the choice of a window around a transit. In each case the matched weights are the ceiling, a well-chosen box is within a known fraction of it, and the size of that fraction says how much processing is worth before a single extra photon is collected.
Still open: what cancels when stars are measured against each other
Everything above measures one star against its own sky. The atmosphere does not stay constant from one exposure to the next: thin cloud, changing extinction and the slow drift of a detector’s response all multiply a star’s counts by a factor that has nothing to do with the star. Measuring the target against other stars on the same frame cancels every such factor at once, and the precision that results is set by what does not cancel — the comparison stars’ own noise, which the ensemble inherits, and the variability of any comparison star that turns out not to be constant.
About the same objects
Not linked from either essay — found by the objects both name.
- The threshold that is not a threshold photon noise · signal-to-noise
What links here
Essays that link to this one from their own argument.
The objects this essay names
Each one links to every other essay that touches it.
Aperture photometryBackground limitedCramer rao boundInverse variance weightingOptimal extractionPhoton noiseThe point-spread functionRead noiseSeeingSignal-to-noiseSky background