Starlight

A magnitude in a band the source never had

A filter passes light at a fixed observed wavelength, and a redshifted source emitted that light at a shorter one. Comparing a distant galaxy with a nearby one through the same filter therefore compares two different parts of two spectra — and the correction between them needs the spectrum, which is what the magnitude was going to be used to find out.

Assumes Magnitudes and Photometric systems.

A magnitude is a number attached to a filter. Change the filter and the number changes, which is why a magnitude has to say which light it was measured in. For a nearby object that is the whole of the story: two observers with the same filter get the same number.

For a distant object it is not. The light arriving in a filter centred at 550 nanometres left the source at 550/(1+z) nanometres, so an observer at redshift one is measuring the source’s ultraviolet flux with a green filter. Comparing that against a nearby galaxy measured in the same filter is comparing the ultraviolet of one object with the green of another.

Four spectra, and 3.9 magnitudes between them at the same redshift. The K-correction — the magnitude that has to be added to compare a redshifted object with a nearby one through the same filter — against redshift, for four power-law spectra. At zero redshift every correction is zero by construction. Beyond that they diverge, because a filter at a fixed observed wavelength samples a different part of the source's own spectrum at every redshift, and how much flux is there depends on the spectrum. A flat spectrum needs no correction at all at any redshift, and the two extremes drawn differ by 3.9 magnitudes by z = 1.2. The circularity is the point: applying the correction requires the spectrum, and the spectrum is what a magnitude is being used to constrain. The dashed line is the way out — observe in a band chosen so that it lands on the rest-frame band of interest, and the spectral term cancels, leaving only the bandwidth stretch.
Fig. 1 The correction that closes the gap, against redshift, for four power-law spectra. At zero redshift every correction is zero by construction. Beyond that they diverge, because how much flux is at the shifted wavelength depends on the spectrum — and a red spectrum and a blue one differ by several magnitudes by a redshift of one.

The correction between them is called a K-correction, it was defined by Hubble in the 1930s, and it is one of those quantities that is simultaneously routine, unavoidable and capable of absorbing the whole of an interesting result if it is wrong.

Where the factors come from

There are two effects and they are usually merged into one symbol, which is a pity because they are quite different.

The bandwidth stretch. A filter of width Δλ\Delta\lambda in the observed frame corresponds to a width Δλ/(1+z)\Delta\lambda/(1+z) in the source’s frame, so it collects light from a narrower slice of the source’s spectrum. That contributes a factor of (1+z)(1+z) and it depends on nothing about the source.

The spectral term. The flux collected is the source’s flux at the shifted wavelength, which depends entirely on what the spectrum is doing there. That contributes a factor that no general argument fixes.

Written together, for a source with spectral flux density FλF_\lambda and a filter response S(λ)S(\lambda),

K=2.5log10[(1+z)F(λ/(1+z))S(λ)dλF(λ)S(λ)dλ].K = -2.5\log_{10}\left[(1+z)\,\frac{\int F(\lambda/(1+z))S(\lambda)\,d\lambda}{\int F(\lambda)S(\lambda)\,d\lambda}\right].

For a power law FνναF_\nu \propto \nu^\alpha the integrals collapse and the whole correction is 2.5(1+α)log10(1+z)-2.5(1+\alpha)\log_{10}(1+z) — which is exactly zero when α=1\alpha = -1. A source whose flux per unit frequency falls as the inverse of the frequency needs no K-correction at any redshift, because the bandwidth stretch and the spectral term cancel exactly.

That special case is worth knowing because it is the boundary between corrections of one sign and corrections of the other, and because real galaxy spectra straddle it.

There is a third factor that is sometimes bundled in and should not be. Cosmological surface brightness dimming and the difference between luminosity distance and other distances are separate effects, with their own powers of one plus the redshift, and they belong to the distance rather than to the bandpass. Merging them into a single symbol makes the accounting shorter and makes the bookkeeping errors invisible; the discipline of keeping the bandpass effects separate from the geometric ones is worth the extra symbols, and it is why the shift that is not a Doppler shift has to be understood before any of this is safe.

It is worth working the arithmetic once, because the smallness of the numbers is deceptive. At redshift 0.1 — nearby, by any modern standard — the light in a 550-nanometre filter left the source at 500 nanometres, which for a galaxy is a shift from the green into the blue-green and across a region where the spectrum of an old population is falling steeply. The correction is already a tenth of a magnitude, which is five per cent in distance and enough to matter for a measurement of the local expansion rate. At redshift 0.5 it is half a magnitude. At redshift one it is more than a magnitude for a red galaxy and rather less for a blue one, and the difference between those two is the whole difficulty.

The circularity, and the way round it

The correction requires the spectrum. The magnitude is being measured because the spectrum is not known — if it were, the magnitude would be redundant.

In practice the circle is broken by templates. A library of observed spectra of nearby objects is redshifted, integrated through the filters, and the closest match to the observed colours is used to compute the correction. That works, and it imports every assumption in the library: that distant galaxies resemble nearby ones, that the library spans the real range, and that the colours determine the spectrum well enough.

For a well-studied class the assumption is defensible. For a supernova used as a standard candle it is checked directly — the spectra of nearby and distant events are compared and found to match — and the checking is a substantial part of why the measurement is believed.

Four spectra, and 2.9 magnitudes between them at the same redshift. The K-correction — the magnitude that has to be added to compare a redshifted object with a nearby one through the same filter — against redshift, for four power-law spectra. At zero redshift every correction is zero by construction. Beyond that they diverge, because a filter at a fixed observed wavelength samples a different part of the source's own spectrum at every redshift, and how much flux is there depends on the spectrum. A flat spectrum needs no correction at all at any redshift, and the two extremes drawn differ by 2.9 magnitudes by z = 2. The circularity is the point: applying the correction requires the spectrum, and the spectrum is what a magnitude is being used to constrain. The dashed line is the way out — observe in a band chosen so that it lands on the rest-frame band of interest, and the spectral term cancels, leaving only the bandwidth stretch.
Fig. 2 The same construction over a wider redshift range and for a narrower range of spectra — the range real galaxies actually span. The corrections still differ by more than a magnitude at the far end, which is to say that misclassifying a galaxy’s type costs more than a factor of two in its inferred luminosity. That is why photometric surveys measure many bands: the colours constrain the spectrum, and constraining the spectrum is what makes the correction possible at all.

The other way round the circle costs nothing and is used wherever the observing plan allows: choose the observed band so that it lands on the rest-frame band of interest. If a galaxy at redshift one is observed in a filter at twice the wavelength of the rest-frame band wanted, then the spectral term cancels identically and only the bandwidth stretch survives.

That is why deep surveys observe in many bands rather than deeply in one, and why the analysis of a high-redshift sample is usually described as being done “in the rest-frame BB band” rather than in whatever filter the photons landed in.

B − V against temperature, computed and measured. B − V against effective temperature. The curve is the colour index of a blackbody, obtained by integrating Planck's law against B and V response functions and subtracting a constant so that the index is exactly zero at 9600 K — the convention that an A0V star has every colour index zero, which is a choice and not a measurement. The 15 points are the main sequence as it is actually measured, and they do not lie on the curve: at the Sun's temperature the blackbody gives 0.446 where the sky gives 0.653, and at M0V 0.995 against 1.40. The model is too blue almost everywhere, and least wrong near 9600 K — which is why the zero point is put where it is. A colour index is a difference of two magnitudes, so it is a difference of two integrals, and changing either filter changes the number.
Fig. 3 The tool the templates are matched with: a colour index against temperature, which is the simplest possible statement of how a spectrum is inferred from photometry. Two magnitudes give one colour, one colour gives one parameter, and one parameter is not a spectrum — it is an index into a one-dimensional family of them. Every K-correction computed from broad-band colours is a lookup in a family of that kind, and its accuracy is bounded by how well the real population is described by however many parameters the survey’s bands can constrain.

The templates themselves deserve a sentence about where they come from, because it is not obvious. The standard libraries are built two ways: empirically, by observing large numbers of nearby galaxies spectroscopically and averaging by type, and theoretically, by combining stellar population synthesis models with prescriptions for dust and for star-formation history. The two disagree in the ultraviolet, which is exactly the rest-frame region a high-redshift optical observation samples, and they disagree because the ultraviolet is dominated by the youngest and rarest stars and by dust — the two things a population model is least able to predict. So the K-correction is best determined in the part of the spectrum where it is least needed, and worst determined where it matters most. That is a recurring shape and it is not a coincidence: the difficult regions are difficult for the same physical reasons in both the model and the observation.

What a K-correction does to a Hubble diagram

The place the correction matters most is the one where it is largest and least checkable: the comparison of distant objects with nearby ones to measure the expansion history.

Distance modulus against redshift, for three universes. The distance modulus μ = 5 log₁₀(D_L/10 pc) against redshift for three universes, all with H₀ = 67.36 km/s/Mpc, with 60 model supernovae drawn from the ΛCDM curve with 0.15 magnitudes of scatter. The point of the figure is how little difference there is: across two decades of redshift the three curves stay within a few tenths of a magnitude, and at z = 0.5 the accelerating and decelerating cases differ by 0.387 mag. A cosmology is not read off this plot. It is read off the residual, which is the next figure.
Fig. 4 The measurement the correction sits inside. Every point at high redshift has had a K-correction applied, computed from a template, and the quantity being extracted — the difference between the observed brightness and what a decelerating universe predicts — is a few tenths of a magnitude. The corrections themselves run to a magnitude or more. A systematic error in them that grows with redshift is indistinguishable from cosmology.

That is not a hypothetical concern and it was taken seriously from the start. The defence has three parts, and all three are about removing the correction rather than computing it better: observe in a band that lands on the rest-frame band, so the spectral term cancels; measure the spectra of the distant objects directly, so the template is checked rather than assumed; and compare like with like, using nearby objects observed through a filter chosen to match.

It is worth noting what would happen if the correction were simply omitted. For a red galaxy the K-correction at redshift one is positive and around a magnitude, so ignoring it makes the galaxy appear a magnitude brighter than it is, which in a distance measurement is a sixty per cent underestimate. The error grows monotonically with redshift, which is exactly the functional form a cosmological signal has, and the two are separable only because the K-correction depends on the object’s spectral type and the cosmology does not. That separation is the whole defence, and it requires that the sample contain a range of types.

There is a version of this trick that is worth naming separately because it is the reason modern surveys look the way they do. If a programme intends to study a population at a range of redshifts, it can choose its filter set so that successive filters correspond to nearly the same rest-frame band at successive redshifts — a filter at 550 nanometres for a nearby object, 770 at redshift 0.4, 1100 at redshift one. Then every object in the sample is measured in the same rest-frame band, the spectral term cancels for all of them, and the K-correction reduces to the bandwidth stretch, which is exact. That is not a data-analysis technique; it is an observing strategy, and it has to be decided before the survey starts. Every distance in the ladder is measured with the last one, and the rungs where the bands were chosen this way are the ones that survived.

What was actually measured

Three measurements bound the problem.

The spectra of distant supernovae. Spectroscopy of high-redshift type Ia events shows the same features at the same relative strengths as nearby ones, to the precision available. That is a direct check of the template assumption, and it is the reason the K-corrections applied to those objects are believed to a few hundredths of a magnitude.

Cross-filter photometry. Observing a distant object in a filter chosen to match the rest-frame band of a nearby comparison, and comparing the two directly, gives a measurement with a much smaller correction. Where both routes are available they agree, which is the check that the template-based correction is not introducing a systematic.

Photometric redshifts, and their failures. Estimating a galaxy’s redshift from its colours alone is the same machinery run backwards, and its failure modes are informative: a red galaxy at low redshift and a blue one at high redshift can have identical colours in a few bands, producing catastrophic outliers at a rate of a few per cent. That rate is a direct measurement of how well colours determine a spectrum, which is exactly the assumption the K-correction rests on.

Two stars at equal brightness, through five filters. Vega (A0V) and the Sun (G2V) drawn at the same total flux — each spectrum normalised so that its area over all wavelengths is identical — with the five Johnson–Cousins responses beneath them and the magnitudes they produce printed at the right. Vega (A0V) reads exactly zero in every band because the zero points are defined from it. The Sun (G2V), radiating precisely as much light in total, reads +0.77 in U, +0.28 in B, −0.17 in V, −0.44 in R, −0.71 in I — a spread of 1.48 magnitudes for two objects of equal brightness. None of those five numbers is the brightness; each is an integral of the spectrum against one piece of glass, and the differences between them are the only reason a colour index exists. The responses are idealised Gaussians at the published effective wavelengths and widths.
Fig. 5 The nearby version of the same problem, and the one that motivates everything above: the same star measured through two different realisations of nominally the same filter. The magnitudes differ, by an amount that depends on the star’s spectrum, and correcting between the two requires knowing that spectrum. A K-correction is this effect with the two bandpasses separated not by a manufacturing difference but by a factor of one plus the redshift.
A colour term of -0.059 magnitudes per magnitude, and one star that will not obey it. Synthetic photometry of blackbodies from 3,000 to 42,000 K through two V bands: the standard one, and a natural system whose effective wavelength is 9 nm longer and whose width is 1.12 times as large. The vertical axis is the difference between the two magnitudes for the same star — not a constant, because a wider redder band collects a different fraction of a hot spectrum than of a cool one. Fitting a straight line against B − V gives a colour term of -0.0585 magnitudes per magnitude and leaves a residual of 3.1 millimagnitudes, which is why a linear transformation is the standard reduction and why it works. The mark off the line is a cool star with molecular absorption bands in the red, drawn from the same blackbody with three synthetic bites taken out of it: it sits 27 millimagnitudes from the fit, 9 times the blackbody scatter. A colour term knows one number about a star and a spectrum has a shape, and that gap is the reason all-sky photometry stops at a per cent while differential photometry on one field reaches a millimagnitude.
Fig. 6 The nearby version of the whole problem, quantified: how far a natural instrumental band sits from the standard band it is supposed to realise, and the colour-dependent transformation between them. The transformation is a polynomial in colour, fitted on standard stars, and it is exact only for stars resembling those standards. A K-correction is the same object with the mismatch generated by cosmology rather than by glass, and with no possibility of observing a standard at the same redshift.

The last of these is the limiting case of the whole argument: a correction to a quantity that has no band at all, which is where the accumulated assumptions become visible because there is nothing left to hide them.

The bolometric correction to V, and where it is smallest. BC = M_bol − M_V against effective temperature, computed as −2.5 log₁₀ of the ratio of the whole radiated flux to the flux through the V band and then shifted so that a 5772 K star has BC = −0.08. The curve has an interior maximum at 6724 K, found on the drawing by root-finding rather than assumed, and it agrees to 1.0% with the narrow-band condition x e^x/(e^x−1) = 4, which contains no filter at all. That condition is worth reading carefully, because the obvious gloss is wrong: it is Wien's displacement with a 4 where the familiar form has a 5, the 4 being the Stefan–Boltzmann exponent, and its root puts 546 nm at the peak of the spectrum per logarithmic interval of wavelength — which is very nearly where the V band sits, at 551 nm. The ordinary per-nanometre Wien peak of a star this temperature is far bluer — about 431 nm — so the band is not sitting on the spectrum's brightest point; it is sitting where a fixed fractional slice of the spectrum carries the largest fraction of the whole. Away from it the correction grows fast and asymmetrically: −0.79 at 3850 K, −0.08 at 5772 K, −2.65 at 30000 K. So a magnitude measured in one band is furthest from the luminosity exactly for the hottest and the coolest stars — the ones whose luminosities matter most, and the ones a single filter sees least of.
Fig. 7 The limiting case of the same idea: the correction from a single band to the total light. A bolometric correction is a K-correction taken to the extreme — the “band” wanted is everything, no instrument measures everything, and the number quoted comes entirely from a model of the spectrum outside the observed range. For a hot star most of the light is in the ultraviolet and the correction is several magnitudes; for a cool one it is in the infrared. The lesson transfers: a correction that is large is a correction that is carrying most of the answer.

Where the picture stops

Three, and the second is the one that has changed the practice.

Emission lines break the template argument. A strong emission line moving into or out of a filter as the redshift changes produces a jump in the correction that no smooth template reproduces. For star-forming galaxies the effect is tens of per cent, it is not smooth in redshift, and it is one of the main sources of structure in photometric-redshift errors.

Dust does the same job and is not separable. Reddening changes a spectrum in a way that mimics a change of type, so an object’s K-correction and its extinction correction are degenerate. Fitting for both requires more bands than fitting for either, and the residual degeneracy is a well-known ridge in the parameter space.

And the correction assumes the object is not evolving. A template library built from nearby galaxies describes what galaxies look like now. Applying it at redshift two assumes galaxies then looked like some galaxies now, and the assumption is testable only where spectra are available — which is at the bright end of the population and not at the faint end where most of the objects are.

A fourth limit is worth adding because it is the one that governs the current generation of surveys. Everything above assumes the filter response S(λ)S(\lambda) is known. For a ground-based survey it is known to a per cent or so and it varies with airmass, with the detector’s temperature, and across the focal plane; for a space mission it is known better and it changes as the optics age. A one per cent error in the shape of a filter’s response produces a K-correction error that depends on the source’s spectrum and therefore on its redshift, which is again the functional form of a cosmological signal. Measuring filter responses in the laboratory, and then verifying them on the sky against stars with known spectra, is a substantial part of the calibration effort of any survey aiming at per cent photometry — and it is the same problem as knowing which light a magnitude was measured in, pushed to the precision cosmology needs.

Why the correction is really a statement about bandpasses

The general point is one this collection meets in several disguises, and it is worth extracting from the redshift context.

A photometric measurement is an integral of a spectrum against a response function. Two such measurements are comparable only if the response functions are the same in the source’s frame, and there are exactly two ways to arrange that: use the same filter on two objects at the same redshift, or use different filters chosen so that they correspond to the same rest-frame band. Everything else requires a model of the spectrum.

That statement contains the colour term between two telescopes, the transformation between photometric systems, and the K-correction as special cases. In each the correction is small when the bandpasses nearly agree and grows with their mismatch; in each it depends on the source’s spectrum; and in each the robust solution is to arrange the observation so that the correction is unnecessary rather than to compute it accurately.

There is a reason that matters more than tidiness. A correction computed from a model has an error that is correlated across every object the model was applied to, so it does not average down over a sample and it does not show up in the internal scatter. A survey of ten thousand galaxies with a one per cent template error has a one per cent systematic, exactly as if it had observed one.

A concluding observation about why this is a rung on the magnitude ladder rather than a footnote to cosmology. The entire difficulty is a consequence of measuring light through a filter, and it exists in miniature at every distance. A brightness becomes a distance only if something is known, and what has to be known is not only the object’s luminosity but the band that luminosity was defined in. A magnitude system is a promise that the band is the same for everybody, and redshift is the one thing that breaks the promise unavoidably — not through any failure of the instruments, but because the source and the observer are not in the same frame. The correction is the price of a comparison between two objects that could never have been observed identically.

End on the one case where the correction is exactly zero and nobody can arrange it. A source whose flux per unit frequency falls exactly as the inverse of the frequency needs no correction at any redshift, because the narrowing of the band and the change of the flux cancel identically. That is not a curiosity: it is the reason radio astronomers, whose sources are often close to that slope, largely ignore the whole subject, and it is the reason optical astronomers cannot. The difference between two fields’ relationship to a correction is entirely a fact about the spectra their sources happen to have.

Notice that the whole difficulty is an artefact of a convention rather than of nature, and that the convention persists for good reasons. Magnitudes exist because photometry was invented before spectroscopy was cheap, and a magnitude in a named band is a compact, comparable, archival quantity that a century of observations can be expressed in. The alternative — quote a flux density at a rest-frame wavelength, computed from a spectrum — requires a spectrum for every object, which for a survey of a hundred million galaxies is not available and will not be. So the correction is the price of a summary statistic that makes a heterogeneous archive comparable at all, and the modern practice of fitting a spectral template to many bands at once is not an abandonment of the convention but a way of estimating the spectrum the convention needs. The band is still the unit; what has changed is that there are enough bands to constrain what falls between them.

Where the ladder goes next

The obvious next rung is the photometric redshift itself: how a spectrum is estimated from a handful of broad-band magnitudes, what the degeneracies are, and why the failures are catastrophic rather than gradual. The one above that is the surface brightness — the one photometric quantity distance does not touch, and which acquires its own factors of one plus the redshift for a completely different reason.