Starlight

Colour is a thermometer, and it reads across the galaxy

A star's colour gives its surface temperature, from two brightness measurements and no other information. It is the cheapest useful measurement in astronomy.

Assumes Magnitudes.

A blacksmith can tell the temperature of iron by its colour: dull red, then orange, then yellow, then white. The judgement is good to within a hundred degrees or so with practice, and it requires no contact with the metal.

Stars do the same thing, for the same reason, and the measurement transfers without modification across a hundred thousand light years. Two brightness measurements through two filters give a temperature, and no other information about the star is needed — not its distance, not its size, not its composition.

Blackbody curves at 3000, 5800, 10000 K. Thermal emission against wavelength, each curve scaled to its own peak so the shift can be seen on one plot. The peak moves to shorter wavelengths as the temperature rises, which is why colour is a thermometer.
Fig. 1 Thermal emission against wavelength at three temperatures, each curve scaled to its own peak so the shift can be seen on one plot. The peak moves to shorter wavelengths as the temperature rises, and the dashed lines mark where Wien’s law puts it.

The curve that only knows one number

Any body in thermal equilibrium radiates with a spectrum that depends on its temperature and on nothing else. Not what it is made of, not how big it is, not what it did yesterday — one number fixes the entire shape.

That is a remarkable claim, and it took until 1900 to justify. Planck’s law gives the emitted power per unit wavelength,

Bλ(T)=2hc2λ51ehc/λkT1,B_\lambda(T) = \frac{2hc^2}{\lambda^5}\frac{1}{e^{hc/\lambda kT}-1},

and its derivation required assuming that energy came in discrete packets, which Planck regarded as a mathematical trick. It was not.

Two consequences carry all the practical weight. The peak wavelength moves inversely with temperature — Wien’s law, λmaxT=2.90×103m⋅K\lambda_{\max}T = 2.90\times10^{-3}\,\text{m·K} — so hotter means bluer. And the total output per unit area rises as the fourth power — Stefan–Boltzmann, F=σT4F = \sigma T^4 — so a small temperature difference is a large brightness difference.

The fourth power is the more startling of the two. A star twice as hot radiates sixteen times as much from every square metre of its surface, which is why the range of stellar luminosities is so much wider than the range of stellar temperatures.

Blackbody curves at 3000, 5800 K. Thermal emission against wavelength, each curve scaled to its own peak so the shift can be seen on one plot. The peak moves to shorter wavelengths as the temperature rises, which is why colour is a thermometer.
Fig. 2 Two curves at the extremes of the common range, area-normalised rather than peak-normalised. The cool star’s output is not merely redder — there is far less of it, because the fourth-power law is doing most of the work.

Turning that into a measurement

Nobody measures a full spectrum to get a temperature. Two filters are enough.

A colour index is the difference between magnitudes in two bands, conventionally BVB - V for blue and visual. Since a magnitude is a logarithm, a difference of magnitudes is a ratio of fluxes, and that ratio is fixed by the shape of the Planck curve — which is fixed by the temperature.

A hot star emits more in blue than in visual, so its BB magnitude is smaller than its VV, and BVB-V is negative. A cool star gives a positive index. The calibration is a monotonic curve: BV=0.3B-V = -0.3 is about 30,000 K, 0.00.0 is about 10,000 K, 0.650.65 is the Sun at 5,772 K, and 1.51.5 is about 3,500 K.

Two numbers, a subtraction, and a lookup. It is cheap enough to do for every star in a survey image at once, which is why colour is the axis of the diagram that organises stellar astronomy — spectra are expensive and colours are free.

The economics of that are worth a number. A spectrum good enough for a temperature requires collecting the star’s light and then spreading it over a few thousand resolution elements, so each one receives a few thousandths of what a photometric measurement would collect; the exposure needed is longer by a comparable factor, and the star has to be observed one at a time. A survey image records every star in the field simultaneously through one filter. Gaia measured colours for about 1.5 billion sources and spectra for about 1% of them, and the ratio is not a choice about priorities. It is what the arithmetic of photon collection permits.

The two filters also do not have to be blue and visual. Any pair that samples different parts of the Planck curve will serve, and the useful pair depends on the temperature being measured: BVB-V saturates above about 10,000 K, where both filters sit far out on the Rayleigh–Jeans tail and the ratio between them stops changing. Hot stars are measured with ultraviolet indices, cool ones with infrared, and the choice of filter pair is a statement about which part of the curve still has slope in it.

Blackbody curves at 4000, 7000, 20000 K. Thermal emission against wavelength, each curve scaled to its own peak so the shift can be seen on one plot. The peak moves to shorter wavelengths as the temperature rises, which is why colour is a thermometer.
Fig. 3 Three curves spanning the range of ordinary stars. Below 400 nm and above 700 the eye records nothing, so the visible band samples a different part of each curve — which is exactly what makes a two-filter ratio informative.

What “the temperature of a star” means

A star is not a blackbody and has no single temperature. Its centre is at fifteen million kelvin and its surface at six thousand, with everything in between.

What the colour measures is the effective temperature: the temperature of the blackbody that would radiate the same total power from the same area. It is a definition,

L=4πR2σTeff4,L = 4\pi R^2 \sigma T_{\text{eff}}^4,

and it is well-defined and useful precisely because the emitting layer is thin. Below it the gas is opaque and photons cannot escape; above it the gas is transparent and there is nothing to emit. The transition happens over a few hundred kilometres out of 700,000, so a star has a surface in the optical sense even though it has no surface in any material sense.

The blackbody approximation works because that layer is very close to equilibrium. It is not exact — the deviations are the spectral lines, and they are where all the chemistry is.

What was actually measured

The colour–temperature calibration is quoted as though it were a law. It is a fit, and the chain that produced it is short enough to state completely.

An effective temperature is defined by a flux and a radius. Getting one without assuming a stellar model therefore requires measuring both, for the same star, independently. The flux is the easy half: integrate the star’s spectrum over all wavelengths, correct for what the atmosphere absorbs and what falls outside the instrument, and the result is the bolometric flux at Earth in watts per square metre. The hard half is the radius, and the only model-free route to it is to resolve the star as a disc and measure its angular diameter.

That can be done, by optical interferometry, for stars that are large and close. A star of one solar radius at ten parsecs subtends about 0.9 milliarcseconds; the CHARA array on Mount Wilson, with baselines up to 331 metres, resolves down to roughly 0.2. The number of stars whose angular diameters have been measured this way to better than 2% is in the low hundreds, and they are overwhelmingly giants and nearby dwarfs — the same population bias that afflicts every direct measurement in stellar astronomy.

For each of them, flux plus angular diameter gives TeffT_{\text{eff}} with no model at all: the flux received divided by the solid angle subtended is the surface flux, and the fourth root of that over σ\sigma is the temperature. Plot those temperatures against the stars’ measured colour indices and the calibration curve is the fit through the result. Everything else — every temperature quoted for every star in every survey — is that curve, interpolated.

Blackbody curves at 5772 K. Thermal emission against wavelength, each curve scaled to its own peak so the shift can be seen on one plot. The peak moves to shorter wavelengths as the temperature rises, which is why colour is a thermometer.
Fig. 4 The Sun’s emission alone, with the visible band marked. Slightly under half the output falls inside it; the rest is infrared and ultraviolet that no visual measurement records. This is why a bolometric correction exists, and why “brightness” always has to specify which brightness.

The Sun is the anchor and it is measured differently, which is worth knowing because the whole scale hangs on it. Its total output comes from radiometers above the atmosphere — the total solar irradiance, 1,361 W/m², known to about 0.5 W/m² — and its angular radius from timing the limb, known to six figures. Those two give Teff=5,772T_{\text{eff}} = 5{,}772 K with an uncertainty of about 0.8 K. No star is known nearly that well, and every stellar temperature is ultimately expressed relative to it.

The gaps, which are the good part

The continuum gives temperature. The lines give everything else.

Fraunhofer catalogued 574 dark lines in the solar spectrum in 1814 without knowing what they were. Kirchhoff and Bunsen identified them in 1859: each element absorbs at its own set of wavelengths, so the dark lines name the elements present in the star’s atmosphere. Astronomy acquired chemistry, at any distance, from a diffraction grating.

The lines also give the temperature a second time, and better. Which lines appear depends on the ionisation and excitation state of the gas, which is far more temperature-sensitive than the continuum shape. Hydrogen lines peak in strength around 9,500 K — weak in hotter stars because the hydrogen is ionised, weak in cooler because it is unexcited. That non-monotonic behaviour is what made the original spectral classification confusing, and Cecilia Payne resolved it in 1925 by showing the sequence was temperature rather than composition. Her thesis also concluded that stars are overwhelmingly hydrogen, which her examiner persuaded her to describe as probably spurious. It was correct.

Lines carry more still. Their Doppler shifts give velocities, which is how unseen companions are found and how the masses that calibrate stellar astronomy are obtained. Their widths give pressures, which separates giants from dwarfs. Their splitting in a magnetic field gives field strengths. One spectrum yields temperature, composition, velocity, surface gravity and magnetism — from a point of light.

Where a 5772 K spectrum meets five filters. A single Planck curve at 5772 K, scaled to its own peak at 502 nm, with the five Johnson–Cousins responses shaded beneath it. Each shaded hump is the spectrum multiplied by that filter's response, so its area is the light the band actually collects, and the percentage beside each name is that area as a fraction of everything the star radiates at all wavelengths: U — 6.9%, B — 12.3%, V — 11.9%, R — 16.3%, I — 13.3%, 60.7% between them. The bands overlap, so those figures do not partition the light; and they cannot add to all of it, because most of a hot star's output is ultraviolet and most of a cool one's is infrared, where none of these filters looks. The responses are idealised as Gaussians at the published effective wavelengths and widths — a real filter's blue edge is steep and its red tail is the detector's, which this cannot show.
Fig. 5 What a colour index actually measures, on the curve it measures it from. Five standard passbands are laid over a 5,772 K spectrum, and an index is the ratio of the light collected in two of them — expressed as a difference because magnitudes are logarithms. Nothing about the star’s distance survives the ratio: both bands are dimmed by the same factor, so the index is a property of the surface alone. Which pair to take is the whole art. B−V straddles the peak here and is sensitive; for a 30,000 K star both bands sit far down the Rayleigh–Jeans tail, where the curve is a straight line of fixed slope and the index stops changing at all.

Where it lands on the diagram

Colour is one axis of the Hertzsprung–Russell diagram, and putting it there is what makes the diagram a physical statement rather than a plot.

The Hertzsprung–Russell diagram. Luminosity against surface temperature, both in solar units and on logarithmic axes, with temperature increasing to the left. The main sequence is computed from the mass–luminosity and mass–radius relations; the dashed diagonals are lines of constant radius.
Fig. 6 Luminosity against surface temperature, with temperature increasing to the left. Colour supplies the horizontal axis; the vertical needs a distance, which is why the diagram took a great deal longer to construct than the measurements it uses.

The horizontal axis is free — two filters. The vertical axis is expensive, because absolute luminosity requires a distance, and distances are the hard part of every quantity in the subject. The asymmetry in cost is why the diagram was first constructed for clusters, where every star is at the same unknown distance and the vertical axis can be left uncalibrated until one cluster is pinned down.

Colour, mass, and the chain between them

Temperature is not a free parameter of a star. It is downstream of the mass, through a chain with no adjustable steps in it.

Luminosity against mass, against a slope of 3.5. Main-sequence luminosity against mass, both in solar units, on logarithmic axes, over the range 0.079 to 63 solar masses. The measured curve comes from the eclipsing binaries and is the same in every drawing of it; what changes here is what it is compared against. The dashed line is a pure power law of exponent 3.5, and the curve crosses it rather than following it — the local slope runs from about 2.3 at the bottom of the range, where the interiors are convective, through nearly 4 near a solar mass where bound-free opacity dominates, to about 3 among the massive stars where electron scattering does. Quoting one exponent across the whole sequence is a convenience and the places it fails are the places the interior physics changes. Because the slope is between three and four across most of the range, a small spread in mass becomes an enormous spread in output: the 63-solar-mass end is 3.0e+9 times brighter than the 0.079-solar-mass end.
Fig. 7 Luminosity against mass. Combined with the mass–radius relation and the Stefan–Boltzmann law, this fixes the surface temperature — so a main-sequence star’s colour is a statement about its mass.

The chain runs: mass fixes luminosity; mass fixes radius; luminosity and radius fix the effective temperature through L=4πR2σT4L = 4\pi R^2\sigma T^4. Nothing is left over. Measuring a main-sequence star’s colour is measuring its mass, indirectly and with a wide error bar, but with no other information required.

That is why the main sequence is a line on the diagram rather than a region. Two independent measurements — colour and brightness — are both functions of one underlying quantity, so the points cannot help but fall on a curve. The residual width is composition, age, rotation and unresolved companions, and separating those is most of the work of stellar astronomy.

The best blackbody ever measured is not a star

The thermometer generalises past the objects it was built for, and the extreme case is the one where it works best.

Stars are mediocre blackbodies. Their spectra are riddled with absorption, their surfaces are structured, and the fit is good to a few percent at best. The finest blackbody spectrum anyone has measured belongs to no object at all: it is the cosmic microwave background, whose spectrum was measured by the FIRAS instrument on COBE in 1990 and found to match a Planck curve at 2.725 K to within 50 parts per million across the whole range sampled. The residuals were smaller than the thickness of the plotted line, and the graph presented at the January 1990 meeting of the American Astronomical Society got a standing ovation, which is not a normal response to error bars.

The reason it is so much better than any star is instructive about what the Planck curve actually requires. A blackbody spectrum is what radiation looks like when it has been in equilibrium with matter for long enough to forget everything else. A stellar photosphere is a boundary — light leaks out of it, so it is never quite in equilibrium. The early universe was opaque throughout, everywhere, for 380,000 years, with nowhere for anything to leak to. The radiation had no boundary to spoil it.

The surprising consequence is that the spectrum survived. The universe has expanded by a factor of about 1,100 since that radiation last touched matter, and every wavelength stretched by the same factor — which is exactly the transformation that takes a Planck curve at one temperature to a Planck curve at a lower one. A blackbody spectrum is the one spectral shape that is preserved by cosmological redshift. Any other shape would have been distorted beyond recognition, and the fact that a thermometer calibrated on iron in a forge reads a number for the whole sky is a coincidence of that stability rather than of anything about the sky.

There is one further consequence of the calibration being anchored on the Sun. Because every stellar temperature is expressed relative to a single star, a revision to the solar value moves the entire scale together — which is convenient, since most published results depend on temperature differences and those are unaffected. It is also a trap for anything that depends on the absolute value, such as a bolometric correction or a model atmosphere’s boundary condition, where a systematic shift propagates in full and does not cancel anywhere.

Two filters that are not the same two filters

A colour index is a difference of magnitudes through two filters, and the number is meaningless until the filters are named. That sounds like bookkeeping and it is a real source of error.

The traditional system uses broad glass filters whose transmission curves were defined in the 1950s by particular pieces of glass in front of a particular photomultiplier. Modern surveys use interference filters with sharper edges and different central wavelengths, and detectors whose sensitivity as a function of wavelength is nothing like a photomultiplier’s. A star’s index in one system is not its index in another, and the difference is not a constant offset: it depends on the star’s own spectrum, because two filters of different shapes weight the same spectrum differently.

So converting between systems requires a transformation, and the transformation is a fit — usually a polynomial in colour, derived from stars observed in both systems. It works well for ordinary stars and badly for anything with an unusual spectrum: a strongly reddened star, a star with emission lines, a cool star with deep molecular bands. Those are exactly the objects most likely to be interesting.

The professional response is to avoid the conversion. A survey defines its own natural system — the filters and detector it actually has — publishes magnitudes in it, and provides the transmission curves so that anyone comparing against a model can integrate the model through the real response rather than through an idealised band. That is more work and it removes a systematic that no amount of care in the observing can.

A photometric measurement is a number attached to an instrument, and the calibration that makes it comparable with somebody else’s is a separate piece of work with its own error budget. The Sun’s index of 0.65 is quoted here in one particular system, and in another it is a different number describing the same star.

Where the model stops

Stars are not blackbodies. The absorption lines remove a substantial fraction of the light, more in some bands than others, and the deficit is worst for cool stars where molecular bands dominate.

Reddening. Interstellar dust scatters blue light preferentially, so a distant star looks cooler than it is — and fainter than its distance implies, which are two errors that partly disguise each other. The correction is essential and is itself derived from colours, by comparing the observed index with the one the spectral type implies.

Not one temperature. Sunspots are 1,500 K cooler than their surroundings, and the colour of the disc as a whole is an average over a structured surface.

Not all thermal. Some emission is not thermal at all — synchrotron radiation, emission lines from hot gas — and applying a colour temperature to those gives a number with no meaning.

The figures share one distortion, and the first caption admits it: each curve is scaled to its own peak. That makes the shift visible and hides the fourth-power law entirely. Drawn at true relative amplitude, the 3,000 K curve would be a barely visible line at the bottom of the plot while the 10,000 K curve filled the frame — the honest picture, and one in which the colour shift would be invisible. No single plot shows both, which is why the second figure exists.

The curve has one parameter, and both what it does to the shape and what a filter set makes of it are worth reading at temperatures either side of the Sun’s.

Blackbody curves at 2500, 4000, 15000 K. Thermal emission against wavelength, each curve scaled to its own peak so the shift can be seen on one plot. The peak moves to shorter wavelengths as the temperature rises, which is why colour is a thermometer.
Fig. 8 Curves at 2,500, 4,000 and 15,000 kelvin. The peak moves as the inverse of the temperature and the height as its fourth power, so the hottest curve is off the top of any linear scale that shows the coolest — which is why the diagram is always drawn logarithmically and why a colour index is a ratio.
Where a 3500 K spectrum meets five filters. A single Planck curve at 3500 K, scaled to its own peak at 828 nm, with the five Johnson–Cousins responses shaded beneath it. Each shaded hump is the spectrum multiplied by that filter's response, so its area is the light the band actually collects, and the percentage beside each name is that area as a fraction of everything the star radiates at all wavelengths: U — 0.7%, B — 2.5%, V — 4.6%, R — 9.9%, I — 12.4%, 30.1% between them. The bands overlap, so those figures do not partition the light; and they cannot add to all of it, because most of a hot star's output is ultraviolet and most of a cool one's is infrared, where none of these filters looks. The responses are idealised as Gaussians at the published effective wavelengths and widths — a real filter's blue edge is steep and its red tail is the detector's, which this cannot show.
Fig. 9 And where a 3,500-kelvin spectrum meets five standard filters. Almost none of the flux is in the blue bands at all, so a colour index built from them measures the far tail of the distribution — which is both why it is sensitive and why it is noisy.

The ladder from here

Later rungs: Planck’s law derived, and the ultraviolet catastrophe it resolved. Wien’s law and Stefan–Boltzmann as its moments. Colour indices and the temperature calibration. Bolometric corrections. The spectral sequence and Payne’s resolution of it. Line formation and curve-of-growth analysis. Doppler broadening, pressure broadening, and Zeeman splitting. Reddening-free indices. And the ways colour lies: unresolved binaries, dust, and metallicity, each of which shifts the index by more than the measurement error.

Auguste Comte wrote in 1835 that the chemical composition of the stars was something humanity would never know. The spectroscope was turned on the Sun twenty-four years later.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 43 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

BlackbodyColour indexEffective temperatureStellar spectraWien's law