Starlight

A temperature and a gravity that trade against each other

A stellar spectrum contains the star's temperature, its surface gravity and its composition, and no single feature in it contains only one of the three. Every diagnostic is a band in the parameter plane rather than a point, and the answer is where the bands cross — which makes choosing diagnostics that disagree in their sensitivities the whole of the art.

Assumes Spectra and Line formation.

A spectrum is asked for three things at once: how hot the star is, how strong the gravity is at its surface, and what it is made of. The three are not independently visible. Every line in the spectrum responds to all three, and the analysis is a matter of finding the combination that reproduces everything.

That would be unremarkable if the responses were sufficiently different. They are not. Raising the temperature and raising the gravity have similar effects on many observable quantities, so a spectrum that is fitted well by one pair is fitted nearly as well by another — and the “nearly” is what the error bars are made of.

Two diagnostics, two bands, and a crossing to 49 kelvin. The plane of effective temperature against surface gravity, with the constraints from two spectroscopic diagnostics drawn as bands. The wings of a hydrogen line are broadened by collisions, so they respond steeply to the gravity and weakly to the temperature: a narrow, steep band. An ionisation balance — requiring that the same element give the same abundance from its neutral and its singly ionised lines — responds to both, and its band is much shallower. Neither diagnostic determines either quantity on its own. Where the two cross is the answer, and the size of the crossing region is set by the band widths divided by the difference of the slopes — so two diagnostics that respond similarly give a long, thin, nearly useless error region however precise each one is. Choosing diagnostics that disagree in their sensitivities is the whole of the art.
Fig. 1 The plane of temperature against surface gravity, with the constraints from two diagnostics drawn as bands. The wings of a hydrogen line are broadened by collisions, so they respond steeply to the gravity and weakly to the temperature. An ionisation balance responds to both, and its band is much shallower. Neither determines either quantity on its own; where the two cross is the answer, and the size of the crossing region is set by the band widths divided by the difference of the slopes.

The situation is not a failure of the data. A high-resolution spectrum of a bright star contains an enormous amount of information — thousands of lines, each measured to a per cent — and the difficulty is not that the information is scarce but that most of it points in the same direction. Adding a thousand more lines of the same kind narrows a band without changing its orientation, and the crossing region stays elongated.

What each diagnostic actually measures

Hydrogen line wings. The far wings of the Balmer lines are broadened by the linear Stark effect — the perturbation of a hydrogen atom’s levels by the electric fields of nearby ions and electrons. The broadening therefore measures the charged-particle density, which for a star in hydrostatic equilibrium is set by the pressure, which is set by the gravity. For cool stars the wings are an excellent gravity indicator and a poor thermometer; above about 8,000 K the hydrogen ionises and the sensitivity reverses.

Ionisation balance. The same element observed in two ionisation stages must give the same abundance. The ratio of the two populations is set by the Saha equation, which depends steeply on temperature and, through the electron pressure, on gravity. Requiring consistency therefore defines a locus in the plane rather than a point.

Excitation balance. Lines of the same species arising from levels of different excitation energy must also give the same abundance. That condition depends almost entirely on temperature, through the Boltzmann factor, and hardly at all on gravity — which makes it the closest thing to a clean thermometer the analysis has.

Photometric colours. A colour index is a crude thermometer that is nearly independent of gravity, which is exactly what is wanted, and it is contaminated by reddening, which is not. Colour is a thermometer with a well-known systematic attached.

Two diagnostics, two bands, and a crossing to 49 kelvin. The plane of effective temperature against surface gravity, with the constraints from two spectroscopic diagnostics drawn as bands. The wings of a hydrogen line are broadened by collisions, so they respond steeply to the gravity and weakly to the temperature: a narrow, steep band. An ionisation balance — requiring that the same element give the same abundance from its neutral and its singly ionised lines — responds to both, and its band is much shallower. Neither diagnostic determines either quantity on its own. Where the two cross is the answer, and the size of the crossing region is set by the band widths divided by the difference of the slopes — so two diagnostics that respond similarly give a long, thin, nearly useless error region however precise each one is. Choosing diagnostics that disagree in their sensitivities is the whole of the art.
Fig. 2 The same construction for a red giant rather than a dwarf. The bands are in the same relative orientation and everything has moved to lower gravity, where a further problem appears: the hydrogen wings are weak in a giant’s spectrum because the pressure is low, so the steep band is also the poorly measured one. The crossing is worse determined for exactly the stars whose gravities span the widest range, which is why giant parameters are so much less certain than dwarf ones.

It is worth noticing that these four diagnostics come from three different kinds of physics — pressure broadening, ionisation equilibrium and excitation equilibrium — and that is precisely why they are used together. Diagnostics drawn from the same physics respond to the same combination of parameters and add nothing. The practical rule that follows is one an analyst can act on: when the errors are not shrinking, the missing ingredient is a diagnostic with a different physical origin rather than more of the same measurement.

There is a fifth diagnostic that deserves mention because it is the one that has transformed the field, and it is not spectroscopic at all. The frequency of maximum power in a star’s acoustic oscillations scales with the surface gravity divided by the square root of the temperature, almost independently of anything else — so a light curve long enough to resolve the oscillations gives a gravity to a hundredth of a dex with no atmosphere model in it. That is a vertical line of extraordinary sharpness through the plane in the previous figure, and stars with both measurements are the benchmarks everything else is calibrated against.

Why the crossing angle is the whole story

The precision of a two-parameter solution is not set by how well each diagnostic is measured. It is set by how well each is measured divided by how differently they respond.

Two bands crossing at a right angle give a compact region whose size is the band widths. Two bands crossing at a shallow angle give a long thin region whose length is the band widths divided by the sine of the angle, which can be an order of magnitude larger. In the limit of parallel bands the region is infinite: two diagnostics that respond identically contain the information of one.

That is a purely geometric statement and it has a practical consequence that is easy to state and hard to act on: improving a diagnostic is worth less than adding a differently-oriented one. Halving the width of the steep band shrinks the crossing region by a factor of two; adding a diagnostic at right angles to both can shrink it by ten.

Two diagnostics, two bands, and a crossing to 24 kelvin. The plane of effective temperature against surface gravity, with the constraints from two spectroscopic diagnostics drawn as bands. The wings of a hydrogen line are broadened by collisions, so they respond steeply to the gravity and weakly to the temperature: a narrow, steep band. An ionisation balance — requiring that the same element give the same abundance from its neutral and its singly ionised lines — responds to both, and its band is much shallower. Neither diagnostic determines either quantity on its own. Where the two cross is the answer, and the size of the crossing region is set by the band widths divided by the difference of the slopes — so two diagnostics that respond similarly give a long, thin, nearly useless error region however precise each one is. Choosing diagnostics that disagree in their sensitivities is the whole of the art.
Fig. 3 The same two diagnostics measured twice as well. The crossing region has shrunk in proportion — halving both widths halves the temperature uncertainty — and it has kept its elongated shape, because the shape is set by the angle rather than by the widths. Better data give a smaller version of the same degeneracy, and the direction of the remaining uncertainty is unchanged. That direction is what has to be quoted alongside the answer, and almost never is.

There is a corollary about how such results should be reported. A two-parameter solution has a covariance, not two error bars, and the covariance is where nearly all the information about the shape of the constraint lives. Quoting a temperature of 5,750 ± 60 K and a gravity of 4.40 ± 0.08 dex describes a rectangle; the actual constraint is a long ellipse tilted across it, and the difference matters for anything computed from both numbers. Publishing the covariance costs one extra number and is rare.

It is also worth noticing what the geometry says about which quantity is best determined. The steep band constrains gravity well and temperature poorly; the shallow band constrains a combination of both. Their intersection therefore determines gravity better than temperature whenever the steep band is narrow, and the reverse when it is not — so the relative precision of the two answers depends on the star’s spectral type through the strength of the hydrogen wings, and it is not a property of the method that can be quoted once. A survey covering a range of types has parameter uncertainties whose ratio varies across the sample, which is a nuisance for anything that averages over the sample.

The third parameter, and the fourth

The picture above has two axes and the real problem has more.

Metallicity enters everything. The number of electrons donated by metals sets the electron pressure, which affects the ionisation balance and the continuous opacity; the line strengths used for the balances depend on the abundance directly. In practice temperature, gravity and metallicity are solved for together, and the plane above is a slice through a three-dimensional degeneracy.

Microturbulence is worse, because it is not a physical parameter at all. Lines in real stars are broader than thermal motion alone allows, and the standard fix is a fudge velocity added in quadrature. It is fitted by requiring that strong and weak lines of the same species give the same abundance, and it is correlated with everything — a change in microturbulence can be traded against a change in temperature over a substantial range.

The result is that a “spectroscopic parameter set” is a point in a four-dimensional space determined by four consistency conditions, and the conditions are not orthogonal. Different codes, different line lists and different weighting schemes land in different places along the degenerate directions, which is why two competent analyses of the same spectrum routinely differ by a hundred kelvin in temperature and 0.2 dex in gravity while each quotes an uncertainty a third of that.

Hydrogen is half ionised by 9,555 K. Ionisation fractions from the Saha equation at an electron pressure of 20 N/m², against temperature. Each species' stages sum to one at every point, which is checked rather than assumed. Hydrogen crosses half-ionised at 9,555 K and is 99% ionised by 12,527; calcium, whose first ionisation potential is 6.113 eV against hydrogen's 13.598, is already 99% singly ionised at 6,000 K, where hydrogen is 100.00% neutral. Those two curves are the reason the same photosphere shows strong Ca II and weak Balmer at one temperature and the reverse at another — with the abundances fixed and only the temperature moving. The partition functions are held at their ground-state weights, which is the usual first approximation and shifts these curves by a few hundred kelvin rather than by their shape.
Fig. 4 The physics under one of the two bands: the fraction of an element in each ionisation stage, against temperature. The steepness of these curves is what makes the ionisation balance a sharp diagnostic, and the fact that they shift with electron pressure — and therefore with gravity — is what makes it a diagonal band rather than a vertical one. Everything about the shape of the constraint in the parameter plane is visible here as a dependence of one curve on a second variable.
The wings, on a logarithmic depth scale, out to 6 half widths. Depth as a fraction of the depth at line centre, against distance from line centre in units of each profile's own half width, so that all three pass through a half at one and separate only afterwards. The Gaussian falls as 2^(−x²), the Lorentzian as 1/(1+x²) — asymptotically x⁻², a power law where the other is an exponential — and the rotational profile is exactly zero past 1.28. The collisional wing is 3.0 times the thermal one at two half widths, 49 times at three, 3744 at four and 1.27·10⁶ at five: six orders of magnitude at five half widths, on profiles whose equivalent widths are identical to 0.1%. That is why a stellar spectrum's damping wings measure a pressure while its core measures almost nothing, and why 11% of a Lorentzian's equivalent width lies beyond the right-hand edge of this plot.
Fig. 5 Why microturbulence is needed at all: the shape of a line’s wings against the shape its thermal and pressure broadening alone predict. Real lines are broader, the excess is real, and attributing it to an isotropic small-scale velocity is a description rather than an explanation — the actual cause is the convective velocity field, which is neither isotropic nor small-scale. The parameter exists because a one-dimensional model has nowhere else to put a three-dimensional effect, and it is the clearest case in the analysis of a fitted quantity that corresponds to nothing.

One more parameter belongs on the list and is usually swept into the others: the macroturbulence, a second fudge velocity representing large-scale motions that broaden a line without changing its equivalent width. It is degenerate with the projected rotation, which does the same thing with a different profile shape, and separating the two requires a resolution and signal-to-noise that most spectra do not have. So the parameter set is really six, two of which are inventions, and the conditions available to fix them are four. The gap is closed by fixing the inventions at values from a calibration relation fitted to other stars — which is to say by importing an answer.

What was actually measured

The size of the degeneracy is measured rather than argued, in three ways.

Comparison against model-independent parameters. For a handful of stars the temperature is known without spectroscopy — from an interferometric angular diameter combined with a bolometric flux, which gives the effective temperature almost directly — and the gravity is known from an asteroseismic mass and radius or from an eclipsing binary. Comparing spectroscopic parameters against those benchmarks measures the systematic, and the answer is that spectroscopic temperatures are good to about a hundred kelvin and spectroscopic gravities to about 0.1 dex, both several times the quoted internal precision.

The spread between pipelines. Large surveys run several analysis codes on the same spectra. The spread between them is a direct measurement of the systematic from the degeneracy, because the codes differ mainly in how they weight the diagnostics — which is to say in where along the degenerate direction they land.

Asteroseismic gravities. For stars with measured oscillations the gravity follows from the frequency of maximum power almost without a model, to about 0.01 dex. Fixing the gravity at that value and re-solving the spectrum removes one axis of the degeneracy entirely, and the temperature’s uncertainty falls by more than a factor of two — which is the direct experimental demonstration that the temperature error was mostly degeneracy rather than measurement.

A strong line at three gravities, 6000 K. Left, one strong absorption line of Fe computed at three surface gravities and drawn on the same wavelength scale — no normalisation to each profile's own width, which is exactly what the comparison is about. The thermal core is identical in all three, because the Doppler width is 0.00223 nm at 6000 K whatever the star's size. The wings are not: collisional damping is proportional to the density of perturbers, that density is proportional to the gas pressure, and in a grey atmosphere the pressure at the photosphere is proportional to g — so log g = 4.4 carries a damping parameter 794 times that of log g = 1.5. Right, the width at a tenth of the core depth against gravity, measured off those curves: a slope of 0.499, against the half a Lorentzian wing forces. That is the second dimension of a spectral classification. The first is the temperature, read from which lines are present; this one is the pressure, read from how wide they are, and it is the whole reason a spectrum can be turned into an absolute magnitude and then into a distance.
Fig. 6 The steep diagnostic itself: the pressure-broadened wings of a hydrogen line at several gravities. The core saturates and stops responding, so the information is entirely in the wings — which is why this diagnostic needs a good continuum, and why a continuum placed a few per cent low moves the inferred gravity. The two systematics are not independent, and an error in the continuum shifts the answer along one of the degenerate directions rather than perpendicular to it.
The Balmer maximum is at 9,870 K, and it is a maximum in temperature. What two absorption lines actually count, each normalised to its own maximum. The first curve is the fraction of all hydrogen sitting in n = 2 — the only hydrogen a Balmer line can absorb — which is a Boltzmann factor climbing with temperature multiplied by the neutral fraction falling with it. The product peaks at 9,870 K, where 35.2% of the hydrogen is still neutral and only 8.73·10⁻⁶ of all of it is in n = 2 at all. Below the peak there is plenty of hydrogen and almost none of it excited; above it there is plenty excited and almost none of it neutral. The second curve is the fraction of calcium that is singly ionised, which is what the Ca II K line counts, and it peaks at 6,335 K — cooler, because calcium gives up its first electron at 6.113 eV. At 6,000 K the Balmer curve is at 0.1% of its own maximum while Ca II is near its peak, and calcium is 4.5·10⁵ times rarer than hydrogen in the same gas. A spectrum in which Ca II K is the strongest line is not a spectrum of a calcium star. Reading it as one is precisely the error that had the Sun made of iron until 1925.
Fig. 7 The second half of the same physics, and the one that gives the cleanest thermometer: the population of an excited level against temperature, which is a Boltzmann factor and depends on the gravity hardly at all. That near-independence is what makes excitation balance a nearly vertical band in the parameter plane, and therefore a good complement to the hydrogen wings — the two are close to orthogonal, which is the property the previous sections say is worth more than precision.

Where the picture stops

There are three, and the second is the one that has changed the field.

The models are one-dimensional and in equilibrium. Classical analyses use plane-parallel model atmospheres in local thermodynamic equilibrium. Real atmospheres are three-dimensional, convective and in places far from equilibrium, and correcting for both moves parameters by amounts comparable to the degeneracy — sometimes in the same direction, so the two systematics do not simply add.

Fixing a parameter externally is worth more than measuring it better. This is the practical lesson of the asteroseismic result and it generalises. Any external constraint that removes one axis — a parallax giving the luminosity and hence a gravity, an interferometric diameter giving a temperature, a cluster membership giving a metallicity — is worth more than a large improvement in the spectrum, because it changes the shape of the problem rather than the size of the errors.

And the degeneracy has a direction that matters downstream. An error along the degenerate direction moves temperature and gravity together in a correlated way, and quantities computed from both — a mass, a radius, an age from an isochrone — inherit the correlation. Propagating independent errors on the two parameters, which is what is usually done, understates the uncertainty on some derived quantities and overstates it on others.

A fourth limit is worth recording because it is the reason the systematic does not average away over a survey. The degeneracy has the same orientation for every star of a given type, so an analysis that lands slightly off along the degenerate direction lands off in the same direction for the whole sample. The scatter between stars looks fine, the internal precision looks excellent, and the sample as a whole is displaced — which then appears as a discrepancy against an isochrone, or as a trend of derived age with temperature, or as a metallicity gradient that is really a temperature gradient. The same failure mode as a systematic in any calibration: invisible in the internal scatter and present in full in the answer.

Why this is a rung on the spectrum ladder

The second thing a spectrum says, after composition, is the physical state of the gas that produced it, and this essay is about why that reading is not direct.

The general shape is one the collection meets in every field: an observable that depends on two quantities determines only a combination of them, and separating the two requires a second observable with a different dependence. What is distinctive about the spectroscopic case is the number of quantities — four, in the standard treatment, with only four conditions to fix them — and the fact that the conditions were chosen for convenience rather than for orthogonality. Nobody designed the ionisation balance to be a poor complement to the hydrogen wings; it is simply what the atomic physics offers.

The response the field has converged on is instructive. Rather than looking for better spectroscopic diagnostics, it has gone outside spectroscopy: to interferometry for radii, to asteroseismology for gravities, to parallaxes for luminosities, to eclipsing binaries for masses. Each of those breaks the degeneracy from a direction spectroscopy cannot reach, and each is a much harder measurement made on many fewer stars — which are then used to calibrate the spectroscopic analysis of the many.

That is the same structure as the distance ladder, and it has the same vulnerability — a bias at the anchor propagates to everything calibrated on it: a small number of hard, direct, model-light measurements anchoring a large number of easy, indirect ones. The diagram that sorted the stars was built from spectroscopic classifications long before any of its axes could be measured absolutely, and the modern version of the same enterprise is still calibrating the easy measurement against the hard one.

One observation to end on about the vocabulary, because it hides the problem. The phrase “spectroscopic parameters” suggests a set of measurements read off a spectrum, in the way a radial velocity is. They are not measurements in that sense: they are the output of a fit in which several consistency conditions are imposed simultaneously on a model that is known to be an idealisation. Every number in the set depends on every condition and on the model, and changing the line list changes all of them together. The distinction matters most when such parameters are used as though they were independent inputs to something else — fed into an isochrone to get an age, or into a mass–radius relation to get a mass — because the input errors are correlated and the output error is computed as though they were not.

The honest summary is short. Hydrogen’s lines are strongest where hydrogen is not, and every other line in the spectrum is similarly a statement about the state of the gas rather than about its contents. Untangling the state from the contents is a four-parameter inversion with four non-orthogonal conditions, and its accuracy is set by the least orthogonal pair.

It is worth recording the practical consequence for anybody comparing two published parameter sets. Because the answer is the crossing point of several bands, two analyses using different diagnostics do not merely have different errors — they have differently shaped errors, oriented along different directions in the plane. Two temperatures agreeing to fifty kelvin can conceal gravities differing by a factor of three, and two gravities agreeing can conceal temperatures differing by two hundred. So the honest comparison is between the two-dimensional confidence regions rather than between the marginalised numbers, and the marginalised numbers are what catalogues contain. That is the reason large spectroscopic surveys publish their own homogeneous parameters rather than compiling them from the literature: a set of parameters derived one way is internally comparable, and a compilation of parameters derived several ways is not comparable at all.

One more consequence follows for anybody building a pipeline rather than reading its output: the diagnostics should be chosen for the angle between their bands, not for the tightness of either. A superbly precise diagnostic that happens to run parallel to another adds nothing at all, and a mediocre one crossing at a right angle halves the error region. That is a statement about design rather than about analysis, and it is why the best parameter determinations combine a spectroscopic measurement with a photometric one — the two are individually worse and they cross steeply.

Where the ladder goes next

The rung directly above is the three-dimensional atmosphere: what changes when convection is computed rather than parameterised, why the microturbulence fudge disappears, and how much of the classical parameter offsets it accounts for. Further up sits the inverse problem — determining a composition once the state of the gas is fixed, and what the residual degeneracy does to the abundances.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 9 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

CovarianceDegeneracyEffective temperatureExcitation balanceIonisation balanceMicroturbulencePressure broadeningSaha equationSpectroscopic parametersSurface gravity