Stars

Two scaling relations calibrated on one star

Asteroseismology gives a star's mass and radius from two numbers read off its oscillation spectrum. The two relations are exact in their exponents and approximate in their constants, and the constants were fixed by requiring that the Sun come out right — so an error of a few per cent in one observable is tens of per cent in a mass and nearly a factor in an age.

Assumes Asteroseismology and Stellar evolution.

Asteroseismology has given masses and radii for hundreds of thousands of stars, which is more than every other technique combined by three orders of magnitude. It does so with two numbers extracted from an oscillation spectrum, and the extraction is robust: both numbers are frequencies, both are measured to parts in a thousand, and neither depends on the star’s distance, its brightness, or its spectrum.

The step that is not robust is what happens next. Two numbers become a mass and a radius through two relations whose exponents are exact and whose constants are fitted to a single star.

4 per cent in one observable is 48 per cent in an age. How an error in the calibration of the large frequency separation propagates into the quantities derived from it. The two scaling relations are exact in their exponents, so a fractional error in the separation appears as twice that in the radius, four times in the mass, and — because a main-sequence lifetime falls as roughly the two-and-a-half power of the mass — ten times in an age. At the 4 per cent level, which is about what the theoretical corrections to the relation amount to for a red giant, that is 15 per cent in mass and 48 per cent in age. Nothing about the seismology is uncertain at that level; the frequencies are measured to parts in a thousand. What is uncertain is the constant of proportionality, and it is uncertain because it was calibrated on one star.
Fig. 1 How an error in the calibration propagates. The relations give a fractional error in the large separation appearing as twice that in the radius, four times in the mass, and — because a main-sequence lifetime falls as roughly the two-and-a-half power of the mass — ten times in an age. Nothing about the seismology is uncertain at that level. What is uncertain is the constant of proportionality, and it is uncertain because it was calibrated on one star.

The situation is unusual in observational astronomy and worth naming for what it is: a technique whose measurement precision exceeds its accuracy by two orders of magnitude, with the entire gap sitting in two constants. The interior read from a comb of frequencies is a genuinely direct probe; what this essay is about is the last step, in which that probe’s output is compressed into two numbers and multiplied by a constant taken from somewhere else.

Where the two relations come from

The large separation. The high-order acoustic modes of a star are nearly evenly spaced in frequency, and the spacing Δν\Delta\nu is the inverse of twice the sound travel time from the centre to the surface. Since the sound speed goes as the square root of the temperature, and the temperature profile is set by hydrostatic equilibrium, the travel time scales as the square root of the mean density. So

ΔνρM/R3.\Delta\nu \propto \sqrt{\rho} \propto \sqrt{M/R^3}.

That relation is not an approximation in any deep sense; it is an asymptotic result that follows from the wave equation, and its accuracy for a real star is limited by how far the star departs from the homology the derivation assumes.

The frequency of maximum power. Oscillations are excited by convection and damped by it, and the envelope of the excited modes peaks at a frequency νmax\nu_{\rm max}. The scaling here is on much weaker ground: it is argued that νmax\nu_{\rm max} tracks the acoustic cut-off frequency of the atmosphere, which is proportional to g/Teffg/\sqrt{T_{\rm eff}}, so

νmaxMR2Teff.\nu_{\rm max} \propto \frac{M}{R^2\sqrt{T_{\rm eff}}}.

That argument is plausible, it is supported empirically across five orders of magnitude in νmax\nu_{\rm max}, and it has no derivation of the same standing as the first.

Solving the pair gives Mνmax3Δν4Teff3/2M \propto \nu_{\rm max}^3 \Delta\nu^{-4} T_{\rm eff}^{3/2} and RνmaxΔν2Teff1/2R \propto \nu_{\rm max}\Delta\nu^{-2}T_{\rm eff}^{1/2}. The fourth power in the mass is where the trouble is.

14 radial orders of the Sun, at 135.1 μHz apart. The p-mode spectrum of the Sun — 1 solar mass in 1 solar radius — from the asymptotic relation with its second-order term, drawn as 14 radial orders of ℓ = 0, 1 and 2 under a Gaussian envelope centred on ν_max = 3,090 μHz. Two numbers are marked and they do very different work. The large separation, 135.1 μHz, is the spacing between consecutive ℓ = 0 modes and fixes the mean density. The small separation, 9.00 μHz, is 2.7 pixels on this axis — it fixes the age, and it is why the échelle diagram exists rather than being a convenience. The vertical axis is the measurement: each mode moves the surface by about 20.0 cm s⁻¹ at the peak, and brightens it by a few parts per million, which is why this was impossible before a decade-long velocity series. Each mode is one line: its true width is set by its lifetime and is far below a pixel here.
Fig. 2 The observable itself: a power spectrum with a comb of nearly evenly spaced peaks under a broad envelope. The spacing of the comb is the large separation and the centre of the envelope is the frequency of maximum power. Everything in this essay is what happens after those two numbers are read off — the reading is the easy part, and it is done to a fraction of a per cent from a few months of photometry.

The asymmetry between the two relations is worth dwelling on because it decides where the effort goes. The first is asymptotic theory: it follows from the wave equation for high-order acoustic modes, its derivation makes assumptions that can be checked, and its departures from exactness can be computed from a model. The second is a scaling argument with an empirical constant. So the two ingredients of every seismic mass have completely different epistemic status, and the weaker one enters the mass with a cube.

There is a useful way to see why the exponents come out as they do. The two observables are, in effect, a mean density and a surface gravity: the first relation is a density and the second is g/Tg/\sqrt{T}. A density is M/R3M/R^3 and a gravity is M/R2M/R^2, so extracting MM and RR from the pair means solving two power laws whose exponents differ by one in RR and not at all in MM. That near-parallelism is exactly the ill-conditioning of the previous essay’s crossing angle, in a different setting: two constraints that respond similarly to one variable determine it poorly, and the amplification factors of four and two are the numerical statement of how similarly.

The exponents amplify, and they amplify differently

The propagation is exact and it is worth having in front of one.

A fractional error ff in Δν\Delta\nu gives 2f-2f in the radius and 4f-4f in the mass. A fractional error in νmax\nu_{\rm max} gives +f+f in the radius and +3f+3f in the mass. A fractional error in the temperature gives f/2f/2 and 3f/23f/2.

So a per cent error in the large separation is a four per cent error in the mass, and — since a main-sequence lifetime scales as roughly M2.5M^{-2.5} — a ten per cent error in an age. The relation is a lever with a long arm.

The consequence for practice is that the radius is the trustworthy output and the mass is not. Seismic radii agree with interferometric radii to a per cent or two; seismic masses agree with dynamical masses to five or ten per cent; seismic ages are quoted with uncertainties of twenty per cent and are probably worse.

1.5 per cent in one observable is 16 per cent in an age. How an error in the calibration of the large frequency separation propagates into the quantities derived from it. The two scaling relations are exact in their exponents, so a fractional error in the separation appears as twice that in the radius, four times in the mass, and — because a main-sequence lifetime falls as roughly the two-and-a-half power of the mass — ten times in an age. At the 1.5 per cent level, which is about what the theoretical corrections to the relation amount to for a red giant, that is 6 per cent in mass and 16 per cent in age. Nothing about the seismology is uncertain at that level; the frequencies are measured to parts in a thousand. What is uncertain is the constant of proportionality, and it is uncertain because it was calibrated on one star.
Fig. 3 The same propagation at a smaller calibration error, which is roughly where the relations stand for main-sequence stars. Even there the age error reaches fifteen per cent, from an observable measured to a tenth of a per cent. The gap between what is measured and what is delivered is entirely in the exponents, and no improvement in the photometry closes it.
Where each mode turns back: ℓ = 0 through the centre, ℓ = 300 in the outer 2 per cent. Why a set of frequencies is a depth profile and a single frequency is not. An acoustic wave travelling into a star meets a rising sound speed and is refracted back; it turns where its horizontal phase speed matches the local sound speed, which happens at c(r)/r = 2πν/√(ℓ(ℓ+1)). The horizontal axis is the angular degree on a logarithmic scale and the vertical axis is the fractional radius of that turning point, drawn at 2000, 3090, 4000 microhertz. The ordering is the content. A radial mode, ℓ = 0, has no horizontal phase speed at all and passes straight through the centre. Degrees one and two turn deep in the core. By ℓ = 300 the mode is trapped in the outer 2 per cent and knows nothing about anything below. So a frequency measured to a part in ten thousand constrains an average of the interior weighted in a way the mode itself decides, and measuring thousands of modes of different degree gives thousands of differently weighted averages — which is a solvable inverse problem, and is how the base of the convection zone was located at 0.713 of the radius rather than assumed. The sound speed here is a polytrope's rather than a tabulated solar model's, so the curve is the right shape and the wrong star in its outer tenth, where the real Sun is convective and this one is not. What the picture cannot show is the frequency dependence at fixed degree, which is weaker but not negligible: a higher-frequency mode of the same degree turns slightly deeper, and the three curves separating toward the right is that effect.
Fig. 4 What the frequencies are actually sensitive to, and why compressing them into two numbers throws so much away: each mode’s sensitivity to the sound speed at each depth. Different modes probe different regions, and the ensemble constrains a profile rather than a mean. The scaling relations use only the average spacing and the envelope centre, which is two numbers out of a spectrum containing dozens of independent frequencies — an enormous compression, made because it requires no model, and paid for in exactly the way this essay describes.

There is a corollary worth acting on. Because the radius is the better-determined of the two outputs, any argument that can be made with a radius should be made with a radius rather than with a mass. A planet’s radius measured from a transit depth needs the star’s radius and not its mass; a surface gravity needs both but combines them in a way that partly cancels; a mean density needs neither separately. Reformulating a question so that it asks for the well-determined combination is worth more than any improvement in the data, and it is available surprisingly often.

Why the constant is not one

The proportionalities above become equations by inserting the solar values, which amounts to assuming that a star with the Sun’s mean density has the Sun’s large separation. That is true to the extent that the star is homologous to the Sun, and it is less true the further the star is from being a solar analogue.

Two departures matter.

The mass distribution. A red giant has an enormously centrally condensed structure, with a dense degenerate core and a vast tenuous envelope. Its mean density is the same quantity as a dwarf’s mean density and its sound-travel time is not related to it in the same way, so the constant in the first relation shifts — by a few per cent for a giant, in a way that depends on its evolutionary state.

The surface. The outer layers of a star are where the model is worst: convection is inefficient there, the treatment is one-dimensional, and the frequencies computed from a model are systematically too high by several microhertz. That is the surface effect, it affects all modes in a way that depends on frequency, and it biases the fitted Δν\Delta\nu.

Both are corrected by comparison against theoretical models, which reintroduces the model dependence the technique was prized for avoiding. Corrections of two to four per cent in Δν\Delta\nu are typical for giants, and the four in the exponent turns that into ten to sixteen per cent in mass.

The same comb folded at 135.1 μHz — three ridges and their curvature. Frequency against frequency modulo Δν, for 14 radial orders of the Sun. Folding at 135.1 μHz stacks the orders into three near-vertical ridges, one for each degree, and that is what makes Δν a fact about the data rather than a fitted parameter: get it wrong and the ridges lean. What is left over is the curvature — the ℓ = 0 ridge wanders 6.7 per cent of Δν across the drawn range, which is the departure from the asymptotic relation and the part of the spectrum that knows about the star's outer layers. The small separation is here too and here it is visible: 9.00 μHz between the ℓ = 0 and ℓ = 2 ridges, 39 pixels on this axis against 39 — the same quantity that is three pixels wide in the unfolded spectrum. A constant offset of 27 μHz has been subtracted before folding so that no ridge wraps round the edge.
Fig. 5 Where the surface effect shows up: the échelle diagram, in which the spectrum is folded at the large separation so that each ridge is a sequence of modes of one angular degree. A perfectly homologous star gives vertical ridges; a real one gives ridges that curve, and the curvature at high frequency is the surface effect. That curvature is what makes the fitted large separation depend on which modes were included in the fit, which is one of the two ways the calibration becomes a choice.

There is a third departure that is smaller and instructive: the relations are usually written with a temperature in them, and the temperature comes from spectroscopy or photometry with its own systematic. A hundred kelvin — which is about the accuracy of a spectroscopic temperature — is 1.7 per cent, which enters the mass with a power of three halves, giving 2.6 per cent. So a technique advertised as distance-independent and model-independent carries a spectroscopic systematic in its mass at the level of a few per cent, and there is no version of the relations that avoids it.

It is worth being explicit about what “calibrated on the Sun” means operationally, because it sounds like a minor convention. The solar values of the two observables are known to better than a tenth of a per cent, so the calibration introduces no measurement error at all. What it introduces is an assumption of homology: the claim that a star of a different mass, composition and evolutionary state relates its observables to its mass and radius by the same constants the Sun does. That assumption is testable only by finding stars with independent masses, and the number of such stars with detected oscillations is a few dozen. So the calibration rests on one star for its value and on a few dozen for its transferability, against a catalogue of hundreds of thousands.

What was actually measured

The relations are tested wherever a mass or radius is available independently, and the tests are the reason the size of the problem is known.

Eclipsing binaries with oscillating components. A handful of systems have a red giant showing oscillations in a binary with a measurable dynamical mass. The seismic masses come out systematically high by about five per cent before correction and agree to within a few per cent after the model-based corrections are applied — which is a validation of the corrections rather than of the raw relations.

Interferometric radii. For bright nearby dwarfs the angular diameter is measurable directly, and combined with a parallax it gives a radius with no seismology in it. Seismic radii agree to about a per cent, which is the strongest evidence that the first relation is sound.

Cluster members. Giants in a cluster share an age and a distance, so their seismic masses should agree with each other and with the cluster’s turn-off mass. They do, to within the scatter, and the comparison constrains the mass scale to a few per cent.

And the distance test. A seismic radius plus an effective temperature gives a luminosity, which with an apparent magnitude gives a distance — a route entirely independent of parallax. Comparing that against measured parallaxes tests the radius scale, and it is one of the ways the parallax zero point itself has been checked.

The plane the two relations make, and how badly it is conditioned in mass. The large separation against the frequency of maximum power, both logarithmic, for four stars whose masses, radii and temperatures are stated and whose frequencies are computed from them — the Sun at 3,090 and 135 μHz; a subgiant at 671 and 39.0 μHz; a red-clump star at 34 and 4.06 μHz; a red giant at 4.43 and 0.862 μHz. Over the plane are lines of constant radius, at 1, 3, 10, 30 solar radii, and of constant mass, at 0.8 and 2. The two families are not at right angles and that is the finding: the constant-mass lines run at a slope of about 0.77 and the constant-radius lines at 0.5, so a factor of 2.5 in mass moves a star only 0.11 decades across the plane. Mass enters the inversion at the quarter power of the observables and radius at the first, which is why an asteroseismic radius is good to a few per cent and an asteroseismic mass to something nearer ten. Every point here was put on the plane by the relations and then read back off it: solving the two equations for mass and radius returns the stated values to 7.8e-16.
Fig. 6 The relations drawn as the scalings they are, across the range of stars they are applied to. Five orders of magnitude in the frequency of maximum power separate a red giant from a main-sequence dwarf, and the relations are used across all of it — calibrated at one point, near the middle in logarithm and nowhere near the middle in structure. That the extrapolation works as well as it does is the surprising part; that it needs percent-level corrections at the ends is not.
Three clusters, three ages, one diagram. Isochrones for populations of 100 Myr, 1 Gyr, 12 Gyr, each drawn as the main sequence up to its own turnoff and then the post-main-sequence track of the turnoff mass. The turnoffs are at 5.52 M☉, 2.20 M☉, 0.94 M☉, from t = 10¹⁰ M/L with this file's own mass–luminosity relation. Nothing here is a track along which a star moves. Every point is a different star of a different mass, all the same age, and the bend is simply where the population runs out of stars that have had time to leave. That is why a cluster has an age and a field star does not: the bend needs a population, and one star is not one. The oldest globular clusters sit near the 12 Gyr line, and in the 1990s the same construction gave them ages of 16 to 18 Gyr against a universe measured at 10 — a two-standard-deviation contradiction that was resolved from the distance side, by Hipparcos, and not from this one.
Fig. 7 What the masses and radii are for. Placing a star on a set of isochrones with a known mass and radius fixes its age far better than a position in temperature and luminosity alone, because the isochrones are widely separated in radius and closely spaced in temperature. That is the whole argument for seismic ages, and the propagation in this essay is the argument against believing them at the quoted precision.

Where the picture stops

Three limits stand out, and the second is the one being worked on.

The second relation has no derivation. The scaling for νmax\nu_{\rm max} rests on an argument about the acoustic cut-off that is dimensionally right and quantitatively unestablished. It works empirically over an enormous range, which is evidence that something is right about it, and there is no theory that says the constant should be the same for a giant as for a dwarf.

Individual frequencies are better than the scalings. Fitting the actual mode frequencies against a stellar model — rather than compressing them into two numbers — uses far more information and gives masses and ages several times more precise. It also requires a model, so it trades one dependence for another, and it is only possible for stars with high enough signal-to-noise to resolve individual modes.

And the ages are the weakest output and the most wanted. Galactic archaeology needs ages for large samples, seismology is the only technique that can supply them in bulk, and the propagation above means that a ten per cent age is close to the best that scaling relations can do — which is not good enough to distinguish the formation epochs of the Galaxy’s components.

And a fourth, and it is about the population the technique reaches. Oscillations are detectable when their amplitude exceeds the photometric noise, and the amplitude scales steeply with luminosity — so seismology works easily on red giants, with difficulty on solar-type dwarfs, and not at all on anything hotter than about the Sun’s temperature, where convective envelopes are too thin to excite modes. The sample is therefore selected by evolutionary state, which correlates with mass and age, and any population statistic built from it inherits that selection. What a survey could have seen has to be modelled before what it found means anything, and seismic samples have an unusually sharp and unusually well-understood selection boundary.

Why a calibrated relation is a different object from a derived one

The general point is one this collection meets repeatedly, and it is worth stating cleanly.

A relation with derived exponents and a fitted constant behaves quite differently from one that is derived throughout. Its differential use is excellent: comparing two stars, the constant cancels and the exponents do the work, so relative masses and relative radii are far better determined than absolute ones. Its absolute use inherits everything about the calibrating object, and there is only one calibrating object.

That is the same structure as a Coulomb logarithm calibrated against simulations and as a mixing length calibrated on the Sun: a piece of physics whose form is understood and whose normalisation is imported. Such a relation is enormously useful and it is not a measurement in the sense a parallax is, and the distinction matters most when the result is being compared against a prediction that shares the same calibration.

The escape, where one is available, is to find a second calibrator of a different kind. For the seismic relations that is what the eclipsing binaries and the interferometric radii provide, and the programme of finding more such benchmarks is the main line of work in the field — not better photometry, of which there is already enough, but a handful of stars measured a completely different way.

One remaining observation about why the technique is nevertheless transformative. Before seismology, a field star’s mass was not measurable at all — it was inferred from a position in a colour–magnitude diagram, which requires a distance and a model and gives a factor. Seismology gives it to five or ten per cent, for hundreds of thousands of stars, from photometry alone. That a five per cent mass is a poor measurement by the standards of an eclipsing binary is true and beside the point: the only stars whose masses were known numbered a few hundred, and they were not a representative sample of anything. A worse measurement of a much larger and fairer sample is a different kind of instrument, and most of what has been learned about the Galaxy’s stellar populations in the last fifteen years came from it.

A final observation on how the field has responded, because the response is instructive. Rather than trying to derive the second relation, the effort has gone into building a set of benchmark stars — objects with masses and radii from eclipsing binaries, interferometry, or cluster membership, and with detected oscillations — against which the relations are calibrated empirically as a function of evolutionary state. That converts a theoretical problem into an observational programme, it works, and it changes the character of the technique: the relations become an interpolation across a calibrated grid rather than a piece of physics. The distance ladder has the same architecture, and it has the same vulnerability — a systematic in the benchmarks propagates to everything, undiluted, and cannot be found from within.

One further consequence is worth carrying, because it decides how such masses should be used in a population study. The calibration is a single multiplicative constant applied to every star, so an error in it moves every seismic mass in the same direction by the same fraction — it is a shared systematic rather than a random one. A sample of ten thousand seismic masses therefore has a statistical precision far better than its accuracy, and averaging does not help at all. That matters most for the quantities derived by comparing a sample against a model: a galactic age distribution built from seismic ages inherits the calibration’s error as a rigid shift of the whole distribution, which looks like a real feature and is not. The remedy is the usual one and it is being pursued — anchor the constant on stars whose masses are known by other means, which for this purpose means eclipsing binaries and interferometrically resolved stars, and there are a few dozen of them.

End on what makes the relations worth using despite all of this. They require two numbers read off a power spectrum and nothing else — no distance, no spectrum, no model of the star’s interior — and they deliver a mass and a radius for any star with a long enough light curve. Nothing else in stellar astrophysics does that, which is why an approximate constant fitted on one star has been applied to hundreds of thousands.

Where the ladder goes next

The rung directly above is the surface effect: what it is physically, why one-dimensional models get the outer layers wrong, and how the empirical corrections are constructed. The one above that is what replaces the scalings — fitting individual mode frequencies against models, and what that costs in model dependence for what it buys in precision.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Benchmark starError propagationFrequency of maximum powerLarge separationRed giantScaling relationSolar calibrationStellar massSurface effectSystematic error