Stars

A length nobody derived, fitted to one star

Convection in a star is turbulent, three-dimensional and impossible to compute inside an evolution code. What is used instead is one number — how far a blob of gas travels before dissolving — fixed by requiring that a model of the Sun come out with the Sun's radius, and then applied to every star ever modelled.

Assumes Energy transport and Hydrostatic equilibrium.

A star’s outer third is a boiling fluid. Energy arrives at the base of that region as radiation and leaves the top as radiation, and in between it is carried by the bulk motion of gas — hot material rising, cool material sinking, on scales from the size of a granule to the depth of the whole convection zone.

Computing that flow is a three-dimensional, compressible, radiative hydrodynamics problem with a Reynolds number of 101210^{12} — a harder calculation than any of the fluid problems this collection has met, and one that has to be embedded inside a calculation of the whole star’s evolution. It cannot be done inside a stellar evolution code, which has to integrate a star’s structure through a million time steps.

A free parameter worth 88 kelvin across its plausible range. The effective temperature a stellar model predicts, against mass, for three values of the mixing-length parameter. The parameter has no derivation: it is the distance a convective blob is supposed to travel before dissolving, in units of the local pressure scale height, and it is fixed by requiring that a model of the Sun reproduce the Sun. The three curves span 88 kelvin, which at fixed luminosity is a radius difference of 1.5 per cent — comparable to the precision with which radii are now measured by interferometry and by eclipsing binaries. Every stellar age, every isochrone and every mass inferred from a position in the temperature–luminosity plane depends on the value chosen, and there is no reason beyond convenience to expect the solar value to apply to a red giant or to a metal-poor dwarf.
Fig. 1 What is used instead: one parameter, the distance a convective blob travels before dissolving, in units of the local pressure scale height. The three curves span the range of values in current use, and they differ by nearly ninety kelvin in effective temperature at a solar mass — which at fixed luminosity is a radius difference comparable to the precision with which radii are now measured.

So a quantity that decides the radius, the effective temperature and the evolutionary track of every star with a convective envelope — which is every star cooler than about 7,000 kelvin, and the outer layers of most others — is a single fitted number. A star is held up by its own weight is a statement that needs no free parameters; how the heat gets out is not.

What the parameter stands in for

The mixing-length theory is a piece of dimensional analysis dressed as a physical model. A blob of gas is displaced upward, it is hotter than its surroundings because the star’s temperature gradient is steeper than the adiabatic one, so it is buoyant and accelerates. After travelling a distance \ell it dissolves and gives up its heat.

The convective flux follows from the excess temperature, the velocity and the density. The velocity follows from the buoyant acceleration acting over \ell. And \ell is set to αHp\alpha H_p, with HpH_p the local pressure scale height and α\alpha a number of order one that nothing in the theory determines.

There is no derivation of α\alpha, there is no reason it should be the same at every depth in one star, and there is no reason it should be the same in two different stars. It is a fitting parameter with the dimensions of a length, and everything about the outer layers of every stellar model depends on it.

It is worth being clear about the scale of the approximation. The theory replaces a flow with structure across ten orders of magnitude in length — from the depth of the convection zone down to the viscous dissipation scale — with a single length and a single velocity, both local. It is the crudest kind of closure, it has no small parameter to expand in, and it is not obviously better than several alternatives that were proposed and abandoned. What recommends it is that it is one number and that it fits.

Why it controls the radius

The reason α\alpha matters at all is that convection in a star’s deep interior is almost perfectly efficient. The gas carries so much heat that only a minute excess over the adiabatic gradient is needed, so the temperature profile there is adiabatic whatever α\alpha is, and the answer is insensitive.

Near the surface it stops being efficient. The density falls, the heat capacity per unit volume with it, and the gas has to be substantially superadiabatic to carry the flux. How superadiabatic depends on α\alpha — a longer mixing length carries the flux more easily and needs less of a gradient.

That superadiabatic layer is thin, and it sets the entropy of the whole adiabatic interior below it, which sets the radius at which the star’s structure matches its surface boundary condition. So a parameter describing a few hundred kilometres of a star’s outer skin fixes the radius of the entire object.

A free parameter worth 44 kelvin across its plausible range. The effective temperature a stellar model predicts, against mass, for three values of the mixing-length parameter. The parameter has no derivation: it is the distance a convective blob is supposed to travel before dissolving, in units of the local pressure scale height, and it is fixed by requiring that a model of the Sun reproduce the Sun. The three curves span 44 kelvin, which at fixed luminosity is a radius difference of 0.8 per cent — comparable to the precision with which radii are now measured by interferometry and by eclipsing binaries. Every stellar age, every isochrone and every mass inferred from a position in the temperature–luminosity plane depends on the value chosen, and there is no reason beyond convenience to expect the solar value to apply to a red giant or to a metal-poor dwarf.
Fig. 2 A narrower range of the parameter — the range a single calibration is usually quoted to — showing that even a tenth in α is forty kelvin. Interferometric radii for nearby dwarfs are now good to about a per cent, which is fifteen kelvin at fixed luminosity, so the measurement is more precise than the parameter is known. That is the state that turns a convenient fudge into a limiting systematic.

There is a corollary that makes the sensitivity easier to remember. The luminosity of a low-mass star is set by its core, which does not care about α\alpha; the radius is set by the envelope, which does. Since L=4πR2σTeff4L = 4\pi R^2\sigma T_{\rm eff}^4, fixing LL and changing RR changes TeffT_{\rm eff}, and the change appears as a horizontal shift of the whole evolutionary track in the temperature–luminosity plane. That is why an age read off the bend of a cluster’s main sequence depends on the parameter: the bend’s position in temperature moves with it, and matching an observed sequence to a model sequence is matching a temperature.

There is a second reason the outer layers dominate, and it is geometric rather than thermodynamic. The superadiabatic region is a few hundred kilometres thick in a star seven hundred thousand kilometres in radius — a part in two thousand. Its thermodynamic role is to set the entropy of everything below it, and entropy is conserved along an adiabat, so a change made in that thin skin propagates unchanged through the entire interior. A star is unusually vulnerable to its own surface in a way that most physical systems are not, and the mixing-length parameter is the lever that acts there. The same leverage makes a spectroscopic gravity so sensitive to the treatment of the outer layers, for the same reason.

The calibration, and what it assumes

The standard procedure is to build a model of one solar mass with the solar composition, evolve it for 4.57 billion years, and adjust α\alpha — together with the initial helium abundance — until the model reproduces the Sun’s radius and luminosity.

That gives a value near 1.8 for most codes, and the value is code-dependent because it absorbs whatever else that code gets wrong about the outer layers: the equation of state, the low-temperature opacities, the treatment of the atmosphere.

Two assumptions are then made without much comment. That the calibrated α\alpha applies to stars of other masses, other metallicities and other evolutionary states. And that the same α\alpha applies at every depth within one star.

Neither is obviously true and both are now testable, because three-dimensional simulations of stellar surface convection can be run for a grid of stellar parameters and the effective α\alpha read off by matching a one-dimensional model to the simulation’s entropy profile.

A free parameter worth 78 kelvin across its plausible range. The effective temperature a stellar model predicts, against mass, for three values of the mixing-length parameter. The parameter has no derivation: it is the distance a convective blob is supposed to travel before dissolving, in units of the local pressure scale height, and it is fixed by requiring that a model of the Sun reproduce the Sun. The three curves span 78 kelvin, which at fixed luminosity is a radius difference of 1.4 per cent — comparable to the precision with which radii are now measured by interferometry and by eclipsing binaries. Every stellar age, every isochrone and every mass inferred from a position in the temperature–luminosity plane depends on the value chosen, and there is no reason beyond convenience to expect the solar value to apply to a red giant or to a metal-poor dwarf.
Fig. 3 The regime where the answer matters most: low-mass stars, whose radii have been measured in eclipsing binaries and found to be several per cent larger than models predict. A change in α of the size the simulations suggest moves the models in the right direction, and so do magnetic fields and spots, and disentangling the three is an open problem. This is the clearest case in stellar physics of a well-measured discrepancy with three plausible explanations and no way yet to choose between them.

Notice that the second assumption is known to be false in a specific way. The pressure scale height varies by orders of magnitude through a convection zone, so =αHp\ell = \alpha H_p varies with it — but the physical eddy size does not, near the top, where the granules have a size set by the depth at which they form rather than by the local scale height. So a constant α\alpha is certainly wrong in the layer that matters most, and it is used because a depth-dependent α\alpha would require knowing the depth dependence, which is the thing that was too hard to compute in the first place.

What the simulations found

The three-dimensional calculations give a definite answer and it is not a constant.

The effective α\alpha varies systematically across the temperature–gravity plane, by about ten per cent from the Sun to a metal-poor turn-off star and by more towards the giants. The variation is smooth, it is reproducible between independent simulation codes, and it goes in the direction of lower α\alpha for hotter and lower-gravity stars.

Ten per cent in α\alpha is about twenty kelvin at the Sun and more elsewhere, which is small compared with the older uncertainties and large compared with modern measurements. Adopting a varying α\alpha from the simulations changes stellar radii by a per cent or two and cluster ages by a few per cent — not a revolution, and not negligible either.

The simulations also say something the one-dimensional theory cannot express: the convective flux is carried by a small number of fast downdrafts rather than by a symmetric up-and-down exchange, which is a qualitatively different flow from the one the mixing-length picture describes. That the parameterisation works as well as it does, given that its physical picture is wrong in that specific way, is the genuinely surprising part.

Which gradient is steeper, and where. The temperature gradient radiation would need in order to carry the whole luminosity, against the gradient a rising blob of gas follows, through an n = 3 polytrope with a Kramers opacity and an energy generation going as ρT^4. Wherever the first exceeds the second the layer is unstable and convects. This model puts the outer boundary at 89 per cent of the radius; helioseismology measures the base of the Sun's convection zone at 71.3 per cent, and the gap between the two numbers is the price of a polytrope with one opacity law rather than a solar model. The curve also rises above the adiabatic value inside 4 per cent of the radius, which is a convective core — it appears when the energy generation is concentrated enough, and it is the structure a star burning by the CNO cycle has.
Fig. 4 The criterion the whole apparatus sits underneath: the comparison of the radiative and adiabatic temperature gradients that decides whether a region convects at all. That comparison is unambiguous and needs no free parameter — it is what decides that the Sun’s outer third boils and its core does not. Everything uncertain is downstream of it, in the question of how much heat the convection carries once it has started.
How long the light takes to get out. The time for energy to diffuse from the centre of a sphere the size of the Sun, against the mean free path of a photon inside it. A random walk of step ℓ covers a distance ℓ√N, so escaping a radius R takes (R/ℓ)² steps and (R/ℓ)²·ℓ/c of time — inversely proportional to the mean free path. At the centimetre or so a solar interior allows, that is tens of thousands of years, against 2.3 seconds for the same distance in a straight line.
Fig. 5 The alternative transport mechanism, for scale: a photon’s random walk out of a star, which takes tens of thousands of years and is what convection is competing against. Where radiation can carry the flux, it does, and the structure is determined without any free parameter at all. Convection takes over precisely where radiation cannot cope — where the opacity is high and the gradient would otherwise be too steep — and that is why the uncertain physics is confined to the outer layers of cool stars rather than being everywhere.

The third of these is the one that connects the parameter to something a telescope sees directly rather than to a model’s output, and it is worth ending the sequence on it.

The limb is 40% as bright as the centre, and that is a temperature gradient. Left: a stellar disc shaded by the grey-atmosphere law I(μ)/I(1) = (2 + 3μ)/5, in 26 steps, with μ = cos θ read from the centre outwards. Right: that law against μ, with linear laws at the measured solar coefficients from 400 to 1600 nm. The grey law's coefficient is exactly 3/5 and its very limb is exactly 2/5 of the central brightness — both read off the drawn curve rather than quoted — because the Eddington–Barbier relation makes the emergent intensity at angle μ the source function at optical depth τ = μ, and in radiative equilibrium that source function is linear in τ. The limb is not a cooler part of the star. A sight line entering at the edge reaches unit optical depth higher up, where the gas is cooler, so what the darkening measures is the run of temperature with depth; a star with an isothermal atmosphere would show a uniform disc, and one with a steeper gradient a darker limb. The measured coefficients fall from 0.9 at 400 nm to 0.35 at 1600 — the same gradient seen through a less steep Planck function — which is why a radius measured from a transit or a fringe null has to say which colour it was measured in.
Fig. 6 The one place the convection can be seen directly rather than parameterised: the solar surface, where granulation is resolved and its velocities, sizes and lifetimes are measured. Everything a mixing-length parameter is trying to summarise is visible there in detail — and the detail says the flow is asymmetric, intermittent and dominated by narrow downdrafts, which is not the picture the parameterisation is built on. The theory works despite describing the wrong flow, which is worth remembering when it is used somewhere the flow cannot be observed.

What was actually measured

Three classes of measurement bear on the parameter, and they disagree in a way that is itself informative.

Interferometric and eclipsing-binary radii. For stars with directly measured radii, comparing against models constrains α\alpha directly. The results are consistent with the solar value for solar-type stars and require lower values for hotter ones, in the direction the simulations predict.

The solar convection zone’s depth. Helioseismology measures the depth of the Sun’s convective envelope to a fraction of a per cent, and it is one of the constraints a solar model must satisfy. It is not very sensitive to α\alpha — the depth is set by the opacity and the composition — but it does constrain the combination of α\alpha and helium abundance, which is why the two are calibrated together.

Red giant temperatures. The effective temperature of a red giant branch is almost entirely set by α\alpha, because the envelope is convective and enormously extended. Requiring that models reproduce the observed giant branches of clusters gives α\alpha values that differ from the solar one, and the differences are the most direct evidence that a single value is inadequate.

A 1 solar-mass star from the main sequence to the tip, with the clock on it. The post-main-sequence track of a 1 solar-mass star: luminosity against surface temperature, temperature increasing to the left, both logarithmic. The main-sequence and subgiant durations are stated, at 10 and 2 billion years; the 0.99 billion years to climb the giant branch is not stated but integrated, because the shell converts hydrogen at exactly the rate the luminosity demands and the core-mass–luminosity relation is therefore a clock. The three sum to 12.99 billion years. The nine dots along the giant branch are 100 Myr apart and they crowd at the foot: the last doubling of the radius, from 61 to 122 solar radii, takes 4.8 Myr of the whole climb. That is why an observed colour–magnitude diagram shows a crowded main sequence and a sparse giant branch — the diagram is a histogram of durations, not a census of stars.
Fig. 7 Why the giant branch is the sensitive case: an evolutionary track whose position in temperature is set almost entirely by the convective envelope. A star on the giant branch has a structure in which nearly all the mass outside the core is convective and nearly adiabatic, so the surface temperature is fixed by the entropy of that adiabat — which is fixed by the superadiabatic layer at the top — which is fixed by α. The whole branch slides horizontally when the parameter changes, which is what makes it the best available laboratory and the worst available systematic.

There is also a fourth result worth stating because it undermines the parameter’s interpretation rather than its value. Comparing simulations run with different numerical resolutions and different codes shows that the effective α\alpha recovered depends slightly on how the matching to the one-dimensional model is done — on which quantity is required to agree, and at what depth. That is not a numerical error; it is a statement that “the mixing length” is not a well-defined property of the flow, only of the procedure used to summarise it. A parameter defined by a fitting procedure inherits the procedure, which is the same observation as the one about a Coulomb logarithm’s cut-off being a choice rather than a quantity.

One more consequence of the simulations deserves recording, because it is a genuine physical result rather than a calibration. The flow they show is not a symmetric exchange of rising and falling blobs. It is a broad, slow, warm upflow occupying most of the area and a set of narrow, fast, cold downdrafts occupying very little of it — a topology forced by the steep density stratification, since a rising parcel expands and a sinking one is compressed. The mixing-length picture has no way to express that asymmetry, and the asymmetry is what determines the overshooting at the base of the zone, the mixing of material below it, and the excitation of the oscillations the previous essay’s technique depends on.

Where the picture stops

The picture stops in three places, and the third is where the field is heading.

The parameter absorbs other errors. A calibration that adjusts α\alpha to fit the Sun will compensate for an error in the opacities, the equation of state or the atmospheric boundary condition by shifting α\alpha. So the calibrated value is not a measurement of a mixing length; it is a measurement of one number’s worth of everything the code gets wrong near the surface.

Rotation and magnetic fields are not in it. A magnetised, rotating convection zone carries flux differently, and the effect is in the direction of inflating a star. That is one of the leading explanations for the radius discrepancy in low-mass stars, and it is degenerate with α\alpha by construction — both act by changing the efficiency of convection.

And the replacement is coming from a different direction. Rather than improving α\alpha, the current work replaces the surface boundary condition of a one-dimensional model with a patched-on three-dimensional simulation. That removes the parameter’s most important role — setting the entropy of the adiabat — and leaves it doing much less. It is expensive, it is being done on grids rather than star by star, and it is the likely end of the parameter’s forty-year career as the largest single uncertainty in stellar structure.

A fourth deserves stating because it changes what can be claimed from a comparison. Because α\alpha absorbs the code’s other surface errors, two codes with different opacities and different atmospheres will calibrate to different α\alpha and then agree with each other about the Sun by construction. Their disagreement about a red giant is therefore a measurement of the difference between their absorbed errors, not of anything about convection — and quoting the spread between codes as an uncertainty on a stellar age counts that difference as if it were physical. The same objection applies to comparing two spectroscopic analyses that share a calibration: agreement between two methods that were tuned on the same object is not evidence.

Why one fitted number is worth an essay

The general shape is one this collection meets from several directions, and the repetition is the point.

A theory with derived structure and one fitted constant is enormously powerful and is not a measurement. The Coulomb logarithm in dynamical friction is one; the constants in the seismic scaling relations are another; the mixing length is a third. In each case the exponents and the functional form come from physics, the normalisation comes from a fit, and the result is used far outside the range the fit was made in.

Such a construction behaves well differentially and poorly absolutely, it hides other errors by absorbing them, and its failures show up as small systematic discrepancies rather than as visible breakdowns. That last property is the dangerous one: a model with a well-tuned fudge factor fits everything slightly and nothing exactly, and there is no single observation that refutes it.

What eventually resolves such a parameter is never a better fit. It is a calculation from below that does not need it — a simulation, a first-principles derivation, a measurement of the thing itself. That is what the three-dimensional convection calculations are, and it is worth noticing how long it took: the parameter was introduced in 1958 and the calculations that measure it became routine around 2010.

One remaining observation about what the parameter’s persistence says. It has been known to be inadequate since it was introduced, its replacement has been technically possible for over a decade, and it is still what nearly every published stellar model uses. The reason is not inertia: it is that a stellar evolution calculation has to run through millions of time steps and a three-dimensional simulation of one instant costs more than the whole track. So the parameterisation survives because of an arithmetic constraint on computation rather than because anybody believes it, and it will be replaced when the grids of simulations are dense enough to be interpolated rather than when the physics improves. That is a common shape in computational science and it is worth naming, because the resulting models carry an assumption that everybody involved knows to be wrong and that nobody can currently avoid.

Be precise about what the solar calibration does and does not transfer. Fixing the parameter on the Sun makes one model reproduce one star’s radius at one age, and that is a single constraint on a single number — it says nothing about whether the number should be the same for a star twice as massive, half as metal-rich, or in a completely different evolutionary state. The three-dimensional simulations that have been run precisely to answer that find a parameter that varies systematically across the temperature–gravity plane, by twenty per cent or so between the Sun and a red giant. Twenty per cent is small against the range the parameter is sometimes given and large against the precision that seismology now delivers, so the calibration’s constancy has stopped being a harmless simplification and become a measurable error.

Where the ladder goes next

The rung directly above is the surface boundary condition: how a three-dimensional simulation is attached to a one-dimensional interior model, what has to match at the join, and what is left of the mixing length afterwards. The rung after that is the radius discrepancy in low-mass stars — a well-measured disagreement with three candidate causes, all of which act through the efficiency of convection.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

ConvectionEffective temperatureEvolutionary trackFree parameterMixing lengthSolar calibrationStellar radiusSuperadiabatic gradientSystematic errorThree-dimensional simulation