The iron clock has no single delay
Assumes Chemical evolution, Supernovae and Binary stars.
The ratio of α elements to iron in a star’s atmosphere is one of the most useful clocks in astronomy, and it is drawn almost everywhere with a single delay built into it. The argument that introduced the knee went like this: core-collapse supernovae come from massive stars that live a few million years and make oxygen, magnesium and silicon with a little iron; Type Ia supernovae come from white dwarfs, arrive about a billion years later, and make iron with almost no α elements. So a system’s [α/Fe] runs flat at the massive-star value, then bends down at whatever metallicity the system had reached when the first Type Ia exploded.
“About a billion years later” was a simplification, and it was drawn as a step: no Type Ia supernovae before a billion years, all of them after. The actual distribution of delays has been measured, in two independent ways, and it is not a step. It is a power law, with as many explosions in each decade of delay as in any other, from forty million years to the age of the universe.
Two ways to measure a delay
A Type Ia supernova carries no label saying when its progenitor formed. Its delay has to be inferred statistically, by comparing where and when supernovae happen with where and when stars formed.
By host age. Take a large sample of galaxies whose spectra have been decomposed into stellar populations of different ages — how much mass formed less than 400 million years ago, how much between 400 million and 2.4 billion, how much before. Count the supernovae each galaxy has hosted over a survey. The rate in each galaxy is the sum over its age bins of the mass in that bin times the rate per unit mass at that delay, and with thousands of galaxies the rates per unit mass can be solved for. The answer has three features: a large rate at short delays, a rate still well above zero in galaxies with no star formation for billions of years — elliptical galaxies host Type Ia supernovae — and a fall between them close to .
By cosmic history. The volumetric rate of Type Ia supernovae at each redshift is the cosmic star-formation history convolved with the delay-time distribution. The star-formation history is measured independently from the light of galaxies, so the supernova rate at several redshifts constrains the convolution kernel. That kernel also comes out close to .
The two methods share almost nothing — one uses nearby galaxies and population synthesis, the other deep-field supernova searches and a cosmic average — and they agree. The integral of the distribution is measured too: about one Type Ia supernova for every thousand solar masses of stars formed, over a Hubble time.
Why a power law, and why this one
A power law has no characteristic time, which is itself a clue: the mechanism that sets the delay must spread over many orders of magnitude without a preferred scale.
Two classes of progenitor have been proposed. In the single-degenerate picture, a white dwarf accretes from a normal companion star until it approaches the Chandrasekhar mass and ignites; the delay is roughly the companion’s lifetime plus the time it takes to transfer enough mass, and it tends to cluster at particular ages set by which companions can transfer mass stably. In the double-degenerate picture, two white dwarfs in a close binary spiral together by emitting gravitational waves and merge; the delay is the time to form both white dwarfs plus the inspiral time.
The inspiral time for a circular binary goes as the fourth power of the separation:
Now suppose white-dwarf binaries emerge from their common-envelope phase with separations spread evenly in logarithm — constant, which is the distribution of separations observed for binaries of many kinds and has been called Öpik’s law since the 1920s. A distribution flat in is flat in , because . And a distribution flat in is exactly .
So the measured delay-time distribution is what a population of merging white dwarfs with Öpik separations produces, with no parameters adjusted. That does not prove double-degenerate progenitors — other channels could produce similar shapes, and the minimum delay of about forty million years is set by the lifetime of the most massive stars that make white dwarfs rather than by any inspiral — but it is the most direct connection in the subject between a measured cosmological distribution and a law of gravitational radiation. The same dependence shrinks the orbits of binary pulsars at exactly the rate general relativity predicts.
A knee that is a bend
Replace the step in the α-element model with the measured distribution and the knee changes character.
The power law’s prompt explosions — a tenth of them within about seventy million years — begin adding iron while the system is still extremely metal-poor, so the ratio starts to decline almost at once. But they add iron slowly, a constant number per decade of time, so the decline is gradual. By the time a step-function knee would appear, the power-law track has already fallen partway, and it continues to fall for a dex of metallicity further.
That makes the position of the knee a less simple measurement than it was drawn to be. With a single delay the knee is at the metallicity the system reached at that delay, so it measures the star-formation rate directly. With a power law there is no single point, and “the knee” becomes a matter of definition — the first departure from the plateau, or the steepest part of the decline, or the metallicity at which the ratio has fallen halfway — and each definition responds differently to the star-formation history.
The ordering survives, and that is the most important thing the figures say. The Galactic bulge, which formed fast, bends at high metallicity; the Milky Way’s thick disc at intermediate metallicity; dwarf spheroidal galaxies, which formed slowly and inefficiently, at [Fe/H] below −1, some below −1.5. Every one of those orderings holds with a power-law distribution, because a faster-forming system reaches any given metallicity sooner and so has had less time for delayed iron whatever the shape of the delay. What changes is the translation from a knee’s metallicity to a star-formation timescale, which becomes a model calculation with the distribution inside it.
The rate across cosmic time
The second way of measuring the distribution is itself a figure worth drawing, because it shows the convolution explicitly.
A convolution with a delay does two things to a history: it shifts the peak to later times, and it smooths the curve by the width of the kernel. The power law does less shifting than a Gaussian at three billion years because so many of its explosions are prompt, and more smoothing than a step because its tail extends across the whole of cosmic time. Today’s Type Ia rate is fed partly by stars formed in the last few hundred million years and partly by stars formed at the peak of cosmic star formation ten billion years ago, and the power law says those contributions are comparable.
The measured Type Ia rate rises by about a factor of two to three from the present to redshift one and then flattens, which is the power-law curve’s shape. Beyond redshift 1.5 the rates come from a few dozen supernovae found in the deepest fields of space telescopes, and the uncertainties there are large enough that a broad Gaussian is not firmly excluded by the cosmic method alone. It is the combination with the host-age method, which constrains short delays well, that pins down the power law.
Iron as a supernova’s signature
A Type Ia supernova — the explosion told apart from a collapsing core by the lines its spectrum lacks — makes about six-tenths to seven-tenths of a solar mass of iron-peak elements, mostly as radioactive nickel-56 that decays through cobalt to iron; a core-collapse supernova makes about a tenth of that. The light curve of a Type Ia is powered by exactly that decay, and its peak brightness measures the nickel mass, which is why these explosions can be standardised at all. The iron that eventually ends up in stars’ atmospheres is, for most elements of the iron peak, mainly the ash of these explosions.
That is what makes the delay-time distribution a chemical quantity. The rate per unit mass sets how fast iron accumulates; the delay sets when. Over a Hubble time about one Type Ia per thousand solar masses of stars formed, each making about 0.7 solar masses of iron, gives roughly 0.7 × 10⁻³ of all the mass that formed stars returned as iron — comparable to what the ten or so core-collapse supernovae per thousand solar masses of the same stars supply. Delayed iron therefore roughly doubles what the prompt explosions made, which is why the α-to-iron ratio of the Sun is lower than that of the oldest stars by a factor of two to three, and why that factor is set as much by the normalisation of the distribution as by its shape.
Iron where no stars are forming
The long tail of the distribution has consequences that are easy to overlook because they happen in galaxies that have stopped doing anything else.
An elliptical galaxy — a member of the red sequence that has almost nothing between it and the blue cloud — that finished forming stars eight billion years ago still hosts Type Ia supernovae, at a rate per unit stellar mass a few times lower than a star-forming spiral’s. On the power law that is exactly as expected: the stars formed at the start are now at delays of eight to ten billion years, and a distribution still delivers about a twentieth of its explosions per factor of 1.3 in delay there. Those supernovae explode into hot, thin gas rather than into a star-forming disc, and their iron has nowhere to go but the galaxy’s halo — or, in a cluster, the intracluster gas that holds most of the cluster’s ordinary matter.
The iron in that gas is measured from X-ray lines, at about a third of the solar abundance, spread over hundreds of kiloparsecs, and it is a problem. Adding up the iron that the cluster’s stars could have made — core-collapse supernovae at their birth plus Type Ia supernovae with the field galaxies’ delay-time distribution and normalisation — falls short of the iron observed, by a factor of two or more. Either clusters’ stars produced more Type Ia supernovae per unit mass than stars elsewhere, which the supernova rates measured in clusters do suggest, or the initial mass function was different, or more stars were stripped into intracluster light than are counted. The integral of the distribution, in other words, may not be universal, and the most iron-rich places in the universe are where it is least likely to be.
A standard candle with an age in it
The delay-time distribution also reaches into cosmology, by a route that runs through the brightness that has to be standardised before a Type Ia supernova is a distance.
Standardisation corrects each explosion’s peak brightness by its light-curve width and colour. After that correction a residual remains that correlates with the host galaxy: supernovae in massive or old hosts are slightly brighter, by a few hundredths of a magnitude, than those in young ones. If that residual is a property of the progenitor’s age — if white dwarfs that explode at long delays produce slightly different explosions from those at short delays — then the delay-time distribution matters directly for distances.
The reason is that the mix of delays changes with redshift. A supernova at redshift one exploded when the universe was six billion years old, so it cannot have had a delay of ten billion years; the power law’s long tail is truncated at high redshift and fully present today. A population of supernovae whose standardised brightness depends on delay therefore has a mean brightness that drifts with redshift, and a drift of a few hundredths of a magnitude between redshift zero and one is the size of the signal that distinguishes a cosmological constant from a slowly varying dark energy. Whether the host correlation is an age effect, a dust effect or a metallicity effect is currently argued, and the answer determines whether the distribution drawn at the top of this page is a nuisance parameter in the equation of state of the universe.
Where the delay matters for a disc
In a disc whose gradient is built ring by ring, each ring has its own history, and each ring’s iron arrives with the same distribution of delays after its own star formation. Inner rings that formed stars fast bend at high metallicity and outer rings that formed slowly bend low, so a disc does not have one α-element knee: it has a family, one per radius.
Stars near the Sun, which are a mixture of stars born across the disc, therefore do not trace any single track. The two sequences seen in the solar neighbourhood’s [α/Fe] against [Fe/H] — a high-α sequence and a low-α sequence, separated by a gap — have been explained by two distinct episodes of gas accretion, by a thick disc formed fast and a thin disc formed slowly, and by the superposition of many rings mixed together by migration. The delay-time distribution enters all three explanations identically; what separates them is the history of star formation and the movement of stars, which is why the gap is still argued about.
What the figures cannot show
The chemistry is a single-zone calculation with instantaneous mixing: every supernova’s iron is assumed to be spread through the whole system immediately. In a real galaxy iron from a single explosion is mixed into a few hundred thousand solar masses of gas over tens of millions of years, and in the earliest, most metal-poor stars the abundance pattern of an individual supernova can show through. The figures also treat the yields as fixed, and Type Ia yields depend on the explosion mechanism — whether the white dwarf is at the Chandrasekhar mass or well below it — which changes manganese and nickel much more than iron, and is one of the few chemical handles on which progenitor channel dominates.
And the distribution is taken to be universal. It may not be: the fraction of stars in close binaries, the efficiency of common-envelope ejection and the white-dwarf mass distribution all depend on metallicity, and a distribution measured in present-day galaxies of near-solar metallicity is being applied to systems a hundred times poorer in metals.
A distribution with no timescale draws no feature of its own
A single delay is a timescale, and a timescale produces a feature — a knee, a break, an edge — at the moment it is reached. A distribution with no timescale produces no feature of its own; any sharp feature in the output must come from somewhere else. When a measured knee is sharp and the delay distribution is a power law, the sharpness is information about the star-formation history or about mixing, not about the supernovae.
The same reading applies wherever a response is a convolution. A disc’s gradient history seen in old stars is the birth gradient convolved with migration, and the flatness of the result is a property of the kernel. A cosmic supernova rate is a star-formation history convolved with a delay. In both cases the kernel can be measured independently, and in both the temptation is to read a feature of the output as a feature of the input.
Still open: the first stars’ yields
The earliest stars formed from gas with no metals at all, and their supernovae — possibly including pair-instability explosions of stars hundreds of times the Sun’s mass, which leave no remnant and make a distinctive abundance pattern — enriched the gas that the oldest surviving stars formed from. No such first-generation star has been found, and perhaps none survives. What is left is the abundance pattern of the most metal-poor stars known, a few with iron abundances below a millionth of the Sun’s, whose ratios of carbon to iron and of odd to even elements are the only evidence there is about a population nobody has observed.
The objects this essay names
Each one links to every other essay that touches it.
Alpha enhancementAlpha kneeChemical evolutionConvolutionCosmic star formation historyDelay time distributionDouble degenerateGravitational wavesSingle degenerateType ia supernovaeWhite dwarf