A birth rate measured from light nothing young emitted
Assumes Star formation, Initial mass function and Extinction.
“This galaxy forms three solar masses of stars per year” is one of the most-quoted sentences in extragalactic astronomy, and almost nothing in it was measured.
What was measured is a flux — usually the Hα recombination line, or the ultraviolet continuum, or the far-infrared. Each of those is produced by stars in a particular mass range, and none of the ranges contains most of the mass — because most stars are small and most of the light is not, and the two facts are the same fact. The conversion from the one to the other is an integral over the initial mass function, extended down through masses that contribute nothing to the light being counted and are not observed in any galaxy but this one.
None of this is a criticism of the practice. A rate quoted this way is reproducible, comparable between galaxies, and correct to within a factor that is known and roughly constant. What it is not is a measurement of a mass, and the difference matters whenever the conversion factor is the thing that varies.
What each indicator counts
The three standard indicators do not measure the same thing, and their differences are as useful as their agreement.
Hα. Massive stars ionise the gas around them; the gas recombines; and a fixed fraction of recombinations produce an Hα photon. The line luminosity is therefore proportional to the rate at which ionising photons are being produced, which is proportional to the number of stars above about fifteen solar masses. Those stars live under ten million years, so Hα measures the star-formation rate now — averaged over about ten million years and no longer.
The ultraviolet continuum, at around 1,500 ångströms, comes from stars above about three solar masses, which live for a few hundred million years. So it measures a longer average.
The far infrared measures the absorbed bolometric output of the whole population, reprocessed by dust, and its timescale depends on which stars are supplying the heating — which is to say it depends on the dust geometry. In a young star-forming region the heating is entirely from massive stars and the timescale is short; in a quiescent disc the general stellar population heats the dust, and the far-infrared luminosity has almost nothing to do with recent star formation at all. That component is called the cirrus, it can be most of a quiescent galaxy’s infrared output, and separating it is a persistent difficulty.
A fourth indicator avoids dust entirely and is worth mentioning for the contrast. The radio continuum from a star-forming galaxy comes from synchrotron emission by cosmic rays accelerated in supernova remnants, so it counts massive stars a few tens of millions of years after they formed. It is transparent, it correlates with the far infrared to a remarkable tightness across five orders of magnitude, and nobody has a fully satisfying account of why the correlation is so good — which is a reason to use it and a reason to be careful.
The conversion, and what is inside it
Turning a luminosity into a rate requires knowing how much light a population of a given total mass produces, which requires a model of a population.
That model — stellar population synthesis — takes an initial mass function, a set of evolutionary tracks, and a set of model atmospheres, and computes the spectrum of a coeval population as a function of age. Integrating over a constant star-formation history gives a steady-state luminosity per unit rate, and that number is the calibration constant.
The calibration constant is therefore a piece of theory, not a piece of data, and it has changed by tens of per cent over the decades as the ingredients were revised. A quoted star-formation rate carries the vintage of the population-synthesis model that produced it, and comparing numbers from papers twenty years apart without checking is comparing two different quantities.
Every ingredient of it is uncertain, and one of them is unmeasurable. The evolutionary tracks for massive stars depend on rotation, on mass loss and on binary interaction, all of which are active research problems. The atmospheres of hot stars are complicated by winds. And the mass function below about half a solar mass, which contains a third of the mass and none of the light, cannot be observed in any galaxy where the stars are not individually resolved.
There is a further dependence that is easy to miss, because it does not appear in the calibration constant as a stated ingredient at all. A massive star’s ionising output is set by its surface temperature, and a metal-poor star of a given mass is hotter and more compact than a metal-rich one — its envelope is less opaque and its wind strips less of it away. So the number of hydrogen-ionising photons produced per solar mass of stars formed rises as the metallicity falls, by something like a factor of two between solar metallicity and a tenth of it, and by more at the extremes reached in the most metal-poor dwarfs.
The consequence is that one calibration applied across a sample converts light into mass at the wrong rate for the metal-poor members of it, always in the same direction: it overstates their star-formation rate, because it credits them with the number of massive stars that solar-metallicity stars would have needed to produce the photons observed. Dwarf galaxies are the metal-poor population, so the systematic is not scattered through the sample but concentrated in one corner of it.
That is the same corner in which the next section’s disagreement lives, and the two effects point opposite ways: a metallicity too low for the calibration inflates the Hα rate, while a burst caught after its peak deflates it. Anything measuring the deficit is measuring their difference, which makes the observed deficit a lower bound on the burstiness rather than a measurement of it.
Dust, which is not a correction
The ultraviolet indicator has a difficulty the others do not: a large and variable fraction of the light it counts never escapes. The modern practice is to measure both and add them, which is a genuine improvement over correcting one with an assumed attenuation law. It works because the infrared side is nearly assumption-free — absorbed energy has only one place to go — while a correction applied to the ultraviolet depends on the geometry of the dust rather than on any property of the galaxy.
For the most vigorous star-forming galaxies over ninety per cent of the light has been reprocessed, so the ultraviolet measurement is a measurement of a residue. For dwarf galaxies with little dust the reverse holds and the infrared is negligible. Neither regime is a correction to the other.
The number that comes out, and its shape across cosmic time
The single most-used product of all this machinery is the star-formation history of the universe: the rate per unit comoving volume, plotted against redshift.
Its shape is well established and slightly surprising. The rate rises from the present day back to a redshift of about two, peaks there at roughly ten times the present value, and declines beyond — so most of the stars now in existence formed between about eight and eleven billion years ago, and the universe has been winding down for the whole of the second half of its life.
That curve is assembled from ultraviolet surveys at high redshift, infrared and submillimetre surveys at intermediate redshift, and Hα and radio work nearby, each with its own calibration and its own obscuration correction. The agreement between the methods where they overlap is at the level of a factor of about two, which is both a genuine achievement and a fair statement of how well the quantity is known.
The peak is the part that most depends on the machinery. It sits at exactly the redshift where obscuration is worst and where the ultraviolet alone would give a much lower answer, so its height is essentially a statement about how well the reprocessed component has been recovered. The shape either side is more robust than the amplitude.
Why the indicators disagree, and what the disagreement measures
Comparing indicators on the same galaxies produces a systematic pattern: the ratio of Hα to ultraviolet falls in low-mass galaxies. For years that was treated as a calibration problem. It is now generally read as a physical result.
The alternative readings were checked and did not survive. A varying mass function would produce the pattern, and would also produce a corresponding pattern in the ionised gas’s own diagnostics, which is not seen. A dust-correction error would produce it, and would depend on inclination, which it does not.
If star formation happens in bursts rather than steadily, then an indicator averaging over ten million years and one averaging over three hundred report different things, and the distribution of their ratio across a population measures how bursty the star formation is. A galaxy massive enough to contain many independent star-forming regions averages over them and looks steady; a dwarf galaxy with a handful does not.
The same reading applies to a puzzle at the other end of the mass scale. Galaxies at high redshift with apparently enormous rates are often found to have modest gas reservoirs, which would give them lifetimes of tens of millions of years — and that is only a paradox if the rate is a steady state. Read as the peak of a burst caught at maximum, with the gas depleting rapidly, it is not.
So a nuisance became an instrument. The scatter between two indicators, which had been an embarrassment, is a measurement of the duty cycle of star formation in systems too distant to resolve.
A second way to make the same deficit
Burstiness is one reading of the falling ratio and it is not the only one that survives the checks above. The other is a statistical property of the mass function itself, and at low rates it is not optional.
A galaxy forming stars at a hundredth of a solar mass per year has produced a hundred thousand solar masses of stars within the ten-million-year window Hα responds to. One star in four hundred is above the ionising threshold by number, and the threshold stars are rare enough by mass that a hundred thousand solar masses contains only a few dozen of them. A few dozen is not a large number, and the ionising output within that few dozen is dominated by its largest two or three — so the Hα luminosity of such a galaxy is set by a handful of objects drawn at random from a steep distribution.
Drawn at random, that quantity has a distribution rather than a value, and the distribution is strongly skewed. Most draws fall below the mean; the mean is carried by the uncommon draw that happens to contain a very massive star. A survey measuring many such galaxies therefore recovers something near the median rather than the mean, and the median is the lower of the two. The ultraviolet, counting stars down to three solar masses and averaging over thirty times as long, is sampled well enough that the same effect is negligible.
The two explanations predict different things and are separated on those predictions rather than on plausibility. Sampling predicts a scatter that grows as the rate falls in a computable way, and a ceiling — no galaxy can produce more ionising photons than a fully sampled population of its rate would give. Burstiness predicts excursions in both directions, including galaxies caught mid-burst with more Hα than their ultraviolet permits. Those over-luminous outliers are observed, which is the evidence that burstiness is real; the skew of everything else is the evidence that sampling matters as well.
Both act on the same conversion, and both are largest exactly where the resolved checks of the next section cannot reach.
The practical response has been to stop quoting a rate for an individual faint galaxy at all, and to quote the distribution for a population instead. That is an honest retreat: the quantity the light constrains is a distribution, and a single number extracted from it was never the measurement, only its most convenient summary.
The one place the whole chain can be checked
There is exactly one setting in which a star-formation rate can be measured without any of this machinery: a region close enough that the individual stars are resolved and counted.
In nearby star-forming regions the young stars can be identified, placed on a colour–magnitude diagram, assigned masses from evolutionary tracks, and counted — so the mass in young stars is measured rather than inferred, and the mass function is observed down to the hydrogen-burning limit rather than assumed. Doing the same on the integrated light of the same region and comparing is the only end-to-end test the method has.
The tests that have been done broadly agree, at the level of tens of per cent, which is reassuring and is a weaker statement than it sounds. The nearby regions where the comparison is possible are all in this galaxy, at roughly solar metallicity, forming stars at modest rates. They do not test the regime the calibration is mostly used in — metal-poor dwarfs, and starbursts a thousand times more intense — and there is no prospect of testing it there directly.
So the conversion is calibrated where it is least needed and applied where it cannot be checked. That is not unusual in this subject; it is the same structure as a distance ladder, where each rung is calibrated on the one below and the extrapolation is where the interest lies. The discipline is the same too: state the assumption, and quote the number that would move if it were wrong.
What the rate is then used for
The rate is rarely wanted for its own sake. It is wanted as one side of a comparison. A last point about the conversion, because it is the one most often left implicit. Every number in this chain is a mean over an assumed population, and the population that a particular galaxy actually formed is a draw from it. For a galaxy forming stars at a solar mass a year the draw is large enough that the mean is the answer; for one forming stars at a thousandth of that, the same mean is a statement about an ensemble the galaxy is a single member of. The conversion factor does not become wrong at low rates — it becomes a distribution, and quoting its centre as though it were a measurement is the error the next section is about.
Where the picture stops
The mass function may not be universal. Everything above assumes one initial mass function everywhere. There is now reasonable evidence that the most massive early-type galaxies have a bottom-heavy one and that extreme starbursts may have a top-heavy one, and if either is right then rates measured in those systems are wrong by factors, in opposite directions.
Binaries change the massive-star output substantially. A large fraction of massive stars exchange mass with a companion, which strips envelopes, spins up the accretor and produces hot stripped stars that are strong ionising sources long after a single-star model says the ionising output should have ceased. Population synthesis including binaries gives noticeably different calibrations, and which to use is not settled.
The rate is an average and the object is not homogeneous. Every number here is integrated over a whole galaxy, and a galaxy is a set of regions with different gas densities, different metallicities and different dust geometries. A single rate for an object thirty kiloparsecs across is a statistic in the same sense that a mean opacity is a statistic — useful, and weighted in a way that has to be remembered.
And the escape fraction is not zero. The Hα method assumes every ionising photon is absorbed by gas in the galaxy and converted. Some escape, and some are absorbed by dust before reaching any hydrogen — both of which make the rate an underestimate by an amount nobody can measure directly for an individual galaxy.
One more upper limit shows how much of the ionising budget rests on stars nobody has counted.
Where this ladder goes next
Later rungs on this anchor: population synthesis in detail, and how much of the calibration is stellar physics rather than statistics; the burstiness measurement and what the indicator-ratio distribution says about the duty cycle; radio continuum and X-ray indicators, which avoid dust entirely and introduce their own assumptions; the initial mass function’s universality and the evidence against it; and the resolved measurements in the Local Group, which are the only place any of these conversions can be checked star by star.
What links here
Essays that link to this one from their own argument.
- A mass function corrected by an age stars
- The count theory predicts, and the inference it costs galaxies
- A cloud that cannot become a star galaxies
- A coincidence that is a factor of fourteen cosmology
- A disc that turns slower than its mass requires galaxies
- It ends when the walls meet cosmology
- Red, gas-poor, and still spiral-shaped galaxies
The objects this essay names
Each one links to every other essay that touches it.
BurstinessCalibration factorDust correctionIndicator timescaleInitial mass functionIonising photonsThe Kennicutt–Schmidt lawObscured star formationRecombination lineSpecific star formation rateStar formation rateStellar population synthesis