Most stars are small, and most of the light is not
Assumes The mass–luminosity relation and Stellar lifetimes.
Ask what a typical star is and there are three defensible answers, differing by a factor of two hundred.
By count, a typical star is a red dwarf of about a quarter of a solar mass, and there are more stars below the Sun’s mass than above it by a factor of twenty. By mass, the typical star is nearer half a solar mass. By light — meaning: pick a photon leaving a freshly formed population at random and ask which star emitted it — the typical star is above ten solar masses, a class that makes up one star in several thousand and burns out in a few million years — which is why the biggest stars take the galaxy with them.
All three are integrals of the same function against different weights. The disagreement between them is not a subtlety. It is the reason a galaxy’s mass cannot be read off its brightness without an assumption, and the assumption is one of the largest uncertainties in extragalactic astronomy.
The function
The initial mass function is the number of stars formed per unit mass interval, and over most of its range it is a power law:
Salpeter measured in 1955 from the solar neighbourhood, counting stars in a volume, correcting for the ones that had already died, and correcting for the ones too faint to see. The modern form is broken: above half a solar mass, flattening to about 1.3 between 0.08 and 0.5, and flattening further below the hydrogen-burning limit into the brown dwarfs.
The exponent 2.35 is the number that does the work, and the reason is arithmetic. The number of stars per logarithmic interval goes as — falling. The mass per logarithmic interval goes as — falling slowly. And the light per logarithmic interval goes as — rising steeply.
The crossing
The clearest way to see the consequence is cumulatively: what fraction of each quantity comes from stars heavier than a given mass.
That number — half the light from above 58 solar masses — is worth being careful about. It is a statement about a zero-age population, every star still on the main sequence, which is a state that lasts a few million years. It is the right number for a starburst and the wrong number for a galaxy, and the difference between them is the next section.
The one integral that ages
Nothing about the mass function changes with time. The stars change: the heavy ones leave the main sequence and stop contributing to the light integral, and they are exactly the stars the light integral was made of.
The steepness of that curve is the practical problem. A factor of two error in the assumed age of a stellar population is a factor of two or three in the inferred mass, which is larger than most of the effects such masses are used to study. And the age is not a single number: a real galaxy contains stars of every age, so its mass-to-light ratio is an integral over a star-formation history, which is a function nobody observes and which runs out of gas before the galaxy runs out of time.
Why the exponent is what it is
There is no accepted derivation of , and it is worth saying so plainly rather than gesturing at turbulence.
The leading account is that a molecular cloud is fragmented by supersonic turbulence into a spectrum of density peaks, that the mass spectrum of those peaks is set by the turbulent power spectrum, and that some roughly constant fraction of each core’s mass ends up in a star. That produces a power law with an exponent in the right neighbourhood, and it produces a characteristic mass — the break near half a solar mass — from the point at which a core’s thermal Jeans mass matches its turbulent one.
What that account does not do is predict 2.35 rather than 2.1 or 2.6, and it does not explain the most consequential observed fact about the function, which is that it looks the same everywhere it has been measured.
What was actually counted
Salpeter’s measurement is worth following through, because every modern determination is the same three corrections applied more carefully.
He started with the observed luminosity function of the solar neighbourhood: how many stars per cubic parsec at each absolute magnitude, from a volume-limited sample. That is a count, and it is the only directly observed quantity in the chain.
First correction: luminosity to mass. The mass–luminosity relation converts the axis, and its slope enters as a Jacobian — — so the steepness of the relation is doing arithmetic work here quite apart from its role in the light integral.
Second correction: present to initial. The observed function is of stars that are still here. Stars above about one solar mass formed early in the Galaxy’s history have died, so the count at high mass is short by the fraction of the Galaxy’s age over the star’s lifetime. Undoing that requires a star-formation history, which is assumed constant — and the correction is a factor of ten at ten solar masses. The high-mass end of the mass function is not counted; it is reconstructed.
Third correction: completeness. Faint stars are missed. A volume-limited sample of red dwarfs requires knowing distances to objects too faint for parallax before Gaia, so the bottom end was estimated from proper-motion samples with a kinematic distance model.
The result is a function whose two ends are each dominated by a different correction and whose middle is where the data are. That is not a criticism — every subsequent determination has the same structure — but it explains why the exponent has stayed at for seventy years rather than converging: the uncertainty is not statistical.
Universality, and the reason to doubt it
The mass function has been measured in the solar neighbourhood, in open clusters, in globular clusters, in the Magellanic Clouds, and in a handful of star-forming regions. Within the uncertainties it is the same in all of them, across a range of metallicity of about a factor of a hundred and a range of density of rather more.
That is a remarkable result and it is also a weak one, because all of those environments are nearby, and the places where it matters most are not.
So every extragalactic mass measurement assumes a low-mass end that has been measured only here. The suspicion that it might differ — that early, metal-poor, or intensely star-forming environments might produce a top-heavy function — has been raised repeatedly, and the evidence remains indirect: the ratio of light to dynamical mass in massive elliptical galaxies, the abundance ratios in old populations, the counts of high-redshift galaxies — and, at one remove, the rate at which the elements heavier than helium were made, since almost all of them come from the top of this function.
The recent evidence, from stellar-population fitting of gravity-sensitive absorption lines, actually points the other way for the most massive ellipticals: bottom-heavy, with more low-mass stars than the local function has. If that is right, the masses of those galaxies have been underestimated by a factor of about two.
Where the top of the function is spent
The high-mass end supplies almost nothing to the mass budget and almost everything else. It is worth listing what, because the list is the reason a per-cent-level population is worth arguing about.
The ultraviolet. Stars above about fifteen solar masses supply essentially all the photons capable of ionising hydrogen, so every measurement of a star-formation rate from an emission line is a count of those stars, converted by an assumed mass function into a total mass formed. Change the exponent by 0.3 and the conversion factor changes by nearly a factor of two.
The heavy elements. Core-collapse supernovae come from stars above about eight solar masses — the ones whose carbon exists because of a resonance and whose iron cores cannot be supported once they exceed the mass a cold star cannot exceed, and they produce the oxygen, neon, magnesium and silicon that make up most of the mass of the non-hydrogen universe. The yield of a population is an integral over the top of the mass function against a yield curve, so the abundance pattern of a galaxy is a fossil record of its mass function.
The mechanical energy. Stellar winds and supernovae inject momentum into the interstellar medium, which is what regulates star formation in most models. That energy budget is dominated by the same one star in three thousand.
And the neutron stars and black holes, whose number per unit mass formed is again an integral over the same end.
So the mass function’s steep tail is the input to four quite separate subjects, and each of them is sensitive to it in a different power. A quantity measured to thirty per cent locally is propagated into conclusions about the entire history of chemical enrichment, which is a fair summary of why the universality question refuses to go away.
What the picture cannot show
Two things.
The first is the low-mass end, which the figures draw down to 0.08 solar masses and stop. Below that hydrogen does not ignite and the objects are brown dwarfs, which cool and fade continuously and therefore have no mass–luminosity relation at all — their brightness depends on age as much as on mass, so the diagram that sorted the stars has no place to put them. They contribute mass and no light, so they are exactly what an inferred mass-to-light ratio cannot see, and their total contribution is estimated rather than measured. Current estimates put them at a few per cent of the stellar mass, and the estimate is a prediction of the same fitted function.
The second is the top. The figures cut off at 100 solar masses because that is roughly where the observed upper limit sits, and the limit itself is a measurement of something else — the brightness at which radiation pressure exceeds gravity. Whether a genuine upper mass limit exists at 150 or 300 solar masses is unsettled, and the answer barely affects the mass integral and substantially affects the light one, which is the pattern of this entire essay.
What this rung establishes
The site has an essay saying that mass decides a star’s luminosity, its temperature and its lifetime. This one says what happens when that relation is applied to a population rather than to a star, and the answer is that it does not commute with averaging: the mean of the luminosities is nothing like the luminosity of the mean.
That is a general hazard and it has a general name — a nonlinear function of an average is not the average of the function — but it is unusually severe here, because the nonlinearity is a power of 3.5 acting on a distribution spanning three decades. Any quantity that is an integral over the mass function against a steep weight is dominated by a part of the distribution that is barely populated, and is therefore both uncertain and volatile.
Two of the three integrals over the function are worth reading with the mass limits moved, because the answers depend on limits that no observation actually fixes.
The one place the function is counted rather than inferred
The solar-neighbourhood determination described above reconstructs the high-mass end rather than counting it, because the stars have died. There is one class of object where nothing has died yet and the counting is direct: a star-forming region a few million years old.
In such a region every star ever formed is still present, including the most massive, so no correction for stellar deaths is needed. The nearest examples are close enough that individual objects can be resolved down to and below the hydrogen-burning limit, which is the part of the function that dominates the mass and that nothing else reaches.
Two difficulties replace the ones that were removed.
The first is that a pre-main-sequence star’s mass is not read off a mass–luminosity relation, because it is not on the main sequence. It is still contracting, its luminosity is falling as it does so, and its position on a colour–magnitude diagram gives a mass only through evolutionary tracks. Those tracks disagree with each other at low mass by tens of per cent, because they depend on convection, on the accretion history and on the starting condition, none of which is well constrained. A mass function measured this way inherits that disagreement directly.
The second is binaries. A pair of stars too close to resolve is counted as one object of the combined brightness, so the measured function is a function of systems rather than of stars. Since roughly half of stars are in multiples and the fraction depends on mass, correcting from one to the other is a substantial operation — it steepens the low-mass end and it requires a multiplicity fraction that is itself measured from the same regions.
There is a third complication that is not a measurement problem. A young cluster is dynamically active, and the lightest members are the first to be lost as the cluster relaxes and evaporates. So a cluster old enough for its tracks to be reliable is old enough to have lost part of the population being counted, and a cluster young enough to be intact has the least reliable masses.
The direct count and the reconstructed one are measuring different things by different routes and agree within their errors, which is why the function is believed and also why its uncertainty has not shrunk.
Where the break might come from
The essay said above that there is no accepted derivation of the slope, and that remains true. There is a better-founded account of the break — the characteristic mass near half a solar mass — and it is worth stating because it explains the universality that the slope does not.
A cloud fragments while it can cool. As a fragment collapses it heats up, and if it can radiate that heat away it stays cold and continues to collapse and to fragment further. Fragmentation therefore continues until the gas becomes opaque to its own cooling radiation, at which point the collapse becomes adiabatic, the temperature rises steeply, and no smaller fragment can form.
The mass at which that happens can be estimated from the condition that the radiated luminosity equals the compressional heating, and it comes out at a few thousandths of a solar mass — the opacity limit for fragmentation, and it sets the bottom of the range rather than the break.
The break itself comes from a different balance: the mass at which a fragment’s thermal Jeans mass, computed at the temperature the gas actually sits at, matches the mass turbulence is delivering. That temperature is set by the balance between heating from cosmic rays and the ambient radiation field and cooling by dust and molecular lines.
The important feature of that balance is how weakly it depends on composition. The cooling rate depends on the abundance of dust and metals, and so does the heating; over a wide range the two shift together and the equilibrium temperature barely moves. The Jeans mass depends on temperature to the three-halves power and on density to the minus one-half, so a temperature that hardly moves gives a characteristic mass that hardly moves.
That is the strongest available argument for why the function looks the same in environments differing by two orders of magnitude in metallicity, and it also predicts where it should stop being true: at metallicities low enough that dust cooling fails entirely, which is the regime of the first stars and is the one place nobody can look.
And the ageing integral drawn for a single slope, since the mass-to-light ratio’s evolution is the one prediction the function makes that can be checked against a real population.
One more reading contrasts the standard slope with a steeper one over the same integrands.
Where the ladder goes next
The rung above is population synthesis proper: convolving the mass function with a star-formation history and a set of evolutionary tracks to produce a predicted spectrum, which is how every galaxy’s mass, age and metallicity are actually estimated. The rung below — and it is the one that would settle the most — is the physics of fragmentation, which would say where the characteristic mass comes from and therefore whether it should be expected to move at low metallicity.
What this makes readable
Essays that name this one as a prerequisite.
About the same objects
Not linked from either essay — found by the objects both name.
- A mass function corrected by an age initial mass function · mass-to-light ratio · stellar population
- The mass that is not the light initial mass function · mass-to-light ratio · stellar population
What links here
The 8 of 14 essays linking to this one that name the most of the same objects.
- The count theory predicts, and the inference it costs galaxies
- A birth rate measured from light nothing young emitted galaxies
- The gas runs out before the galaxy does galaxies
- The same curve, two galaxies galaxies
- A clock with no fuel in it stars
- A smaller star puts the valley lower exoplanets
- An exponent that is a slope, not a law stars
- The cloud that cannot hold itself up galaxies
The objects this essay names
Each one links to every other essay that touches it.
Brown dwarfInitial mass functionMass luminosity relationMass-to-light ratioPopulation synthesisPower lawSalpeter slopeStar formation rateStellar lifetimesStellar population