Stars

Most stars are small, and most of the light is not

The mass function is steep and the mass–luminosity relation is steeper, so counting stars and measuring their output are integrals of the same function dominated by opposite ends of it. Every mass inferred from a brightness passes through that mismatch.

Assumes The mass–luminosity relation and Stellar lifetimes.

Ask what a typical star is and there are three defensible answers, differing by a factor of two hundred.

By count, a typical star is a red dwarf of about a quarter of a solar mass, and there are more stars below the Sun’s mass than above it by a factor of twenty. By mass, the typical star is nearer half a solar mass. By light — meaning: pick a photon leaving a freshly formed population at random and ask which star emitted it — the typical star is above ten solar masses, a class that makes up one star in several thousand and burns out in a few million years — which is why the biggest stars take the galaxy with them.

All three are integrals of the same function against different weights. The disagreement between them is not a subtlety. It is the reason a galaxy’s mass cannot be read off its brightness without an assumption, and the assumption is one of the largest uncertainties in extragalactic astronomy.

One population, three integrals, two ends. The same mass function weighted three ways, each normalised so that the area under it is one, against mass on a logarithmic axis — so a share of the page is a share of the total. The light is a zero-age population's, every star still on the main sequence, which is the only age at which the top of the range is present at all. The steep curve on the left is the number of stars, which peaks at 0.08 M☉ and falls away because the function is steeper than m⁻¹; the middle one is the mass they carry; the one on the right is the light they emit, computed from this file's own main-sequence relation and peaking at 55.0 M☉ — 682 times further up the axis. The population that is counted and the population that is seen are two different populations. Above 8 M☉ there are 0.63% of the stars, 21% of the mass and 99% of the light. The slopes are α = 1.3 above 0.08 M☉, α = 2.3 above 0.5 M☉, and the low-mass end is where the honesty runs out: it has never been measured in another galaxy, and every mass inferred from a luminosity assumes it.
Fig. 1 The three integrals, each normalised so the area under it is one, against mass on a logarithmic axis — so a share of the page is a share of the total. Number peaks at the bottom of the range; mass a little above it; light at 55 solar masses, five hundred times further up the axis. The population that is counted and the population that is seen are two different populations. Above eight solar masses there are 0.03 per cent of the stars, 1.3 per cent of the mass, and most of the light. The light here is a zero-age population’s, every star still on the main sequence, which is the only age at which the top of the range is present at all.

The function

The initial mass function is the number of stars formed per unit mass interval, and over most of its range it is a power law:

dNdmmα.\frac{dN}{dm} \propto m^{-\alpha}.

Salpeter measured α=2.35\alpha = 2.35 in 1955 from the solar neighbourhood, counting stars in a volume, correcting for the ones that had already died, and correcting for the ones too faint to see. The modern form is broken: α2.3\alpha \approx 2.3 above half a solar mass, flattening to about 1.3 between 0.08 and 0.5, and flattening further below the hydrogen-burning limit into the brown dwarfs.

The exponent 2.35 is the number that does the work, and the reason is arithmetic. The number of stars per logarithmic interval goes as m1α=m1.35m^{1-\alpha} = m^{-1.35} — falling. The mass per logarithmic interval goes as m2α=m0.35m^{2-\alpha} = m^{-0.35} — falling slowly. And the light per logarithmic interval goes as m2α×L/mm0.35×m2.5=m2.15m^{2-\alpha}\times L/m \approx m^{-0.35}\times m^{2.5}= m^{2.15} — rising steeply.

Luminosity against mass, against a slope of 3.5. Main-sequence luminosity against mass, both in solar units, on logarithmic axes, over the range 0.079 to 63 solar masses. The measured curve comes from the eclipsing binaries and is the same in every drawing of it; what changes here is what it is compared against. The dashed line is a pure power law of exponent 3.5, and the curve crosses it rather than following it — the local slope runs from about 2.3 at the bottom of the range, where the interiors are convective, through nearly 4 near a solar mass where bound-free opacity dominates, to about 3 among the massive stars where electron scattering does. Quoting one exponent across the whole sequence is a convenience and the places it fails are the places the interior physics changes. Because the slope is between three and four across most of the range, a small spread in mass becomes an enormous spread in output: the 63-solar-mass end is 3.0e+9 times brighter than the 0.079-solar-mass end.
Fig. 2 The relation that makes the third integral rise. Luminosity against mass on logarithmic axes, with the slope between three and four across most of the range: mass decides everything, by a power of three and a half. Set that against a mass function falling as m2.35m^{-2.35} and the product still rises — the mass–luminosity relation wins, and it wins by more than a whole power of the mass. That single inequality is the reason the three curves in the opening figure point in different directions.

The crossing

The clearest way to see the consequence is cumulatively: what fraction of each quantity comes from stars heavier than a given mass.

The fraction contributed by everything heavier than a given mass. Cumulative shares of number, mass and light above each mass, from the same function, for a population whose stars are all still on the main sequence. Half the stars are heavier than 0.24 M☉ and half the light comes from stars heavier than 57.9 M☉ — a factor of 244 between the two medians, in one population. The mass-to-light ratio of the whole is 0.00048 in solar units, and it is that number, not any star's, that turns a galaxy's brightness into a mass. The three curves cross nothing and separate everywhere, which is why "a typical star" is a phrase with three different answers. What the figure assumes and cannot check is that the function is universal: it is measured in the solar neighbourhood and in a handful of clusters, and applied to galaxies at redshift 6.
Fig. 3 Half the stars are heavier than 0.24 solar masses. Half the light comes from stars heavier than 58 — a factor of 244 between two medians of one population. The mass-to-light ratio of the whole is 0.00048 in solar units, which is to say that a freshly formed population is about two thousand times brighter per unit mass than the Sun. “A typical star” is a phrase with three answers, and which one is meant is decided by the weight in the integral rather than by anything about stars.

That number — half the light from above 58 solar masses — is worth being careful about. It is a statement about a zero-age population, every star still on the main sequence, which is a state that lasts a few million years. It is the right number for a starburst and the wrong number for a galaxy, and the difference between them is the next section.

The one integral that ages

Nothing about the mass function changes with time. The stars change: the heavy ones leave the main sequence and stop contributing to the light integral, and they are exactly the stars the light integral was made of.

The one integral that ages. Mass-to-light ratio against age for a population formed in a single burst with this mass function, on logarithmic axes. Nothing about the mass function changes along this curve — what changes is which stars are still on the main sequence, and the light integral is dominated by exactly the stars that leave it first. From 0.0029 at a million years to 15 at 13.0 Gyr, a factor of 5323. The marked ages carry their turnoff masses: 10 Myr at 14 M☉, 100 Myr at 5.5 M☉, 1 Gyr at 2.2 M☉, 10 Gyr at 1.0 M☉. A galaxy's mass is inferred from its light through this number, so a mass is an age assumption before it is a measurement — and the mass here is the initial mass, remnants included at their birth weight, which is the approximation this figure makes.
Fig. 4 The consequence, on logarithmic axes. Mass-to-light ratio against age for a population formed in a single burst: from 0.003 at a million years to 15 at thirteen billion, a factor of over five thousand, with the mass function held fixed throughout. The marked ages carry their turnoff masses — 14 solar masses at 10 Myr, 5.5 at 100 Myr, 2.2 at 1 Gyr, 1.0 at 10 Gyr — which are the same turnoffs an isochrone fit reads. A galaxy’s mass is inferred from its light through this number, so a mass is an age assumption before it is a measurement.

The steepness of that curve is the practical problem. A factor of two error in the assumed age of a stellar population is a factor of two or three in the inferred mass, which is larger than most of the effects such masses are used to study. And the age is not a single number: a real galaxy contains stars of every age, so its mass-to-light ratio is an integral over a star-formation history, which is a function nobody observes and which runs out of gas before the galaxy runs out of time.

Why the exponent is what it is

There is no accepted derivation of α=2.35\alpha = 2.35, and it is worth saying so plainly rather than gesturing at turbulence.

The leading account is that a molecular cloud is fragmented by supersonic turbulence into a spectrum of density peaks, that the mass spectrum of those peaks is set by the turbulent power spectrum, and that some roughly constant fraction of each core’s mass ends up in a star. That produces a power law with an exponent in the right neighbourhood, and it produces a characteristic mass — the break near half a solar mass — from the point at which a core’s thermal Jeans mass matches its turbulent one.

What that account does not do is predict 2.35 rather than 2.1 or 2.6, and it does not explain the most consequential observed fact about the function, which is that it looks the same everywhere it has been measured.

What was actually counted

Salpeter’s measurement is worth following through, because every modern determination is the same three corrections applied more carefully.

He started with the observed luminosity function of the solar neighbourhood: how many stars per cubic parsec at each absolute magnitude, from a volume-limited sample. That is a count, and it is the only directly observed quantity in the chain.

First correction: luminosity to mass. The mass–luminosity relation converts the axis, and its slope enters as a Jacobian — dN/dm=(dN/dMV)(dMV/dm)dN/dm = (dN/dM_V)(dM_V/dm) — so the steepness of the relation is doing arithmetic work here quite apart from its role in the light integral.

Second correction: present to initial. The observed function is of stars that are still here. Stars above about one solar mass formed early in the Galaxy’s history have died, so the count at high mass is short by the fraction of the Galaxy’s age over the star’s lifetime. Undoing that requires a star-formation history, which is assumed constant — and the correction is a factor of ten at ten solar masses. The high-mass end of the mass function is not counted; it is reconstructed.

Third correction: completeness. Faint stars are missed. A volume-limited sample of red dwarfs requires knowing distances to objects too faint for parallax before Gaia, so the bottom end was estimated from proper-motion samples with a kinematic distance model.

The result is a function whose two ends are each dominated by a different correction and whose middle is where the data are. That is not a criticism — every subsequent determination has the same structure — but it explains why the exponent has stayed at 2.35±0.32.35 \pm 0.3 for seventy years rather than converging: the uncertainty is not statistical.

Universality, and the reason to doubt it

The mass function has been measured in the solar neighbourhood, in open clusters, in globular clusters, in the Magellanic Clouds, and in a handful of star-forming regions. Within the uncertainties it is the same in all of them, across a range of metallicity of about a factor of a hundred and a range of density of rather more.

That is a remarkable result and it is also a weak one, because all of those environments are nearby, and the places where it matters most are not.

One population, three integrals, two ends. The same mass function weighted three ways, each normalised so that the area under it is one, against mass on a logarithmic axis — so a share of the page is a share of the total. The light is a zero-age population's, every star still on the main sequence, which is the only age at which the top of the range is present at all. The steep curve on the left is the number of stars, which peaks at 0.08 M☉ and falls away because the function is steeper than m⁻¹; the middle one is the mass they carry; the one on the right is the light they emit, computed from this file's own main-sequence relation and peaking at 55.0 M☉ — 682 times further up the axis. The population that is counted and the population that is seen are two different populations. Above 8 M☉ there are 0.19% of the stars, 13% of the mass and 99% of the light. The slopes are α = 2.35 above 0.08 M☉, and the low-mass end is where the honesty runs out: it has never been measured in another galaxy, and every mass inferred from a luminosity assumes it.
Fig. 5 Salpeter’s original single power law, without the low-mass break, drawn the same way. The number integrand now rises without limit towards the bottom of the range rather than turning over — which is what a pure m2.35m^{-2.35} does, and it is why Salpeter’s own paper was careful to state a lower cutoff. The break at half a solar mass is where most of the mass is decided, and it is precisely the part of the function that has never been measured outside the Local Group: no telescope resolves individual half-solar-mass stars in another galaxy.

So every extragalactic mass measurement assumes a low-mass end that has been measured only here. The suspicion that it might differ — that early, metal-poor, or intensely star-forming environments might produce a top-heavy function — has been raised repeatedly, and the evidence remains indirect: the ratio of light to dynamical mass in massive elliptical galaxies, the abundance ratios in old populations, the counts of high-redshift galaxies — and, at one remove, the rate at which the elements heavier than helium were made, since almost all of them come from the top of this function.

The recent evidence, from stellar-population fitting of gravity-sensitive absorption lines, actually points the other way for the most massive ellipticals: bottom-heavy, with more low-mass stars than the local function has. If that is right, the masses of those galaxies have been underestimated by a factor of about two.

Where the top of the function is spent

The high-mass end supplies almost nothing to the mass budget and almost everything else. It is worth listing what, because the list is the reason a per-cent-level population is worth arguing about.

The ultraviolet. Stars above about fifteen solar masses supply essentially all the photons capable of ionising hydrogen, so every measurement of a star-formation rate from an emission line is a count of those stars, converted by an assumed mass function into a total mass formed. Change the exponent by 0.3 and the conversion factor changes by nearly a factor of two.

The heavy elements. Core-collapse supernovae come from stars above about eight solar masses — the ones whose carbon exists because of a resonance and whose iron cores cannot be supported once they exceed the mass a cold star cannot exceed, and they produce the oxygen, neon, magnesium and silicon that make up most of the mass of the non-hydrogen universe. The yield of a population is an integral over the top of the mass function against a yield curve, so the abundance pattern of a galaxy is a fossil record of its mass function.

The mechanical energy. Stellar winds and supernovae inject momentum into the interstellar medium, which is what regulates star formation in most models. That energy budget is dominated by the same one star in three thousand.

And the neutron stars and black holes, whose number per unit mass formed is again an integral over the same end.

So the mass function’s steep tail is the input to four quite separate subjects, and each of them is sensitive to it in a different power. A quantity measured to thirty per cent locally is propagated into conclusions about the entire history of chemical enrichment, which is a fair summary of why the universality question refuses to go away.

What the picture cannot show

Two things.

The first is the low-mass end, which the figures draw down to 0.08 solar masses and stop. Below that hydrogen does not ignite and the objects are brown dwarfs, which cool and fade continuously and therefore have no mass–luminosity relation at all — their brightness depends on age as much as on mass, so the diagram that sorted the stars has no place to put them. They contribute mass and no light, so they are exactly what an inferred mass-to-light ratio cannot see, and their total contribution is estimated rather than measured. Current estimates put them at a few per cent of the stellar mass, and the estimate is a prediction of the same fitted function.

The second is the top. The figures cut off at 100 solar masses because that is roughly where the observed upper limit sits, and the limit itself is a measurement of something else — the brightness at which radiation pressure exceeds gravity. Whether a genuine upper mass limit exists at 150 or 300 solar masses is unsettled, and the answer barely affects the mass integral and substantially affects the light one, which is the pattern of this entire essay.

What this rung establishes

The site has an essay saying that mass decides a star’s luminosity, its temperature and its lifetime. This one says what happens when that relation is applied to a population rather than to a star, and the answer is that it does not commute with averaging: the mean of the luminosities is nothing like the luminosity of the mean.

That is a general hazard and it has a general name — a nonlinear function of an average is not the average of the function — but it is unusually severe here, because the nonlinearity is a power of 3.5 acting on a distribution spanning three decades. Any quantity that is an integral over the mass function against a steep weight is dominated by a part of the distribution that is barely populated, and is therefore both uncertain and volatile.

Two of the three integrals over the function are worth reading with the mass limits moved, because the answers depend on limits that no observation actually fixes.

The fraction contributed by everything heavier than a given mass. Cumulative shares of number, mass and light above each mass, from the same function, for a population whose stars are all still on the main sequence. Half the stars are heavier than 0.24 M☉ and half the light comes from stars heavier than 62.5 M☉ — a factor of 264 between the two medians, in one population. The mass-to-light ratio of the whole is 0.00042 in solar units, and it is that number, not any star's, that turns a galaxy's brightness into a mass. The three curves cross nothing and separate everywhere, which is why "a typical star" is a phrase with three different answers. What the figure assumes and cannot check is that the function is universal: it is measured in the solar neighbourhood and in a handful of clusters, and applied to galaxies at redshift 6.
Fig. 6 The cumulative contributions from the hydrogen-burning limit to a hundred and twenty solar masses. The number is dominated by the bottom and the light by the top, and both integrals converge — which is the only reason a mass-to-light ratio exists at all.
One star in 313 makes essentially all the ionising light. Three cumulative fractions against stellar mass, for a broken power-law initial mass function with slopes 1.3 and 2.3 breaking at 0.5 solar masses. Each curve says what share of one quantity is produced by stars heavier than the mass on the axis, and the three do not resemble one another. Only 0.26 per cent of the hydrogen-ionising photons come from stars below 15 solar masses, because the ionising output of a star climbs by five orders of magnitude between eight and twenty. One star in 313 is above that mass, and between them those stars hold 19 per cent of the mass. Those two numbers are the leverage in every star-formation rate quoted from an Hα line. What is measured is the light of a handful of very massive stars; what is reported is the mass of a whole population; and the number in between is an integral over a part of the mass function that no extragalactic observation reaches. The medians are marked but should be read with care, and the reason is visible in the curves: the mass-weighted median at 1.51 solar masses is a property of the population, while the light-weighted one at 88 is a property of where the plot stops — halving the upper mass limit moves it to 68. An integrand that rises with mass has its median wherever the axis ends.
Fig. 7 The ionising output with the upper limit at three hundred solar masses rather than a hundred. One star in three hundred and thirteen makes essentially all of it, and moving the upper limit changes that fraction substantially — so the ionising budget of a galaxy depends on the rarest stars in it.

The one place the function is counted rather than inferred

The solar-neighbourhood determination described above reconstructs the high-mass end rather than counting it, because the stars have died. There is one class of object where nothing has died yet and the counting is direct: a star-forming region a few million years old.

In such a region every star ever formed is still present, including the most massive, so no correction for stellar deaths is needed. The nearest examples are close enough that individual objects can be resolved down to and below the hydrogen-burning limit, which is the part of the function that dominates the mass and that nothing else reaches.

Two difficulties replace the ones that were removed.

The first is that a pre-main-sequence star’s mass is not read off a mass–luminosity relation, because it is not on the main sequence. It is still contracting, its luminosity is falling as it does so, and its position on a colour–magnitude diagram gives a mass only through evolutionary tracks. Those tracks disagree with each other at low mass by tens of per cent, because they depend on convection, on the accretion history and on the starting condition, none of which is well constrained. A mass function measured this way inherits that disagreement directly.

The second is binaries. A pair of stars too close to resolve is counted as one object of the combined brightness, so the measured function is a function of systems rather than of stars. Since roughly half of stars are in multiples and the fraction depends on mass, correcting from one to the other is a substantial operation — it steepens the low-mass end and it requires a multiplicity fraction that is itself measured from the same regions.

There is a third complication that is not a measurement problem. A young cluster is dynamically active, and the lightest members are the first to be lost as the cluster relaxes and evaporates. So a cluster old enough for its tracks to be reliable is old enough to have lost part of the population being counted, and a cluster young enough to be intact has the least reliable masses.

The direct count and the reconstructed one are measuring different things by different routes and agree within their errors, which is why the function is believed and also why its uncertainty has not shrunk.

Where the break might come from

The essay said above that there is no accepted derivation of the slope, and that remains true. There is a better-founded account of the break — the characteristic mass near half a solar mass — and it is worth stating because it explains the universality that the slope does not.

A cloud fragments while it can cool. As a fragment collapses it heats up, and if it can radiate that heat away it stays cold and continues to collapse and to fragment further. Fragmentation therefore continues until the gas becomes opaque to its own cooling radiation, at which point the collapse becomes adiabatic, the temperature rises steeply, and no smaller fragment can form.

The mass at which that happens can be estimated from the condition that the radiated luminosity equals the compressional heating, and it comes out at a few thousandths of a solar mass — the opacity limit for fragmentation, and it sets the bottom of the range rather than the break.

The break itself comes from a different balance: the mass at which a fragment’s thermal Jeans mass, computed at the temperature the gas actually sits at, matches the mass turbulence is delivering. That temperature is set by the balance between heating from cosmic rays and the ambient radiation field and cooling by dust and molecular lines.

The important feature of that balance is how weakly it depends on composition. The cooling rate depends on the abundance of dust and metals, and so does the heating; over a wide range the two shift together and the equilibrium temperature barely moves. The Jeans mass depends on temperature to the three-halves power and on density to the minus one-half, so a temperature that hardly moves gives a characteristic mass that hardly moves.

That is the strongest available argument for why the function looks the same in environments differing by two orders of magnitude in metallicity, and it also predicts where it should stop being true: at metallicities low enough that dust cooling fails entirely, which is the regime of the first stars and is the one place nobody can look.

And the ageing integral drawn for a single slope, since the mass-to-light ratio’s evolution is the one prediction the function makes that can be checked against a real population.

The one integral that ages. Mass-to-light ratio against age for a population formed in a single burst with this mass function, on logarithmic axes. Nothing about the mass function changes along this curve — what changes is which stars are still on the main sequence, and the light integral is dominated by exactly the stars that leave it first. From 0.0047 at a million years to 18 at 13.0 Gyr, a factor of 3884. The marked ages carry their turnoff masses: 10 Myr at 14 M☉, 100 Myr at 5.5 M☉, 1 Gyr at 2.2 M☉, 10 Gyr at 1.0 M☉. A galaxy's mass is inferred from its light through this number, so a mass is an age assumption before it is a measurement — and the mass here is the initial mass, remnants included at their birth weight, which is the approximation this figure makes.
Fig. 8 The mass-to-light ratio against age for the Salpeter slope alone. It rises by more than an order of magnitude over ten billion years, because the stars carrying the light die and the stars carrying the mass do not — which is why an old population’s mass cannot be read off its brightness without a model of its age.

One more reading contrasts the standard slope with a steeper one over the same integrands.

One population, three integrals, two ends. The same mass function weighted three ways, each normalised so that the area under it is one, against mass on a logarithmic axis — so a share of the page is a share of the total. The light is a zero-age population's, every star still on the main sequence, which is the only age at which the top of the range is present at all. The steep curve on the left is the number of stars, which peaks at 0.08 M☉ and falls away because the function is steeper than m⁻¹; the middle one is the mass they carry; the one on the right is the light they emit, computed from this file's own main-sequence relation and peaking at 55.0 M☉ — 682 times further up the axis. The population that is counted and the population that is seen are two different populations. Above 8 M☉ there are 0.060% of the stars, 4.3% of the mass and 98% of the light. The slopes are α = 2.35 above 0.08 M☉, α = 2.7 above 0.5 M☉, and the low-mass end is where the honesty runs out: it has never been measured in another galaxy, and every mass inferred from a luminosity assumes it.
Fig. 9 The number and light integrands for the Salpeter slope and for one steeper by a third. The crossing between them moves and the qualitative statement does not: the number is always dominated by the smallest stars and the light by the largest, because the two integrands have opposite signs of slope whatever the exponent.

Where the ladder goes next

The rung above is population synthesis proper: convolving the mass function with a star-formation history and a set of evolutionary tracks to produce a predicted spectrum, which is how every galaxy’s mass, age and metallicity are actually estimated. The rung below — and it is the one that would settle the most — is the physics of fragmentation, which would say where the characteristic mass comes from and therefore whether it should be expected to move at low metallicity.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 14 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Brown dwarfInitial mass functionMass luminosity relationMass-to-light ratioPopulation synthesisPower lawSalpeter slopeStar formation rateStellar lifetimesStellar population