Galaxies

The count theory predicts, and the inference it costs

A luminosity function is measured. A stellar mass function is inferred, one galaxy at a time, through a ratio that runs by a factor of six from the bluest galaxies to the reddest — so the conversion changes the shape and not merely the units.

Assumes Luminosity function and Stellar evolution.

The luminosity function has a knee in it, and the knee is where galaxy formation stops being efficient. That is a statement about mass, made with a measurement of light, and the step between the two is the subject here.

No theory of galaxy formation predicts a luminosity. What a simulation produces is a mass of stars, and turning it into a magnitude requires a stellar population model, a dust model and a filter — three assumptions added to a prediction. The alternative is to move the data instead, and convert every observed galaxy’s luminosity into a mass. That is what is done, and the conversion is not a rescaling of an axis.

The count theory predicts, and the inference it costs. The galaxy stellar mass function: galaxies per cubic megaparsec per dex of stellar mass, both axes logarithmic. Two Schechter components share a characteristic mass of 10^10.66 M☉ — one of slope -0.35 carrying the quenched galaxies at the knee, one of slope -1.47 carrying the star-forming ones below it — and the dashed line is the single component a luminosity function is usually fitted with. Integrated over the range drawn it gives 0.0487 galaxies per cubic megaparsec holding 2.22·10⁸ solar masses of stars, of which 51 per cent sits above the knee. This function is not measured. What is measured is a luminosity function; turning one into the other needs a mass-to-light ratio for every galaxy in the sample, and that ratio is not a constant — it runs by a factor of about five from the bluest galaxies to the reddest, so the conversion moves the red end of the distribution further than the blue end and changes the SHAPE rather than the units. A stellar mass function is a luminosity function plus a stellar population model, and the second half is where its disagreements live.
Fig. 1 The galaxy stellar mass function: two Schechter components sharing a characteristic mass of 1010.66M10^{10.66}\,M_\odot, one of slope 0.35-0.35 carrying the quenched galaxies at the knee and one of slope 1.47-1.47 carrying the star-forming ones below it. Integrated over the range drawn it gives 0.0487 galaxies per cubic megaparsec holding 2.22×1082.22\times10^{8} solar masses of stars, of which 51 per cent sits above the knee. The two-component shape is itself the evidence — a single Schechter function is too low at the faint end, and the second component is there because the population is two populations.

Two functions, not one

The most visible difference from the luminosity function is that this one needs two Schechter components and the luminosity function is usually fitted with one.

That is not a failure of the single fit. It is a consequence of the mass-to-light ratio being different for the two populations, so that converting from light to mass stretches the red galaxies further than the blue ones and separates two distributions that overlapped in luminosity.

Which means the two-component form carries information. The high-mass component has a shallow slope and holds the quenched, red, passively evolving galaxies; the low-mass one is steep and holds the star-forming ones. The colour bimodality says the same thing in a different plane, and the fact that both planes require two components for the same galaxies is the strongest argument that the split is real rather than a fitting artefact.

The transition between them happens near the knee, which is consistent with everything else about the knee: it is where the galaxies stop growing.

What one number per galaxy costs

The conversion runs through a mass-to-light ratio, and the standard way to get one is from a colour. An old red population has more mass per unit light than a young blue one, because the massive stars that dominated the light have died and the low-mass stars that carry the mass have not.

The relation used is empirical and short:

log(M/Li)=0.70(gi)0.68,\log(M_\star/L_i) = 0.70\,(g-i) - 0.68 ,

calibrated against stellar population synthesis models fitted to full spectra, and carrying a scatter of about a tenth of a dex that nothing in a broadband colour removes.

A factor of five, read off one colour. The stellar mass-to-light ratio against colour, on the relation the mass function is built with: log(M⋆/Lᵢ) = 0.7(g − i) − 0.68, with 460 galaxies scattered about it by 0.1 dex. Across the colour range drawn the ratio runs by a factor of 5.9 — a red galaxy holds 5.9 times the mass per unit light that a blue one of the same luminosity does — which is why converting a luminosity function into a mass function is not a shift of the axis. The red end moves further than the blue end, and the knee moves with it. The scatter is the other half. Applied to a count whose logarithmic slope is -1.25, a 0.1 dex scatter inflates the count at a fixed mass by 4.2 per cent — the Eddington bias, an inflation rather than a smearing, because more galaxies scatter up from the steep side of the distribution than down from the shallow one. Where it matters is the other end. At a few times the knee the exponential has taken over and the local slope is near -3.25; the same scatter then inflates the count by 32 per cent, and it steepens without limit beyond that. The brightest galaxies in a mass function are the ones the mass-to-light ratio's scatter has most inflated, and the correction is a property of the scatter rather than of the galaxies. What no colour can supply is the thing the ratio actually depends on, which is the distribution of stellar ages in the galaxy; two populations with the same g − i and different histories differ in M⋆/L by more than this scatter, and the relation absorbs that difference into a number it calls noise.
Fig. 2 The relation, with 460 galaxies scattered about it by 0.1 dex. Across the colour range drawn the ratio runs by a factor of 5.9 — a red galaxy holds nearly six times the mass per unit light that a blue one of the same luminosity does. That is why the conversion changes the shape. The red end of the distribution moves further than the blue end, the knee moves with it, and the two populations that overlapped in luminosity separate in mass.

The factor of six is the important number and it is easy to underrate. A luminosity function and a mass function are often drawn as though they were the same curve on relabelled axes, and if the ratio were constant they would be. It is not constant, it varies systematically with the quantity being plotted, and a systematic variation along the axis is precisely what deforms a distribution.

Where the two functions differ most

It is worth being concrete about which parts of the curve move, because the answer is not uniform and the parts that move most are the parts most quoted.

The knee moves furthest. Galaxies at the knee are a mixture of red and blue with the red ones dominating the mass, so the characteristic mass sits at a mass-to-light ratio near the red end of the range while the characteristic luminosity was set by a mixture. The two characteristic values are therefore not related by the mean ratio, and computing one from the other with an average is wrong by tens of per cent.

The faint end moves least. Low-mass galaxies are almost all blue and almost all star-forming, so the ratio is nearly constant across them, and the faint-end slope of the mass function is close to the faint-end slope of the luminosity function. That near-invariance is useful: it means the steepest and least well measured part of the curve is the part least corrupted by the conversion.

And the bright tail moves in a way that depends on the scatter rather than on the mean, which is the subject of the next section and is the largest of the three effects.

Those three statements together are why a mass function is fitted independently rather than derived from a published luminosity function. There is no transformation to apply; the conversion has to be made galaxy by galaxy, before any function is fitted at all.

The scatter is not noise

The second half of the ratio’s cost is its scatter, and the scatter does something a symmetric error usually does not: it biases the answer in one direction.

Applying a scatter of width σ\sigma in log-mass to a count falling with a logarithmic slope γ\gamma multiplies the count at a fixed mass by exp(12γ2σ2ln210)\exp(\tfrac12\gamma^2\sigma^2\ln^2 10). The exponent has no sign in it, so the effect is always an inflation — a scatter on a falling distribution promotes more galaxies into a bin than it demotes out of it, because there were more below than above.

That is the Eddington bias, and its size depends on where on the function it is applied.

In the power-law region, with a slope near 1.25-1.25, a tenth of a dex inflates the count by 4.2 per cent — negligible against everything else. At the exponential cut-off, a few times the knee, the local slope steepens past 3-3, and the same scatter inflates the count by about a third. Further out it steepens without limit.

So the brightest galaxies in a mass function are the ones the scatter has most inflated, and the correction is a property of the measurement rather than of the galaxies. Any comparison between an observed high-mass tail and a simulated one has to apply the observational scatter to the simulation rather than removing it from the data, because the simulation has no scatter and the data cannot be un-scattered.

A factor of five, read off one colour. The stellar mass-to-light ratio against colour, on the relation the mass function is built with: log(M⋆/Lᵢ) = 0.7(g − i) − 0.68, with 460 galaxies scattered about it by 0.2 dex. Across the colour range drawn the ratio runs by a factor of 5.9 — a red galaxy holds 5.9 times the mass per unit light that a blue one of the same luminosity does — which is why converting a luminosity function into a mass function is not a shift of the axis. The red end moves further than the blue end, and the knee moves with it. The scatter is the other half. Applied to a count whose logarithmic slope is -1.25, a 0.2 dex scatter inflates the count at a fixed mass by 18.0 per cent — the Eddington bias, an inflation rather than a smearing, because more galaxies scatter up from the steep side of the distribution than down from the shallow one. Where it matters is the other end. At a few times the knee the exponential has taken over and the local slope is near -4; the same scatter then inflates the count by 446 per cent, and it steepens without limit beyond that. The brightest galaxies in a mass function are the ones the mass-to-light ratio's scatter has most inflated, and the correction is a property of the scatter rather than of the galaxies. What no colour can supply is the thing the ratio actually depends on, which is the distribution of stellar ages in the galaxy; two populations with the same g − i and different histories differ in M⋆/L by more than this scatter, and the relation absorbs that difference into a number it calls noise.
Fig. 3 Twice the scatter, and a steeper place to apply it. The relation is unchanged and the points spread over twice the range; the inflation at the shallow slope rises to 17 per cent and at the cut-off to a factor of three and a half. A disagreement about the most massive galaxies in the universe can therefore be a disagreement about how well a colour predicts a mass-to-light ratio, which is a question about stellar populations and not about galaxy formation at all.

What a colour cannot know

The relation absorbs into its scatter a thing that is not random, and this is the honest limit of the whole method.

A galaxy’s mass-to-light ratio depends on the distribution of ages in its stellar population — how much of its mass was formed when. A colour is one number and an age distribution is a function, so the mapping from the second to the first is many-to-one, and galaxies with identical colours can have genuinely different ratios.

The worst case is well known and is called outshining. A galaxy that formed almost all of its stars ten billion years ago and then made a small burst of new ones recently will be blue, because the young stars dominate the light, while nearly all of its mass is in the old population the colour cannot see. The inferred mass can be too low by a factor of two or more.

A factor of five, read off one colour. The stellar mass-to-light ratio against colour, on the relation the mass function is built with: log(M⋆/Lᵢ) = 0.7(g − i) − 0.68, with 460 galaxies scattered about it by 0.05 dex. Across the colour range drawn the ratio runs by a factor of 9.5 — a red galaxy holds 9.5 times the mass per unit light that a blue one of the same luminosity does — which is why converting a luminosity function into a mass function is not a shift of the axis. The red end moves further than the blue end, and the knee moves with it. The scatter is the other half. Applied to a count whose logarithmic slope is -1.25, a 0.05 dex scatter inflates the count at a fixed mass by 1.0 per cent — the Eddington bias, an inflation rather than a smearing, because more galaxies scatter up from the steep side of the distribution than down from the shallow one. Where it matters is the other end. At a few times the knee the exponential has taken over and the local slope is near -3.25; the same scatter then inflates the count by 7 per cent, and it steepens without limit beyond that. The brightest galaxies in a mass function are the ones the mass-to-light ratio's scatter has most inflated, and the correction is a property of the scatter rather than of the galaxies. What no colour can supply is the thing the ratio actually depends on, which is the distribution of stellar ages in the galaxy; two populations with the same g − i and different histories differ in M⋆/L by more than this scatter, and the relation absorbs that difference into a number it calls noise.
Fig. 4 The relation across a wider colour range with a smaller scatter, which is what a measurement with excellent photometry and no population problem would look like: the ratio now spans a factor of nine and the points lie almost on the line. The scatter drawn here is the photometric one, and that is not the limiting term. The systematic uncertainty from the population’s unknown age distribution is larger than the scatter in every published version of this relation, and it is not reduced by measuring the colour better.

The escape is more information: a full spectrum instead of a colour, or many bands instead of two, or a near-infrared luminosity, which is dominated by old stars and therefore tracks the mass more directly. Each helps and none closes the gap, because all of them still depend on the initial mass function.

The count theory predicts, and the inference it costs. The galaxy stellar mass function: galaxies per cubic megaparsec per dex of stellar mass, both axes logarithmic. Two Schechter components share a characteristic mass of 10^10.66 M☉ — one of slope -0.35 carrying the quenched galaxies at the knee, one of slope -1.7 carrying the star-forming ones below it — and the dashed line is the single component a luminosity function is usually fitted with. Integrated over the range drawn it gives 0.239 galaxies per cubic megaparsec holding 3.05·10⁸ solar masses of stars, of which 38 per cent sits above the knee. This function is not measured. What is measured is a luminosity function; turning one into the other needs a mass-to-light ratio for every galaxy in the sample, and that ratio is not a constant — it runs by a factor of about five from the bluest galaxies to the reddest, so the conversion moves the red end of the distribution further than the blue end and changes the SHAPE rather than the units. A stellar mass function is a luminosity function plus a stellar population model, and the second half is where its disagreements live.
Fig. 5 The same knee with a steeper and more abundant low-mass component — a faint-end slope of 1.7-1.7, which is towards the upper end of what deep surveys allow. The stellar mass density rises and the count rises far more, since the number diverges at that slope while the mass does not. The two ends of this function answer different questions and are measured by different surveys: the knee by a wide shallow survey with enough volume to contain rare massive galaxies, the faint end by a narrow deep one, and the two are joined in the middle by an overlap that is never as large as either would like.

That join is where most of the disagreement between published mass functions lives. Two surveys agreeing on the knee and disagreeing on the faint-end slope produce functions that differ by a factor of three at 108M10^{8}\,M_\odot and by nothing at all at 101110^{11}, and their integrated mass densities differ by rather little, because almost none of the mass is at the faint end. The same asymmetry between counting and weighing runs through the whole of this subject.

The assumption nothing measures

Underneath all of this sits a quantity that no extragalactic observation constrains at all, and it enters every mass in this essay multiplicatively.

The stellar mass of a galaxy is inferred from its light, and the ratio of the two depends on how many faint stars accompany each bright one. The light comes from stars above about two solar masses; the mass comes from stars below one. Nothing in the integrated light of a galaxy says how many of the second there are per unit of the first.

So the initial mass function is assumed, and it is assumed to be the same everywhere. A different assumption shifts every stellar mass in the universe by a constant factor — a Salpeter function gives masses about 1.7 times a Chabrier one — and that factor cancels out of most comparisons and survives in the ones that matter, such as the ratio of stellar mass to halo mass.

And there is evidence that it is not universal. The most massive elliptical galaxies show spectral features suggesting an excess of low-mass stars relative to a Chabrier function, which would raise their masses by up to a factor of two — exactly at the knee, and exactly in the population whose mass function is the most-compared quantity in the field.

The count theory predicts, and the inference it costs. The galaxy stellar mass function: galaxies per cubic megaparsec per dex of stellar mass, both axes logarithmic. Two Schechter components share a characteristic mass of 10^10.89 M☉ — one of slope -0.35 carrying the quenched galaxies at the knee, one of slope -1.47 carrying the star-forming ones below it — and the dashed line is the single component a luminosity function is usually fitted with. Integrated over the range drawn it gives 0.0618 galaxies per cubic megaparsec holding 3.77·10⁸ solar masses of stars, of which 50 per cent sits above the knee. This function is not measured. What is measured is a luminosity function; turning one into the other needs a mass-to-light ratio for every galaxy in the sample, and that ratio is not a constant — it runs by a factor of about five from the bluest galaxies to the reddest, so the conversion moves the red end of the distribution further than the blue end and changes the SHAPE rather than the units. A stellar mass function is a luminosity function plus a stellar population model, and the second half is where its disagreements live.
Fig. 6 The same function with every mass raised by 0.23 dex — the shift between a Chabrier and a Salpeter initial mass function. The shape is untouched and the whole distribution slides along the axis, which is the honest picture of what that assumption does. It is a nuisance parameter that cannot be measured from the data being fitted, and its effect is exactly degenerate with a change in the knee, so a paper reporting a characteristic mass without naming its initial mass function has reported a number with a factor of 1.7 missing from it.

The integral that is the point

A mass function is most often used not as a curve but as one number: its integral, weighted by mass, which is the stellar mass density of the universe.

The value drawn here is 2.22×1082.22\times10^{8} solar masses per cubic megaparsec, and it is one of the few quantities in extragalactic astronomy that can be compared directly against an entirely independent accounting. The cosmic star formation history integrates to a stellar mass formed, and after subtracting the fraction that has been returned to the interstellar medium by winds and supernovae, the two should agree.

For a long time they did not. The integral of the star formation history exceeded the measured stellar mass density by a factor approaching two — a discrepancy large enough that one of the two measurements had to be wrong, and small enough that nobody could say which.

The gap has closed from both sides. The star formation rates came down as the dust corrections improved; the masses went up as deeper imaging found the faint outskirts of galaxies that shallower surveys had cut off. What resolved it was not a new idea but two measurements each moving by twenty per cent, which is the ordinary way a factor-of-two disagreement ends and is worth remembering when the next one appears.

The residual disagreement is at the level of the initial mass function’s own uncertainty, which is to say that the comparison is now testing the same assumption on both sides and can no longer discriminate.

The count theory predicts, and the inference it costs. The galaxy stellar mass function: galaxies per cubic megaparsec per dex of stellar mass, both axes logarithmic. Two Schechter components share a characteristic mass of 10^11.00 M☉ — one of slope -0.35 carrying the quenched galaxies at the knee, one of slope -1.47 carrying the star-forming ones below it — and the sum is what is measured. Integrated over the range drawn it gives 0.0693 galaxies per cubic megaparsec holding 4.86·10⁸ solar masses of stars, of which 51 per cent sits above the knee. This function is not measured. What is measured is a luminosity function; turning one into the other needs a mass-to-light ratio for every galaxy in the sample, and that ratio is not a constant — it runs by a factor of about five from the bluest galaxies to the reddest, so the conversion moves the red end of the distribution further than the blue end and changes the SHAPE rather than the units. A stellar mass function is a luminosity function plus a stellar population model, and the second half is where its disagreements live.
Fig. 7 The same decomposition with a characteristic mass a third of a dex higher and the single-component comparison removed, so that the two Schechter terms can be read against each other directly. The knee’s position is the most-quoted number this function has and the least robustly measured, because it sits where the exponential begins and an exponential’s foot is where a fit is least constrained. Moving it by a third of a dex changes the stellar mass density by tens of per cent and changes the faint end not at all — which is why two surveys can agree on every galaxy they both observe and publish different characteristic masses.

The sensitivity of the knee is worth one more sentence because it explains a pattern in the literature. Published characteristic masses for the local stellar mass function span about 0.3 dex, and the spread is not random: surveys using shallower imaging systematically find a lower knee, because they miss the faint outer envelopes of the most massive galaxies and therefore underestimate their masses. A galaxy’s light profile falls off slowly enough that a surface-brightness cut removes a real and mass-bearing part of it, and the part removed is largest for the largest galaxies.

So the knee is biased by the same effect that biases the faint-end slope, acting on different objects. Both are consequences of a survey having a surface-brightness limit as well as a flux limit, which is the second selection every galaxy catalogue has and the one least often stated.

Nowhere in the chain is a galaxy weighed

The luminosity function is measured, from a redshift survey with a well-defined selection and a completeness correction. That part is photometry and volume, and it is solid.

The colours are measured, to a few hundredths of a magnitude.

The mass-to-light relation is calibrated, against stellar population synthesis models fitted to spectra of galaxies whose masses were themselves derived from the same class of model. There is no independent measurement of a galaxy’s stellar mass anywhere in the chain — no galaxy has ever been weighed by counting its stars — so the calibration is internal.

The one external check is dynamical. A galaxy’s rotation or velocity dispersion gives a total mass within some radius, and subtracting an estimate of the dark matter leaves a stellar mass. The agreement is at the level of a factor of 1.3, which is both reassuring and about the size of the assumptions being tested.

Strong lensing offers a second, and it is the cleanest test available. A lensing mass is a pure geometry, depending on nothing but the deflection angles and the distances, so an elliptical galaxy acting as a lens gives a total mass inside the Einstein radius with no stellar population model anywhere in it. Subtract a dark matter profile and what remains is a stellar mass that owes nothing to a colour.

Those measurements are what produced the evidence for a varying initial mass function: the lensing masses of the most massive ellipticals exceed their colour-derived stellar masses by more than any plausible dark matter contribution allows, unless the low-mass stellar population is richer than assumed. It is a handful of galaxies against a survey of hundreds of thousands, and it is the only place the assumption has been tested rather than adopted.

A fitting form, an extrapolation, and an error that is not symmetric

The two Schechter components are a fitting form, not two identified populations. Galaxies are not labelled, and the decomposition into a shallow high-mass component and a steep low-mass one is a description of the shape rather than an assignment of objects. A galaxy near the knee belongs partly to each.

The faint end below about 108M10^{8}\,M_\odot is an extrapolation. Surveys complete enough to measure a mass function that low cover volumes too small to contain a representative sample, and the ones with large volumes are not complete that faint. The steep slope drawn there is a fit constrained mostly from above.

And the drawn scatter is symmetric in log-mass, which the real one is not. Outshining biases masses low rather than scattering them, so the real distribution of errors has a tail on one side, and an asymmetric error produces a different Eddington correction from the symmetric one computed here.

Why the trade was worth making

It is worth being explicit about why a measured quantity is given up for an inferred one, since the essay has spent most of its length on what the inference costs.

A luminosity is not a conserved quantity. A galaxy’s light varies by a factor of several over a few hundred million years as its star formation rises and falls, and two galaxies with identical histories observed at different moments have different luminosities. Its stellar mass only ever goes up, slowly, and by an amount that is the integral of everything that has happened to it.

So the mass function is the quantity that connects to a history, and the luminosity function is a snapshot of a population’s current activity. A theory predicts an integral and a telescope measures a derivative, and one of the two has to be converted.

The conversion has been made in the direction that puts the uncertainty on the observation rather than on the theory, and that choice is worth examining, since it is not the only one available. Converting the models into luminosities — forward modelling — puts the same population synthesis into the prediction instead, and is increasingly what is done, because a model can carry its full star formation history through the synthesis rather than compressing it into one colour.

The shape this argument has elsewhere

The structure — a measurable quantity that is not the one the theory speaks about, and a conversion with a systematic in it — is one of the most common in the subject.

A velocity semi-amplitude gives a planet’s minimum mass and the true mass needs an inclination nothing in the velocities supplies. A cluster’s richness gives its mass through a relation calibrated on a subsample. A line width gives a galaxy’s rotation speed, through an inclination correction and an assumption about the gas’s distribution.

In each case the honest statement is the same: the quantity reported is a measurement multiplied by a model, and quoting the measurement’s error bar alone understates the uncertainty by whatever the model is worth. What distinguishes them is whether the model can be checked independently, and for the stellar mass-to-light ratio the answer is barely.

What changes if the ratio is measured instead of inferred

There is one class of galaxy where the mass-to-light ratio is not inferred from a colour, and the comparison is instructive.

A galaxy whose stars can be resolved individually — the nearest dwarfs, and the Magellanic Clouds — can be weighed by counting. Every star is placed on a colour–magnitude diagram, the diagram is fitted with a set of stellar isochrones, and the total mass follows from the numbers in each mass bin plus an extrapolation below the detection limit.

The extrapolation is the same initial mass function assumption as before, so the method does not escape that. What it does escape is the compression of a whole history into one colour, and the comparison between the two methods on the same galaxy is therefore a direct test of the outshining problem.

The answer is that the colour-based ratio is systematically low for galaxies with recent bursts, by up to a factor of two in the worst cases, and correct to a few tens of per cent for galaxies with smooth histories. Which is exactly what the outshining argument predicts, measured rather than argued, and it is available for a few dozen galaxies out of the hundreds of thousands the mass function is built from.

Still open: what the mass function is being compared to

Every use of this function is a comparison, and the comparison is almost never with another observation.

It is compared with the output of a simulation, and the simulation produces galaxies inside dark matter halos whose own mass function is a calculation rather than a measurement. So the real question is how the two functions relate to each other — how much stellar mass a halo of a given mass has managed to assemble — and that ratio is the quantity every theory of feedback is a theory of.

Drawing the two counts on one pair of axes settles it. They are not the same shape, and they disagree at both ends.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

CompletenessEddington biasFaint end slopeInitial mass functionLuminosity functionMass-to-light ratioQuenchingSchechter functionStellar mass functionStellar population