A sample brighter than the population it came from
Assumes Distance ladder and Magnitudes.
A brightness is a distance only if something is known, and the something is usually an absolute magnitude assumed for a class of object.
A standard candle is a class of object assumed to have a known absolute magnitude. The known magnitude comes from calibrating the class on nearby examples whose distances are known some other way, and then the class is used further out.
The step that goes wrong is the calibration, and it goes wrong for a reason that has nothing to do with the objects. A survey detects what is bright enough to detect. At any given distance the faint members of the class fall below the limit and the bright ones do not, so the sample used for the calibration is not a sample of the class — it is a sample of its bright tail, and the tail gets brighter the further out the survey looks.
The effect is named after Gunnar Malmquist, who quantified it in 1920 for a problem in stellar statistics, and it has been rediscovered under other names in nearly every subfield since. It is not subtle, it is not controversial, and it is still routinely absent from analyses that need it — because it does not announce itself. Every distance in the ladder is measured with the last one, and a bias in one rung is a multiplicative factor on everything above.
Two biases with one name
The literature calls two different things Malmquist bias and they are worth separating, because the corrections are different and applying the wrong one is common.
The classical, or homogeneous, Malmquist bias is about a sample as a whole. Take a class of objects with a Gaussian luminosity function of width , distributed uniformly in a Euclidean space, and observe them to a limiting apparent magnitude. The mean absolute magnitude of everything detected is brighter than the population’s by exactly magnitudes — a coefficient that is and comes entirely from the volume element growing as the cube of the distance.
That is the version derived by Malmquist in 1920 and it is a statement about a sample, not about an object. Applying it as a correction to an individual star is meaningless, and it is done.
The second, or inhomogeneous, Malmquist bias is about individual objects at a known distance. If a set of objects is selected at a distance where the survey is incomplete, their mean absolute magnitude is brighter by an amount depending on how far into the luminosity function the limit has cut. That is what the figure above shows, it depends on distance, and it is what actually corrupts a calibration.
The two coincide only for the specific geometry the first assumes. For a sample drawn from a cluster, a survey with a complicated selection function, or a class whose luminosity function is not Gaussian, the classical coefficient is simply the wrong number.
It is worth being explicit about where the classical coefficient comes from, because its form explains its limits. In a Euclidean universe the number of objects within a distance grows as the cube of that distance, so at any apparent magnitude there are far more objects at the far edge of the detectable volume than near its centre. Those distant ones are detectable only if they are luminous. Integrating the luminosity function against that volume weighting gives an exponential tilt towards bright objects, and for a Gaussian luminosity function an exponential tilt shifts the mean without changing the shape — by exactly . Everything about that derivation depends on the cube: change the geometry, change the coefficient.
What it does to a distance
The consequence for the distance ladder is direct and it has a sign.
A calibration that returns an absolute magnitude that is too bright, applied to a distant object of the same class, gives a distance modulus that is too small — the object is inferred to be closer than it is. Every distance measured with the calibration comes out short, and every derived quantity that scales with distance comes out wrong in the corresponding direction: luminosities too low, sizes too small, and the expansion rate too high.
The size is not negligible. For a class with a luminosity scatter of half a magnitude the classical bias is 0.35 magnitudes, which is seventeen per cent in distance. For a scatter of a magnitude it is 1.4 magnitudes and a factor of two.
That dependence on the square of the scatter is the single most useful thing to know about the effect. It means the bias is negligible for a tight relation and severe for a loose one, and it means that halving a class’s intrinsic scatter reduces the bias fourfold rather than twofold. Most of the history of the distance ladder is a search for standard candles with small , and that search is usually motivated by precision; the reduction in bias is the larger gain.
There is a second consequence that is less often stated and is arguably worse. The bias is not merely an offset — it changes with distance, so it distorts the shape of any relation fitted across a range of distances. A Hubble diagram built from a magnitude-limited sample curves, because the mean absolute magnitude of the sample brightens with distance, and the curvature has the same sign as the one an accelerating universe would produce. Distinguishing the two requires modelling the selection to a precision comparable to the signal, and the history of the subject contains more than one claimed cosmological result that turned out to be a selection function.
The historical case
The effect was found in exactly this way and it caused the largest single revision of the distance scale ever made.
Hubble’s original calibration of the extragalactic distance scale used Cepheid variables in nearby galaxies. Several problems were later found in it — the confusion of two distinct classes of variable star with different period–luminosity relations was the largest — and selection bias was another. The brightest Cepheids in a distant galaxy are the only ones detectable, so the period–luminosity relation was being calibrated on the bright envelope of the relation rather than on its ridge line.
The cumulative effect of all the problems was a revision of the extragalactic distance scale by a factor of about ten between 1929 and the 1960s, and a corresponding reduction in the derived expansion rate. Selection bias was not the whole of that factor and it was a substantial part of it.
The modern version of the same argument is quieter and is still active. Every rung of the ladder is calibrated on a sample that is detected, and at every rung the question of whether the selection has been modelled correctly is one of the largest terms in the error budget.
One detail of the Cepheid case is worth carrying because it shows how the bias hides. The period–luminosity relation has intrinsic scatter, so at a given period the Cepheids in a distant galaxy span a range of luminosities. Detecting only the brightest of them means the relation’s zero point is measured too bright, and — because longer-period Cepheids are more luminous and therefore detectable further out — the bias varies with period, which tilts the fitted slope as well. A tilted slope changes the distance in a way that depends on which periods were observed in which galaxy, which is a systematic that does not even have a single value.
What can be done about it
Four responses, in increasing order of both cost and reliability.
Correct it. Compute the bias from an assumed luminosity function and selection limit and subtract it. This works when the luminosity function is known and the selection is simple, and it fails silently when either assumption is wrong — which is often, because the luminosity function is usually measured from the same biased sample.
Use a volume-limited sample. Restrict the analysis to a distance within which the survey is complete for the faintest member of the class. That removes the bias entirely and it discards most of the data, because the complete volume is small.
Fit in the observed space. Rather than converting magnitudes to absolute magnitudes and averaging, model the observation directly: predict what apparent magnitude each object should have given the relation’s parameters and its distance, apply the survey’s selection function, and compare against what was detected. The bias does not arise, because nothing is being averaged over a truncated distribution.
Use the inverse relation. For a scaling relation — a period–luminosity relation, a Tully–Fisher relation — regressing the predictor on the observable rather than the other way round is much less sensitive to selection on the observable. That is the standard defence in extragalactic distance work and it is not intuitive: the two regressions differ, and which one to use depends on which variable the selection acted on.
The fourth response deserves elaboration because it is the least intuitive and the most used. Fitting on and fitting on give different lines whenever there is scatter, and the difference is not a technicality: the forward fit is unbiased if the selection acted on , and the inverse fit is unbiased if it acted on . For a Tully–Fisher relation the survey selects on magnitude, which is the quantity being predicted, so the inverse fit — regressing line width on magnitude — is the one that survives. That prescription is standard, it is counter-intuitive to anybody trained to put the “dependent” variable on the vertical axis, and getting it backwards produces a distance error of several per cent with no other symptom.
What was actually measured
Three demonstrations, using three different kinds of evidence.
The bias measured against complete samples. Where a class has been surveyed to well below its faint end in a small volume, the luminosity function is known without bias, and the same class observed at greater distance can be compared against it directly. The difference is the bias, and it comes out at the size the truncated-Gaussian calculation predicts.
The distance-dependent trend. This is the same diagnostic that reveals a contaminated crater count and a continuum placed too low, and it is unmistakable and costs nothing to look for: plot the inferred absolute magnitude of a class against the distance at which each member was found. It should be flat. In a magnitude-limited sample it rises, and the shape of the rise is the selection function.
The disagreement between forward and inverse fits. Fitting a scaling relation both ways on the same data gives two different slopes, and the size of the difference is a measurement of how much the selection has bitten. Where the two agree, the sample is effectively complete; where they diverge, the sample is not, and the divergence is quantitative.
Where the picture stops
Three, and the second is the one that generalises furthest.
Real selection functions are not sharp. A survey does not detect everything above a limit and nothing below it; the detection probability falls smoothly over a magnitude or so, and it depends on the sky background, the seeing, the crowding and the object’s morphology. Modelling the selection therefore means modelling the survey, and the accuracy of the correction is bounded by that.
Selection on any correlated quantity does the same thing. Malmquist bias is selection on flux, which correlates with luminosity. Selecting on angular size, on surface brightness, on colour, or on the significance of a detection all produce analogous biases in whatever the selected quantity correlates with. A sample chosen on the significance of its own measurement is the same effect wearing a different hat.
And the correction cannot be applied twice. Because the classical and inhomogeneous versions are different statements about different things, and because some analyses correct for one while quoting a luminosity function derived with the other already removed, double-correction is a real and hard-to-detect error. The diagnostic is the same as the discovery: plot the answer against distance and see whether it is flat.
Why this is a property of the survey rather than of the sky
The general shape is one this collection returns to constantly, and it is worth stating in its cleanest form.
What a survey finds is the product of what is there and what it could have seen. Recovering the first requires dividing by the second, and the second is a property of the instrument, the site, the observing strategy and the analysis — none of which is astronomy. It follows that a survey’s selection function is part of its data, and a survey that did not measure its own selection function has published a catalogue that cannot be turned into a statement about the universe.
It is also why a magnitude has to say which light it was measured in and a catalogue has to say what it could not have seen: both are statements that a number is only interpretable alongside a description of how it was obtained.
That is now standard practice: modern surveys inject artificial sources into their own images and measure what fraction are recovered as a function of magnitude, position and crowding. The resulting completeness curve is published alongside the catalogue, and it is what makes the catalogue usable for anything statistical.
The same discipline applies to exoplanet surveys, where the first exoplanets found were enormous and impossibly close to their stars — a fact about spectrographs rather than about planets. It applies to galaxy counts, to cluster catalogues, and to the census of the solar system’s small bodies. In each case the honest statement is not what was found but what was found divided by what could have been.
A last observation about why this belongs to the distance ladder specifically rather than to statistics. The ladder is a chain of multiplicative calibrations, and its distinguishing feature is that errors at the bottom propagate to the top undiluted. A five per cent bias in the parallax calibration of Cepheids is a five per cent bias in the distance to a galaxy a hundred megaparsecs away, and no amount of data at the top removes it. That is why a brightness becomes a distance only if something is known, and why the something has to be known without the help of anything above it. Selection bias is dangerous here not because it is large but because it is a bias — the one kind of error a chain of measurements cannot dilute.
One structural point is worth adding because it explains why the bias is so persistent despite being understood for a century. The correction depends on the luminosity function’s width at the relevant luminosity, and the width is measured from the same flux-limited sample that needs correcting. So the correction requires knowing the quantity whose measurement it is correcting, and the standard escape is to measure the width in a volume-limited subsample — the nearby objects, where everything is detected — and assume it applies further away. That assumption is exactly the one a population evolving with redshift violates. The result is that the correction is reliable where it is small and uncertain where it is large, which is the general condition of every bias whose size depends on the distribution it is being applied to, and it is why the honest treatments fit the luminosity function and the distances simultaneously rather than correcting one with the other.
The bias also has a name problem worth flagging: the same word is used for the effect on a magnitude-limited sample and for the quite different effect on a distance-limited one, and papers rarely say which they mean. Reading a correction without knowing which of the two it was derived for is how a fix in the right direction ends up the wrong size.
Where the ladder goes next
Later rungs on this ladder start with the second bias in the same family: what happens at the faint end of a steep number-count distribution, where scatter moves more objects up than down and the counts themselves are wrong. The rung after that is the practical machinery — how a survey measures its own completeness by injecting sources it knows about into images it has already taken.
About the same objects
Not linked from either essay — found by the objects both name.
- A count with a knee in it completeness · luminosity function · malmquist bias · volume-limited sample
- The count theory predicts, and the inference it costs completeness · eddington bias · luminosity function
- A line width that is a distance malmquist bias · standard candle
- The same census, taken in two places completeness · luminosity function
What links here
Essays that link to this one from their own argument.
- A mass function corrected by an age stars
- A planet radius is a stellar radius exoplanets
- The distance is not one over the parallax starlight
- A forest with no continuum left cosmology
- A temperature and a gravity that trade against each other starlight
The objects this essay names
Each one links to every other essay that touches it.
CompletenessEddington biasFlux-limited sampleInverse regressionLuminosity functionMalmquist biasSelection functionStandard candleSystematic errorVolume-limited sample