Starlight

The distance is not one over the parallax

A parallax is measured with symmetric errors and a distance is one over it. Inverting a noisy positive quantity is not a change of units — it is a change of distribution, and the one that comes out is skewed, biased outward, and above about twenty per cent error has no mean at all.

Assumes Parallax and Distance ladder.

Trigonometric parallax is the only distance in astronomy measured by geometry alone. The triangle closes, nothing about the star’s physics enters, and the answer is a length in metres. That is why it sits at the bottom of every other distance measurement in the subject.

What is measured, though, is an angle, and what is wanted is its reciprocal. Those two are not the same quantity wearing different units. A measurement with symmetric errors, inverted, produces an estimate with asymmetric errors, a shifted mean, and — at large enough errors — no mean at all.

One over a noisy parallax, at three precisions. The distribution of the distance obtained by inverting a parallax, for a star truly at 100 parsecs measured with fractional errors of 5, 10, 20 per cent. At five per cent the distribution is nearly symmetric and inverting is harmless. At twenty per cent it is strongly skewed: the mean sits at 105 parsecs rather than 100, and the tail runs to distances several times the truth, because a parallax scattered a little towards zero is a distance scattered a long way outward. The asymmetry is a Jacobian and nothing else — the parallax measurement is unbiased and symmetric throughout. Above about twenty per cent the mean of the distribution stops existing at all, because the density falls only as the inverse square of the distance and the integral of d times that diverges.
Fig. 1 The distribution of the distance obtained by inverting a parallax, for a star truly at a hundred parsecs measured at three precisions. At five per cent it is nearly symmetric and inverting is harmless. At twenty per cent the distribution is strongly skewed, its mean sits well beyond the truth, and its tail runs to several times it — because a parallax scattered a little towards zero is a distance scattered a long way outward.

The difficulty is not academic. A modern catalogue contains parallaxes for nearly two thousand million stars, of which the great majority have fractional errors well above the level at which inversion is safe, and the temptation to divide one by the other is enormous — it is one line of code and it produces a number for every entry.

Why the asymmetry appears

The mechanism is the same Jacobian argument that makes a fitted eccentricity refuse to be zero, with a steeper exponent.

The parallax ϖ\varpi is measured with a Gaussian error. The distance is d=1/ϖd = 1/\varpi. Transforming a probability density from one variable to the other multiplies by the derivative of the transformation, which here is 1/d21/d^2 — so a Gaussian in ϖ\varpi becomes, in dd, a Gaussian evaluated at 1/d1/d multiplied by an inverse square.

That factor has two consequences and both matter. The density is compressed at small distances and stretched at large ones, so the distribution is skewed outward. And the far tail falls only as the inverse square of the distance, which is slowly enough that the integral of dd times the density diverges: for a fractional parallax error above about twenty per cent, the mean distance does not exist.

That is not a numerical difficulty. It is a statement that the question “what is the expected distance” has no answer, and that any code returning a number for it has silently imposed something the data did not contain.

One over a noisy parallax, at three precisions. The distribution of the distance obtained by inverting a parallax, for a star truly at 200 parsecs measured with fractional errors of 2, 8, 30 per cent. At five per cent the distribution is nearly symmetric and inverting is harmless. At twenty per cent it is strongly skewed: the mean sits at 227 parsecs rather than 200, and the tail runs to distances several times the truth, because a parallax scattered a little towards zero is a distance scattered a long way outward. The asymmetry is a Jacobian and nothing else — the parallax measurement is unbiased and symmetric throughout. Above about twenty per cent the mean of the distribution stops existing at all, because the density falls only as the inverse square of the distance and the integral of d times that diverges.
Fig. 2 The same construction pushed to a thirty per cent error, which is where a great many catalogue entries sit. The distribution is no longer recognisable as an estimate of anything: its peak is well inside the truth, its mean is well outside it, and the two are separated by more than the truth itself. Quoting the peak, the mean or the median gives three answers spanning a factor of two, and there is no principled reason to prefer any of them without saying something about where such stars are expected to be.

It is worth having the numbers. At a fractional parallax error of five per cent the difference between the median of the distance distribution and the truth is under a per cent, and inverting is fine. At ten per cent it is about two per cent. At twenty per cent the median is five per cent high and the mean is much worse. At thirty per cent the distribution is barely a peak at all. The safe boundary quoted in the literature — invert freely below ten per cent, think carefully between ten and twenty, do not invert above twenty — is a summary of that progression rather than a rule with a derivation.

It is also worth being clear that nothing has gone wrong with the measurement. The parallax is unbiased: measure it a thousand times and the average is the true value. The bias appears entirely in the transformation, and it is therefore a property of the question rather than of the data. That distinction decides what to do about it — no improvement in the astrometry removes it, and asking a different question does. The five parameters that come out of one fit are all well behaved; it is the sixth quantity, the one nobody measured, that misbehaves.

What a prior actually does

The way out is to stop asking for a single number and ask for a probability. The parallax measurement is a likelihood; combining it with a statement of where stars are expected to be gives a posterior; and the posterior is the honest answer.

The statement of where stars are expected to be is a prior, and the classical version is the simplest possible one: stars are uniformly distributed in space. That implies the number of stars in a shell grows as the square of the distance, which multiplies the posterior by d2d^2 and cancels the Jacobian exactly.

The resulting correction — the difference between the peak of that posterior and the naive 1/ϖ1/\varpi — is the Lutz–Kelker correction, derived in 1973. Its size depends only on the fractional parallax error: about 0.01-0.01 magnitudes at five per cent, 0.11-0.11 at ten, 0.43-0.43 at twenty, and formally infinite beyond about seventeen.

The correction is real and it is not a general-purpose fix, for a reason worth stating plainly: the uniform prior is wrong. Stars are not uniformly distributed — they are concentrated towards the Galactic plane, they thin out with distance from the Sun in a way that depends on the population, and any sample selected for a purpose has a selection function of its own. A prior that overstates how many distant stars there are overcorrects, and the classical correction, applied to a sample that was selected on parallax in the first place, can be worse than no correction at all.

What an arc of a given length knows about a parallax. The residual an arc cannot explain, against the parallax assumed for it, with the position and proper motion re-solved exactly at every trial value. The star's true parallax is 50 mas and every curve passes through zero there, because there is no noise in the data — what differs is the shape. A 0.25-year arc gives 0.055 mas of residual per mas of error in the parallax; a 3-year arc gives 0.587, 10.6 times as much. A 0.25-year arc is not 12 times worse than a 3-year one for the reason it is 12 times shorter, and the reason is geometric rather than statistical: over a fraction of a year the parallactic displacement is part of one swing, and a straight line in time imitates part of a swing almost exactly. Only when the ellipse closes and starts again is there anything in the data that a proper motion cannot be. That is why a parallax programme is stated in years and not in observations.
Fig. 3 The other reason a parallax’s error bar is not a simple thing: the measurement is one of five parameters fitted jointly, and the parallax is correlated with the others. A star near the ecliptic has a parallax ellipse that has collapsed to a line, so the parallax signature and the proper-motion signature are nearly parallel and the two trade against each other. The quoted uncertainty on such a star is larger, and — more importantly — its error is correlated with its proper motion in a way a single number does not express.

There is a second, more modern prescription that avoids the argument about which prior to use, and it is worth stating because it is what large catalogues now ship. Rather than a uniform density, use an exponentially decreasing space density with a length scale that varies across the sky, fitted from a Galaxy model. That prior is wrong in a different way from the uniform one, and it is wrong by much less, and — the point that matters — its influence on the answer can be measured by changing it and seeing what moves. A posterior that shifts substantially when the prior is varied is a posterior dominated by the prior, which is a useful thing to be told about a distance.

There is also a question of what “the distance” means for a sample rather than for a star, and it has a cleaner answer than the per-star case. A cluster’s distance is much better determined than any of its members’ distances, because the members share it: combining a thousand parallaxes of stars at the same distance gives a parallax precision a thousand times better in the random part, and the inversion is then perfectly safe. That is why cluster distances from parallax are quoted to fractions of a per cent while individual field-star distances at the same range are quoted to tens. The bias is a property of the fractional error, and the fractional error is what averaging fixes — as it does for any counting measurement.

The error that is not random at all

Everything above is about the random error. There is a second error, it is entirely different in character, and for the modern catalogues it is the one that matters.

Every parallax in a catalogue carries a common offset — the instrument’s zero point. It arises because an astrometric mission measures relative positions exquisitely and has no absolute reference; the zero point is fixed by observing objects whose parallax is known to be zero, which in practice means quasars, and by modelling the instrument’s own basic angle. The residual offset for the current catalogue is about 17-17 microarcseconds, with a dependence on magnitude, colour and position that is itself only partly characterised.

17 microarcseconds, and 0.04 magnitudes at a kiloparsec. What a constant offset in every parallax in a catalogue does, expressed as the error it puts into an absolute magnitude, against distance. An offset of -17 microarcseconds is negligible for a nearby star, where the parallax is measured in milliarcseconds, and it is not negligible at a kiloparsec, where the parallax is a millisecond of arc and the offset is nearly two per cent of it. Because the error is additive in the parallax it is multiplicative in the distance and grows in exact proportion to it, so a calibration of a standard candle carried out at a kiloparsec inherits 0.04 magnitudes of systematic that a calibration at a hundred parsecs does not. The offset is not a nuisance to be averaged away: it is identical for every star, so a sample of a million stars has exactly the same error as one star.
Fig. 4 What a constant offset does, expressed as the error it puts into an absolute magnitude. Seventeen microarcseconds is negligible for a nearby star and is not negligible at a kiloparsec, where the parallax itself is a milliarcsecond. Because the error is additive in the parallax it is multiplicative in the distance and grows in exact proportion to it. The offset is identical for every star, so a sample of a million has the same error as one — averaging does nothing at all.

That last property is the important one. A random error of ten per cent on a million stars averages to nothing; a systematic offset of seventeen microarcseconds on a million stars is seventeen microarcseconds. Every calibration built on the catalogue inherits it in full, and the calibrations that matter most — the zero points of the period–luminosity relation, of the red-clump magnitude, of the tip of the red-giant branch — are all built on stars far enough away for the offset to be a substantial fraction of their parallaxes.

40 microarcseconds, and 0.09 magnitudes at a kiloparsec. What a constant offset in every parallax in a catalogue does, expressed as the error it puts into an absolute magnitude, against distance. An offset of -40 microarcseconds is negligible for a nearby star, where the parallax is measured in milliarcseconds, and it is not negligible at a kiloparsec, where the parallax is a millisecond of arc and the offset is nearly two per cent of it. Because the error is additive in the parallax it is multiplicative in the distance and grows in exact proportion to it, so a calibration of a standard candle carried out at a kiloparsec inherits 0.09 magnitudes of systematic that a calibration at a hundred parsecs does not. The offset is not a nuisance to be averaged away: it is identical for every star, so a sample of a million stars has exactly the same error as one star.
Fig. 5 The same construction for a larger offset, which is roughly the uncertainty on the offset itself in the parts of the parameter space where it is least well determined — faint, red stars in crowded regions. The resulting magnitude error at a kiloparsec approaches a tenth of a magnitude, which is five per cent in distance and is comparable to the entire disagreement between the two competing measurements of the expansion rate.

Be precise about why the offset exists at all, because “the instrument’s zero point” makes it sound like a calibration constant somebody forgot to measure. An astrometric satellite measures the angle between two fields of view separated by a large fixed angle, and it builds a global solution from millions of such measurements. What that construction determines superbly is the difference between parallaxes; what it determines poorly is their common level, because a uniform shift in every parallax is nearly degenerate with a small variation in the satellite’s basic angle. So the offset is not an oversight. It is a nearly-null direction of the global solution, and pinning it down requires objects known independently to be at infinity.

What was actually measured

Three measurements bound the problem, and the third is the one that changed practice.

The offset, measured on quasars. Quasars have parallaxes indistinguishable from zero, so their measured parallaxes are a direct sample of the zero point. Their mean gives 17-17 microarcseconds for the catalogue as a whole. The difficulty is that quasars are faint, blue and point-like, and the offset depends on magnitude, colour and position — so the value measured on quasars is not the value applying to a bright red giant.

The offset, measured on stars with independent distances. Eclipsing binaries in the Magellanic Clouds, whose distances are known geometrically to about one per cent, and stars in clusters with well-determined distances, both provide external checks. They give offsets differing from the quasar value by a few microarcseconds, in a direction consistent with the magnitude dependence.

The effect on the expansion rate. The local measurement of the Hubble constant runs through Cepheids whose absolute magnitudes are calibrated on parallaxes. Changing the assumed zero point by ten microarcseconds moves that calibration by about a per cent, which moves the derived expansion rate by a per cent — a fifth of the disagreement with the early-universe value. The zero point is therefore not a technical detail of one catalogue; it is one of the terms in the most-discussed open question in cosmology.

One path, five numbers. Left: the apparent path of a star over 4 years, with a proper motion of 193 mas a year and a parallax of 50 mas, at ecliptic latitude 42°. It is one curve and there is nothing in the sky it can be compared against — the reference stars have paths of their own. Right: the same path with a straight line taken out of it. What is left is an ellipse of semi-major axis 50.0 mas and semi-minor axis 33.5 mas, closed and repeating once a year. The two are separated by their time signatures and by nothing else: proper motion is secular and parallax is annual, in a phase the Earth's position fixes in advance. That is why the five parameters can be told apart at all, and why an astrometric catalogue quotes five rather than two — a position without them is a position at one instant, which is not a direction to anything.
Fig. 6 Where all of it comes from: the five-parameter fit that produces a parallax in the first place. A star’s path across the sky is a straight line with a one-year ellipse on it, and one fit yields the position, the two components of proper motion and the parallax at once. The parallax is the amplitude of one periodic component of that path, and everything in this essay is about what happens after that amplitude has been divided into one.
What 0.5 magnitudes of scatter costs in distance. The fractional distance error against the uncertainty in the absolute magnitude, exactly and to first order. An error of 0.5 magnitudes in the assumed intrinsic brightness is a 26% error in the distance, however precisely the apparent magnitude was measured — the linear approximation, 46% per magnitude, is drawn dashed and holds to within a few per cent over the range shown.
Fig. 7 Why any of this matters downstream, in the currency the rest of the subject uses. A fractional distance error translates into a magnitude error, and the conversion is asymmetric: a distance too large by a factor and one too small by the same factor do not give symmetric magnitude errors. So a distance distribution with a long outward tail becomes an absolute-magnitude distribution with a long faint tail, and the mean absolute magnitude of a sample is pulled by it. Every calibration that averages absolute magnitudes over a sample of stars with parallax distances inherits that pull.

Where the picture stops

Three, and the second is the one most often ignored.

A negative parallax is not an error. For a distant star the measured parallax can come out negative, and roughly half of the entries for the most distant stars in a catalogue do. Those are perfectly good measurements — the true parallax is small and the noise is symmetric — and they carry information. Discarding them, which is the usual reflex, selects on the noise and biases everything that follows; a proper treatment keeps them and works with the likelihood.

A sample selected on parallax is selected on noise. Cutting a catalogue at, say, better than ten per cent fractional parallax error keeps preferentially the stars that scattered towards larger parallaxes, because those have smaller fractional errors. The resulting sample is closer than it should be, and the bias is not removed by any per-star correction. This is the same shape as a flux-limited sample being brighter than its population, with parallax playing the part of flux.

And the correction depends on what the sample is for. The right prior for a random field star, a selected Cepheid, and a member of a known cluster are three different things, and the same parallax with the same error implies three different distances under them. That is not a defect of the method — it is the honest statement that a single noisy measurement does not determine a distance, and that everything else in the answer came from somewhere.

A fourth deserves stating because it will define the next decade. The random errors of the current catalogue will improve as the mission’s time baseline lengthens — parallax precision improves roughly as the square root of the observing span, and proper-motion precision as its three-halves power. The zero point will not improve in the same way, because it is limited by the reference objects and by the instrument model rather than by counting statistics. So the ratio of systematic to random error is rising, and a catalogue that today is random-error-limited for most of its stars will end its life systematic-limited for nearly all of them. Preparing for that means characterising the offset now, on the data that exist, rather than waiting for the final release.

Why the geometry does not save the measurement

There is an irony worth drawing out, because it is the reason this rung exists.

Parallax is prized because it is geometric: no assumption about the star, no calibration, no model. That is true of the measurement and it stops being true at the moment the measurement is converted into a distance. The conversion is nonlinear, the nonlinearity is severe exactly where the data are weakest, and taming it requires a statement about the population — which is precisely the kind of assumption parallax was supposed to avoid.

So the cleanest distance in astronomy is clean only while it is expressed as an angle. Every use of it — a luminosity, an absolute magnitude, a position in a colour–magnitude diagram — requires the inversion, and the inversion requires a prior.

The practical resolution the field has reached is to move the model to the other side. Rather than convert each parallax to a distance and then fit a relation to the distances, fit the relation directly in parallax space: predict what parallax each star should have given the relation’s parameters and its apparent magnitude, and compare against what was measured. The likelihood is then Gaussian in the measured quantity, no inversion happens anywhere, and the bias does not arise. It is more work, it is what modern calibrations do, and it is a good general prescription — do the inference in the space the data were measured in.

A closing observation about the shape of the whole argument, because it recurs throughout the collection. The parallax problem has two errors with opposite characters: a random one that is large per star and averages away, and a systematic one that is small per star and does not. Almost every effort in the field goes into the first, because it is the one that appears in the error bar and the one a bigger telescope fixes. The second is what limits the answer. Every distance in the ladder is measured with the last one, and at every rung the same asymmetry holds — the random errors of the rung below contribute nothing to the top, and its systematic errors contribute in full.

There is one more thing worth saying about the practical consequence for a reader of catalogues. A published distance derived from a parallax should carry three pieces of information: the fractional parallax error, the prior used, and whether the zero point was applied. In current practice the first is nearly always available, the second is sometimes stated, and the third is often not — and applying an already-corrected parallax’s correction a second time is a common and entirely silent error. A brightness becomes a distance only if something is known, and what has to be known here is as much about the reduction as about the star.

One habit is worth stating because it makes most of this go away and is rarely adopted. Wherever possible, do the comparison in the space the measurement was made in: fit a model that predicts a parallax and compare it against the observed parallax, rather than converting the observation into a distance and fitting there. A model of a cluster’s structure, a period–luminosity relation, a Galactic rotation curve — all of them can be written to predict the observable directly, and doing so keeps the errors symmetric and Gaussian, which is what every standard fitting method assumes. The conversion is what breaks those assumptions, and the conversion is almost never necessary. It is done because a distance in parsecs is what a reader wants to see, which is a good reason to quote one at the end and a bad reason to compute with one throughout.

Where the ladder goes next

Later rungs on this ladder start with the selection: what happens to a calibration when the sample was chosen using the same measurement that is being calibrated, and why that bias survives every per-star correction. The rung after that is what the parallaxes are for — the period–luminosity relation whose zero point they set, and which carries the offset onward to every galaxy it is used on.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

CatalogueDistance estimationJacobianLutz kelker biasParallaxParallax zero pointPosteriorPriorSystematic errorVolume prior