Orbits

An eccentricity that cannot be zero

Fit an orbit to noisy data and the eccentricity that comes back is never zero, not even when the orbit is a perfect circle. The reason has nothing to do with the data and everything to do with the fact that a length cannot be negative.

Assumes Orbit determination and The ellipse.

Take a binary star on an exactly circular orbit. Observe it well but not perfectly, fit an ellipse to the observations, and read off the eccentricity. It will not be zero. Repeat the experiment with fresh noise and it will not be zero again, and the average over many repetitions will not be zero either. The most likely answer is one error bar, the average answer is 1.2533 error bars, and neither number depends on anything about the star.

This is not a failure of the fitting. It is a property of asking a question about a quantity that cannot be negative, and it applies to every catalogue of orbits ever compiled.

The best-fitting eccentricity of a circular orbit. What a fitted eccentricity comes out at when the orbit's true eccentricity is 0 and each component of the eccentricity vector carries an error of 0.03. The distribution is not centred on the truth and cannot be: an eccentricity is the length of the vector (e cos ϖ, e sin ϖ), lengths are not negative, and a quantity bounded below by zero whose components scatter symmetrically has a distribution pushed away from the bound. For a circular orbit the most likely fitted value is exactly one error bar, 0.0300 here, and the mean is 0.0376 — 1.2533 error bars, which is √(π/2) and comes from geometry rather than from any property of the data. The practical consequence is a catalogue of small eccentricities that are all measurements of their own error bars, and the fix is not a better fit but a different question: an upper limit rather than a value.
Fig. 1 What a fit returns when the truth is a circle. The distribution of the fitted eccentricity is pushed entirely to one side of the truth, because there is no other side to be pushed to. Its peak sits at exactly one error bar and its mean at √(π/2) of one — a number that comes from the geometry of a plane and contains nothing about orbits, telescopes or stars.

Two numbers that behave, and one that does not

The trouble is visible as soon as the right variables are named. An orbit’s shape and orientation in its own plane are carried by two numbers, and the natural pair is the eccentricity vector:

h=ecosϖ,k=esinϖ.h = e\cos\varpi, \qquad k = e\sin\varpi.

These are the quantities a fit actually determines. They enter the observation equations smoothly, their errors are close to Gaussian, and their correlation is usually mild. Nothing about them is pathological at e=0e = 0: a circular orbit is h=k=0h = k = 0, an ordinary point in an ordinary plane.

The eccentricity is the length of that vector, e=h2+k2e = \sqrt{h^2 + k^2}, and the longitude of perihelion is its direction. Both of those are badly behaved at the origin — the direction is undefined there, and the length is bounded below by zero — and the whole of the effect follows from the second of those two facts.

The same fit asked two ways. Left, the fit as it actually is: a two-dimensional Gaussian in the components of the eccentricity vector, centred on the truth, with its one, two and three error-bar contours drawn. Right, the same fit asked for the eccentricity alone, which is the distance from the origin in that plane. Turning the first into the second means adding up the probability around each circle, and the amount of circle available grows with its radius — so the density is multiplied by something proportional to the radius, and vanishes exactly at the origin. The peak moves away from the centre for a purely geometric reason, and it moves by one error bar. Nothing about the observations changed between the two panels; only the question did.
Fig. 2 The same fit asked two ways. On the left is what the fit is: a two-dimensional cloud in the components, centred on the truth, symmetric, with nothing unusual anywhere in it. On the right is the answer to the question “what is the eccentricity”, which is the distance from the origin in that same cloud. Converting one to the other means summing the probability around each circle, and the amount of circle grows with its radius — so the density is multiplied by something proportional to the radius and vanishes exactly at the centre.

The factor of the radius is a Jacobian. Changing from Cartesian to polar coordinates in a plane replaces dhdkdh\,dk with ededϖe\,de\,d\varpi, and the extra ee is the entire mechanism. A Gaussian in the plane becomes, in the radius, the Gaussian multiplied by ee — which is a Rayleigh distribution, peaking at one standard deviation and having mean σπ/2\sigma\sqrt{\pi/2}.

Nothing has been assumed about the observations except that their errors are symmetric in the two components. The bias is a change of variables.

It is worth pausing on how little physics is in that sentence, because the smallness is the point. The result does not know that the quantity is an eccentricity. Two Gaussian components of equal width, combined into a length, give a Rayleigh distribution whether they are the components of an eccentricity vector, of a proper motion, of a polarisation, or of the offset between two positions on a photographic plate. The coefficient 1.2533 is π/2\sqrt{\pi/2} and it turns up wherever two dimensions are collapsed into one; the corresponding number for three dimensions is 22/π=1.59582\sqrt{2/\pi} = 1.5958, which is why a speed built from three velocity components is biased harder than a shape built from two.

There is a second consequence, less obvious and more damaging in practice: the uncertainty on the fitted eccentricity is also wrong. The formal error that a least-squares fit reports is the width of a parabola fitted to the likelihood at its peak, and near a boundary the likelihood is not a parabola. For a circular orbit the reported error can come out comfortably smaller than the value it accompanies, so an entry reads e=0.038±0.020e = 0.038 \pm 0.020 — a two-sigma detection of an eccentricity that is exactly zero.

One more property of the Rayleigh distribution is worth having, because it explains why the effect went unnoticed for so long. Its mean and its width are the same size: the standard deviation of the fitted eccentricity of a circular orbit is σ2π/2=0.6551σ\sigma\sqrt{2 - \pi/2} = 0.6551\sigma, so a typical entry reads about 1.25 error bars with a scatter of 0.65 of one. Anybody looking at a single orbit sees a small eccentricity with a plausible uncertainty and nothing to complain about. The effect is only visible in the aggregate, where every entry is biased in the same direction and the errors do not average away. That is the signature of a systematic, and it is the reason the discovery was made by somebody looking at a catalogue rather than at a star.

What it looks like in a catalogue

The consequence for a table of orbits is a floor. Above about three error bars, the fitted eccentricity is an honest measurement of the real one. Below that, the fitted value stops tracking the truth and settles onto a constant.

A measurement that flattens onto its own error bar. The eccentricity that comes out of a fit, averaged over noise, against the eccentricity that went in, for an error of 0.03 on each component. The diagonal is what an unbiased measurement would do and the curve joins it above about three error bars, which is the ordinary statement that a well-measured quantity is well measured. Below that the curve peels away and flattens onto 0.0376 — √(π/2) times the error bar, and independent of the truth. Every orbit in this region is reported with the same small eccentricity whatever it actually has, and the number reported is a property of the observations. The shaded band is where a catalogue entry carries no information about the orbit's shape, and the honest entry there is an upper limit.
Fig. 3 The eccentricity a fit reports against the eccentricity the orbit has, averaged over noise. The diagonal is what an unbiased measurement would do, and the curve joins it above three error bars. Below that it peels away and flattens onto 1.2533 error bars — a number that has nothing to do with the orbit at all. Everything in the shaded region is reported with the same eccentricity whatever it actually has.

A catalogue whose orbits have a range of data quality therefore has a range of floors, and the floors are not marked. Two entries reading e=0.04e = 0.04 can mean completely different things: one a real eccentricity measured at ten sigma, the other a circular orbit measured badly. Nothing in the number distinguishes them, and the formal uncertainty does not either, because the formal uncertainty is what the number is made of.

A measurement that flattens onto its own error bar. The eccentricity that comes out of a fit, averaged over noise, against the eccentricity that went in, for an error of 0.008 on each component. The diagonal is what an unbiased measurement would do and the curve joins it above about three error bars, which is the ordinary statement that a well-measured quantity is well measured. Below that the curve peels away and flattens onto 0.0100 — √(π/2) times the error bar, and independent of the truth. Every orbit in this region is reported with the same small eccentricity whatever it actually has, and the number reported is a property of the observations. The shaded band is where a catalogue entry carries no information about the orbit's shape, and the honest entry there is an upper limit.
Fig. 4 The same picture for observations four times better. The shape is identical and the axis has moved: the floor is now at 0.010 rather than 0.038, and orbits that were unmeasurable are now measured. That is the useful half of the statement — the bias falls in proportion to the error, so it is a problem of data quality rather than a barrier — and it is also the dangerous half, because a catalogue assembled from several sources has several floors in it and the shape of the assembled distribution is a map of the observing programmes.

There is an old habit that makes this worse. Fitting programs often bound the eccentricity below at zero and report the boundary value when the fit wants to go past it, which sounds conservative and is not: it converts a symmetric problem into an asymmetric one and piles up entries at exactly zero alongside entries at the floor, producing a bimodal catalogue out of a population that has no structure in it at all. The alternative — letting the fit run in (h,k)(h, k), where there is no boundary to hit — costs nothing and is what modern codes do.

The effect on a distribution is worse than the effect on any one entry. A population of genuinely circular orbits, measured at an error of 0.03, produces a histogram of eccentricities that peaks near 0.03 and has a tail out to 0.1. It looks like a population with small non-zero eccentricities. Fitting a model to that histogram — an exponential, a beta distribution, whatever the fashion is — recovers a scale length that is a property of the survey.

A true eccentricity of 0.055, fitted ten thousand times. What a fitted eccentricity comes out at when the orbit's true eccentricity is 0.055 and each component of the eccentricity vector carries an error of 0.03. The distribution is not centred on the truth and cannot be: an eccentricity is the length of the vector (e cos ϖ, e sin ϖ), lengths are not negative, and a quantity bounded below by zero whose components scatter symmetrically has a distribution pushed away from the bound. For a circular orbit the most likely fitted value is exactly one error bar, 0.0616 here, and the mean is 0.0640 — 1.2533 error bars, which is √(π/2) and comes from geometry rather than from any property of the data. The practical consequence is a catalogue of small eccentricities that are all measurements of their own error bars, and the fix is not a better fit but a different question: an upper limit rather than a value.
Fig. 5 The same machinery with a genuinely eccentric orbit — 0.055, or not quite two error bars. The distribution has become asymmetric rather than one-sided, its peak has moved to 0.047 and its mean to 0.063, and it is still not centred on the truth. The bias does not switch off at some threshold; it decays, and at two sigma it is still a fifth of an error bar. Reporting the peak, the mean or the median gives three different answers here, and none of them is the truth.
The best-fitting eccentricity of a circular orbit. What a fitted eccentricity comes out at when the orbit's true eccentricity is 0 and each component of the eccentricity vector carries an error of 0.012. The distribution is not centred on the truth and cannot be: an eccentricity is the length of the vector (e cos ϖ, e sin ϖ), lengths are not negative, and a quantity bounded below by zero whose components scatter symmetrically has a distribution pushed away from the bound. For a circular orbit the most likely fitted value is exactly one error bar, 0.0120 here, and the mean is 0.0150 — 1.2533 error bars, which is √(π/2) and comes from geometry rather than from any property of the data. The practical consequence is a catalogue of small eccentricities that are all measurements of their own error bars, and the fix is not a better fit but a different question: an upper limit rather than a value.
Fig. 6 A circular orbit measured two and a half times better. Everything about the shape is unchanged — the same peak at one error bar, the same mean at 1.2533 of one — and the whole figure has simply shrunk towards the origin. That is the useful property of the bias and the reason it is a nuisance rather than a catastrophe: it scales with the error, so it is removed by better data rather than by cleverer analysis, and a survey can be designed against it in advance by asking what eccentricity floor it needs and working backwards to a velocity precision.

What was actually measured

The effect was named in 1971, by Lucy and Sweeney, in a paper about spectroscopic binaries. Their concern was practical: the catalogues of the day were full of orbits with eccentricities between 0.02 and 0.1, and the theory of tidal circularisation predicted that short-period binaries should be circular. Either the theory was wrong or the eccentricities were.

Their test was the right one and it is still used. Fitting a circular orbit costs two fewer free parameters than fitting an eccentric one, so the improvement in the sum of squared residuals from allowing an eccentricity can be compared against what noise alone would produce — an F-test with two and N6N-6 degrees of freedom. Applied to the catalogue, it found that a large majority of the small published eccentricities were not significant at all. The theory was fine.

The measurement that makes this quantitative today comes from the exoplanet radial-velocity surveys, where the same statistic is applied to thousands of orbits rather than dozens. Fitting the same data with a circular model and an eccentric one, and comparing the two properly, moves a substantial fraction of the published low-eccentricity planets to circular — and the fraction depends steeply on the signal-to-noise of the velocity curve, exactly as the Rayleigh argument requires. The published eccentricity distribution of planets near a few days’ period changed shape when this was done, and the new shape agreed with the tidal theory that the old one had appeared to contradict.

A period below which every orbit is round. Orbital eccentricity against period for binaries in four clusters of 0.125, 0.625, 6, 4 billion years, with the eccentricities drawn from one seeded distribution and then damped by exp(−age/τ), where τ rises as the 5.333 power of the period. Each cluster shows the same thing: below a boundary period nothing survives eccentric, above it the original distribution is untouched, and there is almost nothing in between because the timescale is so steep. The boundary is a clock. It moves as the three-sixteenths power of the age, which the figure checks against the drawn curves, and the calibration puts it at 6.5 days at 125 million years, 8.8 at 625 million and 12.5 at four billion — against measured cut-offs near 7.2, 8.5 and 12.5 days in the Pleiades, the Hyades and M67. The boundaries are also read back off the plotted points rather than trusted, and required to move outward with age. This is the cleanest measurement of tidal dissipation in ordinary stars that exists, and its cleanliness comes from the ages: a cluster's age is read off its main-sequence turn-off and owes nothing whatever to the tide being measured.
Fig. 7 The distribution the bias distorts, and the physical edge it is being used to locate: eccentricity against orbital period for a population of binaries, with the circularisation cut-off marked. The interesting quantity is the period below which the eccentricities are consistent with zero, because that period is a clock. An upward bias on every small eccentricity moves the apparent cut-off to shorter periods and makes the transition look gradual when it is sharp — which is the measurement being corrupted rather than merely a number being wrong.

The same argument has been run on the transiting planets, where the eccentricity is not measured from a velocity curve at all but from the duration of the transit compared with the duration a circular orbit would give. That measurement has its own difficulties, and it also has a different geometry: it constrains ecosωe\cos\omega far better than esinωe\sin\omega, because the duration depends on the planet’s speed at transit and the speed depends on the distance, which is set by the component of the eccentricity along the line of sight. The cloud in the plane is a long thin ellipse rather than a circle, the Jacobian argument still applies, and the coefficient is different again — which is exactly the case the tidy π/2\sqrt{\pi/2} does not cover and the general argument does.

The general shape of the fix is now standard and it has three parts. Fit in the components rather than in the length, so the likelihood is the well-behaved object. Report an upper limit rather than a value when the significance test fails. And when a population is being modelled, model it in the components too — fitting a distribution of ee to a set of biased ee values imports the bias into the answer, while fitting a distribution of (h,k)(h, k) to a set of (h,k)(h, k) values does not.

Say what the significance test does and does not settle, because it is often over-read. An F-test comparing a circular fit with an eccentric one answers a narrow question: could noise alone have produced this much improvement. A failure to reject the circular model is not evidence that the orbit is circular — a badly observed orbit of eccentricity 0.3 will also fail the test — and the correct entry for such an orbit is an upper limit rather than either a value or a zero. Catalogues that record “circular” for every orbit that failed the test have thrown away the distinction between measured-to-be-round and not-measured, and the two are combined in every population study afterwards.

The same trap, wearing other clothes

Once the mechanism is named it turns up everywhere, because a great many measured quantities are positive by construction.

An inclination derived from a projected quantity has the same problem with a different Jacobian. A flux measured against a subtracted background can go negative in the data and cannot go negative in nature, so the reported flux of a faint source is biased upward — which is the same effect as the noise floor that comes from counting photons meeting a hard boundary. A parallax is not sign-definite as measured but a distance is, and inverting a noisy parallax to get a distance is this same argument with a much steeper Jacobian, since the volume element grows as the square of the distance rather than as its first power.

A mass ratio, a velocity dispersion, an amplitude of a periodic signal, an equivalent width: all of them are lengths, or squares, or moduli of something that a fit determines in a form where the noise is symmetric. Each carries the same bias, with a coefficient set by the number of components being combined.

The same fit asked two ways. Left, the fit as it actually is: a two-dimensional Gaussian in the components of the eccentricity vector, centred on the truth, with its one, two and three error-bar contours drawn. Right, the same fit asked for the eccentricity alone, which is the distance from the origin in that plane. Turning the first into the second means adding up the probability around each circle, and the amount of circle available grows with its radius — so the density is multiplied by something proportional to the radius, and vanishes exactly at the origin. The peak moves away from the centre for a purely geometric reason, and it moves by one error bar. Nothing about the observations changed between the two panels; only the question did.
Fig. 8 The mechanism at the stage where it is starting to switch off. The cloud in the plane has moved two and a half error bars from the origin, so most of it now lies on one side and the ring geometry has much less to do; the marginal on the right is nearly symmetric and nearly centred. The transition is entirely about how much of the two-dimensional cloud can see the origin, which is why three error bars is the rule of thumb and why it is a rule of thumb rather than a threshold.

The steepest version of all is a quantity built from a ratio of two noisy things, where the boundary is not at zero but at infinity: the distribution of a ratio has tails so heavy that its mean may not exist. That is not this effect, but it belongs to the same family of complaints, and both are instances of the same warning — error propagation by the usual first-order formula assumes the transformation is nearly linear over the width of the error, and every one of these cases is a transformation that is not.

Where the picture stops

Three of them, and the second is the one most often missed.

It assumes symmetric errors on the components. Real orbit fits produce correlated errors, and the correlation can be severe: a poorly sampled orbit constrains ecosϖe\cos\varpi much better than esinϖe\sin\varpi, so the cloud is an ellipse rather than a circle. The bias is still there and the coefficient is no longer π/2\sqrt{\pi/2} — it has to be computed from the actual covariance, which is the object an orbit determination really produces and the object most catalogues do not publish.

It is about the estimator, not about the data. No amount of care in the fitting removes it, because nothing is being done wrongly. The observations contain exactly the information they contain; the bias is introduced when that information is summarised by a single number for a quantity whose allowed values have a boundary. A published posterior distribution has no bias, because a distribution is not an estimate.

And a prior is not a fix, it is a choice. Multiplying by a prior that prefers small eccentricities pulls the answer back down and can be tuned to remove the bias exactly — for a population with the eccentricity distribution the prior assumed. Applied to a different population it introduces a bias of the opposite sign. This is not an argument against priors; it is the observation that the honest version of “the eccentricity is small” is a statement about a population, and it has to be inferred rather than assumed.

A fourth deserves stating because it cuts the other way. Everything here makes a small eccentricity look larger, so it can never manufacture a large one. An orbit reported at 0.4 with an error of 0.03 is eccentric, and no amount of Jacobian argument will make it round. The bias is confined to the region within about three error bars of the boundary, and outside that region a fitted eccentricity is exactly what it appears to be. Knowing where the affected region is, and that it has a hard outer edge, is most of what a reader of a catalogue needs.

Why this belongs to orbit determination rather than to statistics

The reason to file this argument with the problem of getting an orbit out of a handful of sightings rather than with error analysis in general is that orbit determination is where the boundary is structural. Most measured quantities are positive by accident of what they measure. An eccentricity is positive because it is the modulus of a vector that is genuinely two-dimensional, and the two dimensions are physically real: the shape and the orientation of the ellipse are separate facts about the orbit, and the fit determines them jointly.

There is a dynamical reason the pair is the right one as well, and it is the reason the same variables appear in the equations that move the elements. Gauss’s equations for de/dtde/dt and dϖ/dtd\varpi/dt both carry a factor of 1/e1/e and both become singular on a circular orbit; the equations for dh/dtdh/dt and dk/dtdk/dt do not, and neither does anything computed from them. The coordinates that are well behaved for a fit and the coordinates that are well behaved for an integration are the same coordinates, which is a stronger statement than either half on its own.

That is also what makes the fix natural rather than a patch. Working in (h,k)(h, k) is not a statistical trick; it is a return to the coordinates the dynamics is written in. The same pair appears in the secular theory, where the eccentricity vectors of the planets are what oscillate and the eccentricities themselves are moduli that go to zero and come back for no dynamical reason at all. A planet whose eccentricity vector passes near the origin has an eccentricity that dips to zero and rises again, and reading that as an event rather than as a coordinate artefact is the same mistake in a different setting.

The lesson is one this collection keeps arriving at from different directions: the coordinates a quantity is measured in and the coordinates it is reported in are rarely the same, and everything difficult lives in the change between them. An eccentricity is not a hard thing to measure. It is a hard thing to report.

There is a last observation worth making about the shape of the whole argument, because it recurs. Nothing here was discovered by finding a discrepancy in the sky. It was discovered by noticing that a theoretical prediction — short-period binaries are circular — disagreed with a catalogue, and asking which of the two to doubt. The answer turned out to be neither the theory nor the observations but the arithmetic in between, and the way it was settled was by constructing the distribution the noise alone would produce and comparing it with the one the catalogue showed. That is the same manoeuvre as asking what a survey could have seen before asking what it did see, and it is the single most reliable way of telling a fact about the universe from a fact about the instrument.

Where the ladder goes next

The rung above this one is the covariance itself — what an orbit fit actually produces, and why publishing six numbers and six error bars discards most of it. The rung after that is the same boundary problem with a steeper Jacobian and much higher stakes: the distance implied by a parallax, where the quantity being inverted is noisy, the volume element grows as a square, and the resulting bias reaches into every distance in astronomy.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

CovarianceEccentricityEccentricity vectorError propagationJacobianMarginal distributionMaximum likelihoodPositive definite quantityRayleigh distributionUpper limit