The observed sky

Five numbers from one wiggle

A star's path across a plate is a straight line with a one-year ellipse laid on it. Nothing measures either alone — one fit yields five parameters at once, and they are separable only because their time signatures differ.

Assumes Parallax and Aberration.

The phrase “measuring a parallax” describes something nobody has ever done. There is no observation whose output is a parallax. What is observed is a sequence of positions of a point of light, and a parallax is one of five numbers fitted to that sequence — jointly, all at once, each correlated with the others.

That distinction sounds like bookkeeping and is not. It is the reason a parallax programme is stated in years rather than in nights, the reason the first stellar parallax was measured a century after the effect that gave it away, and the reason a catalogue quotes an epoch.

One path, five numbers. Left: the apparent path of a star over 4 years, with a proper motion of 193 mas a year and a parallax of 50 mas, at ecliptic latitude 42°. It is one curve and there is nothing in the sky it can be compared against — the reference stars have paths of their own. Right: the same path with a straight line taken out of it. What is left is an ellipse of semi-major axis 50.0 mas and semi-minor axis 33.5 mas, closed and repeating once a year. The two are separated by their time signatures and by nothing else: proper motion is secular and parallax is annual, in a phase the Earth's position fixes in advance. That is why the five parameters can be told apart at all, and why an astrometric catalogue quotes five rather than two — a position without them is a position at one instant, which is not a direction to anything.
Fig. 1 The whole of the data, and what is in it. On the left, the apparent path of a star over four years: a proper motion of 193 milliarcseconds a year with a parallax of 50 laid on it. It is one curve, and there is nothing in the sky it can be compared against, because every reference star has a path of its own. On the right, the same path with a straight line taken out — an ellipse of semi-major axis exactly the parallax, closed and repeating once a year. The two are separated by their time signatures and by nothing else: one is secular, the other annual and in a phase the Earth’s own position fixes in advance.

The five parameters

The standard astrometric model of a single star is

α(t)=α0+μαt+ϖPα(t),δ(t)=δ0+μδt+ϖPδ(t),\alpha^*(t) = \alpha^*_0 + \mu_{\alpha^*}\,t + \varpi\,P_{\alpha^*}(t),\qquad \delta(t) = \delta_0 + \mu_\delta\,t + \varpi\,P_\delta(t),

with α=αcosδ\alpha^* = \alpha\cos\delta so that both coordinates are angles on the sky rather than one of them being a longitude. Five unknowns: two coordinates at a reference epoch, two components of proper motion, one parallax. The parallax factors PP are not unknowns — they are computed from the Earth’s ephemeris and the star’s direction, and are the same for every star in the field of view up to the direction dependence.

That is the entire model, and its shape decides everything that follows. It is linear in all five parameters, so the fit is a least-squares problem with a closed-form solution; and the five basis functions are 11, 11, tt, tt, and a pair of one-year periodic functions. Whether the parameters can be told apart is the question of whether those basis functions are distinguishable over the span of the observations.

Parallax for a star at 1.3 parsecs. The same star observed from two ends of a baseline. The two sight lines are 1.54″ apart, so the parallax — half of that, the shift seen from a one-astronomical-unit baseline — is 0.769″ for a star 1.3 parsecs away. The definition of the parsec is the distance at which it would be exactly one.
Fig. 2 The geometry the fifth parameter comes from, drawn as the textbook draws it: two sight lines from opposite ends of the Earth’s orbit, and the small angle between them. The triangle is exact and the drawing is not to scale — at the true ratio the two lines would be indistinguishable. What this picture leaves out is that the star is also moving, and that in the real measurement the two positions are not observed at the same time as anything else. The wiggle in the previous figure is this triangle plus that motion, and it is the sum that is recorded.

Why the separation works at all

Two of the five parameters have the same functional form as each other — the two coordinates — and are distinguished only by direction. The interesting pairing is proper motion against parallax.

A proper motion is linear in time. A parallax contribution is periodic with a period of exactly one year and a phase fixed by where the Earth is. Over a full year the periodic term returns to where it started while the linear term does not, and that is what separates them. Over less than a year the periodic term is executing part of one swing, and part of one swing is a shape that a straight line in time can imitate closely.

What an arc of a given length knows about a parallax. The residual an arc cannot explain, against the parallax assumed for it, with the position and proper motion re-solved exactly at every trial value. The star's true parallax is 50 mas and every curve passes through zero there, because there is no noise in the data — what differs is the shape. A 0.25-year arc gives 0.055 mas of residual per mas of error in the parallax; a 3-year arc gives 0.587, 10.6 times as much. A 0.25-year arc is not 12 times worse than a 3-year one for the reason it is 12 times shorter, and the reason is geometric rather than statistical: over a fraction of a year the parallactic displacement is part of one swing, and a straight line in time imitates part of a swing almost exactly. Only when the ellipse closes and starts again is there anything in the data that a proper motion cannot be. That is why a parallax programme is stated in years and not in observations.
Fig. 3 How much an arc of a given length actually knows. For each trial parallax, the four remaining parameters are re-solved exactly by least squares, and what is left is the residual the arc cannot explain. Every curve passes through zero at the true value, because there is no noise in the data — what differs is the shape, and the curvature of the parabola is the information. A three-month arc gives 0.055 milliarcseconds of residual per milliarcsecond of error; three years gives 0.587, eleven times as much. The short arc is not worse in proportion to its length, and the reason is geometric rather than statistical.

The consequence is the practical rule that governs every astrometric mission: the parallax accuracy improves as the square root of the number of observations, as any average does, and it improves separately with the time span, because the span is what breaks the correlation with proper motion. Doubling the number of nights within one season buys 2\sqrt{2}. Extending from one year to five buys far more than 5\sqrt{5}, because the first year is buying separability rather than precision.

What else is in the wiggle

Three effects of the same order have to be removed before the five parameters mean anything, and each was discovered by somebody hunting for parallax and finding it instead.

The same ellipse at 955 times the size, and a quarter of a year out of step. The aberration ellipse (solid) and the parallax ellipse (dashed) for γ Draconis at four ecliptic latitudes, drawn to one scale. Aberration is v/c towards the Earth's own direction of travel, so its ellipse has semi-major axis 20.49551″ for every star in the sky; parallax is 1/d towards the Sun, so its semi-major axis is that star's own 0.02147″ — 955 times smaller, and drawn 955 times smaller here. At the ecliptic pole both are circles; on the ecliptic both collapse to lines; in between both are ellipses of semi-minor axis sin β times the major, and the ratio is identical at every latitude because both effects project the same way. The marks are the same four dates on each: they are a quarter of a year apart between the two, because the aberration displacement follows the Earth's velocity and the parallax displacement follows its position, and velocity leads position by 90° on a circular orbit. That quarter-year is the only thing distinguishing the two phenomena on the sky, and it is why Bradley, looking for the dashed ellipse in 1728, spent months unable to interpret the solid one he had found instead.
Fig. 4 The largest of them, and it is much larger. Aberration displaces every star by up to 20.5 arcseconds — four hundred times the largest parallax there is — in an ellipse of the same annual period. It is independent of distance, which is what makes it removable: the same correction applies to every star in the field, computed from the Earth’s velocity and nothing else. Bradley found it in 1728 while trying to measure a parallax, and it proved the Earth’s motion a century before any star’s distance was known.
Parallax for a star at 60 parsecs. The same star observed from two ends of a baseline. The two sight lines are 33.3 mas apart, so the parallax — half of that, the shift seen from a one-astronomical-unit baseline — is 16.7 mas for a star 60 parsecs away. The definition of the parsec is the distance at which it would be exactly one.
Fig. 5 And at sixty parsecs, where the parallactic ellipse is a hundredth of the aberration ellipse it is drawn inside. That ratio is the reason aberration was discovered first: Bradley set out to measure parallax, found an annual ellipse of the right period and twenty times the expected size, and spent two years establishing that it was a different effect entirely — one that depends on the Earth’s velocity rather than its position, and therefore lags the parallax by exactly a quarter of a year. The phase, not the size, is what separates them.

The third is refraction, which displaces a star towards the zenith by an amount depending on altitude and therefore, at a fixed sidereal time, on the season. The Sun sets before it sets is the same effect at its largest; for a star at moderate altitude it is tens of arcseconds and has an annual component, which is precisely the shape the fit is looking for. It is why the eighteenth-century attempts failed and why the successful measurements of the 1830s all used differential techniques against nearby comparison stars, so that the refraction cancels.

What was actually observed, in 1838

Three parallaxes were announced within a year of one another, by three people using three different instruments, and the differences between them are a useful catalogue of what the fit needs.

Bessel measured 61 Cygni with a heliometer — an objective cut in half so that the two halves can be slid against each other, turning an angular separation into a screw reading. He chose the star because its proper motion was the largest then known, 5.2 arcseconds a year, which is a good proxy for nearness — the same reasoning that makes proper motion a distance indicator when nothing better is available. He observed it against two faint neighbours over eighteen months and published 0.3140.314''; the modern value is 0.2860.286'', so he was within ten per cent.

Henderson measured α\alpha Centauri from the Cape with a meridian circle, which measures absolute positions rather than differential ones and therefore fights refraction directly. He had the data in 1833 and did not reduce it until 1839, by which time Bessel had published. His value was 1.161.16'' against a modern 0.750.75'' — a fifty per cent error, and the instrument is why.

Struve measured Vega with a filar micrometer and got 0.260.26'' against a modern 0.130.13''. His error is a factor of two, and Vega’s parallax is small enough that the measurement was at the edge of what the method could do.

The pattern is worth reading. The differential measurements were good and the absolute one was bad, by a factor of five in fractional error, and that is the whole argument of the previous section made by three people who had not yet had it. It took another 150 years to escape the compromise, and the escape was a satellite.

Relative and absolute

The differential trick that removes refraction introduces the problem that dominated the subject until 1989. Measuring a star’s position relative to faint background stars removes anything common to the field — refraction, plate scale, telescope flexure — and also removes the parallax of the reference stars, which is not zero.

So a ground-based parallax is a relative parallax, and turning it into an absolute one requires knowing how far away the references are, which requires knowing their parallaxes — the same circularity that runs through every rung of the distance ladder. The correction is typically 1 to 2 milliarcseconds and is estimated from a model of the Galaxy rather than measured, which means every ground-based distance carried an error bar with a Galactic model inside it.

What a global solution does instead

The escape from relative parallaxes is to observe two widely separated directions at once. If a satellite measures the angle between two fields 9090^\circ apart, and does so repeatedly as it rotates, the parallax factors in the two fields differ in sign and magnitude, and the resulting system of equations has an absolute solution with no zero point to be assumed.

That is what Hipparcos did in 1989 and what Gaia does now, and it converts the problem from a measurement of a star into the simultaneous solution of a system with a billion sources and tens of billions of observations, in which the reference frame, the satellite’s attitude, and the optical calibration are solved for at the same time as the astrometry. The five parameters per star are still five parameters per star; the difference is that the frame is now solved rather than assumed.

The precision that buys is the thing worth quoting. Hipparcos reached about a milliarcsecond, which is 1000 parsecs at ten per cent. Gaia reaches tens of microarcseconds for bright stars, which is tens of kiloparsecs — far enough to cross the Galaxy.

Parallax for a star at 13 parsecs. The same star observed from two ends of a baseline. The two sight lines are 0.154″ apart, so the parallax — half of that, the shift seen from a one-astronomical-unit baseline — is 76.9 mas for a star 13 parsecs away. The definition of the parsec is the distance at which it would be exactly one.
Fig. 6 The same construction at thirteen parsecs, where the angle is a tenth of what the nearest stars give. Parallax falls as one over the distance, so the whole measurement is a fight against a shrinking signal: at 1.3 parsecs the ellipse is 0.77 arcseconds across and at 13 it is 0.077, and at the distances a space astrometry mission actually works to it is measured in microarcseconds. Nothing about the geometry changes. What changes is that the wobble sinks below the proper motion, below the aberration, and eventually below the instrument, in that order.

What the fit assumes

Every model that is fitted is also a hypothesis, and this one has three assumptions in it that are not always true.

It assumes the star is a point. A resolved binary, an unresolved binary, or a star with a bright spot rotating across its disc all have photocentres that move for reasons the model has no term for, and the fit absorbs that motion into whichever parameter it best resembles. An unresolved binary with a period near one year contaminates the parallax; one with a period much longer than the mission contaminates the proper motion. Gaia’s published solutions carry a goodness-of-fit statistic precisely so that these can be flagged, and about a per cent of bright stars fail it.

It assumes the motion is uniform. Over four years that is excellent, and over the decades separating two catalogues it is not: a star in a binary of thirty-year period has a proper motion that differs between epochs, and comparing two catalogues taken fifty years apart yields an apparent acceleration that is real information about a companion.

And it assumes the reference direction is fixed, which brings the whole problem of the frame back. A systematic rotation of the reference frame at ω\omega appears in every star’s proper motion as a term proportional to ω\omega and in none of the parallaxes, so a frame that is slowly spinning produces a spurious pattern of proper motions across the sky. Measuring that pattern against quasars, and finding it consistent with zero, is one of the standard validations of a modern catalogue — and it is the sense in which an extragalactic frame is not a convenience but a requirement.

The sixth number, which is not in the fit

Five parameters describe the star’s motion across the sky. The motion along the line of sight is invisible to astrometry entirely, and comes from somewhere else. There is one place where the two measurements meet, and it is a rare and beautiful check. A star’s radial velocity changes its distance, which changes its proper motion — the perspective acceleration, or secular acceleration. For Barnard’s star, moving at 10.3 arcseconds a year and approaching at 110 km/s, the effect is about 1.2 milliarcseconds per year per year: detectable, and detected. It is one of the very few cases in which a spectroscopic quantity and an astrometric one predict each other, and the agreement is a check on both.

What the ladder has established

The first rung of this anchor was the triangle and its limit: a geometric distance, honest and short-ranged. This one is about what has to happen before that triangle can be extracted from a sky in which nothing holds still.

The lesson generalises past astrometry. Five quantities are entangled in one observable, and they are recoverable because they have different signatures in time — secular, annual, annual-in-quadrature. That is the same argument that separates a planet’s transit from a star’s variability, and the same argument that separates a rotational line broadening from a gravitational one. What makes a measurement possible is rarely that the signal is large. It is that the signal has a shape nothing else has.

Two properties of a global solution are worth separating before the difficulty below, because they are often run together. One is that the frame is determined rather than assumed, which removes the reference-star correction that limited every ground-based measurement. The other is that the instrument’s own calibration is solved for at the same time, from the data, rather than being measured beforehand and applied. The second is what makes the precision possible and it is also what leaves a residual behind: a calibration solved from the data can only be as good as the data’s ability to constrain it, and there are combinations of instrumental parameters that the observations barely distinguish. Whatever survives in those combinations does not show up as scatter. It shows up as a systematic shared by every source measured the same way, which is exactly the shape of the difficulty described next.

The zero point that is not zero

A global solution removes the need to assume a reference zero point, and it does not remove the zero point. Modern catalogues carry one, it is small, and it is currently a limiting term in the distance scale.

The way it was found is the cleanest possible test. Quasars are so distant that their parallaxes are zero to any precision anyone can reach — so measuring them is a null experiment, and whatever comes out is the instrument’s error. The measured mean came out not at zero but at a small negative value, of order tens of microarcseconds.

A negative parallax offset means every star is placed slightly further away than it is. The size is a few hundredths of a milliarcsecond, which is negligible for a nearby star and is a substantial fraction of the measured parallax for a distant one — precisely the distant ones that matter for calibrating anything.

The awkward part is that the offset is not a constant. It varies with a source’s magnitude, with its colour and with its position on the sky, because it originates in the instrument’s optics and in how the images are processed, both of which depend on those things. So the correction is a fitted function of three variables, derived from quasars where they are available and from other objects of known distance where they are not, and its own uncertainty is several microarcseconds.

The consequence is a systematic that behaves unlike a measurement error. It does not average down over many stars, because it is common to them; it changes if the correction function is revised; and it enters every distance derived from the catalogue in the same direction. Anyone quoting a result that depends on parallaxes at the ten-microarcsecond level now has to say which correction was applied, and results using different corrections are not directly comparable.

A global solution converts a per-field problem into a per-catalogue one, which is a large improvement and not an elimination — and the null experiment that revealed it is the only reason it is known at all.

The lesson is the one the ground-based measurements learned in a different form: a differential technique removes what is common and leaves what is not, and identifying what is not requires an object whose true value is known in advance.

What one visit actually records

The five-parameter fit above treats each observation as a position on the sky, two numbers. A scanning satellite does not deliver that, and the difference explains the design of the whole mission.

The instrument measures precisely in one direction only — along the direction the field is sweeping across the detectors — because that is the direction in which the image’s transit time across a pixel column can be timed. Across that direction the measurement is far coarser, by more than an order of magnitude.

So a single visit is essentially a one-dimensional constraint: it says where the star lies along a particular direction on the sky and says almost nothing about where it lies across it. One such measurement cannot determine a position, let alone five parameters.

What supplies the second dimension is revisiting the same star with the scan running a different way. The satellite’s rotation axis is made to precess slowly about the direction of the Sun, so successive passes over a given star cross it at a range of angles, and the accumulated set of one-dimensional constraints intersects at a point.

That is why the scanning law is a designed object rather than a convenience: it has to guarantee that every part of the sky is visited enough times, at a wide enough spread of position angles, and at epochs spread through the year so that the parallax term is sampled at a range of phases. A region visited many times at one angle is measured well in one direction and badly in the other, and a region visited only in one season has its parallax entangled with its proper motion by the argument of this essay’s third section.

The precision of a catalogue is therefore not uniform across the sky, and the pattern of its variation is the pattern of the scanning law rather than anything astronomical.

The separation of the five parameters depends on how much of the path is observed, and it is worth drawing both the path and the degeneracy at a second setting.

One path, five numbers. Left: the apparent path of a star over 4 years, with a proper motion of 193 mas a year and a parallax of 10 mas, at ecliptic latitude 42°. It is one curve and there is nothing in the sky it can be compared against — the reference stars have paths of their own. Right: the same path with a straight line taken out of it. What is left is an ellipse of semi-major axis 10.0 mas and semi-minor axis 6.7 mas, closed and repeating once a year. The two are separated by their time signatures and by nothing else: proper motion is secular and parallax is annual, in a phase the Earth's position fixes in advance. That is why the five parameters can be told apart at all, and why an astrometric catalogue quotes five rather than two — a position without them is a position at one instant, which is not a direction to anything.
Fig. 7 The same five-parameter path for a star five times further away. The parallactic loop shrinks to a fifth and the proper motion does not, so the path becomes a nearly straight line with a small wiggle on it — and the wiggle is the entire distance measurement.
What an arc of a given length knows about a parallax. The residual an arc cannot explain, against the parallax assumed for it, with the position and proper motion re-solved exactly at every trial value. The star's true parallax is 50 mas and every curve passes through zero there, because there is no noise in the data — what differs is the shape. A 0.5-year arc gives 0.227 mas of residual per mas of error in the parallax; a 5-year arc gives 0.595, 2.6 times as much. A 0.5-year arc is not 10 times worse than a 5-year one for the reason it is 10 times shorter, and the reason is geometric rather than statistical: over a fraction of a year the parallactic displacement is part of one swing, and a straight line in time imitates part of a swing almost exactly. Only when the ellipse closes and starts again is there anything in the data that a proper motion cannot be. That is why a parallax programme is stated in years and not in observations.
Fig. 8 And what arcs of half a year, a year and a half, and five years permit. A baseline shorter than a year cannot separate the parallax from the proper motion at all, because over a fraction of an orbit both are approximately linear — which is why an astrometric mission’s first data release carries only two of the five numbers.

Where the ladder goes next

The rung above is the binary problem: a star with an unseen companion has seven parameters rather than five, and the extra two describe a periodic wobble that is neither secular nor annual. Astrometric detection of exoplanets is the same fit with a longer period in it, and the difficulty is precisely that a period near one year is degenerate with the parallax itself.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

AberrationAstrometric solutionAstrometryDegeneracyEpochParallaxProper motionRadial velocityReference frameSpace velocity