The observed sky

A time scale that is finished after it is used

The world's reference time is not kept by any clock. It is an average of about four hundred and fifty, weighted by how steady each has been, published a month late, and then recomputed a year later against the handful of clocks that define the second — so the best available time for any moment is known only long after the moment has passed. The best clocks are now so good that averaging them requires knowing each one's height to a centimetre.

Assumes Timescales, Radiometric navigation and Oblateness.

Astronomy keeps several kinds of second, and the one that underlies all the others is Terrestrial Time: the time a perfect clock would keep if it sat on the geoid, the surface of constant gravitational potential that mean sea level traces. Its rate is defined by a convention that fixes how it relates to the coordinate times of the Earth and the solar system, and its unit is the SI second, defined since 1967 by a transition frequency of the caesium atom.

A definition is not a clock. No single clock keeps Terrestrial Time, and none could, because every clock drifts, suffers noise, sits at some height above or below the geoid, and eventually fails. What is distributed as the world’s time is an average of many clocks, computed after the fact, and the way the average is built says a good deal about what a time scale is.

How steady each kind of clock is, by how long it is averaged. The Allan deviation — the typical fractional difference in frequency between successive averages of length τ — against τ, for representative noise models of caesium beam, hydrogen maser, caesium fountain, optical lattice clock. Each falls as τ^−½ while random frequency noise dominates, then levels off at a floor set by slower, correlated noise, and a clock with a systematic drift eventually rises again. A caesium beam is still improving at the longest averaging drawn, reaching 10⁻¹⁴; a hydrogen maser reaches its best, 7.3·10⁻¹⁶, after about 27 hours; a caesium fountain is still improving at the longest averaging drawn, reaching 1.2·10⁻¹⁶; an optical lattice clock is still improving at the longest averaging drawn, reaching 10⁻¹⁸. No clock is best at every averaging time — the maser is steadiest over hours and wanders over weeks, the fountain is noisier over minutes and more accurate over months — and a time scale is built by using each where it is best. The curves are typical shapes, not the specification of any particular instrument.
Fig. 1 The fractional frequency instability of four kinds of clock against the time over which their frequency is averaged. Each improves as the averaging lengthens, until slower noise sets a floor; the hydrogen maser is steadiest of the three older kinds over hours and begins to wander over days; the fountain is noisier over minutes and better over months; the optical clock is a hundred times better than any of them at every averaging time.

Steadiness depends on how long it is measured over

The natural way to describe a clock’s quality would be a single number — its error — and it does not work, because a clock’s error depends on the interval over which it is judged. The standard measure is the Allan deviation: take the clock’s average frequency over successive intervals of length τ, and ask how much successive averages differ. It is a function of τ, and its shape identifies the kinds of noise the clock suffers.

White frequency noise — random, uncorrelated fluctuations in the clock’s rate — averages away, and makes the Allan deviation fall as the inverse square root of τ. Flicker noise, whose origin in most clocks is not fully understood but which is ubiquitous in physical oscillators, does not average away and makes a floor. A systematic drift in frequency, such as the slow change in a hydrogen maser’s cavity as it ages, grows linearly with the interval and makes the curve rise again. Every real clock’s Allan deviation is some combination of these shapes.

The consequence is that clocks cannot be ranked. A hydrogen maser is superb over hours: its frequency, averaged over a day, is steady to a part in a thousand million million. Over weeks it drifts, and a caesium beam, far noisier over a day, is steadier over a month. A caesium fountain — atoms launched upward through a microwave cavity and allowed to fall back through it, so that the interaction time is a second rather than milliseconds — is noisier than a maser over minutes and far more accurate over months, because its frequency is set by the atom and not by an aging cavity. Each clock is best somewhere, and a time scale uses each where it is best.

An average that stops improving

Averaging clocks improves the result, as averaging any independent measurements does, and the improvement has a limit that is worth understanding.

Averaging clocks helps until they share something. The instability of an average of N hydrogen masers over 10 days, against N, when each clock alone is steady to 1.3·10⁻¹⁵ and the ensemble shares a common noise of 3·10⁻¹⁶ — from the comparisons that link the clocks, from their shared environment, or from the reference they are steered to. Independent noise averages away as one over the square root of N, the dashed line. Shared noise does not average at all, and the ensemble can never be steadier than it. The two are equal at about 17 clocks, and adding clocks beyond that buys almost nothing: an ensemble of a thousand is barely better than one of a hundred. International atomic time averages about four hundred and fifty clocks in eighty laboratories for exactly this reason — not because more would not help, but because what limits it is the links between them and the few clocks that define its accuracy.
Fig. 2 The instability over ten days of an average of N hydrogen masers, each steady to about a part in a thousand million million alone, when the ensemble shares a common noise three times smaller. The average improves as one over the square root of N until the shared noise dominates, at about seventeen clocks, and barely improves after that.

Independent noise in N clocks averages down as 1/N1/\sqrt{N}. Noise the clocks share does not average at all: if every clock in the ensemble is affected equally by something — a change in the laboratory’s temperature, an error in the links that compare clocks in different places, the reference they are all steered towards — the average is affected equally too. The ensemble’s instability therefore falls with N until the independent part has shrunk to the size of the shared part, and then it flattens. Beyond that point more clocks buy almost nothing.

International Atomic Time, the atomic time scale from which the world’s civil and scientific time is derived, is computed at the International Bureau of Weights and Measures from about four hundred and fifty clocks in some eighty laboratories. Most are commercial caesium beams and hydrogen masers. They are compared with each other, across continents, by timing signals from navigation satellites and by two-way exchanges through communication satellites, and it is those links — the propagation of a signal through the atmosphere and ionosphere between clocks thousands of kilometres apart — that contribute much of the shared noise. A better link is worth more than another clock.

Averaging clocks helps until they share something. The instability of an average of N hydrogen masers over 10 days, against N, when each clock alone is steady to 1.3·10⁻¹⁵ and the ensemble shares a common noise of 10⁻¹⁶ — from the comparisons that link the clocks, from their shared environment, or from the reference they are steered to. Independent noise averages away as one over the square root of N, the dashed line. Shared noise does not average at all, and the ensemble can never be steadier than it. The two are equal at about 157 clocks, and adding clocks beyond that buys almost nothing: an ensemble of a thousand is barely better than one of a hundred. International atomic time averages about four hundred and fifty clocks in eighty laboratories for exactly this reason — not because more would not help, but because what limits it is the links between them and the few clocks that define its accuracy.
Fig. 3 The same ensemble with the shared noise reduced threefold, as better comparison links would do. The point at which adding clocks stops helping moves from seventeen to about a hundred and sixty, and the ensemble’s floor falls by the same factor of three. The improvement came from the links, not from the clocks.

Comparing clocks that are an ocean apart

The ensemble exists only because its clocks can be compared, and the methods of comparison are ingenious in a way that explains why they, rather than the clocks, often limit the result.

The commonest is common view. Two laboratories each record the arrival time of the same navigation-satellite signal at the same moment, against their own clocks. Each measurement contains the satellite clock’s error and the propagation delay; the satellite clock’s error is identical in both and cancels in the difference, and the delays, through similar paths, nearly cancel. What survives is the difference between the two laboratories’ clocks, good to about a nanosecond per comparison and much better averaged over days. The second method, two-way transfer, sends signals in both directions through a communication satellite at the same time, so that the path delay, being nearly the same both ways, cancels when the two directions are subtracted. The third, newest and best, sends light through optical fibres between laboratories; over a thousand kilometres of stabilised fibre, two optical clocks can be compared to parts in ten to the nineteenth, far beyond anything a satellite link can do. Such links now connect several European laboratories, and they are the reason optical clocks in different countries can be compared at their full accuracy — when the fibres exist.

A weighted average, a month late

The average is not an equal one. Each clock is weighted by how steady it has been recently, judged by comparing it with the ensemble over the preceding months, so that a clock that has been behaving well counts for more and one that has started to wander is down-weighted before it can pull the average with it. The weights are capped so that no single clock, however good, can dominate — a precaution against a clock that has been excellent for a year and then fails in a way its record gave no hint of.

The computation is done once a month, over data from the preceding month, and the result is published as a table of the differences between the average and each laboratory’s own realisation of time. So International Atomic Time does not exist in real time. At any moment, what the laboratories distribute are their own local time scales, each steered to follow where the international average is predicted to be, and the actual average is computed a few weeks later and published as a correction to each. The world’s time is, strictly, an extrapolation, and the value it extrapolates towards is known only afterwards.

A free-running average of clocks has a further problem: nothing ties its rate to the definition of the second. Every clock in the ensemble could drift in the same direction — they are made by a few manufacturers, and share design choices — and the average would drift with them. The ensemble is therefore steered, gradually, towards the frequency of the primary standards: the dozen or so caesium fountains, and now some optical clocks, in national laboratories that realise the SI second from first principles, with every known systematic effect evaluated and corrected. The primary standards are few, and they do not run continuously; each contributes evaluations of its frequency over periods of weeks. They set the accuracy of the time scale. The ensemble sets its stability. The combination is better than either.

Terrestrial Time, a year after the fact

Terrestrial Time as distributed in real time is International Atomic Time plus a constant offset of 32.184 seconds, inherited from the ephemeris time it replaced. That is good to a few parts in ten to the sixteenth, and for almost every purpose it is good enough.

For the most demanding purposes it is not, and a better realisation is computed retroactively. Once a year, the International Bureau recomputes the whole of Terrestrial Time using every primary-standard evaluation available, including ones made after the fact, and publishes it as TT(BIPM) with a year’s date attached. The recomputed scale differs from the real-time one by tens of nanoseconds, accumulated over decades, and each annual version revises the previous one slightly.

Each version carries the year of its computation in its name, and a paper that needs the best time for an observation made in some year cites the version it used, since a later version will give a slightly different answer for the same moment. The reference time for a past event is therefore not a fixed number but the output of a computation that is repeated as better data about the clocks of that time become available — a situation more familiar from ephemerides, which are refitted as observations accumulate, than from anything called a clock.

So the best available time for any moment is not the one distributed at that moment, and it is not final even when it is published. Pulsar timing, which compares the arrival times of pulses from rotating neutron stars against a terrestrial clock over decades, uses the retroactive scale because a drift of a few parts in ten to the sixteenth, uncorrected, would appear in the pulsars’ timing residuals as a signal. The pulsar timing arrays that search for the gravitational waves from pairs of supermassive black holes need the clock error to be modelled explicitly, and one of the things they measure, in principle, is the error in the terrestrial time scale itself — a common signal in every pulsar, with a particular pattern on the sky, distinguishable from the gravitational waves they are looking for.

Civil time is a compromise with a spinning Earth

What the world’s clocks display is neither International Atomic Time nor Terrestrial Time but Coordinated Universal Time, which ticks at the atomic rate and is kept within nine-tenths of a second of the Earth’s actual rotation by inserting leap seconds. The Earth’s rotation is slowing, irregularly, and since 1972 thirty-seven seconds have been inserted, so UTC is now thirty-seven seconds behind atomic time. The decision taken in 2022 to stop inserting them by 2035 accepts a growing difference between clock time and the Sun, exactly the choice a calendar makes when it lets its dates drift from the seasons rather than correcting them often.

Satellite navigation keeps its own time too. Each navigation system maintains a system time steered to a national realisation of UTC but without leap seconds, and each satellite carries atomic clocks whose rates are offset before launch to compensate for the relativistic effects of their orbit — the correction that shows a clock’s rate depends on where it is, applied in practice to every receiver on the Earth. The navigation systems are also among the most important users and distributors of the international time scale, because they carry it to every receiver and are the main route by which the world’s laboratories compare their clocks.

When a clock is good enough to feel its own height

The gravitational redshift makes a clock’s rate depend on the potential it sits in, and near the Earth’s surface that means on its height.

The height at which a clock's own uncertainty is a height. The fractional change in a clock's rate from raising it through a height h near the Earth's surface, gh/c², against h, with the uncertainties of three kinds of clock drawn across it. A metre changes the rate by 1.1 × 10⁻¹⁶. A caesium beam cannot tell a height difference smaller than 0.9 km; a fountain, 1.8 m; an optical lattice clock, 0.9 cm. So the best clocks now resolve the gravitational potential of their own site to about a centimetre of height, and comparing two of them in different laboratories requires the difference in their heights on the geoid to be known that well — which conventional levelling across a continent does not deliver. Terrestrial Time is defined by the rate of a clock on the geoid, and realising it with optical clocks turns clock comparison into geodesy.
Fig. 4 The fractional change in a clock’s rate from a change in its height, against the height, with the uncertainties of three kinds of clock drawn across it. A metre of height is 1.1 parts in ten to the sixteenth. A caesium beam cannot resolve a height difference of less than about 900 metres, a fountain about two metres, an optical lattice clock about a centimetre.

Raising a clock by a metre makes it run faster by gh/c2gh/c^2, about 1.1 parts in ten to the sixteenth. For a caesium beam, uncertain to parts in ten to the thirteenth, that is irrelevant: the clock cannot tell whether it has been carried up a mountain. For a fountain, uncertain to a few parts in ten to the sixteenth, the laboratory’s height above the geoid has to be known to a metre or so, which is easy. For an optical clock uncertain to a part in ten to the eighteenth, the height has to be known to a centimetre.

That changes the relation between clocks and geodesy. Terrestrial Time is defined on the geoid, so every clock’s rate must be corrected by its own height above the geoid — its gravitational potential difference, strictly — and at a centimetre that is not a number levelling surveys routinely deliver across a continent. Two optical clocks in two countries cannot be compared at their full accuracy unless the geoid between them is known at the centimetre level. Turned around, comparing them measures the potential difference directly: the clocks become instruments of geodesy, and experiments have already used transportable optical clocks to measure height differences between laboratories to a few centimetres, by comparing their rates. The geoid itself is shaped by the Earth’s flattening and its uneven interior, and a network of such clocks is a new way of mapping it.

At that precision the solid Earth moves too. The tides that the Moon and Sun raise lift the ground, and every clock on it, by tens of centimetres twice a day, and the changing potential changes the clocks’ rates by parts in ten to the seventeenth. Optical clocks see the Earth breathe.

Seven decades of improvement

The pressure towards these refinements comes from the clocks, which have improved at an extraordinary and steady rate.

Seven decades of better clocks. The fractional uncertainty of the best clocks of each period, to the order of magnitude, against the year. From the first caesium standard in 1955 to the optical lattice and ion clocks of the 2020s the improvement has been about a factor of ten every 7 years — nine orders of magnitude in seventy years. The second has been defined by caesium since 1967, and the best clocks are now optical ones a hundred times more accurate than the caesium standards that realise the definition, which is why a redefinition of the second in terms of an optical transition is being prepared. Each factor of ten has made a new term matter: at 10⁻¹⁶ the height of the laboratory, at 10⁻¹⁸ the tides of the solid Earth, which raise and lower every clock by tens of centimetres twice a day.
Fig. 5 The fractional uncertainty of the best clocks of each period, to the order of magnitude. From the first caesium standard in 1955 to the optical clocks of the 2020s the uncertainty has fallen by nine orders of magnitude, about a factor of ten every eight years.

The first caesium standard, built in 1955, was good to about a part in a thousand million. The second was redefined in terms of caesium in 1967, when the atomic clocks had become better than the astronomical definition they replaced — the ephemeris second, tied to the Earth’s orbital motion, which in turn had replaced the rotational second tied to a spinning Earth that turned out to be a poor clock. Caesium beams improved to parts in ten to the fourteenth; fountains, in the 1990s, to parts in ten to the sixteenth. Optical clocks, which count the oscillations of light rather than of microwaves, at frequencies a hundred thousand times higher, passed the fountains in the 2010s and now reach parts in ten to the eighteenth.

The second is still defined by caesium, and so the best clocks in the world are now a hundred times more accurate than the definition they are measured against. A redefinition in terms of an optical transition is being prepared, and one of the obstacles is exactly the one in the last section: an optical second realised in one laboratory and compared with another requires the potential difference between them at a level that geodesy is only beginning to supply.

What the figures leave out

The Allan deviation curves are representative shapes, not the performance of particular instruments; individual clocks of each kind differ by factors of several, and the best of each kind is much better than the typical one that contributes to the international average. The ensemble figures treat all clocks as identical and the shared noise as a single number, where the real ensemble is heterogeneous — masers, beams and fountains mixed — and the shared noise has several sources with different spectra. The actual algorithm that computes International Atomic Time accounts for this with a frequency-prediction model for each clock, and it is considerably more elaborate than an average.

And every figure treats time as something measured at a place. Distributing it — to a navigation satellite, to a telescope, to a spacecraft far from the Earth — is a separate problem, of measuring the delay of a signal whose path length is itself uncertain, and for the most precise applications the transfer is the limit rather than the clock.

Still open: a second that is also a height

When the second is redefined in terms of an optical transition, the clocks that realise it will be accurate to parts in ten to the eighteenth, and a time scale built from them will be steady to a centimetre of height. The definition of Terrestrial Time on the geoid then becomes a definition that needs the geoid known to a centimetre everywhere a contributing clock sits, and the geoid is not known that well. One possibility is that the clocks themselves define it — that a network of optical clocks, compared through optical fibres and satellite links, becomes the reference for the Earth’s gravitational potential rather than a user of it. Whether time and height are then two measurements or one, and which institution maintains the result, is an open question about the future of both metrology and geodesy.