Exoplanets

A forecast that fails on a schedule

A transiting planet perturbed near a resonance keeps a clock that wanders, and a straight-line ephemeris fitted to part of the wander predicts the next transit with a confidence the wander does not deserve. The error is not noise and does not average down; it grows on the pair's super-period, it is many times the formal uncertainty within a year, and how soon it appears depends on which stretch of the wander happened to be observed. When a model that includes the known perturber still fails, the failure has a period, and the period is a planet.

Assumes Transit-timing and Transit-timing.

A transit is only useful if someone is watching when it happens. A planet found by a space telescope is followed up by instruments on the ground and in orbit that book their time months in advance, and every booking rests on a forecast: a transit time, extrapolated from the ones already seen, with an uncertainty attached. For a planet on an undisturbed orbit the forecast is a straight line and its uncertainty grows slowly and honestly, in proportion to how far ahead it reaches.

For a planet whose transits run late and early because a neighbour is pulling on it, the straight line is the wrong model, and its uncertainty is a statement about a planet that does not exist. The forecast does not merely get worse; it fails, on a schedule the neighbour sets, and the size of the failure has nothing to do with how precisely the transits were timed.

An ephemeris fitted to 120 days, 43 minutes wrong within a year against a band of ±5.1. Transit times of a 6 Earth-mass planet on a 10-day orbit, perturbed by a 14 Earth-mass planet at 15.24 days, integrated for 1460 days and compared with a straight-line ephemeris fitted only to the transits in the first 120 days — the shaded window. Inside the window the line fits to 2.1 minutes. Outside it the pair's 317-day super-period carries the transits away from the line, and within a year of the window closing the prediction is 42.8 minutes early of the observed transit, 347 days after the last fitted one. The narrow band is the formal three-sigma uncertainty of the same line for a timing precision of 0.5 minutes per transit, which at that date is ±5.1 minutes: the error is 8.4 times the band. A statistical uncertainty assumes the residuals are noise, and these are a signal, so the band describes a planet that does not exist.
Fig. 1 Transit times of a 6 Earth-mass planet on a 10-day orbit, perturbed by a 14 Earth-mass planet at 15.24 days, integrated for four years and compared with a straight-line ephemeris fitted only to the first 120 days — the shaded window. Inside the window the line fits to 2.1 minutes. Within a year of the window closing the prediction is 42.8 minutes early, 347 days after the last fitted transit, and the formal three-sigma band of the same fit, for 30-second timing, is ±5.1 minutes there: the error is 8.4 times the band.

The wander the line is fitted to

The planets in these figures are an ordinary near-resonant pair: the outer orbit takes 1.524 times as long as the inner, just wide of the 3:2 commensurability. Integrated on its own and set against the best straight line through all four years, the inner planet’s transits look like this.

A transit that will not keep time. Transit times of a planet of 6 Earth masses on a 10-day orbit, minus the best straight line through them, from a three-body integration in which the second planet at 15.24 days is the only thing perturbing it. The residual swings by ±10.9 minutes and repeats over 317 days — the super-period 1/|3/P₂ − 2/P₁|, which is long precisely because the pair is near the 3:2 resonance and gets longer the nearer it is. The amplitude is proportional to the perturbing planet's mass, so fitting this curve weighs a planet that may never transit at all. The short jagged component is not noise: it is the chopping signal, one kick per conjunction at the 29.1-day synodic period, sampled once every 10 days by the transits themselves and so barely above the rate at which it can be followed. The integration conserved energy to 1.7e-10.
Fig. 2 The same pair’s inner-planet transit times over 146 orbits, minus the best straight line through all of them. The residual swings by ±10.9 minutes and repeats every 317 days — the super-period 1/|3/P₂ − 2/P₁|, long because the pair is close to the 3:2 resonance. The short jagged component is the chopping at the pair’s 29.1-day synodic period, sampled once every 10 days by the transits. The integration conserved energy to two parts in ten billion.

Two features of that curve decide everything that follows. The first is its period. At 317 days, a single super-period is longer than a typical season of ground-based observations and far longer than the few months over which a newly found planet’s ephemeris is first measured. A fit to a short window therefore sees not a sinusoid but a fragment of one, and a fragment of a sinusoid is very well described by a line with the wrong slope.

The second is its amplitude, eleven minutes each way. That is small beside a 10-day period — a hundredth of a per cent — and large beside the half-minute to which a good transit can be timed. The planet’s clock is overwhelmingly regular and not regular enough, which is exactly the combination that makes a straight-line fit look excellent and forecast badly.

Why the error grows and does not wander back

A straight line fitted to a fragment of a sinusoid has two faults. Its intercept absorbs where on the wave the fragment sat, which does no lasting harm. Its slope absorbs how fast the wave was rising or falling across the fragment, and that is converted into an error in the period.

A period error does not stay put. A transit’s predicted time is the fitted time of the first one plus the fitted period multiplied by the number of orbits since, so an error of a few seconds in the period becomes an error of a few seconds times the number of orbits, and grows without limit. The wander itself swings back after half a super-period; the line does not follow it. That is why the opening figure’s residuals rise steadily through four years, with the super-period’s swing riding on top of a climb, rather than oscillating about zero.

The same shape turns up wherever a period is estimated from a short arc of a longer motion. An orbit determined from a short stretch of observations carries almost all of its uncertainty along the track, because the thing a short arc measures worst is how long the orbit takes, and that error accumulates into position one orbit at a time. The transit ephemeris is the same problem in one dimension: a clock whose rate was measured over too short an interval to average out its own fluctuations.

The size of the period error is modest by any ordinary standard, which is what makes it dangerous. The 120-day fit’s worst forecast error comes some forty orbits after the middle of its window, so the forty-three minutes it accumulates there correspond to a period that is wrong by roughly a minute out of fourteen thousand — a fractional error of less than one part in ten thousand, from a fit whose residuals inside the window were two minutes. A period quoted to six significant figures is being quoted to about its fifth, and the sixth digit is where the neighbour lives.

Nor is the fitted period a poor estimate of anything in particular. It is an excellent estimate of the rate at which the planet’s transits were advancing during those 120 days, which is a real, physical rate: the planet’s orbit at that phase of the exchange really was that long. An ephemeris is a table that is a fit, and a fit reports the model that best describes the data it was given; what it cannot report is how much of the planet’s future lies outside the family of models it was allowed to choose from.

A band that describes a different planet

The narrow band in the opening figure is not a strawman. It is the uncertainty the fit reports, computed in the standard way from the scatter a timing precision of half a minute would put on each transit, and at the date of the worst error it is ±5.1 minutes. Nothing about that calculation is wrong for the model it assumes, which is a planet whose transits scatter independently about a straight line.

The planet in the figure does not do that. Its departures from a line are correlated from one transit to the next, because they are samples of a smooth wave. A least-squares fit that treats correlated departures as independent noise is overconfident in a specific way: it believes it has as many independent pieces of information as transits, when it has something closer to one piece of information per super-period.

An ephemeris fitted to 120 days, 43 minutes wrong within a year against a band of ±20.4. Transit times of a 6 Earth-mass planet on a 10-day orbit, perturbed by a 14 Earth-mass planet at 15.24 days, integrated for 1460 days and compared with a straight-line ephemeris fitted only to the transits in the first 120 days — the shaded window. Inside the window the line fits to 2.1 minutes. Outside it the pair's 317-day super-period carries the transits away from the line, and within a year of the window closing the prediction is 42.8 minutes early of the observed transit, 347 days after the last fitted one. The narrow band is the formal three-sigma uncertainty of the same line for a timing precision of 2 minutes per transit, which at that date is ±20.4 minutes: the error is 2.1 times the band. A statistical uncertainty assumes the residuals are noise, and these are a signal, so the band describes a planet that does not exist.
Fig. 3 The same fit to the same 120 days, with the formal band computed for a timing precision of 2 minutes per transit — a typical figure for a transit timed from the ground. The prediction error a year later is the same 42.8 minutes, because the transits themselves are the same; the band is four times wider, ±20.4 minutes, and the error is still 2.1 times it. Worse timing makes the forecast look less wrong without making it more right.

That figure contains a result that is easy to misread. Degrading the timing precision widens the band and leaves the error untouched, so the ratio of error to band improves — from 8.4 to 2.1 — and a naive check of whether the forecast failed “significantly” would find it had failed less. A forecast’s honesty is not a property of its error bar. The band here scales with the noise, the error scales with the planet’s neighbour, and the two are unrelated quantities that happen to be expressed in the same units.

The prudent forecast in practice is not a straight line with an inflated band but a model that contains the perturber. When the neighbour transits too, its period is known, the super-period is known, and a two-planet integration fitted to the transits predicts the wander rather than being surprised by it. Every well-characterised timing system is forecast that way. The difficulty is that the ephemeris is needed first — to schedule the observations that would reveal the wander — and the first ephemeris of any planet is a line.

A longer window buys less than it seems

The obvious remedy is to wait: fit more of the wander before forecasting.

An ephemeris fitted to 300 days, 19 minutes wrong within a year against a band of ±1.5. Transit times of a 6 Earth-mass planet on a 10-day orbit, perturbed by a 14 Earth-mass planet at 15.24 days, integrated for 1460 days and compared with a straight-line ephemeris fitted only to the transits in the first 300 days — the shaded window. Inside the window the line fits to 12.0 minutes. Outside it the pair's 317-day super-period carries the transits away from the line, and within a year of the window closing the prediction is 18.5 minutes early of the observed transit, 327 days after the last fitted one. The narrow band is the formal three-sigma uncertainty of the same line for a timing precision of 0.5 minutes per transit, which at that date is ±1.5 minutes: the error is 12.1 times the band. A statistical uncertainty assumes the residuals are noise, and these are a signal, so the band describes a planet that does not exist.
Fig. 4 The same pair with the straight line fitted to the first 300 days — nearly a whole super-period. Inside the window the line now fits to 12.0 minutes, much worse than before, because it is being asked to follow most of a swing. A year after the window closes the prediction is 18.5 minutes early, less than half the error of the 120-day fit; the formal band has shrunk faster, to ±1.5 minutes, and the error is now 12.1 times it.

The forecast improved and its honesty got worse. A window that spans nearly a full super-period averages most of the wave out of the slope, so the period error falls and the extrapolation drifts more slowly. At the same time the fit has twenty-nine transits instead of twelve, spread over a longer lever arm, and its formal uncertainty collapses. The residuals inside the window, twelve minutes against a claimed precision of half a minute, are the warning, and a fit that reports a reduced chi-squared of several hundred is announcing that its error bar is meaningless. The warning is available only if someone looks at the residuals rather than the band.

The general statement is that a linear ephemeris for a planet with timing variations is only as good as the fraction of a super-period it spans, and a window of several super-periods is needed before the line’s slope converges on the planet’s mean period. For a pair near resonance that can be years, and the pairs nearest resonance — with the largest, most easily detected timing variations — have the longest super-periods and so the slowest convergence.

Chains of planets make it worse in a way no single pair does. In a system where four or five planets sit near successive commensurabilities, as they do in a chain that could not have been assembled in place, every adjacent pair contributes its own super-period, and the planets in the middle of the chain feel two at once. Their timing wanders at several incommensurable periods simultaneously, so there is no window length that spans a whole cycle of all of them, and the fitted slope keeps changing as the window grows. The forecasts for such systems are only ever made from full dynamical fits, and those fits have themselves been redone each time new transits arrived, because the masses they return depend on how many beats of each super-period the data contain.

The reverse is also true. A planet with no neighbour near any commensurability has timing variations too small to matter, and its linear ephemeris is exactly as good as its error bar. The forecasts that fail are therefore not scattered at random through a catalogue; they are concentrated in the most dynamically interesting systems, which are the ones follow-up observers most want to catch.

The calendar decides how soon

Two observers fitting the same planet over windows of the same length, at different times, do not get the same forecast.

The same pair, fitted at three phases of its wander. Three straight-line ephemerides for the same 10-day planet, each fitted to 120 days of transits taken at a different point of its 317-day super-period, and the error each makes in the 365 days after its window closes. Fitted over days 0–120: at most 5.3 minutes wrong in the first 60 days, 42.8 at worst within 365; Fitted over days 80–200: at most 15.7 minutes wrong in the first 60 days, 44.3 at worst within 365; Fitted over days 160–280: at most 4.7 minutes wrong in the first 60 days, 50.3 at worst within 365. The planets, the masses and the timing are identical in all three; only the stretch of the wander that happened to be observed differs. Over a whole year every forecast meets the full swing and fails by a similar amount. What the phase decides is how soon: a window that caught the wander while it was bending extrapolates a slope that is already wrong, and one that caught it while it ran straight keeps its accuracy for weeks longer. The error of a forecast from transit times is therefore a property of the calendar of the observations as well as of the system.
Fig. 5 Three straight-line ephemerides for the same planet, each fitted to 120 days of transits at a different point of its super-period, and the error each makes in the year after its window closes. Fitted over days 0–120, the forecast is at most 5.3 minutes wrong in its first 60 days; over days 80–200, 15.7 minutes; over days 160–280, 4.7 minutes. Within the full year every forecast meets the whole swing and fails by a similar amount: 42.8, 44.3 and 50.3 minutes at worst.

Over a whole year the three forecasts fail about equally, because a year is longer than the super-period and each extrapolation meets every phase of the wave. What the fitted phase decides is how soon. A window that caught the wander while it was bending — near the top or bottom of a swing — extrapolates a slope that is already turning wrong, and is off by a quarter of an hour within two months. A window that caught it running nearly straight through its mean extrapolates a slope that stays roughly right for a while, and holds to five minutes over the same two months.

For a follow-up programme the difference is practical. Observations in the two months after an ephemeris is published fall inside a transit window for one of those forecasts and outside it for another, for the same planet with the same timing precision. A missed transit is then commonly attributed to weather or to an instrument, when the cause is the calendar of the original observations.

When the model with the neighbour fails too

A two-planet model predicts the wander of a two-planet system, and when the system has a third planet the prediction fails in a way that is more informative than any success.

What a two-planet solution leaves behind when there are three. Residual transit times of a 10-day planet after subtracting everything a two-planet model can account for: a straight line, the near-resonant sinusoid with its first harmonic at the 305-day super-period of its pairing with the 15.24-day planet, and the chopping at the pair's 29.1-day synodic period with two harmonics. Integrated with only those two planets, what is left is 0.87 minutes rms, the terms of higher order that the analytic model leaves out. Integrated with a 8 Earth-mass third planet added at 6.5 days, which never transits, the same subtraction leaves 2.51 minutes rms, and the leftover has a period: fitted at the 131-day super-period of the third planet's own pairing with the transiting one, it has an amplitude of 2.93 minutes, against 0.06 at the same period in the two-planet residual. The two-planet fit does not fail by giving a wrong answer. It fails by leaving a structured residual, and the structure is the measurement of the planet that was not in the model.
Fig. 6 Residuals after subtracting everything a two-planet model accounts for — a straight line, the near-resonant sinusoid at the fitted 305-day super-period with its harmonic, and the chopping at the 29.1-day synodic period with two harmonics. With only the two planets present, what is left is 0.87 minutes rms. With an 8 Earth-mass third planet added at 6.5 days, which never transits, the same subtraction leaves 2.51 minutes rms, and the leftover has a period: at 131 days, the super-period of the third planet’s own pairing with the transiting one, it carries 2.93 minutes, against 0.06 at the same period without the third planet.

The figure’s two curves have the same cause and different contents. The two-planet residual is the part of a two-planet system’s own signal that the analytic model leaves out — terms of higher order in the masses and the eccentricities — and it has no single period. The three-planet residual is dominated by one period that none of the modelled terms has, and a period that is not in the model is a statement about a body that is not in the model. From that period and the transiting planet’s own, the list of candidates for where the extra body sits can be drawn up, and a velocity search can be pointed at the right periods.

The detail in the caption that the super-period fitted from the data is 305 days, not the 317 days the pair’s starting periods predict, is not an error in either. Orbital elements are osculating — each is the orbit a body would follow if every perturbation vanished at that instant — and a near-resonant pair’s instantaneous periods differ from their long-run averages by a fraction of a per cent. Over four super-periods that fraction becomes a visible phase drift, and a model built from the osculating values would itself leave a structured residual. The periods that matter for forecasting are the mean ones, and they are only available from the data.

A residual of that kind can only be read once everything known has been taken out of it, and taking it out is itself the craft. The subtraction in the figure is the timing version of the series of known perturbations that is subtracted from a planet’s observed positions before anything new is claimed: each term removed is a modelled effect of a known body, and whatever survives every subtraction is attributed to something unknown. The claim is only as good as the completeness of the terms. A two-planet model that omitted the chopping, as the first version of this subtraction did, left a residual of nearly two minutes from the two known planets alone — larger than the signal of a small third planet — and would have attributed the known system’s own conjunctions to an imaginary neighbour.

The whole construction has a famous ancestor. The positions of Uranus refused to follow their tables in the 1830s and 1840s, and the residual had the shape a further planet would produce. Urbain Le Verrier and John Couch Adams computed where that planet had to be, and Neptune was found in 1846 within a degree of the prediction. A decade later the same method was turned on the unexplained advance of Mercury’s perihelion, and predicted a planet inside Mercury’s orbit that was never found, because that residual was not a planet but a correction to Newton’s law of gravity. A structured residual is evidence that the model is missing something. It is not, by itself, evidence of what.

What was actually measured

Kepler-9, the first system in which timing variations were measured, is two giant planets near a 2:1 commensurability whose transits wandered by tens of minutes over the telescope’s first months. Its early ephemerides, extended by a few years, mispredicted later transits by amounts larger than their quoted uncertainties, and only a dynamical fit containing both planets forecast them correctly. The same experience recurred across the Kepler catalogue once other telescopes began re-observing its planets: for systems with strong timing variations, the transit windows computed from Kepler-era linear ephemerides had drifted by hours within a decade, and dedicated networks now maintain ephemerides for planets that future missions intend to observe, precisely because a stale forecast is a missed transit.

The third-planet signature has been seen as well. Systems in which a two-planet fit to the transit times left a periodic residual have been followed by velocity surveys, and some of those residuals turned out to be non-transiting planets at one of the periods the timing allowed. Others have turned out to be stellar activity, whose spots distort individual transit shapes and shift their fitted times, and whose rotation period can masquerade as a periodic timing term.

What the forecast leaves out

The perturber is on a fixed orbit. Over four years its orbit precesses and its eccentricity exchanges slowly with the transiting planet’s, and the figures include that exactly because they are integrations; a forecast based on a two-planet analytic model does not, which is part of the 0.87-minute floor.

The timing noise is uniform and independent. Real transit times from different instruments carry different precisions and different systematic offsets, and a fit that combines them without allowing for an offset between instruments produces a slope error of its own.

The system is coplanar. A mutual inclination between the two planets makes the transiting planet’s impact parameter drift, which changes its transit durations and, at the precision of a few seconds, the fitted mid-times. That is a further way for a forecast to fail, and none of the figures contains it.

Still open: how long a forecast should be trusted before the neighbour is known

A newly found planet has a linear ephemeris and, usually, no way of knowing whether it has a near-resonant neighbour until its timing has been followed for longer than the super-period it would have. The question a follow-up programme actually faces is therefore statistical: given the census of timing variations among planets of this kind and size, how wide should the transit window be at each date ahead, so that the chance of missing the transit is kept below some level? The census that would answer it is biased towards the systems whose variations were large enough to find, which are the ones with the longest super-periods and the worst forecasts, and correcting for that bias is the same problem as a prediction whose expiry date is itself uncertain.

About the same objects

Not linked from either essay — found by the objects both name.

The objects this essay names

Each one links to every other essay that touches it.

Chopping signalEphemerisMean motion resonanceN-body integrationOsculating elementsSuper-periodSynodic periodSystematic errorThree-body problemTransit-timing variation