Orbits

Elements that do not stay constant

Six numbers fix an orbit for all time, and the phrase is only true in a universe containing two bodies. Add a third and the six start moving — some of them wandering and returning, one or two of them drifting in one direction forever, and the difference between those two behaviours is the whole of celestial mechanics after Newton.

Assumes Orbital elements and The three-body problem.

Six numbers fix an orbit for all time, and that sentence has an unstated condition attached to it which is false of every real system: that nothing else is pulling. The moment a third body exists, the six stop being constants. What they become instead is more interesting than either a constant or a mess.

The osculating semi-major axis of a perturbed orbit. The semi-major axis a test particle would have if the perturber vanished, computed from its position and velocity at every step of an integration over 26 orbits of the perturber. It is not constant: a short-period ripple rides on a slow trend, and only the trend accumulates.
Fig. 1 The semi-major axis of a test particle over twenty-six orbits of a perturber one thousandth the mass of the primary, computed from the particle’s position and velocity at every step of an integration. It is not constant and it is not chaotic either. There is a ripple that returns to where it started each time round, and underneath it a slow trend that does not — and the two have entirely different consequences, because only one of them accumulates.

What an element is, once it stops being constant

The trick that makes the whole subject possible is to keep using the elements anyway, and to redefine what they mean.

At any instant a body has a position and a velocity. Those six numbers determine a unique Keplerian orbit — the ellipse with the primary at a focus — the one it would follow from now on if every other body vanished. Its elements are the osculating elements, from the Latin for kissing: the conic that touches the true path at this instant, matching it in position and in velocity but not in curvature, because the perturbation is still acting.

So a perturbed orbit is described as a Keplerian orbit whose elements are slowly varying functions of time. That is a change of variables and nothing more; no approximation has been made yet. Its value is that the new variables are nearly constant, and a differential equation for a nearly constant quantity is far better behaved than one for a rapidly changing one. Position and velocity oscillate by their full magnitude every orbit. The semi-major axis in the figure above moves by three parts in a thousand.

An orbit at i = 42°, Ω = 35°, ω = 55°. The three orientation elements. The orbit is generated in its own plane and rotated by the standard sequence, so the inclination, the node line and the argument of periapsis are the angles that produced the drawn curve rather than labels applied to it.
Fig. 2 The six, in the geometry they describe. Size and shape are the semi-major axis and the eccentricity; orientation is the inclination, the longitude of the ascending node and the argument of periapsis; and the sixth places the body along the path. Under perturbation each acquires its own equation of motion, and they are not equally interesting: the first two say how much energy and angular momentum the orbit holds, and the middle three say only which way it is pointing.

Lagrange and Gauss both wrote down the resulting equations, in different forms, and they share a structure worth stating without the algebra. Each element’s rate of change is proportional to a derivative of the disturbing function — the extra potential the third body contributes — with respect to one of the other elements. The system is coupled and messy, and it has one saving property, which the conservation of areal speed is not enough on its own to supply: the disturbing function can be expanded as a sum of cosines of angles that are integer combinations of the two bodies’ orbital positions.

That expansion sorts every term into two piles, and the sorting is the point.

Two kinds of term, and only one of them matters

A term whose angle runs at some multiple of the orbital frequencies is periodic. It makes the element wobble, the wobble returns to zero every time the angle comes round, and after a hundred orbits it has contributed nothing.

A term whose angle does not depend on either body’s position in its orbit — only on the slowly turning orientations — is secular. It does not return. Its contribution accumulates linearly, and given long enough it dominates everything else however small it is.

The osculating eccentricity of a perturbed orbit. The eccentricity a test particle would have if the perturber vanished, computed from its position and velocity at every step of an integration over 26 orbits of the perturber. It is not constant: a short-period ripple rides on a slow trend, and only the trend accumulates.
Fig. 3 The eccentricity of the same integration — the quantity that decides how far up the walls of the body’s own effective potential it climbs. The ripple is larger relative to the mean than it was in the semi-major axis, and the trend is in the other direction. This is the general pattern, and it has a deep reason: the semi-major axis has no secular term at first order in the perturbing mass, a result of Laplace’s and Lagrange’s that is the closest thing celestial mechanics has to a stability theorem. Energy is exchanged and given back. Shape and orientation are not.

The eighteenth century spent a great deal of effort on the question of which pile a given observation belonged to, because the answer decided whether the solar system was stable. Halley found in 1695 that Jupiter’s mean motion appeared to be accelerating and Saturn’s decelerating, by comparing modern observations with the Chaldean eclipse records and with Ptolemy’s. Left to run, that drift would have taken Jupiter into the Sun. Euler and Lagrange both attacked it and neither settled it.

Laplace settled it in 1785, and the answer was that the term is periodic with a period of about 900 years. Two frequencies that nearly share a small integer ratio produce a long beat, which is the mechanism behind the saros as well, and here the beat is a wobble in a planet’s longitude rather than a repeating eclipse. It is large — some 21 arcminutes in Jupiter’s longitude and 49 in Saturn’s, which is why it was noticed at all — and it is large because the two planets’ mean motions are nearly in the ratio 5 : 2, so the combination 2nJ5nS2n_{\rm J} - 5n_{\rm S} is a small number, and a small number in a denominator makes a big coefficient. It has been called the great inequality ever since. What looked like the solar system falling apart was a very slow oscillation seen over less than a tenth of a cycle.

Small denominators, which is where the method breaks

That near-commensurability is the crack in the whole enterprise. The expansion produces denominators of the form j1n1+j2n2j_1 n_1 + j_2 n_2, and there are infinitely many integer pairs, so some of those denominators are arbitrarily close to zero. The coefficient blows up. The series that was supposed to be a small correction stops being small.

Poincaré showed in 1890 that this is not a repairable defect of the technique: the series of classical perturbation theory are in general divergent, and the three-body problem has no convergent expansion of that kind. What the twentieth century added was the KAM theorem, which says that most orbits survive anyway — the ones whose frequency ratios are sufficiently irrational — while the ones near a resonance do something else entirely.

The Kirkwood gaps. Asteroid numbers against semi-major axis, with the resonant radii marked. Each gap sits where the orbital period is a simple fraction of Jupiter's, and each of those radii is computed from the harmonic law rather than placed by eye.
Fig. 4 What “something else entirely” looks like in the asteroid belt. The histogram is the observed distribution of semi-major axes; the marked positions are where the period would be a simple ratio of Jupiter’s. At those places perturbation theory’s denominators vanish, the eccentricity is pumped up over a few million years instead of oscillating, and the body’s perihelion is dragged into the inner solar system where a planet removes it. The gaps are what is left — the one place in the belt where the small correction is not small.

The perturbation with a closed form

One important case is tractable to the end, and it is not a third body at all. It is also the second time this collection has met a torque acting on a spinning oblate object and producing a slow, unstoppable turn of an orientation: the precession of the equinoxes is the same secular term with the roles of primary and perturber exchanged. A primary that is not a sphere has a potential that is not GM/r-GM/r, and the leading departure is measured by a single number, J2J_2, the coefficient of the quadrupole term. For the Earth it is 1.0826×1031.0826 \times 10^{-3}: small, and by far the largest perturbation acting on anything in low orbit.

Averaged over an orbit, an oblate primary produces no secular change in the semi-major axis or the eccentricity — the same protection as before — and two clean secular rates in the orientation. The node moves at

Ω˙=32J2(Rp) ⁣2ncosi,\dot\Omega = -\frac32 J_2 \left(\frac{R_\oplus}{p}\right)^{\!2} n \cos i,

and the argument of periapsis at a rate proportional to (5cos2i1)(5\cos^2 i - 1).

Node regression against inclination, at 800 km. The rate at which an orbit's node moves under the Earth's oblateness, against the orbit's inclination, for a circular orbit at 800 km. It is proportional to cos i, so a polar orbit does not drift at all and a retrograde one drifts the other way — and at 98.6° the drift is exactly one turn a year, which is what makes a sun-synchronous orbit possible.
Fig. 5 Node regression against inclination for a circular orbit at 800 kilometres. The cosine is the whole shape: a prograde orbit’s node moves westward, a polar orbit’s does not move at all, and a retrograde orbit’s moves east. The marked point is where the rate is one full turn per year — 0.9856 degrees a day — which happens at an inclination of 98.6°, and an orbit there keeps a fixed angle to the Sun for its whole life.

That marked point is the reason nearly every Earth-observation satellite flies retrograde. A sun-synchronous orbit crosses the equator at the same local solar time on every pass, so the illumination in an image taken in March is the illumination in an image taken in September and the two can be differenced. Landsat, the Sentinels, and every operational weather satellite in polar orbit are there. The orbit is not chosen for its inclination; the inclination is solved for, from a perturbation that a two-body treatment says does not exist.

Node regression against inclination, at 20200 km. The rate at which an orbit's node moves under the Earth's oblateness, against the orbit's inclination, for a circular orbit at 20200 km. It is proportional to cos i, so a polar orbit does not drift at all and a retrograde one drifts the other way — and at this altitude the largest rate available anywhere on the curve is 0.067° a day, short of the 0.986 a sun-synchronous orbit needs, so no inclination gives one.
Fig. 6 The same relation at the altitude of a navigation satellite, 20,200 kilometres. The rate has collapsed by more than two orders of magnitude, because J2J_2’s effect falls as (R/a)7/2(R/a)^{7/2} once the mean motion’s own dependence on aa is included — so the same planet’s oblateness dominates the design of a low orbit and is a nuisance correction for a high one. The largest rate available anywhere on this curve is 0.067 degrees a day, against the 0.986 a sun-synchronous orbit needs: at this altitude the orbit does not exist at any inclination, and the generator says so rather than solving for one.

The other rate has a marked point of its own. Setting 5cos2i1=05\cos^2 i - 1 = 0 gives i=63.4°i = 63.4°, at which the apsidal line does not turn. The Molniya orbits used by the Soviet Union for communications at high latitudes are highly eccentric, with apogee over the northern hemisphere, and they were flown at 63.4° for exactly this reason: at any other inclination the apogee would migrate south within a year or two and the orbit would stop doing its job. A number that is a root of a quadratic in cos2i\cos^2 i became an engineering constant.

What was actually measured

Uranus’s longitude, and it was wrong by two arcminutes.

Uranus was discovered in 1781 and has an orbital period of 84 years, so by 1820 something under half an orbit had been observed — but it had also been recorded, unrecognised, as a star on at least twenty occasions going back to 1690, which extended the arc to more than a full revolution. Every attempt to fit a single Keplerian orbit with the known planets’ perturbations included failed. Bouvard’s tables of 1821 fitted the modern observations and missed the older ones; refitted to all of them, they missed the modern ones by an amount that grew to about 2 arcminutes by 1845.

Two arcminutes is a fifteenth of the Moon’s diameter, and it is not a small quantity in positional astronomy — meridian circles of the period were good to a second of arc or so. The residual was real, systematic, and had the shape of a perturbation from something further out.

The inversion was performed independently by John Couch Adams and Urbain Le Verrier. It is a genuinely hard problem, because the perturbing body’s mass, distance and position along its orbit all have to come out of the shape of the residual curve rather than being given. Le Verrier’s prediction reached Johann Galle in Berlin on 23 September 1846; Neptune was found that night, within a degree of the predicted place, by a search that took under an hour because Galle had a star chart good enough to notice an object that was not on it.

The number to keep is the ratio. A residual of 22' in a body’s position, arising from a perturbation of order 3×1053\times10^{-5} of the Sun’s pull, was inverted to give a planet’s location to 1°. Perturbation theory is not a correction applied to observations; here it is the instrument.

The same inversion, in another star’s system

The Neptune method is not a piece of history. It is in current use, and the systems it is used on are ones where nothing but the perturbation is visible.

A transiting planet with no companion crosses its star at strictly even intervals, and the interval is the period the harmonic law relates to the orbit’s size. A second planet perturbs the first, its orbital elements acquire exactly the periodic and secular terms described above, and the transits arrive early and late on a cycle. The signal is a residual in a time of arrival rather than in a position on the sky, and it is inverted for the mass and orbit of a body that may never transit at all — which is Le Verrier’s problem with the observable changed and the arithmetic unchanged.

Two features of the exoplanet version make it sharper than the solar-system one. The amplitudes are enormous by comparison, because the systems found this way are compact and often near resonances, where the small divisors that ruin the series also amplify the signal: timing residuals of hours are routine where Uranus’s residual was two arcminutes. And the perturbations are the only mass measurement available for a system with no radial-velocity detection, so the whole mass scale rests on them. A resonant chain of planets is weighed entirely by how its members disturb each other.

The same effect runs the other way for the bodies that were pushed rather than merely nudged. A giant planet found at a hundredth of an astronomical unit cannot have formed there, and the secular terms of this essay are the machinery by which the disc it formed in moved it.

The method, before it was about orbits

The move made at the start of this essay — take the solution of a solvable problem, then allow its constants to become functions of time — is not an astronomical trick. It is a general technique for differential equations, and recognising it as one explains both why it works and where it fails.

Given an equation whose homogeneous version has a known solution with arbitrary constants in it, the constants can be promoted to unknown functions and substituted back. The result is a new system of equations for those functions, which is no easier in general than the original problem — except that when the extra term is small, the new unknowns change slowly, and a slowly changing unknown is one that can be approximated, averaged, or expanded.

That is the whole of it, and the same construction appears under the name variation of parameters in any course on ordinary differential equations, where it is presented as a way of solving an inhomogeneous linear equation and looks like bookkeeping.

The orbital version is the same operation applied to a nonlinear problem whose solution happens to be known in closed form, which is rare. It is available for the two-body problem and for almost nothing else in mechanics, and that is why celestial mechanics developed a perturbation theory two centuries before most other fields had anything comparable.

The failure mode is also general. The method presumes the new unknowns change slowly, and the expansions used to compute them presume the terms decrease. When a denominator becomes small both presumptions fail together, and the failure is not gradual — the series does not slowly become less accurate, it stops meaning anything. That is the same statement as the small-divisor problem above, made without reference to orbits at all.

The theorem that only covers half the perturbations

The result that the semi-major axis has no secular change at first order is the most reassuring statement in this subject, and it is important to know what it does not cover.

It is a theorem about conservative perturbations — about a third body’s gravity, or a primary’s oblateness, both of which derive from a potential. Energy is exchanged between orbits and returned, and over a long enough interval nothing accumulates.

A dissipative perturbation is not covered and behaves in exactly the opposite way. Atmospheric drag removes energy and nothing gives it back, so the semi-major axis falls monotonically and the protection does not apply. The same is true of tidal dissipation, of the thermal recoil that moves small asteroids, and of gravitational radiation.

So the practical division is not between large perturbations and small ones. It is between those that have a potential and those that do not, and a very small dissipative term outlives an enormous conservative one because only the first accumulates in the quantity that decides whether an orbit survives.

There is a further class that is conservative and still not covered, because the theorem is a first-order result. General relativity contributes a term that produces a secular advance of the pericentre and no secular change in the semi-major axis, so it fits the pattern; the higher-order terms in the ordinary planetary problem do not vanish, and it is those that make the long-term stability of the solar system a question for numerical integration rather than for a proof.

What the picture cannot show

The osculating orbit is a fiction with a precise definition, and the fiction has a cost that the drawings above hide.

An osculating element is not a measurement of anything. Two people using different definitions of the osculating orbit — mean elements averaged over the short-period terms, say, against instantaneous ones — get different numbers for the same body at the same instant, differing by the amplitude of the ripple. Published elements for an artificial satellite are almost always mean elements with the short-period variation already removed, because that is the number that changes slowly enough to be worth transmitting. Feeding mean elements into a formula expecting osculating ones is a standard way of being wrong by a part in a thousand, which for an orbit determination is enormous.

The second thing the picture cannot show is the timescale on which the answer changes. These integrations run for tens of orbits. The secular terms that decide whether a configuration is stable act over millions, and the numerical integrations that address the real question — is the solar system stable — run for billions of years and give a probabilistic answer. Laskar’s integrations put the chance of Mercury’s eccentricity being driven above 0.6 within five billion years at around one per cent, with a small fraction of those cases ending in a collision or an ejection. That is a statement no perturbation series makes and no figure here contains.

One inclination per altitude, and all of them retrograde. The inclination at which a circular orbit's node keeps exact pace with the Sun, against altitude. Every point satisfies −(3/2)J₂(R/p)²n cos i = 0.9856°/day, checked at each drawn point rather than at the ends: 96.33° at 200 km rising to 102.51° at 1600. All of it is retrograde, because a prograde orbit's node moves the wrong way and no prograde inclination can be made to follow the Sun at any altitude. The curve rises because n and (R/p)² both fall with height, so a higher orbit needs more of the cosine to make up the same rate — and it runs out: past 5,974 km the fastest rate available, the one at 180° where cos i = −1, is already slower than the Sun's, and no inclination works at any price. A sun-synchronous orbit sits at 786 km and 98.5°; Landsat sits at 705 km and 98.2°.
Fig. 7 The one place a drifting element is a design requirement. The inclination at which a circular orbit’s node keeps exact pace with the Sun, against altitude: 96.33° at 200 km rising to 102.51° at 1,600, all of it retrograde because a prograde orbit’s node moves the wrong way. The perturbation is used rather than corrected for, and it runs out — past 5,974 km the fastest rate available is already slower than the Sun’s, and no inclination works at any price.

One more element completes the set the averaging treats differently.

The osculating semi-major axis of a perturbed orbit. The semi-major axis a test particle would have if the perturber vanished, computed from its position and velocity at every step of an integration over 26 orbits of the perturber. It is not constant: a short-period ripple rides on a slow trend, and only the trend accumulates.
Fig. 8 The osculating inclination of the same perturbed orbit. It oscillates about a mean that does not drift, unlike the semi-major axis and the eccentricity — which is the general pattern: the elements that carry the energy and the shape drift, and the ones that carry the orientation of the plane mostly do not.

Where the ladder goes next

The natural next rung is the resonant case treated on its own terms, where the small divisor is not a nuisance but the subject: a pair of bodies locked so that the perturbation acts always in the same sense. The rung after it is the secular problem stripped of the fast angles entirely — the Laplace–Lagrange theory, in which the eccentricities and inclinations of a planetary system become a coupled set of oscillators with their own eigenfrequencies, and the question of long-term stability becomes a question about eigenvalues. The third law’s own hidden correction term turns out to be measurable in exactly the systems where these perturbations are not.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 35 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Disturbing functionGreat inequalityNode regressionOblatenessOrbital elementsOsculating elementsPerturbationSecular variationSmall divisorsSun-synchronous orbit