An answer obtained along a path that was not taken
Assumes Two-body relaxation, Hyperbolic orbits and Perturbations.
Three gravitating bodies have no closed-form solution, and a star cluster has a million. Everything anybody knows about how such a system behaves therefore rests on approximations, and the most productive of them is also the most obviously circular: to work out how much a passing mass deflects a body, integrate the force along the path the body would have taken if it had not been deflected.
That should not work. It does, and this essay is about exactly how well and exactly where it fails.
The absence of a free constant is what makes this more than an order-of-magnitude estimate. The shape of the force curve is one over one plus the square of the scaled time, to the three halves, and that function’s integral over the whole line is exactly two.
The derivation, in one line
Put the perturber on a straight line passing the target at distance with relative speed . At time the separation is and the component of the attraction perpendicular to the path is the full attraction times over that separation. So
The parallel component integrates to zero by symmetry: whatever the perturber does to the body on the way in it undoes on the way out. That is the same cancellation that makes a flyby change a spacecraft’s direction without changing its speed in the planet’s frame, and it fails here for the same reason it fails there — the encounter is not really symmetric once the body has moved. That cancellation is exact for a straight path and is the first thing the approximation gives up when the deflection is large.
Three things about the result are worth pausing on. It does not contain the mass of the target, so a grain of dust and a star are deflected identically — which is the equivalence that makes a gravitational field a property of space rather than of what is in it. It falls only as the first power of the impact parameter, not the second — much more slowly than the force does — because a more distant encounter lasts proportionately longer. And it falls as the first power of the speed, so slow encounters do more.
How wrong it is
The circularity can be measured, because the two-body problem this approximates does have a closed-form solution. A body passing another on an unbound orbit follows a hyperbola, and the deflection angle of that hyperbola satisfies a tangent relation in a single dimensionless combination.
The one-directional error is the useful part. An estimate that is wrong by an unknown sign is a nuisance; one that is wrong by a known sign is a bound. Every number in stellar dynamics derived this way is an upper limit on the true deflection, with the excess computable from the same single ratio.
At the ninety-degree impact parameter the estimate is exactly twice the truth. That is not a result but a definition — the ninety-degree radius is where the deflection is a right angle, and a right-angle deflection changes the transverse velocity by exactly the incoming speed rather than by twice it.
Summed over a population, it produces a logarithm
The reason the approximation matters is that nobody wants a single encounter. What is wanted is what happens to a star crossing a cluster, which meets encounters at every impact parameter at once.
Square the kick, because the mean of the kick itself is zero and it is the mean square that accumulates. Multiply by the number of encounters at each impact parameter, which for a uniform field goes as the impact parameter itself. The product goes as the impact parameter to the minus one, which is a constant per logarithmic interval.
That last point is worth a sentence of its own. The divergence in the standard treatment is an artefact of using the approximation outside its range, and correcting the approximation removes it. The hand-placed inner cutoff is a convenience rather than a necessity, and the answer it gives is the right one to within a term that does not depend on anything.
The outer cutoff is a different matter and it does not cure itself. Beyond the size of the system there are no more perturbers, so the sum has to stop somewhere physical, and where it stops is a modelling choice.
The consequence: a time to forget
The sum of squared kicks is the whole content of the relaxation time — the interval after which the accumulated deflections have changed a star’s velocity by of order the velocity itself. There is a second consequence in the same arithmetic. If the perturber is much heavier than the field stars, the kicks it gives are not random with respect to its own motion — it leaves an overdense wake behind itself, and that wake pulls back.
The other failure mode, which is not strength
Everything so far has treated the target as free. Most targets are not: a star in a cluster is on an orbit, a binary has an internal period, a galaxy’s stars have their own frequencies. That changes the problem in a way the deflection formula cannot see.
If the encounter lasts longer than the target’s internal period, the target has time to respond. It moves during the encounter, follows the perturbation, and gives the energy back on the way out. The approximation is not merely inaccurate in that regime — it computes something the system does not keep.
Two computations of the same number, by routes with nothing in common, is the strongest form the site’s own habit takes. One is a quadrature in frequency; the other is an initial-value problem in time; Parseval’s theorem is the only thing that connects them, and a slip in either would show up as a disagreement rather than as a plausible curve.
Why the suppression matters more than it sounds
Adiabatic invariance is usually met as a curiosity about slowly changing pendulums. Here it does real work, because it means an encounter does not heat a system uniformly.
The outer parts of a cluster or a galaxy have long orbital periods, so for them almost any encounter is impulsive: they take the full kick and some of them leave. The core has short periods, so for the same encounter it is adiabatic: it barely notices. A passing perturber therefore strips the outside and leaves the inside intact, which is exactly what is observed of clusters near the Galactic centre and of dwarf galaxies on eccentric orbits.
What is actually measured
The impulse approximation is a piece of mathematics, so the question is not what measurement supports it but what measurements depend on it.
The clearest is the mass function of stars in globular clusters. A cluster loses its lightest stars preferentially, because equipartition drives them to larger orbits where the escape criterion is easiest to meet, and the rate at which that happens is a relaxation rate — computed from the sum above. Comparing an observed present-day mass function with the one the cluster must have been born with is therefore a measurement of how many relaxation times have elapsed, and it is the same bookkeeping that says a cluster boils itself away on a schedule its own density sets. The second is dynamical friction on satellite galaxies, where the predicted sinking time is compared with the observed distribution of satellites around galaxies like the Milky Way. That comparison has an awkward history: the predicted times are sensitive to the Coulomb logarithm’s outer cutoff, which is exactly the part of the calculation that does not cure itself.
The same device, three times over, in this collection
It is worth collecting the places where this one integral has already appeared under other names, because they do not look alike.
The first is the Tisserand parameter. A comet passing Jupiter has its orbit changed enormously and one combination of its elements not at all, and the reason is that in the frame turning with Jupiter the encounter conserves the Jacobi constant. That is a statement about an exact invariant rather than about an approximation — but the reason anybody can use it is that the encounter is short compared with the orbital period, so the before and after can be treated as two separate two-body problems with an instantaneous join. The second is the gravity assist, which is the impulse read as engineering. A spacecraft’s velocity relative to the planet is turned and not changed in magnitude; the change in its heliocentric velocity is the vector difference, and it is bounded by twice the planet’s speed. The third is the tidal shock, which is the differential version and belongs to the next rung. A cluster passing through a galactic disc feels a kick that varies across it, the mean cancels, and what survives is a heating proportional to the square of the gradient rather than of the force.
The generalisations below are about where else the same integral turns up, and what its constants are worth.
The same approximation, without gravity
The device is not gravitational and its name in this field is a local one. The identical calculation — integrate the perturbing force along the unperturbed trajectory, take the transverse impulse, square it, sum over impact parameters — is the Born approximation in scattering theory and the standard treatment of Coulomb collisions in a plasma, and the three were developed independently before anybody noticed they were one thing.
The correspondence is exact rather than analogical. An inverse-square attraction and an inverse-square repulsion differ by a sign that disappears when the kick is squared, so the transverse impulse for an electron passing an ion is the same expression with the gravitational constant times the masses replaced by the product of the charges over the permittivity. The relaxation time of a star cluster and the collision time of a plasma are the same integral with different labels.
The logarithm is where the shared history shows. It is called the Coulomb logarithm in stellar dynamics — a plasma physicist’s name, kept — because that is where the divergence was first confronted and where the cutoffs were first argued about. And the resolution differs between the two fields in a way that says something about both: a plasma has a physical outer cutoff, the Debye length, beyond which the charge is screened and there is genuinely no force. Gravity has no screening and therefore no such length, which is why the outer cutoff in a cluster is the size of the cluster and is a modelling choice rather than a physical scale.
An approximation that appears in three fields under three names is usually approximating something structural, and here what it is approximating is any interaction that falls as the inverse square and acts for a finite time.
What the cutoffs are chosen to be, in practice
The Coulomb logarithm is a logarithm of a ratio and it is therefore forgiving, which is the standard defence of not thinking about it very hard. The defence is weaker than it sounds.
The ratio is usually written as the largest impact parameter over the smallest, and the largest is taken to be the size of the system while the smallest is the ninety-degree radius. That gives values between about six and twelve for real systems — so a factor of two in either cutoff moves the answer by ten per cent, which is tolerable.
The trouble is that the cutoffs are not the same quantity in every application. For the relaxation of a cluster of equal masses, the outer scale is the half-mass radius and the inner is set by the individual stellar masses. For dynamical friction on a massive satellite, the inner scale is the satellite’s own size rather than its ninety-degree radius, because the mass is extended and an encounter closer than that does not see the whole of it. Those two prescriptions differ by orders of magnitude in the ratio, which is a factor of two or three in the logarithm and therefore in the sinking time.
The way the argument was settled was not analytical. Direct N-body simulations, which make no approximation of this kind at all, were run and the effective logarithm was fitted to reproduce them — so the constant in the closed form is now calibrated on the computation the closed form exists to avoid. That inversion is honest and it is worth naming: the analytic result supplies the scaling and the simulation supplies the coefficient, and neither alone is the answer.
Where the energy goes when the mean kick is zero
One feature of the sum deserves separating out, because it explains what relaxation does rather than how long it takes.
The mean kick is zero — encounters are as likely to speed a star up as to slow it down — and the mean square is not. So the effect of many encounters is to spread the velocity distribution rather than to shift it, which is diffusion in velocity space and is why the process is called relaxation rather than drag.
But the diffusion is not the whole of it. A star moving faster than its neighbours meets them at a higher relative speed, and the kick falls as the first power of that speed, so it is perturbed less; a slow star is perturbed more. The result is a systematic drift towards the mean, superimposed on the diffusion, and the two together drive the system to a Maxwellian.
For a population of mixed masses the drift has a further consequence. Equipartition pushes the system towards equal energies rather than equal speeds, so heavy stars slow down and sink while light ones speed up and rise — mass segregation, which is observed in every relaxed cluster and is the same calculation with the perturber and target masses kept distinct.
And equipartition cannot be reached. A self-gravitating system has negative heat capacity, so a core that loses energy to the halo gets hotter rather than colder and loses energy faster. The drift towards equipartition therefore runs away instead of converging, which is the core collapse that every relaxed cluster undergoes and that the impulse sum, read as a timescale, predicts the schedule of.
There is a last piece of bookkeeping implied by all of this, and it is the one that decides a cluster’s fate rather than its rate. Energy is conserved in the encounters themselves, so the heating of one part of a system is the cooling of another and the total is unchanged; what relaxation does is redistribute it, from the ordered energy of the orbits into the random energy of the velocities and from the centre outwards. A system that has finished relaxing has not lost anything. It has simply arranged what it had in the one configuration the encounters cannot change further, and that configuration has no equilibrium — which is why the process does not stop.
It is worth being explicit about what kind of answer the approximation delivers, because it is not a force. Averaged over a population of encounters the mean velocity change is zero — for every star that passes on one side there is one on the other, and the kicks cancel in direction. What does not cancel is the square: the kicks add in quadrature, so the quantity that accumulates is a variance in velocity rather than a velocity. That is why the output of this calculation is a diffusion coefficient, why the relaxation time is defined as the interval over which the accumulated variance equals the square of the orbital speed, and why the same machinery describes a random walk rather than a trajectory. The distinction matters when the approximation is pushed: a first-order calculation whose first moment vanishes is being used entirely for its second, and the conditions under which the second is reliable are not the conditions under which the first would have been.
The suppression has a wide end as well as a narrow one, and it is worth drawing over a range that reaches both.
That width is what makes the correction awkward to apply. A cluster of stars does not have one internal frequency; it has a range spanning the same three or four decades, from the orbital period of a binary at its centre to the crossing time of its outermost members. So a single encounter is impulsive for part of the system and adiabatic for another part, and the energy it deposits has to be integrated over the system’s own distribution of frequencies rather than looked up from a single ratio.
The practical consequence is that the impulse approximation is used with a cutoff rather than with a correction. Everything faster than the encounter is counted at full strength, everything slower is counted at zero, and the error made by that step function is small because the systems in the transition region hold little of the mass. It is a crude thing to do and it works for a reason that has nothing to do with the approximation itself: the mass distribution is steep enough that the boundary is not where the mass is.
That is a pattern worth naming, because it recurs. A calculation that is exact at both ends and wrong in the middle is usable whenever the middle is empty, and much of what is done with the impulse approximation depends on that being true rather than on the approximation being good. Checking that emptiness is the step most often skipped, and it is the step that decides whether the answer is right. A system whose mass sits in the transition region — a cluster of binaries with periods comparable to the encounter time — is one the impulse approximation gets wrong by a factor, and no amount of care with the impulse itself will recover it.
Where the ladder goes
The obvious next rung is the wake and the drag it produces, treated as a force rather than as a heating — where the same integral acquires a velocity dependence that is not monotonic, and the drag turns out to be strongest not when a body is moving fastest but when it has slowed to about the speed of everything around it.
The other direction is what happens when the target is a bound pair rather than a single star. Then the kick is differential across the pair, the leading term cancels, and the energy delivered falls as the fourth power of the impact parameter rather than the second — which changes the whole balance between distant and close encounters, and is why a binary in a cluster is a heat source rather than a victim.
What this makes readable
Essays that name this one as a prerequisite.
- A drag that is strongest in the middle gravitation
- A hole that says mass times time galaxies
- The part of a kick that heats nothing gravitation
- Whether it merges is a ratio of two times galaxies
About the same objects
Not linked from either essay — found by the objects both name.
- A bridge and a tail drawn by one force dynamical friction · impulse approximation
- A drag computed with a logarithm nobody can pin down coulomb logarithm · dynamical friction
- The planet pays, and it shows close encounter · hyperbolic orbit
What links here
Essays that link to this one from their own argument.
- A drag that is strongest in the middle gravitation
- Whether it merges is a ratio of two times galaxies
- The part of a kick that heats nothing gravitation
- A hole that says mass times time galaxies
- Where the chaos comes from gravitation
- A stream is not the orbit it came from galaxies
The objects this essay names
Each one links to every other essay that touches it.
Adiabatic invariantClose encounterCoulomb logarithmCross-sectionDynamical frictionHyperbolic orbitImpulseImpulse approximationPerturbationRelaxation timeTwo-body encounter