A family whose size is a choice
Assumes Asteroid families, Secular theory and Non-gravitational forces.
The first rung of this anchor got a date out of a family: the fragments’ semi-major axes spread apart at a rate that depends on their size, so the cloud is a V whose slope is an elapsed time. That argument takes the membership list as given.
There is no observation that produces one. An asteroid does not carry a label saying which parent body it came from, and the belt is not empty between the families. What exists is a list of orbits, and the assignment of some of them to a family is the output of a clustering algorithm with a free parameter in it.
The metric is a velocity, and that is not a convention
The clustering is done in a distance defined between two orbits:
with the orbital speed. The three coefficients are not chosen. They come from Gauss’s equations for how a velocity impulse applied at an unspecified point on an orbit changes each element, averaged over the orbit — so is, to first order, the speed with which one fragment would have to have left the other in order to be where it is.
That is a strong thing to have. It means the cutoff is a physical quantity with an expected value: a catastrophic collision between bodies of a few tens of kilometres produces fragments with ejection speeds of a few tens to a hundred metres a second, comparable to the escape velocity from the parent body. So a cutoff of forty or eighty metres a second is not an arbitrary knob; it is a statement about the collision.
It is also, in a way the metric conceals, wrong for most of the families it is applied to.
The history is worth a paragraph because the method is a century old and its difficulty has been the same throughout. Kiyotsugu Hirayama noticed in 1918 that the asteroids clump in the elements that stay put rather than in the ones that oscillate, and identified three families by eye from a list of a few hundred orbits. Eye and hand remained the method for sixty years. Automatic clustering arrived in the 1980s with the sample in the thousands, and the plateau criterion arrived with it — not as a theory but as an observation that the count of a real clump is flat over a range of cutoffs and the count of a chance grouping is not. That criterion is still what every published list rests on, and it is a heuristic.
Why semi-major axis is not in that figure
The transverse distance in the last figure leaves out the semi-major axis, and the reason is the whole complication.
A fragment’s semi-major axis does not stay where the collision put it. The Yarkovsky effect — a thermal recoil from the asymmetric re-radiation of absorbed sunlight — pushes it inward or outward according to its spin direction, at a rate that goes inversely with its size, for as long as it exists. Over a hundred million years a kilometre-sized fragment drifts by a few hundredths of an au, which in the metric above is several hundred metres a second.
So an old family is a cigar in the metric: a few tens of metres a second across in eccentricity and inclination, and several hundred long in semi-major axis. Its velocity distance from its own centre is mostly a distance the fragments never travelled at, and any statistic that treats the three coordinates alike is measuring the drift and calling it a dispersion.
The practical response has been to cluster in the two transverse coordinates first and then take a slice in semi-major axis, or to cluster in a space that includes the absolute magnitude so that the V-shape’s own geometry is built into the metric. Both amount to assuming part of the answer, and both are better than not doing it.
There is one further consequence of the cigar shape and it is about the algorithm rather than the physics. Single linkage grows a cluster by absorbing anything within the cutoff of anything already absorbed, so it follows a chain of stepping stones and is perfectly happy with a long thin structure. That is why it finds drift-widened families at all — a method requiring compactness would reject them. The same property is what makes it percolate: a chain is a chain whether the stepping stones are fragments or background. The algorithm’s ability to find the families and its tendency to swallow the belt are the same property.
What a deeper catalogue does
The percolation cutoff is not a property of the family. It is a property of the family and the catalogue.
There is a related trap in comparing family sizes between surveys. A statement that family A has three times the members of family B is a statement about two clustering runs, and if the two families sit at different distances from the Sun then the survey’s completeness differs between them — a given absolute magnitude is a fainter apparent magnitude further out — so part of the ratio is the survey and not the collision. Correcting for that requires knowing the size distribution being cut into, which is one of the things the family is being used to measure.
Reality is worse than the figure, because the density of the belt is not uniform. A family sitting next to another family has a background that is itself clustered; a family near a resonance has a background with a hole in it. The published lists carry these as caveats and the caveats are not quantitative.
It is worth saying what would settle the matter and why it is not available. A membership test that did not depend on proximity at all would break the circularity — something intrinsic to the object rather than to its orbit. Two such tests exist in principle. One is composition, which the next rung is about. The other is a backward integration of the orbits to a common epoch, which works beautifully and only for families young enough that the integration stays deterministic. That limit is the expiry date every prediction in a chaotic system has: the belt’s Lyapunov times are a few hundred thousand to a few million years, so a backward integration is reliable for a handful of million and useless beyond — against the hundreds of millions the interesting families have lived.
What proper elements cost
Underneath everything above is a coordinate system that has to be computed rather than observed.
The proper elements are also not one thing. There are analytic ones, computed from a truncated perturbation series, and synthetic ones, computed by integrating the orbit for a few million years and Fourier-analysing the result. The two agree well where the theory converges and disagree where it does not, which is exactly the regions the theory has trouble with — high inclination, high eccentricity, and near resonances. The published family lists based on the two do not contain the same objects.
Two things follow. The proper elements have their own errors, from the truncation of the perturbation theory and from the object’s own orbit being imperfectly known, and those errors are of order a metre or two per second — small compared with the cutoff, but not negligible for the tightest families. And more seriously, proper elements fail in and near mean-motion resonances, where there is no quasi-integral to compute. A family straddling a resonance is cut in two by the coordinate system rather than by anything physical, and the part inside the resonance leaks away on a timescale of millions of years, which is a real loss and looks like the same thing. A resonance clears a gap whether or not anybody is trying to count what used to be in it.
Why any of it matters outside the belt
It would be possible to read all of this as a specialist’s difficulty with a specialist’s catalogue. It is not, because the families are the input to two arguments that reach much further.
The first is the impact history of the inner solar system. The families are the reservoir from which the near-Earth objects are supplied — fragments drift in semi-major axis until they meet a resonance, the resonance pumps their eccentricity, and they leave the belt. So the size and age distribution of the families is what sets the cratering rate, and a surface dated by counting holes in it is dated against a flux whose supply is estimated this way. A family membership uncertain by a factor of two is a supply rate uncertain by the same factor.
The second is the collisional history of the belt itself. The number of families of each size, against the size of the parent bodies they came from, is a record of how often bodies of each size have been shattered — which is the only observational constraint on the strength of asteroids at kilometre scales that does not come from a laboratory experiment on a rock the size of a fist. Where the mass is and where the light is in a collisional cascade depends on that strength through an exponent, and the exponent is fitted to the families.
What is actually measured
The chain from observation to family is long enough to set out.
An asteroid is discovered as a moving point in a survey image. Enough detections give an orbit, which is a set of osculating elements at an epoch. A numerical integration or an analytic theory then converts those into proper elements, with an accuracy that depends on how long the orbit has been observed and on whether the object is near a resonance. Those proper elements go into the metric above. The clustering runs. A cutoff is chosen by looking for a plateau in a curve of the kind in the first figure. And the resulting list is then used to infer the collision’s age, its energy, and the parent body’s size.
None of the intermediate quantities is checkable against anything else. There is no independent census of how many fragments a collision of a given energy produces, no second method for the age of an old family, and no way to visit one. The internal consistency of the picture is its whole support: the ages come out ordered sensibly against the families’ sizes, the size distributions look like collisional cascades, and the drift rates implied are the ones the thermal recoil predicts from a body’s spin and thermal inertia. That is a good deal of consistency and it is not a measurement of any one thing.
Every quantity in that last sentence depends on the membership list, and the list depends on the cutoff. The age in particular is read off the family’s outer edge in semi-major axis, which is exactly where membership is least certain.
What the numbers actually are
Setting the synthetic figures beside the real catalogue is worth doing, because the orders of magnitude are the thing.
The main belt catalogue holds over a million objects with orbits good enough for proper elements. Recent family analyses identify around a hundred and twenty families, containing between them perhaps a third of all belt asteroids — a third of the belt is collisional debris of identifiable parentage, which is a remarkable statement and one that depends entirely on where the cutoffs were put. The largest families have tens of thousands of members; the smallest that anyone will assert have a few dozen.
The cutoffs actually used run from about twenty metres a second for the tightest families to a hundred and twenty for the most diffuse, chosen family by family rather than globally, on the plateau criterion. Estimated interloper fractions, where they have been estimated at all, run from a few per cent to about thirty. And the number of asserted families has roughly doubled with each major extension of the catalogue, which is the pattern the fourth figure predicts and does not distinguish from genuine discovery.
The generalisation
The structure worth extracting is that single linkage percolates, and that percolation is a threshold rather than a gradual degradation.
Connecting every pair closer than a distance and asking what is connected to what is a classical problem, and its answer is that below a critical density-dependent distance the clusters are local and above it there is one cluster containing almost everything. Between those regimes there is no comfortable middle: the transition is sharp, and it gets sharper as the sample grows. Any method built on single linkage therefore has a parameter it cannot be indifferent to.
There is a second structural lesson, about what it means for a parameter to be chosen by looking at the data. The plateau criterion is defensible — a real clump does produce a flat stretch — but the flatness is not a property of the family alone; it is a property of the family against the local background, and the same family in a different neighbourhood plateaus somewhere else. Choosing a parameter from a feature of the data means the parameter inherits everything that made the feature, and there is no separating the two afterwards. That is the same difficulty an eccentricity that belongs to the whole system rather than to one planet presents in a different guise: a quantity attributed to an object turns out to be a property of the object and its surroundings jointly.
The same shape appears wherever a structure is defined by proximity in a space with a background in it. A galaxy cluster identified by a friends-of-friends algorithm has the same problem with the same cure and the same residual doubt; so does a stellar association identified in position and proper motion. The question to ask of any such catalogue is not how many objects it contains but what happens to that number when the linking length is changed by twenty per cent.
There is a third thing to take from it, which is about how a field lives with a difficulty it cannot remove. Nobody in this subject believes a membership list is exact, and nobody has stopped using them. What has happened instead is that the quantities computed from them have been chosen for robustness: the centre of a family in proper elements is very stable against the cutoff, and so is the family’s existence; the age and the total mass are not. Reading a paper in this field means noticing which of the two kinds of quantity is being quoted, and the good ones say so.
The corollary is about what these lists are for. A family membership list is not a measurement; it is a hypothesis, and it is a good one when the quantity computed from it is insensitive to the cutoff. The age from a V-shape is not, which is why it is the number this essay ends on.
Where the ladder goes next
The next rung takes the interloper problem head on, using the one piece of information the clustering ignores: colour. Fragments of one parent body have the same composition and therefore the same reflectance spectrum, so a broadband colour taken from the same survey that found the object is an independent membership test — and applying it to families identified by clustering removes between five and thirty per cent of the members, depending on the family.
Further rungs on this anchor: the families that are found only in a space including the absolute magnitude, which are old and diffuse and were invisible to the classical method; the halo of a family, which is real and extends far beyond any cutoff; the parent body’s size reconstructed from the fragment size distribution, and the fact that the reconstruction usually accounts for less mass than the family should contain; and the young families, of which a handful are recent enough for the fragments’ orbits to be integrated backwards to a common point in time — the only family ages that are measurements rather than inferences.
What links here
Essays that link to this one from their own argument.
The objects this essay names
Each one links to every other essay that touches it.
Asteroid familyCollisional familyEjection velocityHierarchical clusteringInterloperMain beltPercolationProper elementsSingle linkageYarkovsky effect