Gravitation

One number where two masses were

The hundreds of orbits an inspiralling binary completes inside a detector's band depend on its two masses only through one combination of them. Every pair on that contour radiates an identical sweep, so the early signal — which carries nearly all the signal-to-noise — cannot say which pair it was.

Assumes Gravitational waves, The two-body problem and Relativistic orbits.

The frequency of a gravitational wave from an inspiralling binary climbs, because the orbit shrinks, because energy is leaving. The rate at which it climbs is what is measured, and the standard result for that rate depends on the two component masses only through the combination

M  =  (m1m2)3/5(m1+m2)1/5.\mathcal{M} \;=\; \frac{(m_1 m_2)^{3/5}}{(m_1+m_2)^{1/5}} .

This is the chirp mass, and the fact worth sitting with is not that it appears but that at leading order it appears alone. The two masses enter nowhere else. So two binaries with the same chirp mass and completely different components produce, over the hundreds of cycles that carry most of the detectable signal, the same waveform.

This is not a limitation of the instruments. Two detectors of infinite sensitivity, observing the early inspiral for an arbitrarily long time, would return the same answer: one number, exactly determined, and no information whatever about how it was divided between the two bodies. The degeneracy is in the physics being observed.

One measured number, and every pair of masses that produces it. The plane of the two component masses of an inspiralling binary, with three curves of constant chirp mass across it. The middle one is GW150914's value of 28.7 solar masses, and the chirp mass recovered from the coordinates of the drawn curve varies along its whole length by 7.4e-14 per cent — which is the point: every binary on that line radiates the same frequency sweep at leading order, so the early inspiral cannot tell them apart. Two of them are marked. An equal pair of 33.0 and 33.0 solar masses and a lopsided pair of 63.6 and 18.4 sit on the same contour, and their total masses differ by a factor of 1.24. What separates them is the mass ratio, which enters the phasing only at the first post-Newtonian order, suppressed by the square of the orbital speed in units of the speed of light — small through the hundreds of cycles that carry most of the signal, and appreciable only in the last few, where that speed approaches a third of c. So the chirp mass is a measurement and the individual masses are an inference from the end of the signal, which is exactly the part a detector's high-frequency noise eats first.
Fig. 1 The plane of the two component masses, with three contours of constant chirp mass across it. The middle one is the first detected event’s value. The chirp mass recovered from the coordinates of the drawn curve is constant along its whole length to a part in ten to the eleventh, which is the point: every system on that line radiates the same frequency sweep at leading order. Two of them are marked, an equal pair and a lopsided one, whose total masses differ by a quarter.

Why one combination and not two

The two-body problem has a well-known simplification: it can be replaced by a single body of the reduced mass moving in a fixed potential. Two masses go in and two combinations come out — the total mass MM, which sets the orbital frequency at a given separation, and the reduced mass μ\mu, which sets the inertia of the relative motion. The radiated power depends on the second time derivative of the quadrupole moment, which brings in μ\mu; the orbital frequency at a given separation brings in MM; and the rate of frequency change combines them. When the algebra is done, the two appear in exactly one product, μ3/5M2/5\mu^{3/5}M^{2/5}, which is M\mathcal{M}.

Nothing about that is a coincidence of the algebra. It is a statement that at leading order the inspiral has one dimensionful scale, and the scale is set by one number. The mass ratio is a dimensionless quantity, and dimensionless quantities enter only through corrections.

The same structure appears everywhere a system is observed only through the leading term of an expansion. Kepler’s third law gives the sum of two masses and not either one, for exactly this reason: the leading-order orbit is governed by MM alone, and separating the components requires either a second observable — the two velocity amplitudes, or the wobble of the primary about the barycentre — or a correction term.

Two bodies at a mass ratio of 3 to 1. Both bodies orbit their common centre of mass, on similar ellipses whose sizes are in inverse proportion to the masses — here 3 to 1, so the heavier body's path is 3 times smaller.
Fig. 2 The second observable, in the Newtonian case. Both bodies orbit their common centre of mass on similar ellipses whose sizes are in inverse proportion to their masses, so resolving the two orbits gives the ratio directly. Gravitational-wave detection has no equivalent: nothing is resolved, and the only thing observed is the radiation, which depends on the pair as a unit.
Two parallel lines of slope 11/3, and the gap between them is a mass ratio. The rate at which the frequency rises, against the frequency, for GW150914 and GW170817. In these coordinates the quadrupole sweep is a straight line of slope exactly 11/3, and the only property of the binary that moves it is the chirp mass, which shifts it vertically by 5/3 of a decade for every decade of mass. So the two lines here are parallel and 2.31 decades apart, and that gap is the 24.2-fold difference in chirp mass between 28.72 and 1.18 solar masses — nothing else about either system enters. Two black holes of 36 and 31 solar masses and a pair of equal masses adding to the same chirp mass would draw the same line, which is why the individual masses are always quoted with error bars several times wider than the chirp mass's. The dots mark where each signal entered the detector band, and the horizontal run to the right of each is the whole observation: GW150914, 183 milliseconds; GW170817, 102 seconds.
Fig. 3 The relation in its measured form. The frequency and its rate of change trace a line whose slope is fixed and whose offset is the chirp mass, so a single event’s track on this plane locates it. Two parallel lines mean two chirp masses, and the gap between them is what the measurement returns. Nothing here says anything about the ratio.

What the degeneracy costs

The consequence is visible in every catalogue of detections. Chirp masses are quoted to a per cent or better; individual masses are quoted with uncertainties of tens of per cent, and with the two components strongly anti-correlated — a heavier primary is always paired with a lighter secondary in the posterior, sliding along exactly the contour drawn at the top of this essay.

It is worth naming what that anti-correlation means for how results should be read. A quoted primary mass of, say, 36 solar masses with an uncertainty of 5 is not an independent statement; the secondary’s quoted value is not independent either; and the two cannot be sampled separately to build a population model. Anyone using a catalogue has to work with the joint posterior, and a table of marginal values discards most of what was measured.

For a binary neutron star the effect is extreme. The chirp mass of the 2017 event is known to four decimal places; the two individual masses are quoted as ranges that overlap substantially, and whether the system was two 1.4-solar-mass stars or a 1.6 paired with a 1.2 is not settled by the gravitational-wave data alone.

One measured number, and every pair of masses that produces it. The plane of the two component masses of an inspiralling binary, with three curves of constant chirp mass across it. The middle one is GW170817's value of 1.2 solar masses, and the chirp mass recovered from the coordinates of the drawn curve varies along its whole length by 1.7e-13 per cent — which is the point: every binary on that line radiates the same frequency sweep at leading order, so the early inspiral cannot tell them apart. Two of them are marked. An equal pair of 1.4 and 1.4 solar masses and a lopsided pair of 3.0 and 0.7 sit on the same contour, and their total masses differ by a factor of 1.35. What separates them is the mass ratio, which enters the phasing only at the first post-Newtonian order, suppressed by the square of the orbital speed in units of the speed of light — small through the hundreds of cycles that carry most of the signal, and appreciable only in the last few, where that speed approaches a third of c. So the chirp mass is a measurement and the individual masses are an inference from the end of the signal, which is exactly the part a detector's high-frequency noise eats first.
Fig. 4 The same construction for a neutron-star merger. The contour is far shorter, because both components are constrained by other physics to lie in a narrow band — no neutron star below about a solar mass forms, and none above about two and a half is expected to be stable — so astrophysical priors do most of the work that the waveform cannot. The chirp mass is a measurement and the individual masses are a measurement plus an assumption.

What breaks it, and how late

The degeneracy is exact only at leading order, and the post-Newtonian expansion supplies the correction.

Writing the phase evolution as a series in v/cv/c, the leading term carries M\mathcal{M} alone; the next term, suppressed by (v/c)2(v/c)^2, is where the symmetric mass ratio first appears. That is the whole mechanism, and its timing is the difficulty: early in the inspiral v/cv/c is small, so the correction is negligible, and only in the last few cycles — where v/cv/c approaches a third — does it become appreciable.

It is the same pattern as Mercury’s extra arcseconds, where the leading Newtonian orbit is degenerate with respect to a whole class of theories and the discriminating information sits in a correction of order (v/c)2(v/c)^2. In both cases the interesting physics is in a term that is small precisely because the system is well described by the term in front of it.

So the individual masses are measured from the end of the signal and the chirp mass from the beginning, and the two parts of the waveform are wildly unequal in duration. A stellar-mass black hole binary spends hundreds of cycles in the sensitive band accumulating chirp mass precision, and a handful of cycles near merger accumulating everything else.

GW150914: 33 Hz to 250 Hz in 0.22 seconds. The strain of GW150914 — two black holes — through the last 0.22 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 28.716 solar masses, and drawn at the luminosity distance it returned, 440 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 33 Hz to 250 Hz, and the envelope — the outer curve — grows by a factor of 3.9, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²¹, so the peak here is a fractional length change of about 2.9·10⁻²¹: over the four kilometres of an interferometer arm that is 1.2·10⁻¹⁷ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 5 The signal that arrives, with the frequency climbing through the band. Nearly all the signal-to-noise for a light system is accumulated in the many low-amplitude early cycles rather than in the few loud late ones, because matched filtering accumulates coherently and a long signal has more to accumulate. The information about the mass ratio is concentrated in exactly the part that contributes least to detection.
GW170817: 238 Hz to 1000 Hz in 0.22 seconds. The strain of GW170817 — two neutron stars — through the last 0.22 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 1.185 solar masses, and drawn at the luminosity distance it returned, 40 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 238 Hz to 1000 Hz, and the envelope — the outer curve — grows by a factor of 2.6, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²², so the peak here is a fractional length change of about 3.9·10⁻²²: over the four kilometres of an interferometer arm that is 1.6·10⁻¹⁸ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 6 The contrast, for a much lighter system. The neutron-star merger swept through the band for a hundred seconds and thousands of cycles, against a fifth of a second for the black-hole event — because a lighter binary reaches the same frequency at a wider separation and takes longer to get there. Its chirp mass is correspondingly far better determined, and its individual masses are not, because the last cycles happen at frequencies where the detector is deaf.

A detector’s sensitivity curve therefore does more than set a detection threshold: it decides which parameters are measurable. A band that reaches to higher frequency at the same amplitude buys almost nothing in range and a great deal in mass-ratio precision, because it catches the cycles where the correction lives. That is a different design target from raw sensitivity, and it is why proposed upgrades are argued about in terms of what they measure rather than how far they see.

There is a second effect that breaks the degeneracy, and it is geometric rather than perturbative. If the component spins are misaligned with the orbital angular momentum, the orbital plane precesses, which modulates the observed amplitude and phase. The modulation depends on the mass ratio directly, so a precessing system is much better constrained than a non-precessing one — and the events with the tightest individual masses are the ones that happened to precess.

How the number is actually extracted

The measurement is not a fit to a curve drawn through data points. A gravitational-wave signal is far below the noise in any single sample, and it is recovered by matched filtering: correlating the data against a bank of predicted waveforms and looking for the template that maximises the correlation.

That procedure has a consequence for what is measured. The signal-to-noise accumulates as the square root of the number of cycles the template stays in phase with the data, so the parameter the analysis is most sensitive to is the one that most strongly controls the phase. A fractional error in chirp mass of a part in a thousand, over a thousand cycles, dephases the template by a full cycle and destroys the correlation. A comparable error in the mass ratio does almost nothing until the last few cycles.

The hierarchy of what is well measured is therefore a hierarchy of what the phase depends on, and it is steep: chirp mass, then the effective aligned spin, then the mass ratio, then everything else. The sky position and the distance are measured by a different route entirely — by comparing arrival times and amplitudes between separated detectors — and are correspondingly much cruder.

The template bank has to cover the parameter space finely enough that some template is always within a fraction of a cycle of the truth, which is why the banks contain hundreds of thousands of waveforms and why the computational cost of a search is dominated by the dimensions the phase is most sensitive to.

Why a badly measured mass is still worth having

It would be easy to read the degeneracy as a defect. It is a limitation, and the quantity that survives it is more valuable than either mass separately for the two things gravitational-wave astronomy is mainly used for.

One amplitude, a factor of 2√2 in distance, and the whole error budget. The luminosity distance a fixed strain amplitude implies, against the inclination of the orbit to the line of sight, normalised so that a face-on source sits at GW170817's fitted 40 megaparsecs. A binary seen face-on radiates most strongly towards the observer and an edge-on one least, in the ratio 2√2 — so the same measured amplitude is consistent with everything between 14 and 40 megaparsecs, and the amplitude contains nothing that could choose between them. This is the sense in which a standard siren is absolute but not precise: the chirp mass comes from the sweep and is known to four figures, the distance comes from the amplitude and is known to tens of per cent, and the reason is one angle. Everything that has ever narrowed it has come from outside the waveform — the two detectors' relative amplitudes, the polarisation, and for GW170817 the superluminal motion of its radio afterglow, which fixed the viewing angle independently. The right-hand scale is what the ambiguity does to the answer: at NGC 4993's recession speed of 3017 km s⁻¹, the same signal gives a Hubble constant anywhere from 75 to 213 km s⁻¹ Mpc⁻¹.
Fig. 7 The first use. The amplitude of the wave depends on the chirp mass and the distance; the frequency evolution gives the chirp mass; so the distance follows, with no rungs beneath it. A distance with no ladder under it is what makes a merger a standard siren, and the degeneracy that ruins the individual masses does not touch it, because the same combination appears in both places and cancels.

The second use is population statistics. The distribution of chirp masses across a catalogue of detections is a measurement of how binaries form, and because chirp mass is the best-measured parameter it is the one where structure in the population shows up first. The apparent absence of systems with chirp masses in a particular range — the pair-instability gap — is a statement about stellar evolution read entirely off a quantity that is not a mass of anything. It exists because the most massive stars die in a particular way that leaves no remnant at all above a threshold, so the gap in the observed distribution is a prediction of stellar physics being checked with an instrument that measures spacetime.

The nanohertz version of the same problem

Pulsar timing arrays observe a different band and a different population, and the degeneracy takes a different form there.

Nothing visible in any pulsar, and a quadrupole in the angle between them. Above: 4 millisecond pulsars' timing residuals over 15 years, at the few hundred nanoseconds a good one reaches. Each wanders, and none of them shows anything a reader could call a signal; a gravitational-wave background of amplitude 2.4·10⁻¹⁵ at one cycle per year contributes a common part to all of them that is smaller than each pulsar's own red noise. Below: the correlation between pairs, against the angle on the sky between them. 2211 pairs out of 67 pulsars, binned into 15 angles, against three curves with no free parameters between them. A quadrupolar background gives the Hellings–Downs shape — positive for nearby pulsars, negative near 83°, and back up to exactly half its zero-separation value at 180° because a background looks the same in opposite directions. An error in the observatory clock would give a flat line, because it shifts every pulsar identically. An error in the solar-system ephemeris would give a cosine, because it moves the barycentre in one direction. The drawn points prefer the quadrupole over the flat line by Δχ² = 358. That is the detection: not a waveform, not an event, not a moment — a shape in an angle, accumulated over fifteen years, on data taken for another purpose entirely.
Fig. 8 What a timing array detects: not an individual binary but the superposed background of many supermassive ones, showing up as a correlation between pulsars whose amplitude depends on the angle between them. No individual system is resolved, so no individual chirp mass is measured. What is measured is an integral over the whole population — the number of mergers weighted by chirp mass to the five-thirds — which is one number standing in for a distribution.

The structure is the same one level up. In the band where individual events are resolved, one number stands in for two masses; in the band where they are not, one number stands in for an entire population. In both cases the quantity that survives is the one the radiated power depends on, and inverting it requires an assumption about what is not measured.

What a single number is worth against a distribution

There is one more consequence of measuring M\mathcal{M} rather than m1m_1 and m2m_2, and it is about inference rather than about waveforms.

A population model predicts a joint distribution over the two masses. An observation constrains a curve through that plane. Fitting the model to the catalogue is therefore an exercise in projecting a predicted two-dimensional distribution onto a family of one-dimensional constraints — which works, and which requires many more events than it would if each observation returned a point.

The practical statement is that the number of detections needed to measure a feature of the mass distribution depends on how nearly that feature aligns with the contours. A gap at fixed chirp mass shows up almost immediately. A gap at fixed mass ratio, or a change in the ratio distribution with total mass, requires an order of magnitude more events, because each observation is nearly parallel to the feature being looked for.

That is a general property of degenerate measurements and it is worth extracting from this case: a degeneracy does not merely inflate error bars, it decides which questions the data can answer at all. An orbit determination whose error is nearly all along the track has the same character — the covariance, not the variance, is what says whether a given question is answerable.

The third number, which is also a combination

After the chirp mass, the next-best-determined parameter of a binary is not a mass at all. It is a particular weighted combination of the two spins, and it is degenerate in the same way for the same reason.

Each component may be spinning, and the spin’s component along the orbital angular momentum affects the rate of inspiral: a spin aligned with the orbit delays the merger slightly, an anti-aligned one hastens it. What enters the phase at leading order is the mass-weighted average of those aligned components, usually written as an effective spin.

So a system with a rapidly spinning heavy component and a slowly spinning light one is indistinguishable, in the early inspiral, from one with the average distributed differently. One number stands in for two again.

The astrophysical value of that number is out of proportion to its precision, because it discriminates between formation channels. Two stars that evolved together in a binary should have spins roughly aligned with their orbit, since they inherited their angular momentum from the same disc; two black holes that met by chance in a dense cluster should have spins pointing in random directions. So the distribution of effective spins across a catalogue is a measurement of what fraction of mergers came from each route.

The observed distribution is clustered near zero with a slight positive skew, which is consistent with a mixture and settles nothing. What would settle it is the perpendicular components, which produce precession and are measured far worse.

A quantity that discriminates between hypotheses is worth measuring badly, and the effective spin is the clearest case of that in the subject: nobody cares what any individual system’s value is, and everybody cares about the shape of the histogram.

The remnant, measured a different way

Everything above extracts parameters from the inspiral. There is a second measurement in the same signal, made after the merger, and it is independent of the first.

What is left after two black holes merge is one black hole, distorted, and it settles by radiating. The radiation is a superposition of damped sinusoids whose frequencies and decay times are fixed entirely by the remnant’s mass and spin — a theorem of black-hole physics, and a strong one: no other property of the object enters, because a black hole has no others.

So measuring one frequency and one decay time gives a mass and a spin. Measuring a second mode over-determines them, and the consistency of the two is a test of whether the object really is a black hole of the kind the theory describes.

That gives a check on the whole analysis which is worth more than either half. The inspiral’s parameters predict, through numerical relativity, what the remnant’s mass and spin should be; the ringdown measures them. Comparing the two is the standard consistency test applied to every loud event, and so far the agreement holds within the uncertainties.

The test is limited by how much ringdown there is. A heavy binary merges at a frequency inside the detector’s best band and leaves a long, loud ringdown; a light one merges above the band and leaves almost nothing. So the events that give the best individual masses from the inspiral are the ones that give the worst ringdown, and the reverse.

The signal has two halves that measure the same system by unrelated routes, and the fact that they are best in opposite regimes is the reason a catalogue spanning a wide mass range is worth more than a deep one in a narrow range.

Where the picture stops

Redshift is degenerate with mass. A gravitational wave from a distant source arrives redshifted, and a redshifted waveform from a light binary is indistinguishable from an unredshifted one from a heavier binary. Every mass in a catalogue is a detector-frame mass, and converting it to a source-frame mass requires the redshift — which requires either an electromagnetic counterpart, which almost never exists, or an assumed cosmology. The one event with a counterpart is the reason a standard siren has ever been checked against a ladder rather than merely proposed.

Spin is degenerate with mass ratio. The aligned components of the spins enter the phase at a similar order to the mass ratio, so a system with an unusual mass ratio and one with substantial aligned spin can produce similar waveforms. Breaking that requires either precession or the merger and ringdown, and for high-mass systems the ringdown is a better handle than the inspiral.

Eccentricity has been assumed away. Every waveform above is for a circular orbit, on the argument that gravitational radiation circularises an orbit long before merger. That is right for a binary that formed wide and inspiralled, and wrong for one assembled by a close encounter in a dense cluster, which can enter the band still eccentric. Searching for those requires a different template family, and a circular search misses them.

And the waveform models are approximations. The post-Newtonian expansion diverges near merger and numerical relativity is expensive, so what is used in practice is a hybrid, calibrated against simulations at a finite set of mass ratios and spins. The systematic error from waveform modelling is currently comparable with the statistical error for the loudest events, and will exceed it as detectors improve.

Where this ladder goes next

Later rungs on this anchor: matched filtering itself, and why a template bank has the shape it does; spin-induced precession as the degeneracy-breaker, and what a precessing event gives up; the ringdown and black-hole spectroscopy, where the remnant’s mass and spin are measured from damped oscillations rather than from an orbit; the mass distribution of the population and the gaps in it; and the dark-siren method, which recovers a Hubble constant from many events with no counterpart by using the galaxies in each error volume.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Chirp massDegeneracyDetector sensitivity curveInspiralLuminosity distanceMass ratioMatched filteringMerger ringdownPost-newtonian expansionReduced massSpin induced precessionStandard siren