Gravitation

A distance with no ladder under it

The frequency sweep of an inspiral fixes the chirp mass with no distance in it, and the amplitude then gives the luminosity distance directly, because one expression fixes both. That is a distance measured with nothing calibrated beneath it — and its error budget is one angle.

Assumes Relativistic orbits, Distance ladder and The two-body problem.

Every distance in this collection is measured with the last one. A parallax calibrates a period–luminosity relation, the relation calibrates a supernova, the supernova reaches the far universe, and each rung inherits every error below it plus one of its own. That is what a ladder is, and the whole apparatus exists because there is no way to look at a distant object and read its distance off.

There is now one exception, and it does not fit anywhere on the ladder because it does not rest on anything. A pair of neutron stars spiralling together radiates gravitational waves, and the signal that arrives carries two numbers that are enough on their own: the rate at which the frequency rises gives the mass, and the amplitude then gives the distance — absolutely, with nothing calibrated, because the same expression fixes both.

GW150914: 33 Hz to 250 Hz in 0.22 seconds. The strain of GW150914 — two black holes — through the last 0.22 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 28.716 solar masses, and drawn at the luminosity distance it returned, 440 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 33 Hz to 250 Hz, and the envelope — the outer curve — grows by a factor of 3.9, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²¹, so the peak here is a fractional length change of about 2.9·10⁻²¹: over the four kilometres of an interferometer arm that is 1.2·10⁻¹⁷ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 1 The signal, through the last fifth of a second before two black holes merged. Two things rise together and neither is free to rise on its own: the frequency roughly sevenfold and the amplitude roughly fourfold, because the amplitude goes as the two-thirds power of the frequency and nothing else in it changes over so short a span. The vertical axis is in units of ten to the minus twenty-one, so the peak is a fractional length change of about a part in 102110^{21} — over four kilometres of interferometer arm, a thousandth of the width of a proton.

The law taken as given

The radiation from an orbiting pair is not derived here. It is a consequence of general relativity, it is worked out in the field that owns that derivation, and this essay treats it the way an observational essay treats an equation: as a stated relation whose consequences can be measured against.

The relation is that a slowly inspiralling binary radiates at twice the orbital frequency, at a rate

dfdt=965π8/3(GMc3)5/3f11/3,\frac{df}{dt} = \frac{96}{5}\pi^{8/3}\left(\frac{G\mathcal{M}}{c^3}\right)^{5/3} f^{11/3},

and with an amplitude, for an optimally oriented source at luminosity distance DD,

h=4D(GMc2)5/3(πfc)2/3.h = \frac{4}{D}\left(\frac{G\mathcal{M}}{c^2}\right)^{5/3}\left(\frac{\pi f}{c}\right)^{2/3}.

Only one property of the binary appears in either: the chirp mass

M=(m1m2)3/5(m1+m2)1/5.\mathcal{M} = \frac{(m_1 m_2)^{3/5}}{(m_1+m_2)^{1/5}}.

Two black holes of thirty-six and twenty-nine solar masses and a pair of equal masses with the same chirp mass would sweep along the same curve at the same amplitude. Nothing in the inspiral separates them.

The first number has no distance in it

Look at what is in the sweep and what is not. The rate at which the frequency rises depends on the chirp mass and on the frequency, and on nothing else — not on the distance, not on the orientation, not on how much of the signal was received.

That is why the chirp mass is the best-determined parameter of any gravitational-wave event, usually to three or four significant figures, while the individual masses carry error bars ten times wider. It is read off the timing rather than the amplitude, and timing is the one thing an instrument can be made good at.

Two parallel lines of slope 11/3, and the gap between them is a mass ratio. The rate at which the frequency rises, against the frequency, for GW150914 and GW170817. In these coordinates the quadrupole sweep is a straight line of slope exactly 11/3, and the only property of the binary that moves it is the chirp mass, which shifts it vertically by 5/3 of a decade for every decade of mass. So the two lines here are parallel and 2.31 decades apart, and that gap is the 24.2-fold difference in chirp mass between 28.72 and 1.18 solar masses — nothing else about either system enters. Two black holes of 36 and 31 solar masses and a pair of equal masses adding to the same chirp mass would draw the same line, which is why the individual masses are always quoted with error bars several times wider than the chirp mass's. The dots mark where each signal entered the detector band, and the horizontal run to the right of each is the whole observation: GW150914, 183 milliseconds; GW170817, 102 seconds.
Fig. 2 The same law in the coordinates that make it a straight line. The rate of change of frequency against frequency, in logarithms, is a line of slope exactly 11/3 whose vertical position is the chirp mass — shifted by five thirds of a decade for every decade of mass. Two events forty times apart in mass are two parallel lines two and a third decades apart, and reading a mass off this figure is reading an offset. The dots mark where each signal came up out of the noise: one lasted a fifth of a second, the other a hundred seconds.

A practical note on how the sweep is actually measured, because it is not by watching a curve. The signal is far below the instrument’s noise at any instant; what recovers it is matched filtering — correlating the data against a bank of hundreds of thousands of precomputed waveforms and looking for a correlation that appears in two widely separated detectors within the light-travel time between them. The chirp mass is then the parameter of whichever template matched, which is why it comes out of the analysis with the precision of a frequency measurement.

How long a signal lasts, and what that decides

The time a binary spends between two frequencies follows from the same expression, integrated:

τ=5256(GMc3)5/3(πf)8/3.\tau = \frac{5}{256}\left(\frac{G\mathcal{M}}{c^3}\right)^{-5/3}(\pi f)^{-8/3}.

The M5/3\mathcal{M}^{-5/3} is what separates the two kinds of event completely. A pair of thirty-solar-mass black holes entering a detector’s band at 35 hertz has two tenths of a second left. A pair of neutron stars entering at 24 hertz has a hundred seconds.

That factor of five hundred decides what each event can be used for. A fifth of a second is a few dozen cycles, and everything about the system has to be extracted from them at once. A hundred seconds is tens of thousands of cycles, which is enough for the sky position to be refined while the signal is still arriving — which is how an alert reached optical telescopes in time for them to find a counterpart that faded within a week.

GW170817: 114 Hz to 1000 Hz in 1.6 seconds. The strain of GW170817 — two neutron stars — through the last 1.6 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 1.185 solar masses, and drawn at the luminosity distance it returned, 40 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 114 Hz to 1000 Hz, and the envelope — the outer curve — grows by a factor of 4.3, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²², so the peak here is a fractional length change of about 3.9·10⁻²²: over the four kilometres of an interferometer arm that is 1.6·10⁻¹⁸ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 3 The other kind of signal, drawn on the same terms. Two neutron stars of about one and a half solar masses each, at forty megaparsecs: the chirp mass is a twenty-fourth of the black hole pair’s, so the sweep is far slower and the last second and a half covers a range of frequency the heavier system crossed in a few hundredths. The amplitude is comparable despite the vastly smaller masses, because the source is eleven times nearer — which is the amplitude–distance trade the rest of this essay is about, seen from the other side.

The second number is the distance

Now the amplitude. It contains the chirp mass, which the sweep has already supplied, and the distance, which nothing else has. Rearranging, the distance falls out.

There is nothing to calibrate. The four in the numerator is a four. The relation between the mass and the emitted power is fixed by the theory rather than by a fitted zero point, so unlike every rung of the distance ladder there is no step at which somebody had to measure a nearby example to set a scale.

That is the whole meaning of the phrase standard siren, and it is a better phrase than standard candle: a candle is standard because a population of them was found to be similar, and a siren is standard because the physics says how loud it is.

What was actually observed

The argument became a measurement on 17 August 2017, when a hundred seconds of inspiral from two neutron stars arrived, followed 1.7 seconds later by a short burst of gamma rays and, over the following days, by an optical transient in NGC 4993.

That sequence matters for a reason that has nothing to do with the waves. A distance alone is not a Hubble constant. A Hubble constant needs a distance and a redshift, and a gravitational-wave signal carries no redshift: the source could be anywhere on a sky region tens of square degrees across, and there is nothing in the waveform to say which galaxy it was in.

The optical counterpart identified the host. The host had a spectroscopic redshift already. The pairing gave

H0=708+12 km s1Mpc1,H_0 = 70^{+12}_{-8}\ \text{km s}^{-1}\,\text{Mpc}^{-1},

from a single object, with an error bar of about fifteen per cent — which is not competitive, and which is arrived at by a route that has no rung in common with either of the two methods that disagree.

H₀: nine determinations in two families. Published determinations of H₀, each with its quoted one-sigma interval, sorted into two families — measured locally, calibrated by a ladder, against inferred from z ≈ 1100 through a model. The shaded band behind each family is that family's inverse-variance weighted mean: 72.66 ± 0.75 across 5 of them, against 67.40 ± 0.41 across 4. The difference is 5.26 ± 0.85 km/s/Mpc, which is 6.2 standard deviations, computed here from the quoted errors alone. That number is an upper bound on the significance rather than the significance: the determinations within each family share calibrations, samples and in two cases the same supernovae, so they are not independent, and a correlated pair combines to something wider than the formula used here gives. What the figure does establish is that the split is not one discrepant measurement against a consensus — it is two internally consistent groups, and the grouping is by method rather than by result.
Fig. 4 Why an uncompetitive measurement is still worth having. Two well-established determinations of the Hubble constant disagree by about five standard deviations, and each is internally consistent, so the disagreement is a systematic somewhere and not a statistical fluctuation. Neither party can adjudicate, because each is inside its own method. A siren’s error bar spans both — and a hundred of them would not.

The one thing the amplitude cannot supply

The optimistic version above assumed the binary is seen face-on. It usually is not, and the correction is the entire difficulty.

A binary radiates most strongly along its rotation axis and least strongly in its orbital plane, in the ratio 222\sqrt2 for a detector measuring both polarisations. So a weaker signal is consistent with a more distant face-on source or a nearer edge-on one, and the amplitude contains nothing that could choose between them.

One amplitude, a factor of 2√2 in distance, and the whole error budget. The luminosity distance a fixed strain amplitude implies, against the inclination of the orbit to the line of sight, normalised so that a face-on source sits at GW170817's fitted 40 megaparsecs. A binary seen face-on radiates most strongly towards the observer and an edge-on one least, in the ratio 2√2 — so the same measured amplitude is consistent with everything between 14 and 40 megaparsecs, and the amplitude contains nothing that could choose between them. This is the sense in which a standard siren is absolute but not precise: the chirp mass comes from the sweep and is known to four figures, the distance comes from the amplitude and is known to tens of per cent, and the reason is one angle. Everything that has ever narrowed it has come from outside the waveform — the two detectors' relative amplitudes, the polarisation, and for GW170817 the superluminal motion of its radio afterglow, which fixed the viewing angle independently. The right-hand scale is what the ambiguity does to the answer: at NGC 4993's recession speed of 3017 km s⁻¹, the same signal gives a Hubble constant anywhere from 75 to 213 km s⁻¹ Mpc⁻¹.
Fig. 5 The degeneracy, drawn. For one measured amplitude, the luminosity distance implied by each inclination, normalised so that a face-on source sits at the fitted 40 megaparsecs. The curve spans exactly 222\sqrt2, which is a factor of 2.83 in distance and therefore a factor of 2.83 in the Hubble constant read off it. Everything that has ever narrowed this has come from outside the waveform.

For this event, three things did narrow it. The two detectors’ relative amplitudes and their arrival-time difference constrain the sky position and some of the polarisation. The gamma-ray burst was seen at all, which requires a jet pointed somewhere near the line of sight. And the radio afterglow was subsequently watched to move across the sky at an apparent speed several times that of light — a superluminal motion which, read as relativistic projection, fixed the viewing angle at about twenty degrees independently of everything else. That last measurement cut the distance error roughly in half.

The masses that are not the chirp mass

Two black holes of thirty-six and twenty-nine solar masses and a pair of thirty-three and thirty-two have almost the same chirp mass. Separating them needs something the leading-order inspiral does not have.

What supplies it is the next order. The relativistic corrections to the sweep depend on the mass ratio and on the spins, and they grow as the orbit tightens — so the individual masses are measured from the last few cycles, where the expansion is worst behaved, rather than from the long clean inspiral. That is why the mass ratio is quoted with wide error bars and the chirp mass is not.

The one thing that does have to be calibrated

“Nothing to calibrate” is the phrase this method is sold on, and it is true of the astronomy and false of the instrument. A distance read off an amplitude is only as good as the amplitude, and an amplitude is a number an interferometer reports in metres of arm-length change. That number has an absolute scale, and the scale has to be established by something.

It is established by pushing on a mirror with light. A separate, calibrated laser is aimed at the test mass, and the radiation pressure it exerts moves the mirror by a computable amount — force is power over the speed of light, and the displacement follows from the mirror’s mass and the frequency of the modulation. Sweeping that modulation across the detection band and recording what the interferometer reports establishes the response function from displacement to output, at every frequency, in absolute units.

The chain from there to a distance is short and entirely physical: a laser power measured against a standard, the speed of light, a mirror mass measured on a balance, and the wavelength of the main laser. No astronomical object appears anywhere in it, which is exactly the claim being made — but it is a claim about metrology rather than an absence of calibration.

The achieved accuracy on the absolute amplitude scale is a few per cent, and it enters the distance linearly. For the single-event Hubble constant with its fifteen per cent error bar that is negligible. For a future measurement at the per cent level it is not, and improving it is a laboratory problem: better power standards, better characterisation of the mirror’s response at the frequencies where the calibration lines sit, and better modelling of the small amount of light that scatters back into the beam.

So the standard siren replaces a chain of astronomical calibrations with a chain of laboratory ones, and that is the whole of its advantage. A laboratory calibration can be repeated, transferred between institutions, and compared against a standard held somewhere else; a Cepheid zero point cannot.

The prediction that was already confirmed

None of this was a surprise, because the same expression had already been checked to four significant figures without a single wave being detected.

A binary pulsar radiates too, and radiating removes energy, and removing energy shrinks the orbit and shortens the period. That period change is measurable by counting pulses, which is the most precise measurement in astronomy — and the measured orbital decay of the first binary pulsar agrees with the prediction to about a part in a thousand.

Where the errors are, and where they are not

It is worth being explicit about which parts of the answer share errors with anything else, because that is the whole argument for the method.

The chirp mass shares nothing. It comes from timing in an instrument calibrated by a laser wavelength.

The distance shares nothing with the ladder. Its errors are the calibration of the detectors’ absolute response, the inclination degeneracy above, and the statistics of a signal near the noise. All three are instrumental or geometric, and none of them is a property of any astronomical object.

The redshift shares everything with everything. It is a spectroscopic measurement of a galaxy, corrected for that galaxy’s own motion through the local flow, and the correction is a model.

What the method cannot do

Four limits, and they are all about the sky rather than about the physics.

It needs a host. Without an identified galaxy there is no redshift and no Hubble constant. Of the compact-binary mergers detected so far, exactly one has had an unambiguous host, because black hole mergers emit no light. A statistical version exists — weight every galaxy in the localisation volume by its luminosity and marginalise — and it works, and it needs hundreds of events to reach the precision one good host gives.

The localisation is coarse. Two detectors give an annulus on the sky; three give a patch of tens of square degrees. Finding a transient in that patch before it fades is an observational campaign rather than a pointing.

The reach is short. Amplitude falls as one over distance, so the volume surveyed grows as the cube of the instrument’s sensitivity — which is an argument for building better detectors and not for waiting.

And the redshift is degenerate with the mass. A signal from a distant source arrives redshifted — and a cosmological redshift is a change of scale rather than a speed — so a redshifted chirp is indistinguishable from a chirp of a larger mass nearby. Every quoted mass is therefore a detector-frame mass, and converting it to a source-frame mass needs the redshift — which needs the host. The distance does not suffer from this, because the same redshift enters the amplitude in exactly the way that makes the recovered quantity the luminosity distance rather than a comoving one.

The waveform is the whole measurement, so it is worth drawing over longer stretches for both events — because how long a signal lasts decides how well the chirp mass, and therefore everything else, is measured.

GW150914: 24 Hz to 250 Hz in 0.5 seconds. The strain of GW150914 — two black holes — through the last 0.5 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 28.716 solar masses, and drawn at the luminosity distance it returned, 440 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 24 Hz to 250 Hz, and the envelope — the outer curve — grows by a factor of 4.8, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²¹, so the peak here is a fractional length change of about 2.9·10⁻²¹: over the four kilometres of an interferometer arm that is 1.2·10⁻¹⁷ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 6 Half a second of the black-hole merger. The whole detectable signal is a few dozen cycles, and the chirp mass is measured from how the frequency of those cycles accelerates — a small number of cycles is a poor lever, which is why a black-hole merger’s masses are less well determined than a neutron-star pair’s.
GW170817: 81 Hz to 1000 Hz in 4 seconds. The strain of GW170817 — two neutron stars — through the last 4 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 1.185 solar masses, and drawn at the luminosity distance it returned, 40 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 81 Hz to 1000 Hz, and the envelope — the outer curve — grows by a factor of 5.4, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²², so the peak here is a fractional length change of about 3.9·10⁻²²: over the four kilometres of an interferometer arm that is 1.6·10⁻¹⁸ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 7 And four seconds of the neutron-star merger, which was audible in the band for nearly a hundred. Thousands of cycles pass through the detector, so the chirp mass comes out to five decimal places — and the distance, which depends on the amplitude rather than the phase, does not improve with them at all.

Sirens without a light

One host in a hundred detections is not a programme, and the response has been to find a way of doing without one.

A gravitational-wave event localises its source to a region of sky and a range of distance — a banana-shaped volume of anywhere from tens to thousands of cubic megaparsecs. That volume contains a finite number of galaxies, each with a measured or estimated redshift, and one of them is the host. Any value of the Hubble constant predicts, for the event’s measured distance, which redshifts are consistent; galaxies at other redshifts in the volume are then evidence against that value.

One event that way is nearly useless: the volume holds hundreds of galaxies and the resulting constraint is broad and lumpy. But the true host contributes to every event’s likelihood at the same value of the Hubble constant, while the wrong galaxies contribute at values scattered differently in each event. Stack enough events and the correct value accumulates while the contamination averages down. This is the dark-siren method, and its precision improves as the square root of the number of events rather than as the number.

Two things limit it, and both are about catalogues rather than about waves. The galaxy catalogues covering the relevant volumes are incomplete beyond a few hundred megaparsecs, and the incompleteness is not random — faint galaxies are missing preferentially, and faint galaxies are numerous. What is done instead is to assume the missing hosts follow the same luminosity-weighted distribution as the detected ones, which is a model where a measurement is wanted.

And the events themselves are selected. A detector finds the loud ones, and at fixed intrinsic properties a loud event is a near or face-on one — so the detected population is biased towards small distances and towards face-on orientations, which is Malmquist bias wearing a different hat. The correction is computable, because the detector’s sensitivity as a function of masses, distance and inclination is known from the same calibration described above, and it is applied through a selection function in the likelihood rather than as a correction to individual distances. It is nevertheless the difference between a measurement and a number, and its size grows with the sample it is applied to.

And the degeneracy the chirp mass leaves behind, over the range of component masses the detections actually span.

One measured number, and every pair of masses that produces it. The plane of the two component masses of an inspiralling binary, with three curves of constant chirp mass across it. The middle one is GW150914's value of 28.7 solar masses, and the chirp mass recovered from the coordinates of the drawn curve varies along its whole length by 7.4e-14 per cent — which is the point: every binary on that line radiates the same frequency sweep at leading order, so the early inspiral cannot tell them apart. Two of them are marked. An equal pair of 33.0 and 33.0 solar masses and a lopsided pair of 63.6 and 18.4 sit on the same contour, and their total masses differ by a factor of 1.24. What separates them is the mass ratio, which enters the phasing only at the first post-Newtonian order, suppressed by the square of the orbital speed in units of the speed of light — small through the hundreds of cycles that carry most of the signal, and appreciable only in the last few, where that speed approaches a third of c. So the chirp mass is a measurement and the individual masses are an inference from the end of the signal, which is exactly the part a detector's high-frequency noise eats first.
Fig. 8 Every pair of masses consistent with the chirp mass of a 1.6-and-1.2-solar-mass pair. The curve is a hyperbola and the measurement fixes a point on it, so a chirp mass alone cannot say whether a system is two equal objects or a very unequal pair — and the mass ratio comes from the late inspiral, where the approximation the chirp rests on is failing.

Where this ladder goes next

This rung has taken a signal that carries no image, no spectrum and no position, and got two physical quantities out of it: a mass from a rate of change, and a distance from an amplitude, with nothing calibrated underneath either.

The rung above is the population. Ninety-odd mergers give a mass distribution with structure in it — a gap where pair-instability supernovae should leave one, and objects above the gap that should not exist — and a distribution is a statement about how massive stars end that no individual event could make.

Beside it lies the same measurement made on a completely different instrument: a set of millisecond pulsars used as a galaxy-sized detector, sensitive at nanohertz rather than at hundreds of hertz, and therefore to supermassive binaries rather than to stellar ones.

And below it, the habit that makes this rung worth a place among measurements rather than among mechanisms: a quantity is absolute when the same expression fixes two observables and only one unknown stands between them. The sweep gives the mass; the mass in the amplitude gives the distance; and there is no rung, no calibrator, and no nearby example that had to be measured first.

What this makes readable

Essays that name this one as a prerequisite.

What links here

The 8 of 14 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Chirp massCoalescenceCompact binaryDistance ladderHost galaxyHubble constantInclination degeneracyInspiralLuminosity distanceMatched filteringMultimessengerStandard siren