Gravitation

A distance tangled with an angle

A gravitational wave carries its own distance, with no ladder underneath it and nothing to calibrate. What it also carries, inseparably, is the orientation of the orbit that made it — and one detector cannot tell a nearby binary seen edge-on from one twice as far away seen face-on.

Assumes Gravitational waves and Distance ladder.

A merging binary is the only astronomical object that announces its own distance. The frequency and its rate of change give the masses; the masses give the intrinsic amplitude of the wave; comparing that against the amplitude measured gives the distance, in metres, with no calibrator, no standard candle and no ladder under it.

That is a remarkable thing and it is very nearly useless on its own, because the amplitude does not depend on the distance alone. It depends on the distance divided by a function of the angle between the orbit’s axis and the line of sight, and a single detector measures the quotient.

One amplitude, and every point on this curve fits it. The set of distances and inclinations that produce the same measured amplitude in one detector. A binary seen face-on radiates most strongly along its spin axis, so it can be twice as far away as an edge-on binary and still arrive with the same strain — the curve is (1 + cos²ι)/2 and the factor between its ends is exactly two. Nothing in a single detector's data distinguishes the two ends. What makes this worse rather than merely awkward is that the prior pulls the other way: an isotropically oriented population has half its members beyond 60 degrees, so most binaries really are closer to edge-on, and a posterior that combines a flat likelihood along this curve with that prior returns a distance biased low with an error bar that understates the range. The whole business of standard-siren cosmology is the business of cutting across this curve.
Fig. 1 The set of distances and orientations that produce the same measured amplitude. A binary seen face-on radiates most strongly along its axis, so it can be twice as far away as an edge-on binary and arrive with the same strain. Every point on this curve fits the data exactly, and one detector has nothing to say about which point it is.

The degeneracy is not a defect of the instruments and cannot be engineered away. It is a statement about what a quadrupole radiates, and it would be there for a perfect detector with infinite bandwidth and no noise at all. That distinguishes it sharply from every other limitation on these measurements, all of which improve as the detectors do.

Why the wave knows about the angle

A binary’s two masses orbit in a plane, and the wave they radiate is not isotropic. Along the axis the two mass quadrupoles rotate face-on and the emission is circularly polarised at full strength. In the plane of the orbit, the motion is seen edge-on and only one polarisation survives, at half the amplitude.

Written out, the two polarisation amplitudes are

h+1+cos2ι21D,h×cosι1D,h_+ \propto \frac{1 + \cos^2\iota}{2}\cdot\frac{1}{D},\qquad h_\times \propto \cos\iota\cdot\frac{1}{D},

with ι\iota the inclination of the orbital axis to the line of sight. A detector responds to a linear combination of the two set by its antenna pattern, so what it measures is an effective distance — the real one divided by a factor between one half and one.

The factor of two is not large by the standards of this subject, but it is large compared with everything else in a siren measurement. The masses come out to a per cent or better; the sky position, with three detectors, to tens of square degrees; the distance carries a factor-of-two ambiguity that no amount of signal-to-noise removes, because it is a degeneracy rather than a noise.

Two features of that expression are worth pausing on. First, the plus amplitude never vanishes: even exactly edge-on, a binary is visible at half strength, which is why the degeneracy is a factor of two rather than a factor of infinity. Second, both amplitudes carry the distance in exactly the same way, so no combination of them measures the distance without also measuring the angle. The information about orientation is entirely in their ratio, and the information about distance is entirely in their overall scale, and a single detector measures one number where two are needed.

Two amplitudes whose ratio contains no distance. The two polarisation amplitudes of a circular binary against inclination, in units of the face-on value, together with their ratio. Face-on the two are equal and the wave is circularly polarised; edge-on the cross amplitude vanishes and the wave is linearly polarised. Both amplitudes carry the distance and the ratio does not, so a measurement of the ratio is a measurement of the orientation alone and cuts straight across the degeneracy. The difficulty is that measuring it requires detectors with genuinely different antenna patterns, and two detectors nearly aligned with each other measure almost the same combination. The ratio also changes slowly near face-on — it is still 0.98 at twenty degrees — so it is a poor discriminator exactly where the degeneracy is worst.
Fig. 2 What could break it. The two polarisation amplitudes depend on the inclination differently, so their ratio is an orientation with no distance in it at all — face-on the two are equal and the wave is circularly polarised, edge-on the cross amplitude vanishes and the wave is linear. Measuring that ratio needs detectors whose antenna patterns genuinely differ, and the ratio changes slowly near face-on, which is exactly where the degeneracy is worst.

There is a further complication that is easy to miss: an interferometer’s antenna pattern depends on where the source is on the sky, and the sky position is being fitted at the same time. So the combination of polarisations a given detector measures is itself uncertain, and the distance, the inclination, the sky position and the polarisation angle enter as a tangle of four parameters rather than as two. The pairwise degeneracy in the first figure is a slice through a four-dimensional ridge, and the ridge is what the sampler has to explore. This is why a siren’s distance posterior is computed rather than read off a formula.

The prior makes it worse before it makes it better

There is a second effect, and it acts in the opposite direction to the one intuition suggests.

Orientations are isotropic, so the probability of an inclination is proportional to sinι\sin\iota: half of all binaries lie beyond sixty degrees, and edge-on is far more common than face-on. A prior that prefers edge-on prefers small distances, because an edge-on binary producing the measured amplitude is nearer.

But the detection process pulls the other way. A face-on binary is louder, so it is detectable further away, and the volume surveyed goes as the cube of that distance. Selection therefore favours face-on systems by a factor of about 23=82^3 = 8 in volume, and the detected population is not isotropic even though the underlying one is.

The two effects partly cancel and neither is negligible, so the posterior on the distance is a genuinely two-sided object that depends on both. This is the reason a siren’s distance is quoted with an asymmetric error bar rather than a symmetric one, and the reason the asymmetry is large: a typical single-event distance posterior runs from about 0.7 to 1.6 times its median.

One amplitude, and every point on this curve fits it. The set of distances and inclinations that produce the same measured amplitude in one detector. A binary seen face-on radiates most strongly along its spin axis, so it can be twice as far away as an edge-on binary and still arrive with the same strain — the curve is (1 + cos²ι)/2 and the factor between its ends is exactly two. Nothing in a single detector's data distinguishes the two ends. What makes this worse rather than merely awkward is that the prior pulls the other way: an isotropically oriented population has half its members beyond 60 degrees, so most binaries really are closer to edge-on, and a posterior that combines a flat likelihood along this curve with that prior returns a distance biased low with an error bar that understates the range. The whole business of standard-siren cosmology is the business of cutting across this curve.
Fig. 3 The same curve for a nearer event, at the effective distance the first neutron-star merger was measured at. The shape is identical because it is a property of the emission and not of the source, and everything about the ambiguity scales with the distance rather than being reduced by proximity. Being close buys signal-to-noise, which buys the masses and the sky position, and buys nothing at all on this axis.

It is worth stating the selection effect carefully, because it is a genuine bias in the population rather than an artefact of any one measurement. If the underlying binaries are isotropically oriented, the detected ones are not: the detection horizon is a factor of two further for face-on systems, so the detected sample over-represents them by a factor of eight in volume. That is a large distortion, and any statement about the population — how eccentric they are, what the mass distribution is, whether the spins are aligned — has to be corrected for it, in the same way an exoplanet survey’s yield has to be corrected for what it could have seen.

What the counterpart is worth

The neutron-star merger of August 2017 is the case where the degeneracy was broken from outside, and it is worth following in detail because each step contributed a different amount.

The gravitational-wave data alone gave a distance of 4014+840^{+8}_{-14} megaparsecs — a fractional uncertainty of about a quarter, dominated entirely by the inclination.

An electromagnetic counterpart was found, so the host galaxy was identified and its recession velocity was known. That is the other half of a Hubble measurement and it is what makes a siren useful at all: a distance in megaparsecs and a velocity in kilometres a second, with no rung of any ladder between them.

The counterpart also constrained the inclination independently. The event produced a short gamma-ray burst, and a burst is beamed; the afterglow’s rise and decline, monitored in radio and X-rays for two years, is fitted with a structured jet whose viewing angle is a parameter. That fit gave roughly fifteen to twenty-five degrees from the axis — face-on, in the language of the first figure — and imposing it collapsed the distance posterior.

The published Hubble constant from that event went from about fifteen per cent uncertain on the gravitational-wave data alone to under seven per cent with the jet constraint. Half the error bar was orientation.

One amplitude, a factor of 2√2 in distance, and the whole error budget. The luminosity distance a fixed strain amplitude implies, against the inclination of the orbit to the line of sight, normalised so that a face-on source sits at GW170817's fitted 40 megaparsecs. A binary seen face-on radiates most strongly towards the observer and an edge-on one least, in the ratio 2√2 — so the same measured amplitude is consistent with everything between 14 and 40 megaparsecs, and the amplitude contains nothing that could choose between them. This is the sense in which a standard siren is absolute but not precise: the chirp mass comes from the sweep and is known to four figures, the distance comes from the amplitude and is known to tens of per cent, and the reason is one angle. Everything that has ever narrowed it has come from outside the waveform — the two detectors' relative amplitudes, the polarisation, and for GW170817 the superluminal motion of its radio afterglow, which fixed the viewing angle independently. The right-hand scale is what the ambiguity does to the answer: at NGC 4993's recession speed of 3017 km s⁻¹, the same signal gives a Hubble constant anywhere from 75 to 213 km s⁻¹ Mpc⁻¹.
Fig. 4 The measurement the degeneracy is standing in the way of: a distance from the wave against a recession velocity from the host, whose slope is the Hubble constant with no calibration anywhere in it. One event of this kind is worth an enormous amount precisely because the two axes are independent of everything else in cosmology, and it is worth much less than it should be because the vertical axis carries an ambiguity the horizontal one does not.

The lesson from that sequence is not that counterparts are necessary. It is that the constraint came from a completely different physical process — the geometry of a relativistic jet, measured by watching an afterglow brighten and fade over two years — that happens to share one parameter with the gravitational-wave problem. The inclination is the orbital axis in one measurement and the jet axis in the other, and identifying the two is an assumption, though a good one: the jet is launched along the angular momentum axis of the remnant, which is the orbital axis of what made it.

That assumption is the kind worth flagging. It is almost certainly right, it is not a measurement, and it is carrying a substantial share of the error bar on the best independent Hubble constant available.

The rest of the signal is not degenerate at all

It is worth being clear that the ambiguity is confined. Almost everything else a merger measures is measured extremely well, and the contrast is instructive.

GW150914: 33 Hz to 250 Hz in 0.22 seconds. The strain of GW150914 — two black holes — through the last 0.22 seconds before merger, computed from the quadrupole sweep at the chirp mass its fit returned, 28.716 solar masses, and drawn at the luminosity distance it returned, 440 megaparsecs. Two things rise together and neither is free to rise on its own: the frequency goes from 33 Hz to 250 Hz, and the envelope — the outer curve — grows by a factor of 3.9, because the amplitude goes as f^2/3 and nothing else in it changes over so short a span. The vertical axis is in units of 10⁻²¹, so the peak here is a fractional length change of about 2.9·10⁻²¹: over the four kilometres of an interferometer arm that is 1.2·10⁻¹⁷ metres, a thousandth of the width of a proton. The chirp mass is not fitted to the amplitude at all — it comes from the spacing of these zero crossings, which is why it is the best-determined number in the whole event and why the distance, which does come from the amplitude, is the worst.
Fig. 5 The signal itself, whose frequency and rate of change are what the masses come from. This part of the measurement is a timing problem rather than an amplitude problem, and timing is what these detectors do best: the phase is tracked through hundreds of cycles and the accumulated phase error over the whole signal is a fraction of a radian. Nothing about the orientation enters it.

The masses come from the phase evolution, which is a frequency measurement, and frequency measurements are immune to everything the amplitude is vulnerable to. A detector whose calibration is ten per cent wrong measures the masses correctly and the distance ten per cent wrong.

One measured number, and every pair of masses that produces it. The plane of the two component masses of an inspiralling binary, with three curves of constant chirp mass across it. The middle one is GW150914's value of 28.7 solar masses, and the chirp mass recovered from the coordinates of the drawn curve varies along its whole length by 7.4e-14 per cent — which is the point: every binary on that line radiates the same frequency sweep at leading order, so the early inspiral cannot tell them apart. Two of them are marked. An equal pair of 33.0 and 33.0 solar masses and a lopsided pair of 63.6 and 18.4 sit on the same contour, and their total masses differ by a factor of 1.24. What separates them is the mass ratio, which enters the phasing only at the first post-Newtonian order, suppressed by the square of the orbital speed in units of the speed of light — small through the hundreds of cycles that carry most of the signal, and appreciable only in the last few, where that speed approaches a third of c. So the chirp mass is a measurement and the individual masses are an inference from the end of the signal, which is exactly the part a detector's high-frequency noise eats first.
Fig. 6 The combination the phase actually determines. Two masses enter the leading-order phase evolution only through one number, so the early inspiral gives that number to a fraction of a per cent and the individual masses much less well — a degeneracy of its own, broken later in the signal by higher-order terms. The distinction worth carrying is that this one is broken by more signal and the distance–inclination degeneracy is not.

That is the cleanest statement of the difference between the two kinds of degeneracy this collection keeps meeting. A degeneracy that a better measurement breaks is a nuisance. A degeneracy that a better measurement does not break is a structural feature of what the observable is, and it can only be broken by a different observable.

The same contrast appears in the sky localisation. A single detector localises a source to a ring; two detectors localise it to two arcs on that ring, from the arrival-time difference; three narrow it to a patch. Each addition buys geometry rather than sensitivity, and the pattern is general — a detector the size of a galaxy faces the identical problem at a wavelength eleven orders of magnitude longer, and solves it the same way, by having many independent lines of sight rather than one good one.

What was actually measured

Three numbers anchor the argument.

The factor of two is exact. It is not an empirical statement about detectors: (1+cos2ι)/2(1+\cos^2\iota)/2 takes the value one at ι=0\iota = 0 and one half at ι=90°\iota = 90°, and that is the whole of it. Any single interferometer sensitive to only one polarisation carries exactly this ambiguity.

Three detectors are worth roughly one polarisation ratio. The two LIGO instruments are deliberately near-aligned, so they measure almost the same combination and add signal-to-noise rather than information about the orientation. Virgo is oriented differently, and its addition to the 2017 event is what gave the sky localisation and a weak inclination constraint — enough to exclude edge-on, not enough to pin the angle.

The dark-siren route works and is expensive. Where no counterpart is found, the host can be marginalised over the galaxies in the localisation volume, each weighted by its own redshift. That recovers a Hubble constant without an electromagnetic detection, and it converges as roughly the inverse square root of the number of events. The current dark-siren constraints from tens of events are weaker than the single event with a counterpart.

Two parallel lines of slope 11/3, and the gap between them is a mass ratio. The rate at which the frequency rises, against the frequency, for GW150914 and GW170817. In these coordinates the quadrupole sweep is a straight line of slope exactly 11/3, and the only property of the binary that moves it is the chirp mass, which shifts it vertically by 5/3 of a decade for every decade of mass. So the two lines here are parallel and 2.31 decades apart, and that gap is the 24.2-fold difference in chirp mass between 28.72 and 1.18 solar masses — nothing else about either system enters. Two black holes of 36 and 31 solar masses and a pair of equal masses adding to the same chirp mass would draw the same line, which is why the individual masses are always quoted with error bars several times wider than the chirp mass's. The dots mark where each signal entered the detector band, and the horizontal run to the right of each is the whole observation: GW150914, 183 milliseconds; GW170817, 102 seconds.
Fig. 7 The population these constraints are accumulating over, drawn in the plane the detectors actually measure. Each event contributes a distance with its own orientation ambiguity, and the ambiguities are independent, so they average down — which is the argument for waiting rather than for cleverness. What does not average down is anything systematic in the calibration of the strain amplitude, and that is the reason the detectors’ absolute calibration is itself an experiment.

And there is a fourth number, and it is the one that decides how the field will look in ten years. A single event’s orientation ambiguity contributes about fifteen per cent to a Hubble constant, and independent events average that down as the inverse square root of their number: a hundred events with counterparts would give one and a half per cent, which is decisive against a five-sigma disagreement between an early-universe value and a late-universe one. The rate of neutron-star mergers with detected counterparts is the quantity that decides when, and it is currently about one every few years.

Where the picture stops

There are three, and the second is the one likely to change first.

It assumes the orbit is circular. An eccentric binary radiates in harmonics of the orbital frequency with a different angular pattern, so in principle the eccentricity supplies extra orientation information. In practice every pair arrives circular, because radiation reaction circularises the orbit long before the signal enters the band, so this route is unavailable for exactly the systems being observed.

Higher-order emission breaks it partially, and will break it better. The quadrupole formula is the leading term; the sub-dominant multipoles have their own inclination dependences, and they are detectable in systems with unequal masses. Events with mass ratios far from one already show them, and for those the inclination is measured from the waveform alone. As the detectors improve, this becomes the standard route and the counterpart becomes a bonus.

And the degeneracy is with the luminosity distance, which is not a distance. What a siren measures is the luminosity distance, which in an expanding universe differs from every other distance by factors of the redshift — the same distinction that runs through the whole distance ladder. At forty megaparsecs the difference is negligible and at a gigaparsec it is not, so a siren used for cosmology is measuring a quantity whose relation to the thing wanted is itself cosmology-dependent.

One more limit belongs here because it is about the word rather than the physics. “Standard siren” is an analogy with “standard candle”, and the analogy is inexact in a way that matters. A candle is standard because its intrinsic brightness is assumed known from a calibrated population; a siren’s intrinsic amplitude is computed from the masses that the same signal measures, so nothing is assumed and nothing is calibrated. What the siren does share with the candle is that the observed quantity is an amplitude, and every amplitude in astronomy is a product of a source property and a geometric one. The masses that come out of an orbit as a lower bound are the same trade seen in a different observable: a quantity multiplied by an unknown angle, reported as though the angle were known.

Why the ambiguity is worth understanding rather than deploring

A standard siren is the most direct distance measurement in astronomy. It has no calibration, no metallicity dependence, no extinction correction, and no reliance on any other measurement. Those are exactly the properties that make it valuable in the standoff over the Hubble constant, where the two existing answers disagree by five standard deviations and both are built on long chains of calibration.

An independent method with a factor-of-two ambiguity that averages down is worth a great deal more than a precise method with an unknown systematic. The orientation ambiguity is statistical: it is different for every event, it is understood exactly, and it shrinks as the square root of the number of events. Nothing about it can conspire to shift the answer in one direction.

That is the property that makes the method decisive in the end, and it is worth stating plainly because it is the opposite of how the error bar looks. A siren’s uncertainty is large and honest. The competing methods’ uncertainties are small and rest on assumptions that would move the answer if they were wrong.

It is also worth noticing that the ambiguity has a preferred direction with respect to the thing being measured. Because detected events skew face-on and face-on means further, ignoring the degeneracy entirely — taking the effective distance as the distance — systematically underestimates distances and therefore overestimates the Hubble constant. The size of that error is a factor approaching two in the worst case and around thirty per cent for a typical detected orientation. Nobody makes that mistake deliberately, but it is the shape of the error that any incomplete treatment of the inclination will produce, and knowing its sign is most of what is needed to catch one.

Put the selection effect alongside the degeneracy, because the two combine in a way that biases a population. A face-on binary emits more strongly along the line of sight than an edge-on one, so at a fixed distance a face-on source is louder and more likely to be detected — by a factor that makes the detected population strongly weighted towards face-on orientations. Since face-on is also the orientation that makes a source look nearer than it is, the detected sample is systematically biased towards inferred distances shorter than the true ones unless the selection is modelled. The correction is straightforward in principle: the prior on inclination is not uniform in the angle but weighted by the detection probability. It is easy to get wrong, and getting it wrong shifts every distance in one direction, which is exactly the sort of error that a measurement of the expansion rate cannot absorb.

A final point about what the degeneracy is not. It is not a limitation of the detectors’ sensitivity, and it does not narrow as the instruments improve: a louder signal is measured more precisely along the degenerate direction and no better across it. Only a second observable — a counterpart, a network geometry, a higher harmonic — changes the shape of the answer rather than its size.

Where the ladder goes next

The rung directly above is the sub-dominant multipoles: what a waveform with unequal masses says about its own orientation, and how much of the ambiguity disappears when the emission is no longer purely quadrupolar. The rung after it is the dark-siren statistics — how a distance with no host galaxy is combined with a catalogue of candidate hosts, and what that inference assumes about the completeness of the catalogue.