Gravitation

Nothing in the sky is weighed in kilograms

The Sun's gravitational parameter is known to eleven significant figures. The Sun's mass is known to five. The two statements are about the same object and the difference between them is a constant measured in basements, which is the worst-determined fundamental constant in physics.

Assumes Harmonic law and Shell theorem.

Ask how heavy the Sun is and the answer given is 1.989×10301.989 \times 10^{30} kilograms. Four significant figures, and the fifth is uncertain — not because the Sun is hard to observe, but because nobody can weigh anything to better than about twenty parts in a million using gravity, and the difficulty is entirely in a laboratory on Earth.

Ask instead for the Sun’s gravitational parameter, the product GMGM_\odot, and the answer is 1.32712440018×1020 m3s21.32712440018 \times 10^{20}\ \mathrm{m^3\,s^{-2}}, good to something like a part in 101010^{10}.

The same object, the same physics, six orders of magnitude of difference in how well it is known. Everything in this essay follows from noticing which of those two numbers astronomy actually uses.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — torsion balance, in one form or another, against beam balance, pendulum, atom interferometry. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67435 ± 0.00004 across 11 of them, against 6.67343 ± 0.00009 across 3. The difference is 0.00092 ± 0.00010 10⁻¹¹ m³ kg⁻¹ s⁻², which is 9.3 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 1 Fourteen determinations of the gravitational constant, each with the interval its own authors published, grouped by method rather than by result. The spread across them is about 500 parts per million; the quoted intervals run from 12 to 130. They cannot all be right. This is the state of the least well-known constant in physics, and it has been roughly this state for thirty years.

What an orbit actually reports

Newton’s law contains the product of two things that never appear separately in any orbital calculation. Write the acceleration of a body at distance rr from a mass MM:

r¨=GMr2r^.\ddot{\mathbf r} = -\frac{GM}{r^2}\hat{\mathbf r}.

GG and MM occur together and nowhere else. No measurement of any trajectory — a moon, a planet, a spacecraft, a star in a binary — can separate them, because the equation of motion has no term in which they are apart. What an orbit determines is the product, and the product has its own name and symbol, μ=GM\mu = GM, the standard gravitational parameter. Every gravitational quantity in astronomy is of this kind. Escape speed is 2GM/r\sqrt{2GM/r}. Orbital speed is GM/r\sqrt{GM/r}. The Schwarzschild radius is 2GM/c22GM/c^2. The Hill radius involves only mass ratios. In none of them can GG be removed from MM, and in none of them does anybody want to.

The system built to avoid the problem

Astronomy noticed this early and built its unit system around it.

Gauss, in 1809, defined a constant kk by fixing the Earth’s orbital period, the Earth’s mass and the astronomical unit, and taking k2=GMk^2 = GM_\odot in those units. The consequence was that kk was a defined number, 0.01720209895, rather than a measured one, and every planetary calculation could be carried out to full precision with no reference to any laboratory. Masses were quoted in solar masses. Distances were quoted in astronomical units. Times were quoted in days. In that system the gravitational constant does not appear.

The consequence is a fact worth stating plainly. The mass ratios in the solar system were known to a few parts in a thousand more than a century before anybody had a decent value of GG, because a ratio of masses is a ratio of two μ\mu values and the GG divides out. Newton computed the Sun-to-Jupiter ratio in the Principia and got 1067 against a modern 1047, with no gravitational constant in existence at the time and no concept of one. The unit system was retired in stages: the astronomical unit is now a defined length, exactly 149,597,870,700 metres, fixed by the IAU in 2012 after radar and spacecraft ranging had made the old parallax-based determinations obsolete; and GMGM_\odot is now quoted as a measured quantity in SI units. But the underlying situation is unchanged. Planetary ephemerides are fitted with μ\mu values as parameters, and no ephemeris in the world contains GG.

Why the constant is so hard

The gravitational constant has been measured for two and a quarter centuries and its uncertainty has improved by about three orders of magnitude in that time — while cc became exact, hh became exact, and the electron’s magnetic moment reached twelve significant figures. Four features of gravity explain the gap, and none of them is going away.

It cannot be shielded. Every electromagnetic measurement can be put inside a Faraday cage and its background removed. There is no gravitational cage, so the mass of the experimenter, the water table under the building and the tidal position of the Moon are all present in the signal.

It cannot be modulated. The standard defence against drift and noise is to switch the effect on and off at a known frequency and detect only that frequency. Gravity has no switch. What torsion-balance experiments do instead is move the source masses, which introduces exactly the systematic — a change of geometry — that the measurement is most sensitive to.

It is absurdly weak. The gravitational attraction between two protons is smaller than their electrostatic repulsion by about 103610^{36}. A laboratory measurement of GG works with forces of order 10910^{-9} newtons in the presence of the Earth’s pull on the same apparatus, which is 10910^{9} times larger.

And it needs a mass metrology, not just a force metrology. Every other route to a fundamental constant can be made to depend on a frequency, and frequencies are the best-measured quantities in existence. The source masses must be weighed, their density inhomogeneities characterised, and their positions known to microns. Several of the discrepant results in the figure differ in ways that have been traced to how the source masses’ density was assumed to be distributed.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — one laboratory, four determinations, against every other laboratory. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67425 ± 0.00005 across 4 of them, against 6.67414 ± 0.00005 across 10. The difference is 0.00012 ± 0.00007 10⁻¹¹ m³ kg⁻¹ s⁻², which is 1.6 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 2 The same fourteen determinations with one thing changed: they are sorted by laboratory rather than by method. HUST’s four results against everyone else’s ten, and the two families agree to 1.6 standard deviations — against the 9.3 the method split reports on the identical data. Nothing was re-measured and no interval was adjusted. What the two drawings differ in is the question, and the fact that the answer changes by a factor of six tells you the split does not lie where the first figure puts it. A grouping that produces a large sigma is a hypothesis about the systematic, not a measurement of one.

The result is the state the hero figure shows: not a random scatter around a value, but a set of careful experiments with small quoted errors that disagree with each other. CODATA’s response has been to inflate the recommended uncertainty by a factor of several beyond what a weighted combination would give — an admission that at least one unidentified systematic is at large.

Cavendish did not measure G

The experiment everybody names is the one that did not do this.

Henry Cavendish’s 1798 paper is titled Experiments to Determine the Density of the Earth, and that is what it reports: 5.48 times the density of water, against a modern 5.514. The gravitational constant appears nowhere in it, for the excellent reason that it had not been invented — Newton’s law was written as a proportionality, and the symbol GG with a numerical value attached does not appear in the literature until Cornu and Baille in 1873.

What Cavendish did was compare two attractions. His torsion balance measured the pull of two lead spheres on two smaller ones; the Earth’s pull on the same small spheres was already known, being their weight. The ratio of the two attractions, with the geometry, gives the ratio of the Earth’s mass to the lead spheres’ — and the lead spheres could be weighed. A density came out, not a constant.

That framing is worth recovering, because it is the same trick the rest of this essay is about, run in the other direction. Cavendish did not need GG because he took a ratio; astronomy does not need GG because it takes ratios. Extracting a constant from his result is a modern back-formation, and doing it gives 6.74imes10116.74 imes10^{-11}, which is within one per cent of the current value and better than several nineteenth-century determinations that were trying.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — before 2005, against 2005 and after. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67427 ± 0.00008 across 5 of them, against 6.67418 ± 0.00004 across 9. The difference is 0.00010 ± 0.00009 10⁻¹¹ m³ kg⁻¹ s⁻², which is 1.0 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 3 The same fourteen sorted by date instead: before 2005 against 2005 and after. The two families agree to one standard deviation, which is the weakest split of the four this figure can draw — so the disagreement is not a matter of older experiments being worse. The modern determinations have smaller intervals and scatter just as widely relative to them, which is the whole difficulty: two centuries of improvement have reduced the quoted errors and not the disagreement, and a spread that shrinks no faster than the error bars is a spread that is not statistical.

And it costs astronomy almost nothing

Here is the part that surprises people, and it is the reason this essay belongs in a collection about the sky rather than one about metrology.

A 22-parts-per-million uncertainty in GG propagates directly into the mass of every astronomical body expressed in kilograms. It propagates into essentially nothing else, because astronomy almost never expresses a mass in kilograms. Where the constant does bite is at the joins between astronomy and physics, and the list is short and specific: the number of baryons in a star, which needs a mass in kilograms and a proton mass; the mean density of the universe expressed as a critical density, where GG appears explicitly in 3H02/8πG3H_0^2/8\pi G; the equation of state of a neutron star, where a laboratory nuclear physics must be matched to an astronomical mass. It also bites, quietly, wherever a stellar model is integrated: the interior equations carry GG explicitly, so the pressure that holds a star up and the gradient that decides whether it convects are computed with it. In each of those a 22-ppm uncertainty is negligible against everything else in the calculation — the critical density’s own uncertainty is dominated by H0H_0, which is disputed at the level of eight per cent, or three and a half thousand times worse.

H₀: nine determinations in two families. Published determinations of H₀, each with its quoted one-sigma interval, sorted into two families — measured locally, calibrated by a ladder, against inferred from z ≈ 1100 through a model. The shaded band behind each family is that family's inverse-variance weighted mean: 72.66 ± 0.75 across 5 of them, against 67.40 ± 0.41 across 4. The difference is 5.26 ± 0.85 km/s/Mpc, which is 6.2 standard deviations, computed here from the quoted errors alone. That number is an upper bound on the significance rather than the significance: the determinations within each family share calibrations, samples and in two cases the same supernovae, so they are not independent, and a correlated pair combines to something wider than the formula used here gives. What the figure does establish is that the split is not one discrepant measurement against a consensus — it is two internally consistent groups, and the grouping is by method rather than by result.
Fig. 4 The same drawing for the constant everybody argues about, and the comparison is the point of putting them on one pair of axes. The Hubble tension is a disagreement of about eight per cent between two families of determinations, it comes to five sigma, and it is called a crisis. GG’s determinations disagree by 500 parts per million with quoted errors ten times smaller, which is a worse internal inconsistency by the only measure that matters — the ratio of the disagreement to the claimed precision — and nobody calls it anything. The difference is not statistical. It is that no astronomical measurement needs GG separately from the product GMGM, so a wrong value falsifies nothing, while H0H_0 is the answer to the question it appears in. A discrepancy is a crisis only if something depends on it.

What was actually measured, and when the units changed

The astronomical unit is the hinge of this whole story, because for three centuries it was the one length in the solar system that had to be measured rather than computed.

Every planetary distance was known as a ratio to the Earth’s, from Kepler’s laws and timed observations, to a precision far better than the absolute scale. Fixing that scale required one absolute measurement, and successive attempts are a history of the subject: the transits of Venus of 1761, 1769, 1874 and 1882, whose timings gave the solar parallax to about a part in a thousand; the parallax of the asteroid Eros at its 1930–31 opposition, which reached about 1.5×1041.5\times10^{-4}; and then radar.

The 1961 radar echoes from Venus fixed the astronomical unit to about a part in 10610^{6} overnight, and spacecraft tracking has since improved it to a few parts in 101110^{11}. At that point the AU stopped being a measurement at all and became a defined conversion factor. The last absolute length in solar-system astronomy was retired by a radio pulse and a stopwatch.

The Earth’s own parameter tells the same story one step closer to home. Satellite laser ranging — bouncing pulses off retroreflectors on LAGEOS and its successors, and timing the return to a few millimetres — gives GM=3.986004418×1014 m3s2GM_\oplus = 3.986004418\times10^{14}\ \mathrm{m^3\,s^{-2}}, to about two parts in 10910^{9}. Dividing by GG to obtain the Earth’s mass in kilograms throws away four of those nine digits at a stroke. The Earth is the best-characterised object in the universe and its mass is the worst-known thing about it.

What the picture cannot show

The hero figure computes a difference between its two groups and the difference is not the story. It reports 9.3 standard deviations between the torsion-balance family and the rest, which is what its arithmetic gives; but the same arithmetic applied within the torsion-balance family would report an equally severe inconsistency. The disagreement does not sort by method. It sorts, as far as anybody can tell, by laboratory, and that is precisely what makes it hard to fix.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — torsion balance, time-of-swing, against every other technique, torsion included. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67400 ± 0.00007 across 6 of them, against 6.67428 ± 0.00004 across 8. The difference is -0.00028 ± 0.00008 10⁻¹¹ m³ kg⁻¹ s⁻², which is 3.4 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 5 The split inside the torsion-balance family, which is the one the paragraph above says exists. Time-of-swing determinations against every other technique, torsion included: 3.4 standard deviations, from measurements that all use the same instrument and differ only in how the period is read. A method split that survives restricting to one method is not a method split. Set the four groupings this figure can draw side by side — 9.3, 3.4, 1.6 and 1.0 sigma from one unchanged table — and what they measure is how much a chosen partition can be made to say, which is a caution about the technique rather than a finding about gravity.

Nothing here says GG is variable. A constant that different experiments disagree about is not the same as a constant that changes, and the astronomical constraints on the latter are far tighter than the laboratory ones on the former: lunar laser ranging bounds G˙/G\dot G/G below about 101310^{-13} per year, and binary-pulsar timing does comparably well. The best evidence that GG is constant comes from the sky; the best measurements of its value do not.

And the two headline numbers are not directly comparable. GMGM_\odot at a part in 101010^{10} is the parameter of an ephemeris fitted to ranging data over decades, and its uncertainty is a formal one within a model containing hundreds of other parameters. It is an excellent number and it is not the same kind of object as a laboratory result with an error budget itemised line by line.

And a figure of determinations is not a figure of the truth. Every interval drawn in the hero figure is its own authors’ honest estimate of their own systematics, and the whole content of the picture is that at least one of those estimates is wrong. A plot like this can show inconsistency and can never show which point is at fault; identifying that requires somebody to repeat somebody else’s apparatus, which is expensive, unglamorous and exactly what the field has spent the last decade doing.

Both rungs share a question this essay has dodged. If a mass ratio is what astronomy measures, and the solar mass is the unit, then what is being asserted when a galaxy is said to weigh 101210^{12} solar masses? The answer is that a ratio has been taken between two gravitational parameters twelve orders of magnitude apart, measured by entirely different means, and that the chain connecting them is as long as any distance ladder — which is a different essay and a longer one.

Weighing an ice sheet

There is a class of measurement in which a mass in kilograms is genuinely what is wanted, and it is worth describing because it shows that even there the constant mostly cancels.

A pair of satellites in the same orbit, separated by a couple of hundred kilometres and ranging to each other by microwave link, measure the distance between themselves to a fraction of a micron. As they pass over a region of excess mass, the leading satellite is accelerated first and the separation changes; a hundred kilometres later the trailing one catches the same pull and it changes back.

Integrating those changes over the whole orbit, month after month, maps the Earth’s gravity field and — more usefully — how it changes. The changes are seasonal and secular: groundwater moving, ice sheets shrinking, the crust rebounding from the last glaciation.

The results are quoted in gigatonnes per year, which is a mass rate in kilograms, and the numbers are the primary measurement of how fast the Greenland and Antarctic ice sheets are losing mass.

Now notice what the constant does there. What is measured is a change in the gravitational parameter of a region — a change in GMGM, not in MM — so converting it to a mass requires dividing by GG, and a 22-parts-per-million error in GG is a 22-parts-per-million error in the ice loss. The published uncertainties on those rates are several per cent, dominated by how the signal is separated from the crust’s rebound. The constant contributes a term a thousand times smaller than the smallest thing anyone argues about.

Even the one measurement whose answer must be in kilograms is unaffected, because a mass difference inherits the same fractional error as a mass and the fractional error is negligible against everything else.

The redefinition that left it out

In 2019 the SI base units were redefined so that a set of constants take exact values: the speed of light, the Planck constant, the elementary charge, the Boltzmann constant and the Avogadro constant. The kilogram in particular stopped being a lump of metal in a vault and became a quantity derived from the Planck constant, realised by an instrument that balances a mechanical force against an electromagnetic one.

The gravitational constant was not in the list, and it could not have been.

The redefinition works by tying each unit to a phenomenon whose measurement reduces to a frequency, because frequencies can be measured to eighteen digits against an atomic clock. The kilogram’s realisation is a comparison between mechanical and electrical power, both of which reduce to voltages and velocities that are ultimately frequencies.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — torsion balance, in one form or another, against beam balance, pendulum, atom interferometry. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67435 ± 0.00004 across 11 of them, against 6.67343 ± 0.00009 across 3. The difference is 0.00092 ± 0.00010 10⁻¹¹ m³ kg⁻¹ s⁻², which is 9.3 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 6 The dataset’s own split stated explicitly rather than by default, which is worth doing once because the difference between this figure and the hero is nothing at all. The same fourteen rows, the same two families, the same 9.3 sigma. What that establishes is that the four groupings above are not four datasets: they are four readings of one, and the only thing separating them is a key attached to each row after the measurement was published. A constant that cannot be fixed by definition is a constant whose value depends on which of these pictures somebody believes.

There is no frequency in gravity. Every route to GG requires measuring a force, or a displacement, or a period of a torsion oscillator whose restoring constant must itself be calibrated — and none of them reduces to a frequency comparison against a clock. So GG cannot be fixed by definition without making some other quantity worse, and it remains one of the few constants in the tables that is measured rather than assigned.

The situation has a curious consequence for the mass of the Sun. Now that the kilogram is defined through the Planck constant, expressing a solar mass in kilograms means connecting an astronomical measurement to a quantum-mechanical definition through a constant known to twenty-two parts in a million — and the chain from a planet’s orbit to a photon’s energy passes through the least reliable link in the whole system of units.

G: 14 determinations in two families. Published determinations of G, each with its quoted one-sigma interval, sorted into two families — quoted better than 100 ppm, against quoted worse than 100 ppm. The shaded band behind each family is that family's inverse-variance weighted mean: 6.67420 ± 0.00004 across 11 of them, against 6.67301 ± 0.00048 across 3. The difference is 0.00120 ± 0.00048 10⁻¹¹ m³ kg⁻¹ s⁻², which is 2.5 standard deviations, computed here from the quoted errors alone. The arithmetic is the same one the Hubble figure uses and here it should be distrusted, because the scatter inside each family already exceeds what the intervals allow: eleven torsion-balance determinations spread over 500 parts per million with quoted intervals of 12 to 130 cannot all be right, whatever the difference between the families comes to. That is why the recommended value's uncertainty is expanded far beyond any single experiment's rather than being the weighted combination drawn here — the disagreement is between laboratories using the same method, not between methods.
Fig. 7 And the split that asks whether the confident experiments agree with each other. Sorting by the quoted interval rather than by anything physical — better than a hundred parts per million against worse — puts eleven of the fourteen in one family, and the eleven are the whole problem: they span 500 parts per million while quoting intervals between 12 and 130. The three loose ones sit well away from them and are the only part of the picture behaving as a scatter should. That is what stops GG being fixed by definition. A constant is assigned a value when the experiments claiming the smallest errors converge, and here they are the ones spread furthest apart in units of their own error bars.

The constant that governs the largest structures in the universe is the one the metrologists could not include, and the reason is that gravity does not oscillate at a frequency anybody can count.

One more dataset puts the same disagreement beside a constant everybody agrees is measured well.

H₀: nine determinations in two families. Published determinations of H₀, each with its quoted one-sigma interval, sorted into two families — measured locally, calibrated by a ladder, against inferred from z ≈ 1100 through a model. The shaded band behind each family is that family's inverse-variance weighted mean: 72.66 ± 0.75 across 5 of them, against 67.40 ± 0.41 across 4. The difference is 5.26 ± 0.85 km/s/Mpc, which is 6.2 standard deviations, computed here from the quoted errors alone. That number is an upper bound on the significance rather than the significance: the determinations within each family share calibrations, samples and in two cases the same supernovae, so they are not independent, and a correlated pair combines to something wider than the formula used here gives. What the figure does establish is that the split is not one discrepant measurement against a consensus — it is two internally consistent groups, and the grouping is by method rather than by result.
Fig. 8 The Hubble constant grouped by method. The pattern is the same shape as the gravitational constant’s: measurements that agree within their own technique and disagree between techniques, which is the signature of an unmodelled systematic rather than of bad statistics.

Where the ladder goes next

The obvious rung is the one where the ratio trick runs out: extragalactic astronomy, where masses are quoted in solar masses because there is no alternative, and the solar mass itself has become a unit rather than a measurement. The rung beside it is the geodetic one — the Earth’s own μ\mu, known to two parts in 10910^{9} from satellite laser ranging, against the Earth’s mass, known to 22 parts in 10610^{6}, which is the same disparity as the Sun’s and considerably easier to check.

What this makes readable

Essays that name this one as a prerequisite.

What links here

The 8 of 16 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Astronomical unit (AU)Dimensional analysisGravitational constantKepler's third lawMass ratioRadar rangingSolar massStandard gravitational parameterSystematic errorTorsion-balance