Starlight

Resolution without a mirror

Two telescopes a kilometre apart do not make a kilometre-wide telescope. They measure one number — the Fourier component of the sky at the spatial frequency their separation sets — and an image is what you get by collecting enough of those.

Assumes Angular diameter and Seeing.

The resolution of a telescope is set by its aperture, and the arithmetic is unforgiving. To resolve a stellar disc — a few thousandths of an arcsecond across for the nearest giants — a visible-light telescope would need to be a hundred metres wide. To see the shadow of a black hole in another galaxy, a radio telescope would need to be the size of the Earth — an angle far below what the atmosphere allows a single mirror in any case.

Neither instrument exists and both measurements have been made, because a filled aperture is not what resolution requires. What it requires is the information a filled aperture would have collected, and most of that information can be gathered by a small number of separated pieces.

The transform plane an array of 9 actually samples. Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. The 9 antennas make 36 pairs, each pair measures one point at any instant, and turning the Earth sweeps each of them along an ellipse — so eight hours of tracking turns 36 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian: a real sky forces V(−u,−v) = V(u,v), so half the points are free and the coverage is symmetric through the origin. And every ellipse has axis ratio exactly sin δ = 0.707* at this declination, measured off the longest track as 0.707 — an array is squashed in one direction by where the source is in the sky, and at the equator the tracks collapse to lines whatever the array. What the figure cannot show is the hole in the middle: no baseline is shorter than an antenna is wide, so the largest structures on the sky are simply not measured, and no processing recovers them.
Fig. 1 Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. Nine antennas make 36 pairs, each pair measures one point at any instant, and turning the Earth sweeps each along an ellipse — so eight hours of tracking turns 36 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian, because a real sky forces V(u,v)=V(u,v)V(-u,-v) = V^*(u,v), so half the points are free. And every ellipse has axis ratio exactly sinδ\sin\delta — an array is squashed in one direction by where the source happens to be in the sky, and at the equator the tracks collapse to lines whatever the array.

What a pair of apertures measures

Combine the light from two apertures separated by a vector B\mathbf B and the result is not an image. It is a single complex number: an amplitude and a phase, called the visibility, and it is what the van Cittert–Zernike theorem identifies as the Fourier transform of the sky brightness distribution evaluated at the single spatial frequency B/λ\mathbf B/\lambda.

That theorem is the load-bearing statement and it is taken here as an input rather than derived — it belongs with the physics of coherence, and what belongs here is its consequence, which is a way of thinking about telescopes that inverts the usual one. A filled aperture is not a light bucket; it is a device that measures every baseline within its own diameter, simultaneously, and adds them all together. A two-element interferometer measures exactly one of them, and does it as well as the filled aperture would have.

Two consequences follow at once, and both are counterintuitive on first meeting. The first is that the collecting area and the resolution of an array are entirely independent quantities: the area is the sum of the dishes and the resolution is set by the largest separation, and an array can be built with either one large and the other small. A single dish cannot do that, which is the whole reason the technique exists.

The second is that the visibility of an unresolved source is one, whatever the baseline. A point source has a flat transform, so every pair reports the same amplitude and the same phase; a source that is resolved is one whose visibility has begun to fall, and “resolved” therefore means something precise and measurable rather than “looks like more than a dot”. An interferometer that reports unity on its longest baseline has measured an upper limit on the source’s size and nothing else — which is a perfectly good result and is how most of the compact radio sky is catalogued.

The transform plane an array of 9 actually samples. Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. The 9 antennas make 36 pairs, each pair measures one point at any instant, and turning the Earth sweeps each of them along an ellipse — so eight hours of tracking turns 36 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian: a real sky forces V(−u,−v) = V(u,v), so half the points are free and the coverage is symmetric through the origin. And every ellipse has axis ratio exactly sin δ = 0.707* at this declination, measured off the longest track as 0.707 — an array is squashed in one direction by where the source is in the sky, and at the equator the tracks collapse to lines whatever the array. What the figure cannot show is the hole in the middle: no baseline is shorter than an antenna is wide, so the largest structures on the sky are simply not measured, and no processing recovers them.
Fig. 2 The same nine antennas at 22 GHz rather than 5, and the tracks have grown by the ratio of the wavelengths. Nothing about the array moved: the plane is measured in wavelengths, so a shorter wavelength puts the same physical baseline further out and the resolution improves in proportion. That is the cheapest upgrade an array has — it needs a receiver rather than a hole in the ground — and its cost is that the primary beam shrinks by the same factor, so the field of view falls as the square of the frequency and a survey takes sixteen times as long.

Why the tracks are ellipses

A baseline is fixed to the ground; the sky is not. As the Earth turns, the baseline’s projection onto the plane perpendicular to the line of sight changes, and that projection is what the visibility depends on.

Writing the baseline in equatorial components (Bx,By,Bz)(B_x, B_y, B_z) and the source’s hour angle as HH:

u=BxsinH+BycosH,v=BxsinδcosH+BysinδsinH+Bzcosδ.u = B_x\sin H + B_y\cos H, \qquad v = -B_x\sin\delta\cos H + B_y\sin\delta\sin H + B_z\cos\delta.

Those are the parametric equations of an ellipse, centred at (0,Bzcosδ)(0, B_z\cos\delta), with semi-axes Bx2+By2\sqrt{B_x^2+B_y^2} and that same quantity times sinδ\sin\delta. The declination squashes the track and nothing about the array can undo it. A source at the celestial equator gives a degenerate ellipse — a line — so an east–west array observing an equatorial source measures a one-dimensional slice of a two-dimensional transform — the same collapse of information that an orbit determined from angles alone suffers when the geometry degenerates, and the image it makes is smeared without limit in one direction.

That single fact decides where radio arrays are built and what they can observe. It is also why an array intended for the whole sky is laid out in two dimensions — a Y, a T, a spiral — rather than in the east–west line that would be simplest to survey.

The transform plane an array of 9 actually samples. Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. The 9 antennas make 36 pairs, each pair measures one point at any instant, and turning the Earth sweeps each of them along an ellipse — so eight hours of tracking turns 36 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian: a real sky forces V(−u,−v) = V(u,v), so half the points are free and the coverage is symmetric through the origin. And every ellipse has axis ratio exactly sin δ = 0.985* at this declination, measured off the longest track as 0.985 — an array is squashed in one direction by where the source is in the sky, and at the equator the tracks collapse to lines whatever the array. What the figure cannot show is the hole in the middle: no baseline is shorter than an antenna is wide, so the largest structures on the sky are simply not measured, and no processing recovers them.
Fig. 3 A source near the pole, where sinδ\sin\delta is nearly one and the tracks are almost circular. This is the best geometry an array ever gets: the ellipses are unsquashed, the coverage fills a rough disc, and the synthesised beam comes out nearly round. It is also why the deepest radio fields are at high declination and why an array’s advertised resolution is quoted for a source at the zenith — the number degrades continuously as the source moves south, and by the equator it has become a one-dimensional measurement.
The transform plane an array of 9 actually samples. Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. The 9 antennas make 36 pairs, each pair measures one point at any instant, and turning the Earth sweeps each of them along an ellipse — so eight hours of tracking turns 36 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian: a real sky forces V(−u,−v) = V(u,v), so half the points are free and the coverage is symmetric through the origin. And every ellipse has axis ratio exactly sin δ = 0.342* at this declination, measured off the longest track as 0.342 — an array is squashed in one direction by where the source is in the sky, and at the equator the tracks collapse to lines whatever the array. What the figure cannot show is the hole in the middle: no baseline is shorter than an antenna is wide, so the largest structures on the sky are simply not measured, and no processing recovers them.
Fig. 4 The same array on a source twenty degrees above the equator instead of forty-five. Every track has flattened by the ratio of the sines, sin20°/sin45°=0.48\sin 20°/\sin 45° = 0.48, and the coverage has become an ellipse of arcs rather than a rough disc of them. The resolution along the squashed direction is halved, and the beam is elongated by the same factor — which is why radio images of southern sources from a northern array, and vice versa, are quoted with an elliptical beam and a position angle rather than a single number.

The beam is the transform of the holes

An image is made by inverse-transforming the measured plane. But the plane has only been sampled where the baselines put points, so what is actually transformed is the true sky’s transform multiplied by a sampling function — and by the convolution theorem, the resulting image is the true sky convolved with the transform of that sampling function.

That transform is the synthesised beam, and it is what a point source looks like through the array.

The beam that incomplete sampling produces. A cut through the point-source response of the same array, formed by transforming the sampled plane and nothing else — no sky, no source, no noise. The narrow curve is eight hours of tracking and the broad one is a 12-minute snapshot of the same 36 pairs. Two numbers come out. The main lobe is 1.486″ across against λ/B_max = 1.429″, so the resolution is set by the single longest baseline and by nothing else in the array; and the worst sidelobe falls from 22% of the peak to 8% when the plane is filled in, which is the whole reason for tracking rather than snapping. The sidelobes are not an imperfection of the instrument. They are the transform of the holes, and a point source really is observed with this response; the negative rings are as real as the peak, and deconvolution is an attempt to guess what was in the holes rather than a way of measuring it.
Fig. 5 A cut through the point-source response, formed by transforming the sampled plane and nothing else — no sky, no source, no noise. The narrow curve is eight hours of tracking and the broad one a twelve-minute snapshot of the same pairs. The main lobe is 1.64″ across against λ/Bmax=1.43\lambda/B_{\max} = 1.43″, so the resolution is set by the single longest baseline and by nothing else in the array; the worst sidelobe falls from 26% of the peak to 9% when the plane is filled in, which is the entire reason for tracking rather than snapping. The sidelobes are not an imperfection of the instrument — they are the transform of the holes, and a point source really is observed with this response.

The distinction between the two numbers in that caption is the one worth keeping. The longest baseline sets the resolution; the filling of the plane sets the fidelity. An array of two enormous dishes has superb resolution and a beam with sidelobes at 100%, which makes an image of anything but a point source uninterpretable. An array of many modest dishes has a clean beam and a resolution set by whichever pair happens to be furthest apart.

The transform plane an array of 16 actually samples. Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. The 16 antennas make 120 pairs, each pair measures one point at any instant, and turning the Earth sweeps each of them along an ellipse — so eight hours of tracking turns 120 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian: a real sky forces V(−u,−v) = V(u,v), so half the points are free and the coverage is symmetric through the origin. And every ellipse has axis ratio exactly sin δ = 0.707* at this declination, measured off the longest track as 0.707 — an array is squashed in one direction by where the source is in the sky, and at the equator the tracks collapse to lines whatever the array. What the figure cannot show is the hole in the middle: no baseline is shorter than an antenna is wide, so the largest structures on the sky are simply not measured, and no processing recovers them.
Fig. 6 Sixteen antennas instead of nine, on four arms instead of three. The pair count goes as N(N1)/2N(N-1)/2, so the jump from nine to sixteen takes 36 baselines to 120 — a factor of three and a third for a factor of under two in antennas — and the plane fills in accordingly. The longest baseline has not changed, so the resolution has not changed at all; what has changed is the fidelity, which is exactly the distinction the paragraph above draws. Adding antennas buys a cleaner beam and adding distance buys a narrower one, and no array can do both with the same money.

What was actually measured

Three measurements, in three eras, each of which needed the previous one’s technique to be believed.

A stellar diameter, 1920. Michelson and Pease mounted a six-metre beam with movable mirrors on the 100-inch Hooker telescope, and observed Betelgeuse until the interference fringes vanished. Fringe disappearance is the first null of the visibility curve, and the baseline at which it happened, 3.07 m, gives 47 milliarcseconds. That was the first measurement of the size of any star other than the Sun, and it is still how the answer is obtained a century later.

A galaxy’s structure, 1946 onwards. Ryle’s group in Cambridge developed the technique of moving one antenna between observations and combining them, which turns two dishes into an arbitrarily dense sampling of the plane over the course of weeks. That is aperture synthesis proper, and the Nobel citation in 1974 says so.

A black hole’s shadow, 2019. The Event Horizon Telescope correlated signals from eight stations on four continents at 1.3 mm, giving baselines up to 10,700 km and a beam of about 20 microarcseconds. The coverage was appalling by radio standards — eight stations give 28 pairs, and weather removed several of them — which is exactly why the published image required four independent imaging teams working blind, and why the ring’s diameter is quoted with far more confidence than any detail within it.

The one knob after the observation is over

There is a decision left after all the data are taken, and it is the clearest illustration that an interferometric image is a construction rather than a photograph.

Each measured visibility can be given a weight before the transform, and two choices are standard. Natural weighting gives every measurement equal weight, which maximises sensitivity — most of the samples are on short baselines, so the beam comes out broad and the faintest sources are detected. Uniform weighting gives each region of the plane equal weight regardless of how many samples fell in it, which up-weights the sparse long baselines, sharpens the beam, and throws away signal-to-noise.

The same data therefore yield an image with a 1″ beam and poor sensitivity, or a 2″ beam and good sensitivity, and both are correct. The resolution of a radio image is chosen after the fact, within limits set by the longest baseline at one end and the shortest at the other, and any figure quoting a beam size is quoting a decision as much as an instrument.

Where the model stops

Three limits, and the first two are absolute.

The short baselines are missing. No pair of antennas can be closer than one antenna is wide, so the centre of the plane is a hole, and the largest angular scales on the sky are simply not measured. An interferometer is blind to smooth extended emission — it does not measure it badly, it does not measure it at all — and no processing recovers what was never sampled. The usual remedy is to observe the same field with a single dish and add its measurements in, which is a different instrument answering a different question.

Deconvolution is a guess. Algorithms like CLEAN work by assuming the sky is made of point sources and iteratively removing the beam’s response from the image. That assumption is a prior, it is often reasonable, and it is not a measurement. The honest statement about any synthesis image is that it is one sky consistent with the sampled visibilities, and the ones the holes could hide have been excluded by assumption.

The beam that incomplete sampling produces. A cut through the point-source response of the same array, formed by transforming the sampled plane and nothing else — no sky, no source, no noise. The narrow curve is eight hours of tracking and the broad one is a 60-minute snapshot of the same 36 pairs. Two numbers come out. The main lobe is 1.486″ across against λ/B_max = 1.429″, so the resolution is set by the single longest baseline and by nothing else in the array; and the worst sidelobe falls from 21% of the peak to 8% when the plane is filled in, which is the whole reason for tracking rather than snapping. The sidelobes are not an imperfection of the instrument. They are the transform of the holes, and a point source really is observed with this response; the negative rings are as real as the peak, and deconvolution is an attempt to guess what was in the holes rather than a way of measuring it.
Fig. 7 What an hour buys against a twelve-minute snapshot, on the same nine antennas. The main lobe is unchanged — it is set by the longest baseline and by nothing else — and the sidelobes have come down, which is the entire product of the extra time. Every one of those sidelobes is a place where a real source at the wrong position contributes to the flux measured at the right one, so deconvolution’s difficulty scales with their height. A beam with 26 per cent sidelobes needs a strong prior to invert; one with 9 per cent needs much less of one, and the difference between the two is tracking rather than hardware.

Phase is fragile. The atmosphere, the clocks and the cable lengths all corrupt the phase of a visibility, and the amplitude is far more robust. Much of radio interferometry’s machinery — closure phases, self-calibration — exists to construct quantities that survive that corruption, and the EHT image rests on closure quantities precisely because absolute phase across ten thousand kilometres is unrecoverable. Two pieces of machinery have been named in passing and not explained, and both are load-bearing enough to deserve it: the construction that makes a phase usable at all, and the arrangement that lets antennas on different continents interfere without being connected.

The quantity that survives a broken clock

The essay’s last limitation was that phase is fragile, and the technique that rescues it deserves stating rather than naming, because it is the reason intercontinental interferometry produces images at all.

Take three antennas and measure the three visibility phases around the triangle. Each measured phase is the true phase plus an error contributed by the antenna at each end — an atmospheric delay, a clock offset, a cable length — and each of those errors belongs to one antenna rather than to a pair.

Now add the three measured phases around the closed triangle. Every antenna appears twice, once with each sign, so every antenna-based error cancels exactly. What is left is the sum of the three true phases, uncorrupted, and it is called the closure phase.

The cancellation is exact and it is not statistical: it does not require the errors to be small, or random, or stationary. It requires only that they attach to antennas rather than to baselines, which is what an atmospheric delay and a clock offset both do.

The price is that closure recovers less than all of the phase information. An array of NN antennas has N(N1)/2N(N-1)/2 baselines and therefore that many measured angles, of which N1N-1 are consumed by the unknown per-antenna errors, leaving (N1)(N2)/2(N-1)(N-2)/2 independent closure phases. So the fraction of the phase information that survives is (N2)/N(N-2)/N — a third for three antennas, a half for four, and nine tenths for twenty.

That fraction is why an array’s usefulness grows faster than its number of baselines. A three-element array can measure one closure phase and cannot make an image; a large array loses almost nothing to calibration and can.

A quantity constructed to be immune to a class of error is worth more than a better measurement of the corrupted one, and closure is the clearest example of that construction in observational astronomy.

Correlating a signal nobody sent

The description of an interferometer as combining light from two apertures is right for an optical instrument, where the beams are physically brought together. At radio wavelengths nothing is combined optically, and the difference is worth setting out because it decides what is possible.

Each antenna receives a voltage — a noise-like waveform, because a cosmic radio source emits thermal or synchrotron radiation with no structure in it. That voltage is amplified, mixed down, digitised and timestamped against a hydrogen maser, and then recorded.

The interference happens afterwards, in a computer. The two recordings are cross-correlated as a function of relative delay, and the peak of that correlation is the visibility. Nothing about the process requires the two antennas to have been connected while observing, or even to have known about each other.

That is what makes intercontinental baselines possible. Antennas in Chile, Hawaii, Spain and Antarctica cannot be joined by any cable, and they do not need to be: they need clocks stable enough that the relative timing is known to a fraction of a wavelength period across the observation, and a way of moving the recordings to one place.

The data rates are the practical limit. Sensitivity goes as the square root of the bandwidth, so a modern experiment records several gigabits per second per station, and a few days of observing fills hundreds of terabytes. Those disks are shipped physically to the correlator, which is slow and is still faster than any network to a telescope at the South Pole — where, for part of the year, the disks cannot be flown out at all.

The transform plane an array of 9 actually samples. Every point in this plane is a spatial frequency the array has measured, in thousands of wavelengths. The 9 antennas make 36 pairs, each pair measures one point at any instant, and turning the Earth sweeps each of them along an ellipse — so eight hours of tracking turns 36 measurements into the arcs drawn here. Two properties are structural rather than chosen. The plane is Hermitian: a real sky forces V(−u,−v) = V(u,v), so half the points are free and the coverage is symmetric through the origin. And every ellipse has axis ratio exactly sin δ = 0.707* at this declination, measured off the longest track as 0.707 — an array is squashed in one direction by where the source is in the sky, and at the equator the tracks collapse to lines whatever the array. What the figure cannot show is the hole in the middle: no baseline is shorter than an antenna is wide, so the largest structures on the sky are simply not measured, and no processing recovers them.
Fig. 8 The same nine antennas spread four times further apart, which is the only thing that improves resolution and is also the thing a cable cannot reach across. The tracks move out by the same factor and the centre of the plane empties by it, so the array becomes blind to precisely the large angular scales it could see before. That trade is why intercontinental interferometry produces images of compact sources and nothing else: the recording arrangement makes the long baselines possible, and the long baselines make the short ones missing.

An interferometer is a recording instrument rather than a combining one, and the consequence is that its baseline is limited by the size of the planet and its bandwidth by how fast bits can be written.

The recording arrangement also decides what an array can be built out of: any two antennas with adequate clocks are a valid pair, whatever else they were built for, which is why intercontinental experiments are assembled from telescopes that spend the rest of the year working alone.

The limit the planet imposes

The longest baseline available on the ground is one Earth diameter, and at a given wavelength that is a hard ceiling on resolution. Two ways past it exist and both have been tried.

The first is to shorten the wavelength. Resolution goes as wavelength over baseline, so moving from centimetres to millimetres buys an order of magnitude with the same array. The cost is that the atmosphere becomes opaque: water vapour absorbs strongly at millimetre wavelengths, so the stations have to be high and dry, and the useful weather is a small fraction of the year.

The second is to put an antenna in space. A satellite in a high elliptical orbit provides baselines several Earth diameters long, and two missions have flown — one in the 1990s reaching about three Earth diameters, one later reaching about eight.

The difficulties are instructive. The satellite’s position has to be known to a fraction of a wavelength, which for centimetre observing is centimetres, so the orbit must be tracked far better than an ordinary spacecraft’s. Its clock has to be as stable as a ground maser, or else phase-transferred from the ground continuously. And the data have to be sent down in real time, because a spacecraft cannot carry the recording capacity.

What such an array gains is resolution and what it loses is sensitivity, because one small antenna paired with a large ground dish gives a baseline whose noise is set by the smaller of the two. So space interferometry has measured the sizes of the very brightest compact sources and has not made images of anything faint.

The technique’s resolution is bounded by geometry and its sensitivity by collecting area, and no single change improves both.

Where this ladder goes next

This rung establishes what an array measures and what it does not. The rungs above it follow the technique into the places it has gone.

The nearest is closure: the sum of visibility phases round a triangle of antennas is independent of any per-antenna error, which turns an unusable set of phases into a usable one and is why an image can be made at all across intercontinental baselines. It is a small piece of algebra with an enormous consequence, and it deserves its own rung.

Beyond it lies optical interferometry, where the fringes must be tracked in milliseconds and the light must be piped through delay lines kilometres long to keep two paths equal to within a coherence length. The instruments that do it now measure the orbits of stars around the Galactic centre’s black hole to tens of microarcseconds — which is a mass measurement obtained by the technique in this essay, applied to the smallest angle anybody has yet had a use for.

What this makes readable

Essays that name this one as a prerequisite.

What links here

The 8 of 18 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

Angular resolutionAperture synthesisBaselineDeconvolutionFringeSidelobesSynthesised beamUv planeVan cittert zernikeVisibility