Starlight

The direction a photon count throws away

A photometer records how many photons arrived. It discards a two-component quantity that survives every attenuation on the way, and that quantity carries a magnetic field direction, a grain size, and the shape of an exploding star nobody can resolve.

Assumes Extinction, Photometric systems and Photon noise.

Light arriving from a star carries four independent quantities. A photometer measures one of them.

The other three describe the orientation of the electric field’s oscillation — whether it prefers one plane, whether it rotates, and how strongly. They are usually small, they are difficult to measure, and they survive the journey untouched by anything that merely attenuates. Where a brightness is degraded by distance, dust and atmosphere, a polarisation is not: it is a ratio, and the things that dim the light dim both components of the ratio together.

That makes polarimetry the only channel in optical astronomy that carries a geometry rather than a scalar, and there are objects for which it is the only channel at all.

A peak at 0.55 µm, and therefore a grain size. Interstellar polarisation against wavelength — the Serkowski law, p(λ) = p_max exp[−K ln²(λ_max/λ)] with K = 1.66 λ_max. The heavy curve peaks at 0.550 µm, measured off the drawing rather than read back from the parameter, and falls to half its peak at 1.314 µm on the red side, against the closed form λ_max exp√(ln2/K) = 1.315. That peak wavelength is the measurement. It is set by the size of the grains doing the aligning — bigger grains, longer λ_max — and it is tied to the shape of the extinction curve along the same sight line by R_V ≈ 5.5 λ_max, which gives 3.03 here against the diffuse-medium value of 3.1. The two faint curves are populations peaking at 0.35 µm and 0.75 µm: the same amount of polarisation, distributed differently, and a different dust. Nothing in a photometric measurement of the same star distinguishes them.
Fig. 1 Interstellar polarisation against wavelength — the Serkowski law, p(λ)=pmaxexp[Kln2(λmax/λ)]p(\lambda) = p_{\max}\exp[-K\ln^2(\lambda_{\max}/\lambda)] with K=1.66λmaxK = 1.66\lambda_{\max}. The heavy curve peaks at 0.55 µm, measured off the drawing rather than read back from the parameter, and falls to half its peak at 1.31 µm on the red side against a closed form of 1.31. That peak wavelength is the measurement: it is set by the size of the grains doing the aligning, and it is tied to the shape of the extinction curve along the same sight line by RV5.5λmaxR_V \approx 5.5\lambda_{\max}. The two faint curves are populations peaking at 0.35 and 0.75 µm — the same amount of polarisation, distributed differently, and a different dust.

What the four numbers are

Stokes’ 1852 parameterisation is the one that survived, because its components add when incoherent beams are combined. Written as intensities:

I=total flux,Q=I0I90,U=I45I135,V=IRHILH.I = \text{total flux},\quad Q = I_{0^\circ} - I_{90^\circ},\quad U = I_{45^\circ} - I_{135^\circ},\quad V = I_{\rm RH} - I_{\rm LH}.

The linear polarisation is p=Q2+U2/Ip = \sqrt{Q^2+U^2}/I and its position angle is θ=12arctan(U/Q)\theta = \tfrac12\arctan(U/Q). The factor of a half is not a convention: a rotation of the analyser by 90° exchanges QQ’s two intensities, which is a rotation by 180° in the (Q,U)(Q,U) plane. Polarisation has no sign and no direction, only an orientation, and everything odd about handling it follows from that.

The circular component VV is almost always negligible in astronomy and almost always the interesting one where it is not — it appears in the Zeeman-split lines of magnetic stars and in the synchrotron emission of a few sources, and measuring it well enough to be believed is harder than everything else in this essay put together.

A position angle to 1.7° and a polarisation that is biased upwards. The Stokes plane, with 260 simulated measurements of one star at a true polarisation of 0.6 per cent and a position angle of 62°, each Stokes parameter carrying an independent error of 0.35 per cent. The polarisation is the length of the vector and the position angle is half its azimuth — which is why a rotation of 180° in this plane is a rotation of 90° on the sky, and why polarisation has no sign. Averaging the two components recovers 60.3° against the true 62°. But averaging the lengths gives 0.686 per cent against a true 0.6: a length cannot be negative, so noise can only push it up, and the excess here is 0.112 against the large-signal expectation σ²/2p = 0.102. At zero true polarisation the same effect returns 0.44 per cent from a star that has none, which is the reason a polarimetric detection is quoted in σ and almost never in per cent alone.
Fig. 2 The plane a measurement actually lands in, with 260 simulated observations of a star at 0.6 per cent polarisation and position angle 62°, each Stokes parameter carrying an independent error of 0.35 per cent. Averaging the two components recovers 63.7° against the true 62. Averaging the lengths gives 0.71 per cent against the true 0.6: a length cannot be negative, so noise can only push it up, and at zero true polarisation the same effect returns 0.44 per cent from a star that has none. That is why a polarimetric detection is quoted in standard deviations and almost never in per cent alone.

Where the polarisation comes from

Two mechanisms produce nearly all the linear polarisation seen in optical astronomy, and they have opposite geometries.

Scattering polarises the light perpendicular to the plane containing the source, the scatterer and the observer. For a particle small compared with the wavelength the degree of polarisation is

p(θ)=1cos2θ1+cos2θ,p(\theta) = \frac{1-\cos^2\theta}{1+\cos^2\theta},

exactly zero in the forward and backward directions and exactly one at a right angle. Neither of those is a fitted number: they are what a dipole radiator looks like end-on and side-on.

Zero forwards, zero backwards, and exactly one at a right angle. The degree of linear polarisation produced by a single scattering, against the angle through which the light was turned. The curve is (1 − cos²θ)/(1 + cos²θ): it is exactly zero at 0° and 180° and exactly one at 90°, and neither of those is a fitted number — they are the geometry of a dipole seen end-on and side-on. The lower curves are the same shape divided down by an unpolarised component, for slabs of optical depth 0.2, 1, 3 in which some of the light has scattered more than once; the peak falls to 16 per cent of its single-scattering value at the thickest. The angular position of the maximum does not move, which is why a polarisation map of a reflection nebula locates the illuminating star even when the star is hidden: every position angle is perpendicular to the line back to the source, and the pattern converges on it.
Fig. 3 The scattering curve, with three families beneath it for slabs in which some of the light has scattered more than once. Multiple scattering adds an unpolarised component and divides the whole curve down, taking the peak to a third of its single-scattering value at an optical depth of three — and the angular position of the maximum does not move. That is why a polarisation map of a reflection nebula locates the illuminating star even when the star itself is hidden: every position angle is perpendicular to the line back to the source, so the vectors form a centrosymmetric pattern that converges on it.

Dichroic absorption by aligned grains polarises the light along the direction the grains are aligned by, which is the magnetic field. Interstellar grains are elongated, they spin, and they align with their long axes perpendicular to the field — so they absorb preferentially the component of the electric field along their long axis, and what gets through is preferentially the component along the field.

The mechanism of the alignment took fifty years to settle and is still not entirely closed. Davis and Greenstein proposed paramagnetic relaxation in 1951; the modern account is radiative torques, in which an irregular grain in an anisotropic radiation field is spun up to suprathermal rotation and its angular momentum then aligns with the field. What matters for the observation is that the alignment happens, that it happens with the long axis perpendicular to B, and that the transmitted polarisation is therefore parallel to it.

A map of a field nobody can see

The Galactic magnetic field is a few microgauss. It threads the whole disc, it confines the cosmic rays, it resists the collapse of molecular clouds, and it is invisible to every direct measurement optical astronomy has.

Polarimetry of background stars maps it. Each star gives one number — the position angle of the polarisation its light acquired crossing the intervening dust — and that angle is the field direction projected on the sky, averaged along the line of sight and weighted by the dust. Thousands of such measurements make a vector field.

The picture that came out, first from Hiltner and Hall in 1949 and in modern form from Planck’s submillimetre polarimetry, is a field running predominantly along the Galactic plane, with a coherent large-scale component and a turbulent one of comparable strength. The ratio between those two is a measurement of how magnetically dominated the interstellar medium is, and it is obtained by dispersion of the position angles rather than by measuring any field. The peak of the Serkowski curve is the one number in it that carries physics, and moving it is the whole of what the law is used for.

A peak at 0.35 µm, and therefore a grain size. Interstellar polarisation against wavelength — the Serkowski law, p(λ) = p_max exp[−K ln²(λ_max/λ)] with K = 1.66 λ_max. The heavy curve peaks at 0.350 µm, measured off the drawing rather than read back from the parameter, and falls to half its peak at 1.043 µm on the red side, against the closed form λ_max exp√(ln2/K) = 1.043. That peak wavelength is the measurement. It is set by the size of the grains doing the aligning — bigger grains, longer λ_max — and it is tied to the shape of the extinction curve along the same sight line by R_V ≈ 5.5 λ_max, which gives 1.92 here against the diffuse-medium value of 3.1. The two faint curves are populations peaking at 0.35 µm and 0.75 µm: the same amount of polarisation, distributed differently, and a different dust. Nothing in a photometric measurement of the same star distinguishes them.
Fig. 4 The same law peaked at 0.35 microns instead of 0.55. A peak this far to the blue means the aligned grains are smaller than average, which happens along sight lines through diffuse gas where the largest grains have been destroyed. The wavelength of the peak is a grain size, read off a curve with no imaging of any kind.
A peak at 0.8 µm, and therefore a grain size. Interstellar polarisation against wavelength — the Serkowski law, p(λ) = p_max exp[−K ln²(λ_max/λ)] with K = 1.66 λ_max. The heavy curve peaks at 0.800 µm, measured off the drawing rather than read back from the parameter, and falls to half its peak at 1.648 µm on the red side, against the closed form λ_max exp√(ln2/K) = 1.648. That peak wavelength is the measurement. It is set by the size of the grains doing the aligning — bigger grains, longer λ_max — and it is tied to the shape of the extinction curve along the same sight line by R_V ≈ 5.5 λ_max, which gives 4.40 here against the diffuse-medium value of 3.1. The two faint curves are populations peaking at 0.35 µm and 0.75 µm: the same amount of polarisation, distributed differently, and a different dust. Nothing in a photometric measurement of the same star distinguishes them.
Fig. 5 And peaked at 0.8 microns, which is what a dense cloud gives. Grains there have grown ice mantles and coagulated, so the size distribution shifts upward and the polarisation peak moves with it. The same three-parameter curve therefore distinguishes diffuse from dense material along a line of sight that is entirely unresolved.

A shape from an unresolved point

The second use of polarimetry is the one that produces information no telescope can otherwise obtain.

A spherically symmetric source that scatters its own light produces no net polarisation, however strongly each individual photon is polarised: the contributions from opposite sides of the source cancel exactly. Break the symmetry and they no longer cancel, and the residual is a measure of the departure from spherical.

For a supernova, that is the whole of what is known about its geometry. A supernova at ten megaparsecs subtends a few microarcseconds; nothing resolves it, and nothing will. What is measured is a continuum polarisation of a few tenths of a per cent, and the modelling that turns it into an axis ratio is straightforward enough to be believed: a per cent of continuum polarisation corresponds to an asphericity of about ten per cent.

The results have been decisive. Core-collapse supernovae are systematically aspherical, and the asphericity increases as the ejecta thin and deeper layers are exposed — so the explosion mechanism itself is aspherical rather than the envelope being disturbed on the way out. Type Ia supernovae are nearly spherical in the continuum, at a few tenths of a per cent, which is a constraint on their explosion models that no other observation supplies. There is a third use, smaller in scope and unusually direct. A planet reflecting its star’s light is polarised by the scattering in its atmosphere, and the polarisation varies through the orbit exactly as the scattering angle does — zero at full phase, maximum near quadrature. That is a signal whose shape in time is fixed by geometry alone, and it can be recovered from a light curve in which the planet is never separated from the star.

Why it is rare

Every polarimetric measurement is a difference of two nearly equal counts, and that is the whole of the difficulty.

To measure a polarisation of one per cent to a tenth of its own value requires distinguishing two intensities that differ by one part in a hundred, to one part in a thousand. From counting statistics alone that needs 10610^6 photons; in practice it needs far more, because the systematic errors are larger than the statistical ones and they do not average down. The instrumental answer is to make the measurement differential in as many ways as possible. A dual-beam polarimeter splits the incoming light into two orthogonal polarisations and records both simultaneously on the same detector, so a fluctuation in transparency affects both equally and cancels in the ratio. Rotating a half-wave plate through four positions and combining the eight resulting beams in the right order cancels, to first order, both the difference in the two channels’ efficiencies and the difference between the two detector regions. The instrument is designed so that the quantity of interest survives four cancellations and the systematics do not.

That is why polarimetry is done on a minority of telescopes by a minority of observers, and why a polarimetric result is generally presented with a longer description of its calibration than of its interpretation.

The bias, and how it is handled

Return to the positive bias, because it is the part of polarimetry most often got wrong and it has a clean structure.

The measured p=q2+u2p = \sqrt{q^2+u^2} is the length of a two-dimensional vector whose components carry Gaussian errors. Its distribution is the Rice distribution, and its mean exceeds the true value for every true value. At high signal-to-noise the excess tends to σ2/2p\sigma^2/2p, which is negligible. At zero true polarisation the mean is σπ/2=1.25σ\sigma\sqrt{\pi/2} = 1.25\sigma, which is not.

The practical consequences are worth stating as rules, because they are the difference between a result and an artefact.

A measurement with p/σ<3p/\sigma < 3 is not a detection of polarisation. It is consistent with zero, and quoting the measured pp as a value overstates it.

Debiasing estimators exist — the simplest subtracts σ2\sigma^2 in quadrature — and they are useful in the middle range and misleading at the bottom, where they return zero for a genuine small polarisation as readily as for none.

And the position angle is well behaved where the polarisation is not: θ\theta has no bias at all, because the two components’ errors are symmetric about the truth. A weak polarisation whose angle is consistent across four measurements is more convincing than a strong one measured once, and the angle is the quantity most of the astrophysics is carried by. Both applications share a difficulty that is worth naming once rather than twice: what is measured is an integral along a line of sight, so a polarisation is a weighted average over everything between here and the source, and untangling a foreground from a signal is the same operation in both cases.

Where the model stops

Three limits, and each is a live research question rather than a settled correction.

The alignment efficiency of grains is not known independently. The Serkowski law is empirical; the constant relating its peak to RVR_V is a fit; and the fraction of grains aligned varies with environment, falling in dense cloud interiors where the radiation that drives the alignment cannot penetrate. So a polarisation map of a molecular cloud under-represents the field exactly where the field matters most.

The scattering geometry is degenerate in a way that a single measurement cannot lift. The same net polarisation can come from a moderately aspherical source seen edge-on or a strongly aspherical one seen obliquely, and separating them requires either a time series — as the geometry changes — or a wavelength dependence.

And the interstellar contribution has to be removed before anything intrinsic can be claimed. That is done by measuring nearby stars along the same sight line, which assumes the dust between here and them is the same dust as between here and the target. For a supernova in another galaxy that assumption spans two interstellar media and an intergalactic one, and it is the single largest uncertainty in most published supernova polarimetry.

A slope, not an angle. Above: the polarisation angle of a background source against the square of the observing wavelength, at the five wavelengths a radio survey actually uses. The plane rotates as it passes through magnetised plasma, by an amount proportional to the integral of the electron density times the field along the path — and to λ². The intercept is the angle the source emitted at, which nobody knows, so a measurement at one wavelength contains no information whatever; the slope is the rotation measure, here 42 radians per square metre, and it needs no knowledge of the source at all. Below: the fractional polarisation that survives. A telescope beam covers many lines of sight with slightly different rotation measures, their angles disagree by more at longer wavelengths, and the vector sum collapses — which is why the useful band has a long-wavelength edge that has nothing to do with sensitivity. Divide the rotation measure by the dispersion measure of the same path, 26.8 in the usual units, and the electron density cancels: the mean line-of-sight field is 1.93 µG, obtained without knowing the distance, the density, or where along the path the field was.
Fig. 6 The measurement that needs the direction and nothing else. A magnetised plasma rotates the plane of polarisation by an amount proportional to the square of the wavelength, so measuring the angle at several wavelengths and fitting a line gives the rotation measure — which is the electron density and the line-of-sight field integrated together. Divide by the dispersion measure, which is the electron density alone, and what is left is the mean field along the path. Nothing in a photon count carries that; it is entirely in the angle a total-intensity measurement discards.

Two applications are worth setting out before the ladder, because they are where the technique is currently doing the most work — one on a field that fills the Galaxy, and one on a signal from before there was a galaxy.

Two more settings are worth putting beside those, because each shows a different way the measurement fails.

Zero forwards, zero backwards, and exactly one at a right angle. The degree of linear polarisation produced by a single scattering, against the angle through which the light was turned. The curve is (1 − cos²θ)/(1 + cos²θ): it is exactly zero at 0° and 180° and exactly one at 90°, and neither of those is a fitted number — they are the geometry of a dipole seen end-on and side-on. The lower curves are the same shape divided down by an unpolarised component, for slabs of optical depth 0.1, 0.5, 8 in which some of the light has scattered more than once; the peak falls to 0 per cent of its single-scattering value at the thickest. The angular position of the maximum does not move, which is why a polarisation map of a reflection nebula locates the illuminating star even when the star is hidden: every position angle is perpendicular to the line back to the source, and the pattern converges on it.
Fig. 7 Single scattering against optical depth, with one case far into the multiply-scattering regime. At an optical depth of eight the polarisation is largely washed out: each photon has been turned several times and the position angles average against one another. So a polarisation measurement of an optically thick object measures the surface it escaped from and not the geometry inside it.
A position angle to 0.7° and a polarisation that is biased upwards. The Stokes plane, with 260 simulated measurements of one star at a true polarisation of 1.5 per cent and a position angle of 62°, each Stokes parameter carrying an independent error of 0.35 per cent. The polarisation is the length of the vector and the position angle is half its azimuth — which is why a rotation of 180° in this plane is a rotation of 90° on the sky, and why polarisation has no sign. Averaging the two components recovers 61.3° against the true 62°. But averaging the lengths gives 1.514 per cent against a true 1.5: a length cannot be negative, so noise can only push it up, and the excess here is 0.041 against the large-signal expectation σ²/2p = 0.041. At zero true polarisation the same effect returns 0.44 per cent from a star that has none, which is the reason a polarimetric detection is quoted in σ and almost never in per cent alone.
Fig. 8 The Stokes plane for a star polarised at 1.5 per cent rather than 0.6, with the same noise. The cloud of measurements is now well away from the origin, the position angle is recovered to better than a degree, and the upward bias in the polarisation degree has become negligible. The bias is not a property of the instrument; it is a property of how close the signal is to zero.

Three measurements, three components

The essay has described one probe of the interstellar magnetic field, and it measures one thing: the direction of the field projected on the plane of the sky. A field is a vector, and getting the other components requires two further techniques that share nothing with this one.

Faraday rotation measures the line-of-sight component. A linearly polarised radio wave passing through a magnetised plasma has its plane of polarisation rotated, by an amount proportional to the wavelength squared and to the integral along the path of the electron density times the field component along the line of sight. Observing a background source at several wavelengths and fitting the rotation against wavelength squared gives that integral.

What it returns is a product of two unknowns — the field and the electron density — so converting it to a field strength requires a separate measurement of the density, usually from the dispersion of a pulsar’s pulses along the same sight line. Combining the two gives a density-weighted mean field along the path, with a sign: toward the observer or away.

The Zeeman effect measures the line-of-sight component again, and directly. A spectral line formed in a magnetic field is split, and the splitting is proportional to the field; for the twenty-one-centimetre line of neutral hydrogen the splitting is far smaller than the line width, so what is measured is not a split line but a small circular polarisation whose profile is the derivative of the line’s. It is one of the hardest measurements in radio astronomy and it is the only one that returns a field strength with no other quantity multiplied into it.

So the three techniques divide the vector between them: dust polarisation gives the plane-of-sky direction, Faraday rotation gives the line-of-sight component weighted by electrons, and Zeeman gives the line-of-sight component weighted by neutral hydrogen. Each is sensitive to a different phase of the medium, so combining them is not a matter of adding components of one field — it is a matter of comparing fields measured in different gas.

The three-dimensional field is assembled from measurements that do not overlap, and the assembly is the least secure part of every published field map.

The division above also explains why field strengths quoted for the interstellar medium vary so much between papers: they are frequently measurements of different phases by different techniques, presented as though they were the same quantity.

The same scattering, at recombination

The mechanism at the start of this essay — scattering polarises light perpendicular to the scattering plane — has a second application at a completely different scale, and it is the one the largest experiments in cosmology are built for.

At the moment the universe became transparent, the last photons to scatter off free electrons were scattered from a radiation field that was not perfectly isotropic. A quadrupole anisotropy in the incident radiation produces a net linear polarisation in the scattered light, for exactly the reason a scattering angle of ninety degrees produces one: the two orthogonal components of the incident field are not equal.

So the microwave background is polarised, at a level of a few per cent of its temperature anisotropies, and the pattern carries information the temperature alone does not.

The pattern is decomposed into two kinds. One has no curl and is produced by density perturbations, which is most of what is there. The other has a curl and cannot be produced by density perturbations at all, so its presence would be evidence for something else — gravitational waves from the very early universe being the sought-after source.

The difficulty is the foreground, and the foreground is the subject of this essay. Aligned interstellar grains in this galaxy emit polarised light in the far infrared and microwave, at exactly the frequencies the background is measured at, with a polarisation fraction of several per cent and a pattern that contains both kinds.

Separating them requires observing at several frequencies and exploiting the different spectra — the background’s is a blackbody and the dust’s is a modified one — which is why every experiment of this kind is multi-frequency by design.

A claimed detection of the curl component in 2014 turned out, on the release of better dust maps, to be consistent with Galactic dust alone. The measurement that would say something about the first fraction of a second is limited by the alignment of grains a few hundred parsecs away, and improving it means understanding the dust rather than the background.

Where this ladder goes next

This rung establishes what a polarisation is, the two mechanisms that produce one, and the two failure modes — instrumental and statistical — that make it hard.

Above it lies spectropolarimetry, where the polarisation becomes a function of wavelength and the loops traced in the Stokes plane across a spectral line carry the geometry of the line-forming region.

Beside it lies the circular component, which is nearly always the Zeeman effect and which measures the line-of-sight field strength directly rather than the plane-of-sky direction — the two together giving a three-dimensional field from a measurement made on a point.

And below it, as a caution worth its own rung, lies the whole business of what a small measured quantity means when its estimator cannot be negative. That is not a fact about polarisation; it is a fact about lengths, and it recurs wherever a modulus is measured — in a proper motion, in a lens magnification, in an amplitude of any kind.

What this makes readable

Essays that name this one as a prerequisite.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

The 8 of 11 essays linking to this one that name the most of the same objects.

The objects this essay names

Each one links to every other essay that touches it.

AsphericityGrain alignmentInterstellar dustMagnetic fieldPolarimetric biasPolarisationPosition angleReflection nebulaRice distributionScatteringSerkowski lawStokes parameters