Starlight

An instrument more polarised than the sky

Every oblique reflection polarises. A telescope is a stack of oblique reflections, so it adds a fraction of a per cent of polarisation to everything it looks at — which for most astronomical sources is more than they have themselves, and which is a vector rather than a scale error, so it rotates the answer as well as changing its size.

Assumes Polarimetry and Photometric systems.

Polarimetry measures a fraction of a per cent. Interstellar dust polarises starlight by a few per cent at most and usually much less; a scattering atmosphere gives tenths of a per cent; the alignment signal from a magnetic field in a molecular cloud is at the same level. These are small numbers by the standards of any other astronomical measurement, and they are not small compared with what the telescope contributes.

An oblique reflection from a metal surface reflects light polarised parallel to the plane of incidence differently from light polarised perpendicular to it. Every mirror at a non-zero angle of incidence therefore polarises, and a telescope with a tertiary mirror or a Nasmyth focus can add a per cent or more.

An instrument 1.4 times as polarised as the sky it is measuring. The Stokes plane, with the two linear polarisation parameters as axes. The tight cluster near the centre is a set of stars known to be unpolarised, observed through the same instrument: they should sit at the origin and do not, and their mean is the instrumental polarisation — 0.77 per cent here, which is 1.4 times the real polarisation of the field. The other cluster is the field stars, whose measured values are the sum of their own polarisation and the same instrumental offset. Subtracting one mean from the other recovers 0.59 per cent at the right angle. Two things make this worth doing carefully. The offset is a vector, so leaving it in rotates the measured position angle as well as changing its magnitude — by 31 degrees here — and a calibration that only fixes the scale does not touch that. And the offset depends on where the telescope was pointing, because the reflection angles do, so the standards have to be observed at the same place in the sky and at the same instrument rotation.
Fig. 1 The Stokes plane, with the two linear polarisation parameters as axes. The tight cluster near the centre is a set of stars known to be unpolarised, observed through the same instrument: they should sit at the origin and do not. Their mean is the instrumental polarisation, which here exceeds the real polarisation of the field. Subtracting one cluster’s mean from the other recovers the sky’s own signal.

Polarisation is the piece of information a photon count throws away, and recovering it means measuring a difference between two intensities that are nearly equal. That structure is what makes the technique powerful and what makes it vulnerable: a difference of two large numbers inherits every systematic that affects the two unequally.

Why it is a vector and not a scale

The single most important structural fact about the problem is that instrumental polarisation adds rather than multiplies.

Polarisation is described by two numbers, conventionally QQ and UU, which are differences of intensities measured in perpendicular directions. They combine linearly: light that is the sum of two beams has Stokes parameters that are the sum of theirs. So an instrument that adds its own polarised component adds a fixed vector (Qinst,Uinst)(Q_{\rm inst}, U_{\rm inst}) to whatever arrived.

The degree of polarisation and the position angle are the modulus and the argument of that vector. Adding a vector changes both. An instrumental contribution of half a per cent applied to a source with half a per cent of its own can double the measured degree, halve it, or rotate the angle by ninety degrees, depending on the relative directions.

This is why a photometric-style calibration is useless here. Determining that the instrument transmits ninety per cent of the light corrects a scale; nothing about a scale correction touches an additive offset, and the offset is the whole of the problem.

An instrument that adds 0.20 per cent to a sky that has 0.55. The Stokes plane, with the two linear polarisation parameters as axes. The tight cluster near the centre is a set of stars known to be unpolarised, observed through the same instrument: they should sit at the origin and do not, and their mean is the instrumental polarisation — 0.20 per cent here, which is 0.4 times the real polarisation of the field. The other cluster is the field stars, whose measured values are the sum of their own polarisation and the same instrumental offset. Subtracting one mean from the other recovers 0.58 per cent at the right angle. Two things make this worth doing carefully. The offset is a vector, so leaving it in rotates the measured position angle as well as changing its magnitude — by 10 degrees here — and a calibration that only fixes the scale does not touch that. And the offset depends on where the telescope was pointing, because the reflection angles do, so the standards have to be observed at the same place in the sky and at the same instrument rotation.
Fig. 2 A better-behaved instrument: the same field observed through optics with a third of the polarisation. The two clusters are now well separated, the correction is smaller than the signal, and the residual error after correcting is set by how well the standards’ mean is determined rather than by how large the offset is. That is the design goal — not zero instrumental polarisation, which is unattainable, but an offset small and stable enough that measuring it is easy.

There is a second structural fact that follows from the first and is worth stating separately. Because the offset is a vector, its effect on the measured degree of polarisation is not symmetric: a source whose own polarisation happens to point along the instrumental vector is measured too high, and one pointing against it is measured too low, and one at right angles has its angle rotated with its degree barely changed. So a population of sources with random position angles, measured through an uncorrected instrument, acquires a correlation between degree and angle that no physical mechanism would produce. That correlation is the diagnostic: a plot of measured degree against measured angle for a sample that should have no such relation is the cheapest test of whether the correction is working.

It is worth putting the numbers side by side once. A Cassegrain focus with a symmetric optical train contributes of order a hundredth of a per cent; a Nasmyth focus with a flat tertiary at forty-five degrees contributes several tenths of a per cent, and can exceed a per cent in the blue where aluminium’s reflectivity is most asymmetric. Interstellar polarisation of a star a kiloparsec away is a per cent or two; of a nearby star, a hundredth. Scattering in a circumstellar disc gives a few per cent in the disc and much less when averaged over an unresolved source. So the instrument is comparable to the strongest astrophysical signals and larger than most of them, and the ordering depends on which focus the instrument happens to sit at — a decision made for reasons of mechanical convenience decades earlier.

How it is measured

The calibration is conceptually simple and observationally tedious: observe objects known to have no polarisation, and whatever is measured is the instrument’s.

The standards are nearby stars, chosen because interstellar polarisation is produced by dust and there is little dust within a hundred parsecs. Their polarisations are known to be below a hundredth of a per cent from decades of measurement with many instruments, and the residual disagreement between those measurements is the floor on how well the zero point can be established.

There is a second calibration and it is equally necessary: the efficiency. An instrument does not merely add polarisation, it also fails to transmit all of what arrives, so a source with one per cent of polarisation may be measured as 0.95. That is measured by observing standards of known non-zero polarisation, or by inserting a polariser into the beam, and it is a scale correction — the multiplicative half that the additive half above is not.

A position angle to 1.7° and a polarisation that is biased upwards. The Stokes plane, with 260 simulated measurements of one star at a true polarisation of 0.6 per cent and a position angle of 62°, each Stokes parameter carrying an independent error of 0.35 per cent. The polarisation is the length of the vector and the position angle is half its azimuth — which is why a rotation of 180° in this plane is a rotation of 90° on the sky, and why polarisation has no sign. Averaging the two components recovers 60.3° against the true 62°. But averaging the lengths gives 0.686 per cent against a true 0.6: a length cannot be negative, so noise can only push it up, and the excess here is 0.112 against the large-signal expectation σ²/2p = 0.102. At zero true polarisation the same effect returns 0.44 per cent from a star that has none, which is the reason a polarimetric detection is quoted in σ and almost never in per cent alone.
Fig. 3 What a measurement actually consists of: an ensemble of exposures at different analyser angles, from which the two Stokes parameters are extracted. Every systematic in this essay enters as a shift of this cloud, and the reason the technique is so demanding is visible in the geometry — the quantity wanted is the displacement of the cloud’s centre from the origin, and the cloud’s own size is comparable to it.

Both calibrations have an awkward property in common: they are measured on bright stars and applied to faint ones, and nothing guarantees the instrument behaves identically. Detector non-linearity, charge transfer inefficiency and the different exposure times involved all enter, and the resulting brightness dependence of the polarimetric zero point is one of the least-characterised systematics in the technique. It is the same shape as a photometric calibration transferred from standards to targets, with the difference that a polarimetric measurement is a difference and therefore twice as sensitive to anything that scales.

The parts a zero point does not fix

If instrumental polarisation were a constant vector the story would end there. Three things stop it being constant.

It depends on the pointing. The angles of incidence on the mirrors of an alt-azimuth telescope with a Nasmyth focus change as the telescope moves, so the instrumental vector rotates and changes magnitude through the night. Standards therefore have to be observed at the same altitude and azimuth as the target, or the dependence has to be mapped and modelled.

Crosstalk. A real optical train does not merely add polarisation; it converts one kind into another. Linear polarisation becomes circular and back again on reflection from a metal surface, and the conversion is described by the off-diagonal elements of the instrument’s Mueller matrix. A source with strong circular polarisation and no linear can be measured as linearly polarised, which is a systematic that no unpolarised standard reveals.

Depolarisation. Averaging over a field of view, over a bandpass, or over time reduces the measured polarisation if the position angle varies across whatever is being averaged. That is not an instrumental error in the ordinary sense — the instrument reports the average correctly — but it means the measured value depends on the aperture, the filter and the exposure.

A peak at 0.55 µm, and therefore a grain size. Interstellar polarisation against wavelength — the Serkowski law, p(λ) = p_max exp[−K ln²(λ_max/λ)] with K = 1.66 λ_max. The heavy curve peaks at 0.550 µm, measured off the drawing rather than read back from the parameter, and falls to half its peak at 1.314 µm on the red side, against the closed form λ_max exp√(ln2/K) = 1.315. That peak wavelength is the measurement. It is set by the size of the grains doing the aligning — bigger grains, longer λ_max — and it is tied to the shape of the extinction curve along the same sight line by R_V ≈ 5.5 λ_max, which gives 3.03 here against the diffuse-medium value of 3.1. The two faint curves are populations peaking at 0.35 µm and 0.75 µm: the same amount of polarisation, distributed differently, and a different dust. Nothing in a photometric measurement of the same star distinguishes them.
Fig. 4 Why the bandpass matters: interstellar polarisation has a wavelength dependence with a peak, so a broad-band measurement is an average over a curve rather than a sample of it. The effective wavelength of that average depends on the source’s spectrum, so two stars of different colours measured through the same filter are measured at different points on their own polarisation curves. It is the same structure as a magnitude having to say which light it was measured in, with an extra parameter attached.

A fourth complication belongs with those three and it is astrophysical rather than instrumental: the sky itself is polarised. Moonlight scattered in the atmosphere is strongly polarised, with a degree that depends on the angle from the Moon and reaches tens of per cent, and it enters the aperture along with the source. For a faint target the sky contributes a large fraction of the light in the aperture, so the sky’s polarisation has to be subtracted as a Stokes vector rather than as an intensity — which requires measuring the sky’s polarisation nearby and at the same time, and which is why polarimetry of faint objects is not attempted near a bright Moon.

What was actually measured

Three results establish both the size of the problem and the level to which it can be beaten.

The standards themselves. Catalogues of unpolarised standards, built up over decades, agree at the level of about a hundredth of a per cent. That number is the floor of the technique from the ground, and it is set by the accumulated disagreement between instruments rather than by any single measurement.

The pointing dependence, mapped. Observing standards across the sky at a Nasmyth focus produces a smooth pattern of instrumental polarisation with altitude and azimuth, of amplitude up to a per cent and reproducible from run to run. Modelling it reduces the residual by an order of magnitude, and the residual after modelling is what limits the instrument.

The best achieved precision. Dedicated instruments using rapid modulation — switching between polarisation states hundreds of times a second, so that atmospheric and instrumental drifts are common to both states — reach a few parts per million on bright stars. That is four orders of magnitude below the raw instrumental polarisation, and it is achieved not by measuring the offset better but by arranging that it cancels.

Zero forwards, zero backwards, and exactly one at a right angle. The degree of linear polarisation produced by a single scattering, against the angle through which the light was turned. The curve is (1 − cos²θ)/(1 + cos²θ): it is exactly zero at 0° and 180° and exactly one at 90°, and neither of those is a fitted number — they are the geometry of a dipole seen end-on and side-on. The lower curves are the same shape divided down by an unpolarised component, for slabs of optical depth 0.2, 1, 3 in which some of the light has scattered more than once; the peak falls to 16 per cent of its single-scattering value at the thickest. The angular position of the maximum does not move, which is why a polarisation map of a reflection nebula locates the illuminating star even when the star is hidden: every position angle is perpendicular to the line back to the source, and the pattern converges on it.
Fig. 5 One of the signals the technique is aimed at: polarisation produced by scattering, whose degree depends on the scattering angle and which is what makes a reflection nebula, a planetary atmosphere or a circumstellar disc polarimetrically visible. The quantity to be measured is a per cent or two at best, which is why an instrumental contribution of the same size is not a detail — and why the cases where polarimetry is easiest are the ones where the signal is large enough to survive a mediocre correction.
Where integrating longer stops helping, at V = 10. The same budget against exposure time rather than against brightness, for a star at V = 10. Both random terms fall as the square root of the time — the counting one because a photon count does, scintillation because the atmosphere decorrelates in milliseconds and an exposure averages over very many independent realisations — so on these axes both are straight lines of slope −½. The floor is the term that is not a line. Flat-fielding error, differential colour terms and the imperfect match between the star's spectrum and the comparison's are systematic: they repeat, so averaging leaves them exactly where they were. After 1000 seconds the total is 1.09 times that floor and after 10000 it is 1.01, which is the practical statement that a ground-based night is over long before the photons run out. A space telescope's floor is lower by a factor of ten or more, and that — not aperture, not photons — is what a transit survey buys by leaving the ground.
Fig. 6 What the systematic has to be compared against: the precision photon statistics alone would deliver, against brightness. Polarimetry is a differential measurement between two intensities, so the photon-limited polarimetric precision is roughly the photometric precision divided by the square root of two — which for a bright star is parts per million in minutes. Everything in this essay is about why that number is almost never achieved, and the gap between the two is entirely calibration.

There is a fourth result that is really a warning, and it comes from the history of the subject. Several claimed detections of polarisation in objects of great interest — the alignment of quasar polarisation vectors over large scales, the polarisation of the light from certain supernovae, the circular polarisation of some stellar sources — have been reduced or removed by later observations with better-characterised instruments. In each case the original measurement was at the level of a few tenths of a per cent, which is exactly where the instrumental contribution sits, and in each case the disagreement was traced to calibration rather than to variability. That history is why polarimetric results at that level are reported with the instrumental correction described in detail, and why a paper that does not describe it is not usable.

Where the picture stops

The picture stops in three places, and the second is the one that decides what can be attempted.

The correction is only as good as the standards’ distribution. Unpolarised standards are bright, nearby and unevenly distributed on the sky, so at some pointings there is no suitable standard within a reasonable slew. Interpolating the instrumental model across such gaps is where the residual systematic lives.

Rapid modulation beats calibration, and it costs an instrument. The reason the best polarimeters reach parts per million is that they switch states faster than anything drifts, so the two measurements being differenced share every systematic. That is a design decision made before the instrument is built, and it cannot be retrofitted to data taken with a slow rotating analyser.

And circular polarimetry is much harder than linear. Circular signals in astronomy are typically ten to a hundred times weaker than linear ones, and the crosstalk that converts linear into circular is often larger than the circular signal being sought. A measurement of circular polarisation in the presence of strong linear polarisation is therefore mostly a measurement of crosstalk, which is why the Zeeman signature that measures a magnetic field is one of the most demanding measurements in observational astronomy.

A 3-gauss field moves the line by 0.33% of its width and is measured anyway. Above: the Fe I 6173 Å line, Landé factor 2.5, at a Doppler width of 0.041 Å. The solid curve is the unmagnetised profile; the dashed curve is the same line in a longitudinal field of 3 gauss, which splits it by 1.3·10⁻⁴ Å — 0.33 per cent of its own width — and is drawn on top of it because the two are not distinguishable. The double-lobed curve underneath them is the Stokes V profile of that same field, magnified 100 times: the two σ components are circularly polarised with opposite handedness, so what is lost in the sum survives in the difference, and the difference of two profiles a hair apart is the derivative of one of them. Below: what that buys. The V amplitude is linear in the field, because it is a first derivative; every signature of the same field in the intensity is quadratic, because a symmetric splitting can only broaden. At a polarimetric precision of 10⁻⁴ the first reaches 0.1 gauss; at a line-width accuracy of 0.001 the second reaches 38, a factor of 420 worse. A quantity far too small to resolve is measured because it is the only thing in the signal that carries a sign — and the same argument run the other way says what polarimetry cannot do: a field of mixed polarity inside the resolution element cancels in V and does not cancel in I, so the 3000-gauss field of a sunspot, which does resolve, is measured the other way round.
Fig. 7 The measurement that pushes the technique hardest: the circular polarisation signature of a magnetic field in a spectral line, which is the derivative of the intensity profile scaled by a field strength. The amplitude for a solar-type star’s global field is a fraction of a per cent of the line depth, and every systematic in this essay sits on top of it. That the measurement works at all is a consequence of the signal’s distinctive shape — an antisymmetric profile that no instrumental offset reproduces.

A fourth limit belongs here because it changes what is possible from space rather than from the ground. A space telescope has no sky polarisation, no atmospheric depolarisation and no pointing-dependent gravity flexure, and it has an optical train that cannot be recoated or realigned. So the instrumental polarisation is smaller and far more stable, and the limiting factor becomes the stability of the detector rather than of the optics. That is why the most precise polarimetry of faint objects has come from space and the most precise polarimetry of bright ones from the ground: the two regimes are limited by different terms, and neither instrument is simply better.

There is also a case worth mentioning where the instrumental contribution is deliberately made large. A polarimeter that modulates by rotating a waveplate introduces a known, large, rapidly varying polarisation on purpose, so that the small unknown one is measured as a modulation amplitude rather than as a difference between two exposures. That converts an additive systematic into a phase — and a phase is much easier to keep stable than a level, which is the same reason a frequency measurement beats an amplitude measurement nearly everywhere in this subject.

Why the shape of the signal is the real defence

That last point generalises and it is the most useful thing in this essay.

A constant instrumental offset can be subtracted if it is measured, and measuring it is limited by standards. But a signal with a distinctive structure — a wavelength dependence, a spatial pattern, an antisymmetric line profile — can be separated from an offset without measuring the offset at all, because the offset does not have that structure.

That is why the Zeeman analysis works: the signal is antisymmetric about the line centre and the instrumental contribution is not, so fitting for an antisymmetric component rejects the offset by construction. It is why dust polarisation in a cloud is measured from the pattern of directions rather than from the degree at any one point. And it is why polarimetric imaging of a disc uses the fact that scattered light has a position angle that circles the star — a pattern that an additive constant cannot produce.

The general prescription: when a systematic is additive and constant, look for a signal that varies. The variation need not be large. It has to be a variation the systematic cannot imitate, and identifying one is worth more than any amount of effort spent on the calibration.

The same argument runs through the whole collection — a contaminant with a known dependence is not a contaminant but a second observable, and the design of a measurement is largely the business of arranging for the two to differ in something.

A concluding observation about why this is a rung on the polarimetry ladder rather than an instrumental appendix. Polarimetry is the one common astronomical measurement whose systematic is larger than its signal for most targets, and that inverts the usual order of work. In photometry or spectroscopy the calibration is a correction to a measurement; here the measurement is a correction to the calibration. Everything about how the observations are planned follows from that — standards at the same pointing, modulation faster than any drift, and a preference for signals with structure over signals with amplitude. Two instruments blind in opposite directions is the same principle applied to the physics rather than to the hardware, and both are cases of designing a measurement so that its dominant error has nowhere to hide.

Say which measurements are unaffected by all of this, because the list is short and useful. Anything that depends on a change in polarisation — a variable source measured repeatedly through the same instrument at the same pointing — is nearly immune, because the offset is common to every epoch. Anything that depends on a spatial pattern within one image is similarly protected. What is vulnerable is the absolute polarisation of a single unresolved source measured once, which is unfortunately the most commonly wanted quantity.

One consequence of the vector character deserves stating separately, because it decides what an unpolarised standard is worth. Observing a star known to be unpolarised measures the instrument’s own contribution at that moment and at that pointing — which is exactly what is wanted, and which is only valid for that pointing, because the reflection angles change as the telescope tracks. On an alt-azimuth mount the field rotates with respect to the instrument through the night, so the instrumental vector rotates with respect to the sky while the source’s does not. That is a nuisance and it is also the escape: observing the same source across a large range of parallactic angle modulates the two contributions differently, and fitting for both separates them. The measurement therefore has to be spread across hours rather than concentrated, which is the opposite of what signal-to-noise alone would suggest.

Where the ladder goes next

The rung directly above is the Mueller matrix itself: the full sixteen-element description of what an optical train does to polarised light, how much of it can be measured on the sky, and which elements matter for which observation. The one above that is the modulation scheme — how switching between polarisation states faster than anything drifts converts a calibration problem into a differential measurement.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

CalibrationCrosstalkDepolarisationInstrumental polarisationModulationMueller matrixPolarimetric standardPosition angleStokes parametersSystematic error