Exoplanets

A light curve with a fold in it

A single lens magnifies smoothly. A second mass makes the lens mapping fold, and a fold has an edge — a curve across which two images appear out of nothing and the magnification formally diverges. Crossing it turns the finite size of the source star from a nuisance into a ruler.

Assumes Microlensing, Strong lensing and Detection bias.

The rung below this one drew the single-lens curve: symmetric, achromatic, unrepeatable, and containing one number — the Einstein crossing time — that mixes the lens mass, its distance and its transverse velocity together and cannot be separated into them.

A single point mass is a well-behaved map. It takes each point of the source plane to two image positions, the total magnification is a smooth function of the source’s offset, and the only divergence is at the single point of perfect alignment.

A second mass changes the character of the map, not merely its parameters. The lens equation for two masses is not analytic — it involves the complex conjugate of the image position — and a non-analytic map can fold. Where it does, the Jacobian vanishes, the magnification diverges along a whole curve rather than at a point, and the number of images changes by two across it.

A fold 0.06 Einstein times wide, and two configurations that make the same one. A binary lens of mass ratio 0.003 at a projected separation of 1.5 Einstein radii. Top left: the source plane, with the caustic — the set of source positions at which the magnification is formally infinite — and the track of a background star across it. A single lens has no such curve; it magnifies smoothly and diverges only at one point. A second mass makes the lens mapping fold, and the fold has edges: crossing one, the number of images changes from 3 to 5, because a pair is created out of nothing on the critical curve. The small closed curve near the origin is the central caustic, always there; the larger one at 0.83 Einstein radii is the planetary caustic, and its distance from the origin is s − 1/s, which is where the planet's own image lies. Top right: the central caustic drawn twice, once for s = 1.5 and once for s = 0.667. They are 0.0184 and 0.0169 Einstein radii across and they lie on top of each other. That is not a coincidence of these numbers: to the order that a central-caustic anomaly is measured, a close binary and a wide one with the reciprocal separation produce the same perturbation, so an event with only a central anomaly returns two separations and no way to choose. Below: the light curve along the track. The smooth part is what a single lens of the same total mass would do; the spikes are the two crossings, 0.06 Einstein times apart, so a few hours inside an event lasting a month. The two curves differ in one thing only — the size of the source. A point source diverges at each fold and reaches 27; a source of angular radius 0.006 Einstein radii averages over its own disc and reaches 9 — an eighth of a source radius inside the fold the two are 8 and 4, with the divergence replaced by a rounded shoulder whose width is the source's own diameter. Everywhere else in this collection the finite size of a star is a nuisance that degrades a measurement. Here it is the ruler: the fold is a straight edge of known sharpness sweeping across a disc, so the shape of that shoulder gives the source's angular radius, and dividing it by the crossing time gives the angular Einstein radius — which is the one quantity a light curve otherwise cannot supply.
Fig. 1 The structure a planet adds. Top left: the source plane, with the caustic — the set of source positions at which the magnification is formally infinite — and the track of a background star across it. The small closed curve near the origin is the central caustic, always present; the larger one is the planetary caustic, at s1/ss - 1/s Einstein radii from the origin, which is where the planet’s own image lies. Top right: the central caustic drawn twice, once for a separation ss and once for 1/s1/s, lying on top of each other. Below: the light curve along the track, for a point source and for a source of finite size.

Where a fold comes from

The lens equation maps an image position to a source position. For a single point mass it is a smooth two-to-one map and its Jacobian, which is the magnification’s reciprocal, vanishes on a circle — the Einstein ring — whose image in the source plane is the single point at the origin.

Add a second mass and that circle deforms. The set of image positions where the Jacobian vanishes is the critical curve, and its image in the source plane is the caustic. For a single lens the caustic degenerates to a point; for two masses it is a closed curve with cusps, and for most configurations there is more than one component.

Two facts about a fold matter observationally and neither depends on anything about the lens beyond its being a fold.

The image count changes by two. Approaching a caustic from outside, two images appear on the critical curve, at the same place and with equal and opposite parities. A binary lens has three images outside its caustics and five inside.

The magnification diverges as the inverse square root of the distance from the fold. That is a generic property of a fold catastrophe, not a property of gravity, and it is why a caustic crossing is a sharp rise with a characteristic shape rather than a symmetric peak.

The two images a point lens makes. A source at 0.15 Einstein radii from the lens is split into two images, at 1.08 and -0.93 Einstein radii on either side of it, one outside the ring and one inside. Neither is ever resolved — the ring is a milliarcsecond across — so all that is observed is their combined brightness, which is 6.72 times the source's own. A planet matters when it lands on top of one of them, which is why the method is sensitive at about one Einstein radius and almost nowhere else.
Fig. 2 Where the images actually are for the single-lens case the rung below treats. A point mass makes one image outside the Einstein ring and one inside it, on opposite sides of the lens; as the source moves, the outer image traces most of the brightness and the inner one is faint and close in. A planet’s caustic is a perturbation to one of those image tracks — which is why the planetary anomaly happens at a particular moment in the event rather than at the peak, and why its timing gives the planet’s projected separation.
The two images a point lens makes. A source at 0.35 Einstein radii from the lens is split into two images, at 1.19 and -0.84 Einstein radii on either side of it, one outside the ring and one inside. Neither is ever resolved — the ring is a milliarcsecond across — so all that is observed is their combined brightness, which is 2.99 times the source's own. A planet matters when it lands on top of one of them, which is why the method is sensitive at about one Einstein radius and almost nowhere else.
Fig. 3 What the source’s images are doing while the curve rises. A point lens produces two images, one inside the Einstein ring and one outside, both invisibly small — and the magnification is the sum of their areas. At an impact parameter of 0.35 Einstein radii the images are strongly sheared and the total magnification is near three. Nothing is resolved at any point of this; the entire measurement is the sum of two areas nobody can see, read as a brightness against time.

What a crossing measures

The planetary caustic’s position gives the projected separation ss in Einstein radii, from the relation that the caustic sits at s1/ss - 1/s for a wide binary and at 1/s+s-1/s + s folded inside for a close one. The caustic’s size gives the mass ratio: a planetary caustic’s width goes as q\sqrt{q} and a central caustic’s as qq divided by a function of ss.

So an anomaly lasting a few hours in a month-long event gives two numbers, and neither of them is a mass. What comes out is qq, a ratio, and ss, an angle in units of θE\theta_{\rm E} — and θE\theta_{\rm E} is exactly the quantity a single-lens light curve cannot supply.

That is the position the method has always been in: microlensing measures mass ratios superbly and masses badly. A planet’s mass is qq times the host’s, and the host’s mass is buried in tE=θE/μrelt_{\rm E} = \theta_{\rm E}/\mu_{\rm rel} along with the relative distance and proper motion.

The degeneracy that is not a measurement error

Two configurations produce nearly the same central caustic: a close binary at separation ss and a wide one at 1/s1/s.

This is not a coincidence of particular numbers. Expanded to the order at which a central-caustic anomaly is actually measured, the perturbation from a companion depends on ss only through the combination s1/ss - 1/s — and that combination is invariant under s1/ss \to 1/s up to a sign, which the geometry of the caustic does not resolve. The agreement is exact in a limit and approximate in practice, and it tightens as logs|\log s| grows.

That last point is the sting. The degeneracy is worst exactly where the caustic is smallest, which is where the anomaly is hardest to measure and where a longer or better-sampled light curve does not help. Many published events have two solutions with separations reciprocal to each other, equally good fits, and physically different planets — one inside the snow line and one well outside it.

A degeneracy of this kind cannot be beaten with more of the same data. It has to be broken from outside: by the planetary caustic if the source happens to cross it too, since that one is not degenerate; by higher-order effects in the light curve; or by imaging the lens star directly years later, once it and the source have separated enough.

The nuisance that becomes a ruler

A point source crossing a fold would go infinitely bright. A real source has an angular radius θ\theta_*, and what is observed is the magnification averaged over its disc — so the divergence is replaced by a rounded shoulder whose width in time is the time the source takes to cross the fold.

Everywhere else in this collection a star’s finite angular size is a complication to be removed: it smears a light curve’s ingress, it limits the resolution of an occultation, it blurs an interferometric visibility. Here it is the measurement, and the reason is that the fold is a straight edge of known sharpness sweeping across a disc of unknown size.

The chain is short. The crossing time tt_* and the Einstein time tEt_{\rm E} give ρ=θ/θE=t/tE\rho = \theta_*/\theta_{\rm E} = t_*/t_{\rm E}. The source’s angular radius θ\theta_* is obtained independently — from its dereddened colour and magnitude, through an empirical relation between colour and surface brightness that is calibrated on stars with interferometric diameters. Divide, and

θE=θρ.\theta_{\rm E} = \frac{\theta_*}{\rho}.

That single quantity converts a mass ratio into a mass, because θE\theta_{\rm E} and tEt_{\rm E} together give the relative proper motion, and θE\theta_{\rm E} with a lens distance gives the mass directly.

A magnification, and a spike inside it. The brightness of a background star as a foreground one passes in front of it. The smooth curve is exact: a point mass magnifies a point source by (u² + 2)/(u√(u² + 4)), which peaks at 6.7 for this track. The spike is the planet, of mass ratio 0.00003, lensing one of the two images. Its duration is the Einstein time scaled by √q — about 4 hours against 30 days — so the whole planetary signal is a few hours in an event lasting a month, and it never repeats. The deviation is drawn in the approximation that the planet lenses the image in isolation; a real caustic crossing has structure this smooths over, and its true height is set by the source's size rather than by the geometry.
Fig. 4 The same event with the companion thirty times lighter — an Earth-mass planet rather than a Neptune. The deviation’s duration scales as the square root of the mass ratio, so a thirty-fold lighter planet perturbs the curve for a fifth as long, and its amplitude is undiminished. That is the property that makes microlensing uniquely sensitive to low masses: the signal does not fade with the planet’s mass, it shortens — so what limits the method is cadence rather than photometric precision.

What was actually measured

A microlensing planet detection is a light curve, sampled by a survey that monitors hundreds of millions of bulge stars every fifteen to sixty minutes, with follow-up photometry triggered when an event brightens.

The anomaly is short. For a mass ratio of 10310^{-3} — roughly Jupiter around an M dwarf — the planetary caustic crossing lasts tEqt_{\rm E}\sqrt{q}, about a day out of a month. For an Earth-mass planet, q3×106q \approx 3\times10^{-6} and the anomaly is a couple of hours. Missing it is the normal outcome, and the entire architecture of the field — continuous longitude coverage from several continents, automated alerts, and now a survey cadence fast enough not to need follow-up at all — exists to avoid missing it.

What the fitted parameters are, for a typical published event: tEt_{\rm E} to a few per cent, u0u_0 and the peak time precisely, qq to ten or twenty per cent, ss to a few per cent with a twofold ambiguity, and ρ\rho only when a caustic is crossed — which happens in perhaps a third of planetary events.

The masses that result carry error bars of tens of per cent and are quoted with a distance that is often a probability distribution rather than a number. The census microlensing has produced is a census of mass ratios and separations, and it is turned into one of masses through a model of where the lenses are in the Galaxy.

A magnification, and a spike inside it. The brightness of a background star as a foreground one passes in front of it. The smooth curve is exact: a point mass magnifies a point source by (u² + 2)/(u√(u² + 4)), which peaks at 6.7 for this track. The spike is the planet, of mass ratio 0.001, lensing one of the two images. Its duration is the Einstein time scaled by √q — about 23 hours against 30 days — so the whole planetary signal is a few hours in an event lasting a month, and it never repeats. The deviation is drawn in the approximation that the planet lenses the image in isolation; a real caustic crossing has structure this smooths over, and its true height is set by the source's size rather than by the geometry.
Fig. 5 What is actually on the plot before any of this is extracted. A smooth, symmetric brightening lasting a month, with a spike lasting hours somewhere on its flank. Everything above is recovered from the spike’s timing, duration and shape — and the reason the modelling is hard is that the smooth part is a three-parameter fit and the spike is a five-parameter one whose likelihood surface has several disconnected minima.
Why an Einstein radius is a mass. The geometry, drawn at an angle some ten thousand times larger than the real one so that anything is visible at all. Light from a source directly behind a lens reaches the observer along every path that passes the lens at the same distance, so the image is a ring rather than a point. The ring's angular radius is θ_E = √(4GM/c² · D_ls/D_l D_s), which for a lens of 1.0×10¹² M☉ at these distances is 2.52 arcseconds and encloses 10 kpc at the lens. Rearranged, it is a mass in terms of an angle and three distances — and the mass so obtained is inside a cylinder rather than a sphere, and assumes nothing whatever about the lens being in equilibrium, which is the assumption every other weighing in this collection makes.
Fig. 6 The quantity everything is measured in units of. The Einstein radius is fixed by the lens mass and by the ratio of three distances, and it is the natural angular scale of any lensing configuration — the separation of the images, the size of the caustics and the duration of an event are all set by it. For a solar-mass lens halfway to the Galactic bulge it is about a milliarcsecond, which at the lens is a few astronomical units. That coincidence is the whole reason microlensing finds planets: the scale the geometry picks out happens to be the scale planetary systems are built on.

Cusps, and the anomalies that have no fold in them

A caustic is not made only of folds. Where two folds meet, the curve has a cusp — a point rather than an arc, at which three images merge instead of two — and cusps behave differently enough to be worth separating.

Passing near a cusp without crossing the caustic still produces an anomaly, because the magnification is enhanced in a whole region around a cusp rather than only inside the caustic. That matters practically: a substantial fraction of planetary detections are cusp approaches rather than caustic crossings, and they produce a smooth bump rather than the sharp double-shouldered feature a fold pair gives.

The distinction has consequences for what can be measured. A fold crossing has a sharp edge whose rounding measures the source size, which is what supplies the angular Einstein radius; a cusp approach has no such edge, so the source size is unconstrained and the mass stays unmeasured. The events that are easiest to detect are not the events that yield the most, and a survey’s catalogue is therefore split into a minority with masses and a majority with only mass ratios.

There is a second structure worth naming because it is the source of a persistent ambiguity. Near a cusp, the three merging images have magnifications that satisfy a relation with no free parameters: the sum of the two of one parity equals the third. For a resolved lens that relation is a test — a system violating it has substructure between the images — and for a microlensing event it is one of the few internal consistency checks available on a fit whose likelihood surface has multiple minima.

Cusps also sharpen the close–wide problem rather than resolving it. The central caustic of a close binary and that of a wide one agree in size and shape to the order that matters, and they differ in the arrangement of their cusps — but the difference appears in the light curve only if the source track happens to pass near the cusps that differ, which is a matter of geometry the observer does not control. Some events break the degeneracy and most do not, and which is which is decided by where the source happened to go.

Where the model stops

A binary lens has more parameters than a light curve constrains. Beyond qq and ss there is the angle of the source track relative to the binary axis, the source size, limb darkening, and — for long events — the orbital motion of the lens itself and the parallax from the Earth’s motion during the event. Fits routinely explore a likelihood surface with a dozen dimensions and several isolated islands.

The lens is often not two masses. Triple lenses exist, and a caustic structure from a star with two planets, or from a binary star with one, can imitate a single planet’s anomaly closely enough to be published as one.

And the source is often blended. A bulge field at one arcsecond resolution has several stars per resolution element, so the baseline flux attributed to the source includes light that is not being lensed. Blending is degenerate with the magnification, and it is the reason a colour–magnitude estimate of θ\theta_* is done on the source’s colour recovered from the event rather than on the star’s apparent colour.

The picture cannot show any of the images. Everything drawn in the source plane above is a construction: the images are milliarcseconds apart, the caustic subtends microarcseconds, and the only thing ever measured is a total flux against time. The whole geometry is inferred, and the confidence in it comes from the fact that a fold is a generic mathematical object whose signature could not plausibly be produced by anything else.

A detection that cannot be followed up

Every other planet-detection method produces a target. A transiting planet can be re-observed at another wavelength, its atmosphere probed during a later transit, its mass measured by a spectrograph, its orbit refined over decades. A radial-velocity planet can be watched for a second orbit.

A microlensing planet cannot be any of those things. The alignment that produced the signal will not recur — the lens and the source are unrelated objects passing at a relative proper motion of a few milliarcseconds a year — and once the event is over the system is a faint star in a crowded field with nothing to distinguish it. The planet has been measured once and will never be measured again.

That shapes the field in ways worth stating, because it explains choices that otherwise look strange. It is why the modelling is done so exhaustively on each event: there is no prospect of a second data set to settle an ambiguity, so every degeneracy has to be explored and reported rather than resolved later. It is why the published parameters are quoted as multi-modal distributions rather than as values with error bars. And it is why the field’s results are framed as occurrence rates rather than as catalogues of objects — the individual detections are not follow-up targets, so their value is entirely statistical.

There is one exception and it is the reason large telescopes are pointed at old event fields. The lens and the source separate at a few milliarcseconds a year, so after a decade or two they can be resolved from one another by adaptive optics or from space. Measuring the lens star’s own brightness then gives its mass and distance directly, which converts the event’s mass ratio into a planet mass with none of the modelling above.

That is a strange kind of observation: the decisive measurement of a planet is made ten or twenty years after the only opportunity to observe the planet has passed, on a star that was invisible at the time and is now merely faint. A few dozen events have been resolved that way, and the masses they return have in several cases selected between the degenerate solutions that the light curve alone could not separate. The method’s one route to certainty runs through waiting.

The caustic’s size and the images that produce it are the two halves of the geometry, and each is worth drawing at a second configuration.

A fold 0.03 Einstein times wide, and two configurations that make the same one. A binary lens of mass ratio 0.0006 at a projected separation of 1.5 Einstein radii. Top left: the source plane, with the caustic — the set of source positions at which the magnification is formally infinite — and the track of a background star across it. A single lens has no such curve; it magnifies smoothly and diverges only at one point. A second mass makes the lens mapping fold, and the fold has edges: crossing one, the number of images changes from 3 to 5, because a pair is created out of nothing on the critical curve. The small closed curve near the origin is the central caustic, always there; the larger one at 0.83 Einstein radii is the planetary caustic, and its distance from the origin is s − 1/s, which is where the planet's own image lies. Top right: the central caustic drawn twice, once for s = 1.5 and once for s = 0.667. They are 0.0036 and 0.0036 Einstein radii across and they lie on top of each other. That is not a coincidence of these numbers: to the order that a central-caustic anomaly is measured, a close binary and a wide one with the reciprocal separation produce the same perturbation, so an event with only a central anomaly returns two separations and no way to choose. Below: the light curve along the track. The smooth part is what a single lens of the same total mass would do; the spikes are the two crossings, 0.03 Einstein times apart, so a few hours inside an event lasting a month. The two curves differ in one thing only — the size of the source. A point source diverges at each fold and reaches 19; a source of angular radius 0.006 Einstein radii averages over its own disc and reaches 5 — an eighth of a source radius inside the fold the two are 6 and 3, with the divergence replaced by a rounded shoulder whose width is the source's own diameter. Everywhere else in this collection the finite size of a star is a nuisance that degrades a measurement. Here it is the ruler: the fold is a straight edge of known sharpness sweeping across a disc, so the shape of that shoulder gives the source's angular radius, and dividing it by the crossing time gives the angular Einstein radius — which is the one quantity a light curve otherwise cannot supply.
Fig. 7 The caustic for a planet five times lighter. It shrinks roughly as the square root of the mass ratio, so the probability that a source crosses it falls with the planet’s mass — which is why the microlensing sensitivity curve turns over below an Earth mass and not because the signal becomes unmeasurable.
The two images a point lens makes. A source at 0.05 Einstein radii from the lens is split into two images, at 1.03 and -0.98 Einstein radii on either side of it, one outside the ring and one inside. Neither is ever resolved — the ring is a milliarcsecond across — so all that is observed is their combined brightness, which is 20.02 times the source's own. A planet matters when it lands on top of one of them, which is why the method is sensitive at about one Einstein radius and almost nowhere else.
Fig. 8 And the two images for a much smaller impact parameter. They stretch into arcs on either side of the lens and their combined brightness rises steeply, so the highest-magnification events are the ones most likely to show a planetary perturbation — the images sweep past where the caustic is.

The generalisation

The mathematics here is not specific to gravity. A fold is the simplest of the catastrophes: the generic way a smooth map from one plane to another can fail to be locally invertible, with a square-root divergence in the density of images and a change of two in their number.

The same structure appears wherever a wave or a ray family is focused by a smooth but non-uniform medium. The bright lines on the bottom of a swimming pool are caustics of the water surface; a rainbow is the caustic of refraction through a sphere, and its brightness diverges in exactly the same square-root way at the geometric-optics level; a mirage is the caustic of a temperature gradient. In each case the pattern of the singularity is universal and only its scale depends on the physics.

That universality is what makes the microlensing interpretation robust. The specific caustic shape depends on the two masses and their separation, and getting it wrong gets those wrong. But that there is a curve, that crossing it adds two images, and that the magnification goes as an inverse square root, are consequences of the map being smooth and non-analytic — and no astrophysical alternative reproduces them.

The shortest events, and what they are not

One class of detection deserves separating out, because it is the only one for which the finite-source measurement is not an improvement but a requirement.

An event with no host at all — a single lens whose light curve rises and falls within hours rather than weeks — implies a lens of planetary mass, because the Einstein crossing time scales as the square root of the mass and a day-long event corresponds to something below a Jupiter. Several dozen such events have been reported, and they are read as free-floating planets: objects ejected from the systems that made them, or formed in isolation, wandering the Galaxy unattached.

The reading is only as good as the ruling-out of a host, and a host is ruled out by not seeing an anomaly. That is an argument from absence, and it is weak in a specific way: a wide-separation companion produces a caustic far from the source track and no signal at all, so an event with no anomaly is consistent with a bound planet at a large separation as readily as with a free one. What the non-detections constrain is the separation, not the boundness — a survey can say that no companion lies within a few tens of astronomical units and can say nothing beyond that.

The finite-source effect is what makes such an event worth anything quantitatively. For an event lasting a day the source’s angular size is a substantial fraction of the Einstein radius, so the light curve is visibly rounded rather than pointed, and the rounding gives the ratio of the two. That converts an unremarkable brief brightening into a measurement of an angular Einstein radius, and therefore into a mass–distance relation rather than a single degenerate timescale.

So the shortest events are the ones where the source’s size stops being a correction and becomes the whole measurement, and it is the same fold physics as the rest of this essay operating on a lens with no fold in it at all.

And a longer event at a larger impact parameter, which is the shape most of the survey’s detections actually have.

A magnification, and a spike inside it. The brightness of a background star as a foreground one passes in front of it. The smooth curve is exact: a point mass magnifies a point source by (u² + 2)/(u√(u² + 4)), which peaks at 3.4 for this track. The spike is the planet, of mass ratio 0.001, lensing one of the two images. Its duration is the Einstein time scaled by √q — about 1.9 days against 60 days — so the whole planetary signal is a few hours in an event lasting a month, and it never repeats. The deviation is drawn in the approximation that the planet lenses the image in isolation; a real caustic crossing has structure this smooths over, and its true height is set by the source's size rather than by the geometry.
Fig. 9 A sixty-day event at a third of an Einstein radius. The peak is broad and low, the planetary deviation is a small feature on its flank, and the whole event is over in two months — which is the regime in which a survey has to decide in real time whether to trigger follow-up on something it cannot yet identify.

Where this ladder goes next

Later rungs on this anchor: microlens parallax, in which the Earth’s own motion during a long event distorts the light curve and gives a second mass–distance constraint; space-based parallax, where a satellite an astronomical unit away sees a different light curve and the difference is the geometry; astrometric microlensing, where the image centroid shifts by hundreds of microarcseconds and is now measurable, giving θE\theta_{\rm E} without needing a caustic crossing; free-floating planets, which produce very short events with no host and whose abundance is a direct test of planet formation; and the survey design problem, which is the question of what cadence and which fields maximise the number of anomalies that are actually caught.

About the same objects

Not linked from either essay — found by the objects both name.

What links here

Essays that link to this one from their own argument.

The objects this essay names

Each one links to every other essay that touches it.

Angular einstein radiusBinary lensCausticClose wide degeneracyCritical curveDegeneracyFinite-source effectFold crossingImage multiplicityLens equationMass ratioMicrolensing