Part II · Special Relativity — Chapter 2.1

The Crisis of 1900

Three facts, each of them solid. Any two of them are comfortable together. All three cannot be true at once, and this chapter is the proof.

Where we are

Part I rebuilt mechanics from an action principle, and it worked. One scalar function, varied over paths, gave back Newton's second law. It gave that law back in any coordinates. It handled constraints without ever naming a constraint force. And through Noether it turned every symmetry into a conservation law. The machinery is sound, and nothing in Part II takes any of it back.

Now we point it at light, and it breaks.

Here is the shape of the chapter, stated in advance so you can watch it close. There are three claims. (1) The laws of physics take the same form in every inertial frame, and the transformation relating those frames is x=xvtx'=x-vt, t=tt'=t. That is Galilean relativity. It is older than Newton, and Newton's laws satisfy it exactly. (2) Maxwell's equations describe every electric and magnetic phenomenon anyone had measured by 1890, and they imply that electromagnetic disturbances propagate at a definite speed cc built out of two laboratory constants. (3) Experiment finds no trace of the frame that (1) and (2) together demand must exist.

Each claim is going to be established here rather than asserted. Claim (1) comes by direct calculation in §1. Claim (2) comes from deriving the wave equation out of Maxwell in §2. Claim (3) comes from working out what Michelson and Morley should have seen in §5 and comparing that with what they did see. Section 4 is the hinge. It shows by explicit chain rule that Maxwell's equations change form under the transformation in (1), and that is where the contradiction actually lives.

One thread is picked up from earlier. Chapter 1.1 §3.3 showed that Newton's third law fails for two charges in relative motion: an isolated pair of particles has dPmech/dt0\dd\vv P_{\text{mech}}/\dd t\neq\vv 0, at a computable rate suppressed by v2/c2v^{2}/c^{2}. The only available repair was that the field carries the missing momentum, and that suspicious factor v2/c2v^{2}/c^{2} was left hanging. It is about to reappear as the size of the effect Michelson and Morley went looking for. It is the same v2/c2v^{2}/c^{2}, and the match is not a coincidence. Both are the leading signature of one fact: electromagnetism does not transform the way mechanics does.

What this chapter does not do. It does not derive the Lorentz transformation. That is Chapter 2.2's job, and doing it early would spoil the only honest reason to accept it. That reason is that by the end of §7 you will want some transformation that resolves the contradiction, and there turns out to be essentially one candidate. This chapter ends with two postulates and nothing else.

Tools you'll need  — Chapter 0.3: the binomial series, and what "to leading order in β2\beta^{2}" means quantitatively. Chapter 0.6 §5.2: the multivariable chain rule for a change of coordinates, which is all that §4 is, ground out. Chapter 0.7 §4 (curl), §3 (divergence), §7 (the Laplacian and the second-derivative identities): all four Maxwell equations are written in that language, and §2 manipulates them there. Chapter 0.8 §7.6: the wave equation, its solutions f(xvt)f(x-vt), and the fact that a wave equation's speed is a property of the medium. Chapter 1.1 §3.3: the third law failing for moving charges, and the loose end that failure left.

1 · Galilean relativity, stated precisely

Let's start by looking closely at an idea that has been central to physics for over four hundred years: Galilean relativity. It is that old, and it is correct. To understand the crisis that sparked modern physics, we first need to separate the core principle of relativity from the specific mathematical transformation we use to calculate it. For centuries these two concepts were treated as the exact same thing. Once we pry them apart, the path forward becomes much clearer. Chapter 2.2 is where they finally come apart for good.

1.1 · The transformation

Imagine two observers, each with their own coordinate system, or "frame", of reference. We will call the stationary frame SS and the moving frame SS'. Let's give them both a measuring tape and a clock, and assume they perfectly agree on how to use them.

To keep the maths simple, we'll align their axes and have them start their stopwatches at t=0t=0 exactly when they pass each other. After that moment, the SS' frame slides away in the positive xx direction at a steady speed vv, as measured in SS. This setup is known as standard configuration, and it is the baseline for everything in Part II.

Now, suppose something happens — an "event" — and both observers write down where and when it occurred.

  • The observer in frame SS records the position as xx and the time as tt.
  • Meanwhile, the origin of the moving frame SS' has shifted forward by a distance vtvt. Because of this shift, the moving observer will measure the event's position as xvtx-vt.

We are also going to make a crucial assumption here, and we put it in explicitly so that we can watch it die later: both observers will record the exact same time tt. This leads us directly to the classical Galilean transformation:

x=xvt,y=y,z=z,t=t. x' = x - vt, \qquad y' = y, \qquad z' = z, \qquad t' = t. (2.1.1)

Let's pause and look at that final equation, t=tt'=t. It might look like simple bookkeeping, but it actually contains a massive physical claim: there is a single, universal clock ticking at the exact same rate for everyone in the universe, regardless of how fast they are moving, so that two events either are simultaneous or are not, with no reference to who is asking. Isaac Newton built his mechanics on this idea of absolute time, and he stated it in exactly those terms: absolute time, flowing equably, without relation to anything external. It feels so intuitive to our everyday experience that it is hard to realise it is merely an assumption.

As we dive deeper into relativity, the first three equations of (2.1.1) will survive into Chapter 2.2 with a few tweaks. This universal time equation will completely fall apart.

1.2 · Velocities add

Next, let's see how our observers view motion. Imagine a particle travelling along a path. The stationary observer tracks its position as x(t)x(t). Using our transformation (2.1.1), the moving observer tracks its position as x(t)=x(t)vtx'(t') = x(t) - vt.

To find the particle's velocity, we just need to take the derivative. Since we assumed that time is the same for both observers (t=tt'=t), differentiating with respect to tt' is identical to differentiating with respect to tt:

u    dxdt  =  ddt(xvt)  =  dxdtv  =  uv. u' \;\equiv\; \dv{x'}{t'} \;=\; \dv{}{t}\big(x - vt\big) \;=\; \dv{x}{t} - v \;=\; u - v. (2.1.2)

This equation tells us something very familiar: velocities add or subtract depending on your point of view. If you throw a ball forward at uu' inside a train moving at vv, a person standing on the platform sees the ball moving at u+vu'+v, the speed of the throw plus the speed of the train. Nobody has ever been surprised by this, and for everyday objects it is incredibly accurate.

But notice the hidden mechanics here. The subtraction comes from the shift in space, while the ease of the derivative relies entirely on the assumption that t=tt'=t. Chapter 2.2 keeps the first ingredient and destroys the second. When that assumption breaks later on, this rule for adding velocities will have to change too, and (2.1.2) is what changes.

1.3 · Accelerations do not change, so F=ma\vv F=m\vv a survives

Now differentiate (2.1.2) once more. The relative velocity vv is a constant, since that is what it means for one inertial frame to move uniformly relative to another, so its derivative vanishes:

a  =  dudt  =  ddt(uv)  =  dudt  =  a. a' \;=\; \dv{u'}{t'} \;=\; \dv{}{t}\big(u - v\big) \;=\; \dv{u}{t} \;=\; a. (2.1.3)

In three dimensions the same computation on each component gives r¨=r¨\ddot{\vv r}' = \ddot{\vv r}. Acceleration is a Galilean invariant. Two observers in uniform relative motion disagree about where a particle is and about how fast it is going, and they agree exactly about how it is accelerating.

That single fact is why Newtonian mechanics is compatible with the principle of relativity. Here is the argument. Newton's second law in SS reads F=mr¨\vv F = m\ddot{\vv r}, and we want to know what it looks like in SS', so transform it piece by piece. The mass mm is a property of the body and does not depend on who is watching. The acceleration is unchanged by (2.1.3). So the law in SS' reads F=mr¨\vv F = m\ddot{\vv r}', and the only remaining question is whether the force is the same in both frames.

For the forces Newtonian physics actually uses, it is. Every one of them depends on the particles' relative positions, and possibly on their relative velocities:

rirj=(rivt)(rjvt)=rirj,r˙ir˙j=(r˙iv)(r˙jv)=r˙ir˙j. \begin{aligned} \vv r_{i}' - \vv r_{j}' &= \big(\vv r_{i} - \vv v t\big) - \big(\vv r_{j} - \vv v t\big) = \vv r_{i}-\vv r_{j},\\[4pt] \dot{\vv r}_{i}' - \dot{\vv r}_{j}' &= \big(\dot{\vv r}_{i} - \vv v\big) - \big(\dot{\vv r}_{j} - \vv v\big) = \dot{\vv r}_{i}-\dot{\vv r}_{j}. \end{aligned} (2.1.4)

The boost cancels in both differences. Gravity depends on rirj\abs{\vv r_{i}-\vv r_{j}}, and so do Coulomb's law and Hooke's law. Even a drag force proportional to the relative velocity of body and fluid is untouched. So F=F\vv F' = \vv F for each of them, and

F=mr¨F=mr¨. \vv F = m\ddot{\vv r} \qquad\Longrightarrow\qquad \vv F' = m\ddot{\vv r}'. (2.1.5)

The law has the same shape in the new frame. Not merely a shape that can be corrected into the old one by adding terms, but the identical shape, with the identical symbols meaning the identical things. Compare Chapter 1.1 §4.2, where writing F=ma\vv F=m\vv a in polar coordinates produced two extra terms out of nothing. Nothing like that happens here.

Grind box — the full Galilean group, and what "same form" is allowed to mean

(2.1.1) is one member of a larger family. The complete set of transformations that carry one inertial frame to another, in Newtonian physics, is

r  =  Rr    vt  +  d,t  =  t+s, \vv r' \;=\; R\,\vv r \;-\; \vv v\,t \;+\; \vv d, \qquad t' \;=\; t + s,

where RR is a constant rotation matrix (33 parameters), v\vv v a constant boost velocity (33), d\vv d a constant spatial displacement (33) and ss a constant time offset (11). Ten parameters. They compose and invert among themselves, since the composition of two boosts is a boost with v1+v2\vv v_{1}+\vv v_{2} and the inverse of a rotation is a rotation. So they form a group, the Galilean group. Chapter 6.1 gives that word a definition. For now the content is just that "change of inertial frame" is a closed operation.

Differentiating twice kills d\vv d and vt\vv v t outright and leaves r¨=Rr¨\ddot{\vv r}' = R\,\ddot{\vv r}, so acceleration is not quite invariant under the full group. It rotates, as a vector must. But it rotates the same way F\vv F does, so F=mr¨\vv F = m\ddot{\vv r} still maps to F=mr¨\vv F' = m\ddot{\vv r}'. That is the precise sense in which "the form is unchanged": both sides of the equation transform identically, so the equality is preserved. Hold onto that formulation. It is the exact criterion Chapter 2.4 turns into the definition of a tensor equation, and it is the reason index notation exists.

What is not allowed. Let the relative velocity depend on time, r=rs(t)\vv r' = \vv r - \vv s(t) with s¨0\ddot{\vv s}\neq\vv 0. Then

r¨=r¨s¨(t)mr¨=Fms¨(t), \ddot{\vv r}' = \ddot{\vv r} - \ddot{\vv s}(t) \qquad\Longrightarrow\qquad m\ddot{\vv r}' = \vv F - m\ddot{\vv s}(t),

and an extra term has appeared that is not any interaction between bodies. That is the fictitious force of Chapter 1.1 §1.1, and its presence is exactly why the first law is an existence claim about a restricted class of frames rather than a corollary of the second. The principle of relativity was never a claim about all frames. It is a claim about the inertial ones, and the boost velocity must be constant for the argument of §1.3 to go through.

One more piece of fine print, cheap now and expensive later. Nothing above proves that (2.1.1) is the only transformation with these properties. We derived F=ma\vv F'=m\vv a' from the Galilean transformation. We did not show that a theory respecting the principle of relativity has to use it. That gap is the whole of Chapter 2.2, and the reason nobody looked into it for three hundred years is that t=tt'=t did not feel like an assumption.

1.4 · The principle itself

Now separate the two ideas that §1.1–1.3 ran together.

The principle of relativity

The laws of physics take the same form in all inertial frames. No experiment performed entirely within a uniformly moving laboratory can determine that laboratory's velocity.

This is not Einstein's. It is Galileo's, from 1632, and his statement of it is better than most modern ones. Shut yourself below decks on a large ship with some flies, a bowl of water, and a friend to throw you a ball. You will find that everything proceeds exactly as it did in port. The flies do not pile up at the stern, the ball needs no extra effort thrown forwards rather than backwards, and the drops fall straight into the vessel beneath. ⚑ (Paraphrased from the Dialogo. The reasoning is his, the words are not.) By 1900 nobody proposed to abandon it. It was, and is, one of the best-tested statements in physics.

The principle says something about the relationship between frames. It does not say what the transformation between them is. Those are separate questions, and the failure to notice they were separate is the single reason the crisis took forty years to resolve. Written out:

ClaimStatus in 1900
(P)The laws take the same form in all inertial framesUniversally accepted; verified constantly
(G)Inertial frames are related by (2.1.1), with t=tt'=tAssumed without comment by everyone

Everybody held both. It had never occurred to anyone that they were different claims, because (G) was not perceived as a claim at all. It was perceived as what the words in (P) meant. The resolution of this chapter's crisis is that (P) is right and (G) is wrong. Keep the pair separate from here on, because the rest of the chapter is about the strain between them.

In plain terms 2.1.1

A principle and the transformation that implements it are separate claims, and for three centuries nobody noticed, because the second was never perceived as a claim at all but as what the words of the first meant. The principle is Galileo's, from 1632: shut yourself below decks with some flies and a bowl of water, and everything proceeds as it did in port, so no experiment inside a smoothly moving laboratory can reveal how fast it is going.

The implementation is the familiar one. Subtract the distance the other origin has slid, leave the two directions across the motion alone, and give both observers the same reading for when anything happened. That last instruction carries all the weight, because it asserts a single universal clock, so that two events either are simultaneous or are not, with no reference to who is asking. It is the assumption the previous part promised would have to go.

Mechanics does not mind in the slightest. Differentiate twice and the steady relative velocity vanishes, so both observers agree about acceleration; and every force Newtonian physics uses depends on separations between bodies, which the shift leaves untouched. The law reappears in the moving frame with the identical shape rather than one needing repair by extra terms, which is exactly what the turning basis directions of polar coordinates had inflicted on it.

2 · Maxwell's equations produce a speed

Now the second of the three claims. It has to be derived rather than quoted, because the number that falls out at the end is the whole point.

⚑ Quoted, not derived — Maxwell's equations

The four equations below are the empirical input of this chapter. Each one summarises a body of nineteenth-century laboratory work, from Coulomb, Gauss, Ørsted, Ampère and Faraday. The last term in the last equation is Maxwell's own addition, which Chapter 0.7 Problem 4 showed is forced by charge conservation and the identity (×A)=0\nabla\cdot(\nabla\times\vv A)=0. In vacuum (ρ=0\rho=0, J=0\vv J=\vv 0) and in SI units:

E=0,×E=Bt,B=0,×B=μ0ϵ0Et. \begin{aligned} \nabla\cdot\vv E &= 0, &\qquad \nabla\times\vv E &= -\pdv{\vv B}{t},\\[5pt] \nabla\cdot\vv B &= 0, &\qquad \nabla\times\vv B &= \mu_{0}\epsilon_{0}\,\pdv{\vv E}{t}. \end{aligned}

Nothing else about electromagnetism is assumed anywhere in this chapter. The two constants ϵ0\epsilon_{0} and μ0\mu_{0} are measured, and §2.3 says exactly how. Chapter 2.6 rewrites all four equations as a single manifestly Lorentz-covariant statement, and shows that they were never four equations at all. Chapter 6.3 derives them from a symmetry principle. Here they are input.

2.1 · One identity first

The derivation needs the curl of a curl, which Chapter 0.7 §7 did not have occasion to compute. It is worth deriving rather than looking up, because the pattern of the answer is not guessable.

×(×A)  =  (A)    2A. \nabla\times\big(\nabla\times\vv A\big) \;=\; \nabla\big(\nabla\cdot\vv A\big) \;-\; \nabla^{2}\vv A. (2.1.6)

Here 2A\nabla^{2}\vv A means the Laplacian of Chapter 0.7 §7.5 applied to each Cartesian component separately. That definition is only available because the components are Cartesian, a caveat that becomes important in Chapter 3.3 and is harmless here.

Grind box — (2.1.6) by brute force, one component

Write C=×A\vv C = \nabla\times\vv A, whose components are, from Chapter 0.7 §4,

Cx=yAzzAy,Cy=zAxxAz,Cz=xAyyAx. C^{x} = \partial_{y}A^{z}-\partial_{z}A^{y}, \qquad C^{y} = \partial_{z}A^{x}-\partial_{x}A^{z}, \qquad C^{z} = \partial_{x}A^{y}-\partial_{y}A^{x}.

Take the xx-component of ×C\nabla\times\vv C and substitute:

(×C)x=yCzzCy=y(xAyyAx)    z(zAxxAz)=yxAy+zxAz    yyAxzzAx. \begin{aligned} \big(\nabla\times\vv C\big)^{x} &= \partial_{y}C^{z} - \partial_{z}C^{y}\\[3pt] &= \partial_{y}\big(\partial_{x}A^{y}-\partial_{y}A^{x}\big) \;-\; \partial_{z}\big(\partial_{z}A^{x}-\partial_{x}A^{z}\big)\\[3pt] &= \partial_{y}\partial_{x}A^{y} + \partial_{z}\partial_{x}A^{z} \;-\; \partial_{y}\partial_{y}A^{x} - \partial_{z}\partial_{z}A^{x}. \end{aligned}

Now the trick, and it is the only step with any content. Add and subtract the missing term xxAx\partial_{x}\partial_{x}A^{x}. The first group then completes to a divergence and the second to a Laplacian:

(×C)x=xxAx+yxAy+zxAzx(A)    xxAx+yyAx+zzAx2Ax. \begin{aligned} \big(\nabla\times\vv C\big)^{x} &= \underbrace{\partial_{x}\partial_{x}A^{x} + \partial_{y}\partial_{x}A^{y} + \partial_{z}\partial_{x}A^{z}}_{\textstyle \partial_{x}\left(\nabla\cdot\vv A\right)} \;-\; \underbrace{\partial_{x}\partial_{x}A^{x}+\partial_{y}\partial_{y}A^{x}+\partial_{z}\partial_{z}A^{x}}_{\textstyle \nabla^{2}A^{x}}. \end{aligned}

Both regroupings used Clairaut's theorem (Chapter 0.6 §6.1) to swap yx\partial_{y}\partial_{x} for xy\partial_{x}\partial_{y}, and its hypothesis is continuous second partials. That is the same hypothesis that made (×A)=0\nabla\cdot(\nabla\times\vv A)=0 work in Chapter 0.7 §7.2. The yy and zz components follow by cycling xyzxx\to y\to z\to x, under which every expression above is invariant. Hence (2.1.6). ✓

Why the answer looks like that. ×(×  )\nabla\times(\nabla\times\ \cdot\ ) is a second-order operator, so it must be some combination of ij\partial_{i}\partial_{j} acting on AkA^{k}. Only two such combinations produce a vector: the gradient of the divergence, and the Laplacian of the vector. The identity says the coefficients are +1+1 and 1-1. In Chapter 3.5 this becomes the statement that the operator dd\dd\star\dd splits in exactly this way, and the minus sign is the one that makes the electromagnetic wave equation come out with the right relative sign between space and time.

2.2 · Take the curl of Faraday's law

We want an equation for E\vv E alone. Faraday's law relates E\vv E to B\vv B, and Ampère–Maxwell relates B\vv B back to E\vv E. So apply one to the other and eliminate B\vv B. The operation that lets Ampère–Maxwell in is the curl, because the curl of B\vv B is precisely what that law tells us. Start from Faraday and hit both sides with ×\nabla\times:

×(×E)  =  ×(Bt)  =  t(×B). \nabla\times\big(\nabla\times\vv E\big) \;=\; \nabla\times\left(-\pdv{\vv B}{t}\right) \;=\; -\,\pdv{}{t}\big(\nabla\times\vv B\big). (2.1.7)

The second step swapped a space derivative with a time derivative. That is Clairaut again. The curl is built from x,y,z\partial_{x},\partial_{y},\partial_{z}, none of which is t\partial_{t}, so as long as B\vv B has continuous second partials the two operations commute. It is a small step and worth naming, because in Chapter 3.3 the analogous swap will fail, and the failure will be the curvature of spacetime.

Now substitute Ampère–Maxwell on the right, with ×B=μ0ϵ0E/t\nabla\times\vv B = \mu_{0}\epsilon_{0}\, \partial\vv E/\partial t:

×(×E)  =  μ0ϵ02Et2. \nabla\times\big(\nabla\times\vv E\big) \;=\; -\,\mu_{0}\epsilon_{0}\,\pdv{^{2}\vv E}{t^{2}}. (2.1.8)

Next expand the left-hand side with (2.1.6), using E=0\nabla\cdot\vv E=0, which holds because we are in vacuum with no charge anywhere to source a divergence:

×(×E)  =  (E)=0    2E  =  2E. \nabla\times\big(\nabla\times\vv E\big) \;=\; \nabla\underbrace{\big(\nabla\cdot\vv E\big)}_{\textstyle =\,0} \;-\; \nabla^{2}\vv E \;=\; -\nabla^{2}\vv E. (2.1.9)

Set the two expressions equal and cancel the overall minus sign:

  2E  =  μ0ϵ02Et2.   \boxed{\;\nabla^{2}\vv E \;=\; \mu_{0}\epsilon_{0}\,\pdv{^{2}\vv E}{t^{2}}.\;} (2.1.10)

Look at what has happened. We started with four equations coupling two fields, performed three manipulations, and produced an equation for E\vv E by itself with no B\vv B in it. And the equation is not a new object. It is the wave equation of Chapter 0.8 §7.6, so let us set the two side by side:

2ut2  =  v22ux2,here2Et2  =  1μ0ϵ02E, \pdv{^{2}u}{t^{2}} \;=\; v^{2}\,\pdv{^{2}u}{x^{2}}, \qquad\text{here}\qquad \pdv{^{2}\vv E}{t^{2}} \;=\; \frac{1}{\mu_{0}\epsilon_{0}}\,\nabla^{2}\vv E, (2.1.11)

with three spatial dimensions instead of one, and a vector where the string had a scalar. Now read off the coefficient:

v2  =  1μ0ϵ0  c    1μ0ϵ0.   v^{2} \;=\; \frac{1}{\mu_{0}\epsilon_{0}} \qquad\Longrightarrow\qquad \boxed{\;c \;\equiv\; \frac{1}{\sqrt{\mu_{0}\epsilon_{0}}}.\;} (2.1.12)

Nobody put a speed in. Two constants measured with capacitors and current-carrying wires went in, and a propagation speed came out.

Grind box — the same for B\vv B, and the plane-wave solution

The magnetic field obeys the identical equation. Run the argument the other way. Take the curl of Ampère–Maxwell:

×(×B)=μ0ϵ0t(×E)=μ0ϵ02Bt2, \nabla\times\big(\nabla\times\vv B\big) = \mu_{0}\epsilon_{0}\,\pdv{}{t}\big(\nabla\times\vv E\big) = -\mu_{0}\epsilon_{0}\,\pdv{^{2}\vv B}{t^{2}},

using Faraday in the last step. Now expand the left with (2.1.6) and use B=0\nabla\cdot\vv B=0, which holds always rather than merely in vacuum, since there are no magnetic charges. So

2B  =  μ0ϵ02Bt2, \nabla^{2}\vv B \;=\; \mu_{0}\epsilon_{0}\,\pdv{^{2}\vv B}{t^{2}},

the same equation with the same speed. The two fields propagate together, which they must, since they are two aspects of one thing, as Chapter 2.6 makes literal.

A plane wave, and c=ω/kc=\omega/k. Try a solution travelling along xx with E\vv E pointing along yy, so that E=y^E0cos(kxωt)\vv E = \hat{\vv y}\,E_{0}\cos(kx-\omega t). Then 2E=k2E\nabla^{2}\vv E = -k^{2}\vv E and 2E/t2=ω2E\partial^{2}\vv E/\partial t^{2} = -\omega^{2}\vv E, so (2.1.10) requires

k2=μ0ϵ0ω2ωk=1μ0ϵ0=c. -k^{2} = -\mu_{0}\epsilon_{0}\,\omega^{2} \qquad\Longrightarrow\qquad \frac{\omega}{k} = \frac{1}{\sqrt{\mu_{0}\epsilon_{0}}} = c.

That is Chapter 0.8's relation between frequency, wavenumber and wave speed, now with the speed fixed by two laboratory constants. Notice that it holds for every kk. The speed does not depend on the wavelength, so vacuum is non-dispersive and a pulse of any shape travels undistorted. That is Chapter 0.8's f(xct)f(x-ct), and it is why we see sharp images of distant stars rather than smeared ones.

One honest caveat. (2.1.10) is a consequence of Maxwell's equations, not an equivalent of them. Every solution of Maxwell solves it, and not every solution of it solves Maxwell. Feeding the plane wave back into the original four equations imposes two further conditions we did not need above. First, E=0\nabla\cdot\vv E=0 forces E\vv E\perp propagation direction, so the wave is transverse. Second, Faraday fixes B=x^×E/c\vv B = \hat{\vv x}\times\vv E/c, so B\vv B is perpendicular to both E\vv E and the direction of travel, in phase with E\vv E, and smaller by a factor cc. None of that is needed for this chapter's argument, and all of it matters in Chapter 2.6.

2.3 · Put the numbers in

Now the moment. Here is where ϵ0\epsilon_{0} and μ0\mu_{0} come from, because the provenance is the point.

ϵ0\epsilon_{0}, the permittivity of free space, is fixed by measuring the capacitance of a parallel-plate capacitor. Put a known charge on two plates of known area and separation, measure the voltage, and ϵ0\epsilon_{0} falls out of C=ϵ0A/dC=\epsilon_{0}A/d. This is a tabletop electrostatics experiment involving no light, no motion and no waves. Its value is ϵ0=8.854×1012 C2N1m2\epsilon_{0}=8.854\times10^{-12}\ \mathrm{C^{2}\,N^{-1}\,m^{-2}}.

μ0\mu_{0}, the permeability of free space, is fixed by measuring the force between two parallel current-carrying wires. Run known currents through them, measure the attraction per unit length, and μ0\mu_{0} falls out of F/=μ0I1I2/2πdF/\ell = \mu_{0}I_{1}I_{2}/2\pi d. This is a tabletop magnetostatics experiment involving no light and no waves either. Its value is μ0=4π×107 NA2=1.2566×106 NA2\mu_{0}=4\pi\times10^{-7}\ \mathrm{N\,A^{-2}} = 1.2566\times10^{-6}\ \mathrm{N\,A^{-2}}.

Multiply them together and take the reciprocal square root:

μ0ϵ0=(1.2566×106)(8.854×1012)=1.1126×1017 s2m2,c=1μ0ϵ0=13.3356×109 sm1  =  2.998×108 ms1. \begin{aligned} \mu_{0}\epsilon_{0} &= \big(1.2566\times10^{-6}\big)\big(8.854\times10^{-12}\big) = 1.1126\times10^{-17}\ \mathrm{s^{2}\,m^{-2}},\\[6pt] c = \frac{1}{\sqrt{\mu_{0}\epsilon_{0}}} &= \frac{1}{3.3356\times10^{-9}\ \mathrm{s\,m^{-1}}} \;=\; 2.998\times10^{8}\ \mathrm{m\,s^{-1}}. \end{aligned} (2.1.13)

The measured speed of light, from Fizeau's toothed wheel in 1849 and Foucault's rotating mirror in 1862, was ⚑ about 3.0×108 ms13.0\times10^{8}\ \mathrm{m\,s^{-1}}.

What just happened

A capacitor and a pair of wires told you the speed of light.

Not "a number of the same order". The speed of light, to the precision of the input data, out of two constants that were measured in experiments containing no light at all. There is no route by which optics could have leaked into either measurement. One is a static charge on metal plates, and the other a steady current in a wire. The only way the number can come out right is if light is the propagating disturbance (2.1.10) describes.

⚑ Maxwell drew exactly that conclusion in 1862, in one of the most consequential sentences in physics: we can scarcely avoid the inference that light consists in the transverse undulations of the same medium which is the cause of electric and magnetic phenomena. Optics was a subject two thousand years old, with its own laws, its own instruments and its own practitioners. It stopped being a separate science and became a chapter of electromagnetism. Hertz generated and detected the waves directly in 1887, and radio, X-rays and everything else followed from taking (2.1.12) seriously across the spectrum.

This is what a successful unification looks like from the inside. Not a philosophical reorganisation, but two numbers you measured for unrelated reasons combining into a third number you had already measured for a third reason.

⚠ Why this isn't obvious — is (2.1.13) circular?

A sharp reader will object. In the SI system in force for most of the twentieth century, μ0\mu_{0} was defined to be exactly 4π×1074\pi\times10^{-7}, as part of the definition of the ampere. And since 1983 the metre has been defined by fixing cc exactly. If the constants are defined rather than measured, has anything been discovered?

Yes, and the objection is worth taking seriously, because it locates the physical content precisely. Unit conventions can move a numerical factor from one constant to another. They cannot create a relationship between independently measurable quantities. The invariant statement is this: the ratio of the electrostatic to the electromagnetic unit of charge has the dimensions of a speed, and that speed can be measured with nothing but charges, currents and a balance. Weber and Kohlrausch did precisely that in 1856, discharging a capacitor of measured electrostatic capacity through a galvanometer of measured electromagnetic sensitivity, and got ⚑ 3.107×108 ms13.107\times10^{8}\ \mathrm{m\,s^{-1}}, six years before Maxwell's remark and with no reference to light whatever. That measurement is the physics. It cannot be conjured out of a choice of units, and it is the reason the coincidence was persuasive at the time.

In modern SI the bookkeeping runs the other way. Now cc is exact by definition, μ0\mu_{0} is a measured quantity differing from 4π×1074\pi\times10^{-7} in the ninth decimal place, and (2.1.12) is the relation used to determine it. The equation is doing work in both directions. Only the label "defined" has moved.

Familiar ground — a rate constant predicted from two static measurements

The shape of the argument in this section is one you have run yourself, and having run it is what makes §2.3 feel inevitable rather than lucky.

Take a drug obeying the first-order kinetics of Chapter 0.1's callout. Two of its constants can be measured without ever watching a concentration fall. The volume of distribution comes from a single dilution: give a known dose, let it mix, measure the concentration once, and Vd=D/C0V_{d}=D/C_{0}. The clearance comes from a steady state: infuse at a known rate until the concentration stops changing, and CL=Rinf/Css\mathrm{CL}=R_{\mathrm{inf}}/C_{\mathrm{ss}}. Neither experiment contains a half-life, and neither requires waiting for anything to decay. Yet

dAdt  =  CLAVdt1/2  =  ln2  VdCL, \dv{A}{t} \;=\; -\,\mathrm{CL}\,\frac{A}{V_{d}} \qquad\Longrightarrow\qquad t_{1/2} \;=\; \frac{\ln 2\;V_{d}}{\mathrm{CL}},

gives a time, assembled out of a volume and a flow. That is (2.1.12) with different letters on it: a rate of propagation manufactured out of two constants measured in experiments containing no propagation.

The step that matters is the comparison, not the prediction. Computing t1/2t_{1/2} is bookkeeping. Going out, measuring the terminal slope of a real concentration–time curve, and finding that it agrees is a test. And when it fails to agree, the disagreement is informative rather than embarrassing, because it says a second compartment is hiding, which is the case Chapter 0.8's callout works through in full. Maxwell's step is exactly this comparison, performed once. A capacitor and a current balance predicted 2.998×108 ms12.998\times10^{8}\ \mathrm{m\,s^{-1}}, Fizeau's toothed wheel had already measured 3.0×1083.0\times10^{8}, and the two numbers had no business being the same.

One difference is worth naming, because it is why his conclusion was so much larger than yours. VdV_{d} and CL\mathrm{CL} are properties of a drug and a patient, so agreement validates a model of that pair and says nothing about anything else. ϵ0\epsilon_{0} and μ0\mu_{0} are properties of empty space, so agreement could not possibly be a fact about a particular apparatus, and the only object left available to be identified was light itself. A model check and a unification share their arithmetic and differ entirely in their consequences, and what separates them is what the constants belong to.

In plain terms 2.1.2

Nothing resembling a speed was put into the four equations of electromagnetism, and a speed came out. Apply the law tying a changing magnetic field to a circulating electric one to its partner, which ties a changing electric field to a circulating magnetic one, and the two fields uncouple, leaving one equation apiece. That equation is the one a chain of masses and springs produced back in the toolkit, and reading off its coefficient gives a speed built from two laboratory constants.

The provenance of those constants is the point. One is fixed by putting a known charge on two metal plates and measuring the voltage across them; the other by running known currents through parallel wires and measuring how hard they pull. Neither experiment contains any light, any motion, or any waves.

No channel exists by which optics could have leaked into a static charge on metal or a steady current in a wire, so the only way the number comes out right is if light is the disturbance those equations describe. Maxwell drew that conclusion in 1862, and optics, two thousand years old and with its own laws and its own instruments, stopped being a separate science. This is what a successful unification looks like from the inside: two numbers measured for unrelated reasons combining into a third that had already been measured for a third.

3 · The question nobody could answer: cc with respect to what?

Now the trouble starts, and it starts with a question that is embarrassing in its simplicity.

(2.1.12) gives a speed. A speed is a rate of change of position, and position is measured relative to something. So here is the question: relative to what is the light going at cc?

3.1 · Every other wave answers this question easily

Chapter 0.8 §7.6 built the wave equation from a chain of masses and springs and got v=T/ρv=\sqrt{T/\rho}. Look at what that formula is made of: the tension and the linear density of the string. Both are properties of the medium, so the speed is a property of the medium. It is therefore the speed relative to the medium. If the string is being reeled in while the wave travels along it, an observer at the side of the room measures something else.

Every wave in nineteenth-century physics was like this, without exception:

WaveSpeedRelative to
Transverse wave on a stringT/ρ\sqrt{T/\rho}the string
Sound in airγp/ρ343 ms1\sqrt{\gamma p/\rho}\approx343\ \mathrm{m\,s^{-1}}the air
Ripples on watergλ/2π\sqrt{g\lambda/2\pi} (deep water)the water
Seismic SS-wavesG/ρ\sqrt{G/\rho}the rock

And the consequence is entirely familiar. Stand in a wind of speed uu and sound travelling downwind reaches you at 343+u343+u, upwind at 343u343-u. That is (2.1.2) applied to a wave, and it is measured routinely. The speed in the formula is the speed in the rest frame of the medium, and in any other frame you add the medium's velocity.

3.2 · The wave equation itself makes the point

You do not even need the physical picture. The mathematics says it directly. The wave equation is written in some particular set of coordinates (x,t)(x,t), and its solutions are f(xvt)f(x-vt) and g(x+vt)g(x+vt), which are disturbances moving at ±v\pm v in those coordinates. Suppose the equation holds in one coordinate system with speed vv, and suppose coordinates transform by (2.1.1). Then in a coordinate system moving at uu the same disturbance moves at vuv-u. The number vv in the equation is therefore attached to one preferred frame, namely the frame in which the equation takes that form.

Apply this to (2.1.10). Maxwell's equations, written as they always are written, hold in some frame. In that frame light goes at cc in every direction. In a frame moving at uu through it, light should go at cuc-u in one direction and c+uc+u in the other. So Maxwell's equations appear to single out a preferred frame of reference. That flatly contradicts claim (P) of §1.4, and it does so for reasons that have nothing to do with any experiment.

3.3 · The ether, taken seriously

Given all of §3.1, there was one obvious move, and it was not a stupid one. If light is a wave, it is a wave in something. Name that something the luminiferous ether, and declare Maxwell's equations to hold in the ether's rest frame. The preferred frame of §3.2 is then no longer an embarrassment. It is only the rest frame of a medium, exactly like the air for sound. The principle of relativity survives untouched, because the ether is a physical object, and detecting motion relative to it is no more mysterious than feeling a breeze.

This was not a fudge. It was the only known way for a wave to work, and every other wave anyone had ever studied confirmed it. Refusing to posit a medium in 1880 would have been the strange move, not the sober one. And the hypothesis was productive. It made a sharp prediction: the Earth moves, so there must be an ether wind, so the speed of light must be anisotropic in the laboratory. That prediction is testable, and §5 tests it.

The ether was, however, a demanding object. Its properties were fixed by (2.1.10), and they do not sit comfortably together.

Grind box — what the ether had to be like, quantitatively

It had to be a solid. Light is transverse, since a plane wave has E\vv E perpendicular to the direction of travel, as the previous grind box showed. A transverse wave is a shear disturbance, with neighbouring layers of the medium sliding past one another. Fluids do not resist shear, which is why sound in air is purely longitudinal and why the Earth's liquid outer core transmits no SS-waves. So the ether had to have rigidity. It had to be an elastic solid filling all of space, and the planets had to pass through it without measurable resistance.

And an extraordinarily stiff one. For a transverse wave in an elastic solid the speed is v=G/ρv=\sqrt{G/\rho}, with GG the shear modulus, which is the same square root as T/ρ\sqrt{T/\rho} in Chapter 0.8. Setting v=cv=c fixes neither the stiffness nor the density but their ratio:

Gρ  =  c2  =  8.99×1016 m2s2. \frac{G}{\rho} \;=\; c^{2} \;=\; 8.99\times10^{16}\ \mathrm{m^{2}\,s^{-2}}.

Compare steel: shear modulus 79.3 GPa79.3\ \mathrm{GPa}, density 7850 kgm37850\ \mathrm{kg\,m^{-3}}, so G/ρ=1.01×107 m2s2G/\rho = 1.01\times10^{7}\ \mathrm{m^{2}\,s^{-2}}. And indeed 1.01×107=3.18 kms1\sqrt{1.01\times10^{7}}=3.18\ \mathrm{km\,s^{-1}}, which is the shear-wave speed in steel. The ether's stiffness-to-density ratio therefore had to exceed steel's by a factor

8.99×10161.01×107  =  8.9×109, \frac{8.99\times10^{16}}{1.01\times10^{7}} \;=\; 8.9\times10^{9},

nearly ten billion. You may make the density as small as you like to keep the planets moving, but the ratio is not negotiable, and it is the ratio that has to be explained.

And it had to be undetectable in every other way. Transparent, frictionless, incompressible enough to fill the space between the stars, non-interacting with matter except through this one channel, and possessed of no measurable effect on anything except that light goes through it.

Nineteenth-century physicists were entirely aware of this list, and it bothered them. Kelvin called the ether one of the two clouds on the horizon of physics. But an awkward hypothesis that makes a testable prediction is a perfectly respectable scientific object, and the response was the right one: go and measure the wind.

In plain terms 2.1.3

Sound in air, ripples on a pond, a pulse running down a rope, tremors through rock: every one of them travels at a speed assembled out of properties of the stuff it travels in, and therefore travels at that speed relative to the stuff. Stand in a wind and sound reaches you faster downwind than up, by exactly the wind's speed. So asking what the new speed was measured against was not a foolish question but an obligatory one.

The mathematics says it more sharply than the analogy does. A wave equation is written in some particular set of coordinates and its solutions run at the stated speed in those, so to an observer drifting past at some rate the same disturbance runs at a different one. The equations of electromagnetism therefore single out one frame, contradicting the principle of relativity before any experiment has been performed.

Naming a medium was the sober response rather than a fudge, because it was the only known way for a wave to work. It cost something. The medium had to resist shearing in order to carry a transverse disturbance, so it had to be an elastic solid filling all of space, with a stiffness-to-density ratio ten billion times steel's, through which the planets passed without measurable drag. Nineteenth-century physicists knew that list and it bothered them. An awkward hypothesis making a sharp prediction is still a respectable object.

4 · Maxwell's equations are not Galilean invariant

Section 3 argued from physical analogy that Maxwell's equations must pick out a frame. This section proves it, by computation, with no analogy anywhere. It is the hinge of the chapter.

Strip the problem to its skeleton. Take one Cartesian component of E\vv E and call it ϕ\phi. Suppress yy and zz, so that the Laplacian is 2/x2\partial^{2}/\partial x^{2}. Then (2.1.10) becomes the one-dimensional wave equation:

2ϕx2  =  1c22ϕt2. \pdv{^{2}\phi}{x^{2}} \;=\; \frac{1}{c^{2}}\,\pdv{^{2}\phi}{t^{2}}. (2.1.14)

Here is the question. What does this equation look like in a frame moving at vv along xx? Not what do its solutions look like, but what does the equation look like. If the principle of relativity holds with the Galilean transformation, the answer must be 2ϕ/x2=c22ϕ/t2\partial^{2}\phi/\partial x'^{2} = c^{-2}\,\partial^{2}\phi/\partial t'^{2}, character for character.

4.1 · The chain rule, set up carefully

We have new coordinates x=xvtx'=x-vt and t=tt'=t, and a function ϕ\phi that we may regard either as a function of (x,t)(x,t) or of (x,t)(x',t'). It is one physical field, described two ways. Chapter 0.6 §5.2 gives the rule for converting derivatives: each old derivative becomes a sum over the new ones, weighted by how each new coordinate responds to the old one.

x  =  xxx+txt,t  =  xtx+ttt. \pdv{}{x} \;=\; \pdv{x'}{x}\,\pdv{}{x'} + \pdv{t'}{x}\,\pdv{}{t'}, \qquad\qquad \pdv{}{t} \;=\; \pdv{x'}{t}\,\pdv{}{x'} + \pdv{t'}{t}\,\pdv{}{t'}. (2.1.15)

The four partial derivatives of the new coordinates with respect to the old come straight from x=xvtx'=x-vt and t=tt'=t:

xx=1,xt=v,tx=0,tt=1. \pdv{x'}{x} = 1, \qquad \pdv{x'}{t} = -v, \qquad \pdv{t'}{x} = 0, \qquad \pdv{t'}{t} = 1. (2.1.16)

Take a moment over t/x=0\partial t'/\partial x = 0. It says that changing where you are does not change what time it is. That is the mathematical form of absolute simultaneity. It is the only place t=tt'=t enters this calculation, and it is the entry that Chapter 2.2 will make nonzero. Now substitute (2.1.16) into (2.1.15):

  x  =  x,t  =  t    vx.   \boxed{\;\pdv{}{x} \;=\; \pdv{}{x'}, \qquad\qquad \pdv{}{t} \;=\; \pdv{}{t'} \;-\; v\,\pdv{}{x'}.\;} (2.1.17)

The spatial derivative is untouched, and the time derivative acquires an extra piece. That is the whole asymmetry, and everything below is bookkeeping. The physical reading of the second relation is worth having. Holding your position fixed in the old frame means drifting backwards at v-v in the new one, so the rate of change you measure standing still picks up a term from the motion. It is the convective derivative of fluid mechanics, arrived at without meaning to.

4.2 · Second derivatives, and the term that ruins everything

Grind box — squaring the operators, every step

Space. Apply (2.1.17) twice:

2x2=x(x)=x(x)=2x2. \pdv{^{2}}{x^{2}} = \pdv{}{x}\left(\pdv{}{x}\right) = \pdv{}{x'}\left(\pdv{}{x'}\right) = \pdv{^{2}}{x'^{2}}.

Nothing to do.

Time. Here we must square a sum of two operators. The one thing to be careful about is that operators need not commute, so (A+B)2=A2+AB+BA+B2(A+B)^{2}=A^{2}+AB+BA+B^{2}, and we may not jump to A2+2AB+B2A^{2}+2AB+B^{2} without checking. Write A=/tA=\partial/\partial t' and B=v/xB=-v\,\partial/\partial x':

2t2=(tvx)(tvx)=2t2    vtx    vxt  +  v22x2. \begin{aligned} \pdv{^{2}}{t^{2}} &= \left(\pdv{}{t'} - v\pdv{}{x'}\right)\left(\pdv{}{t'} - v\pdv{}{x'}\right)\\[5pt] &= \pdv{^{2}}{t'^{2}} \;-\; v\,\pdv{}{t'}\pdv{}{x'} \;-\; v\,\pdv{}{x'}\pdv{}{t'} \;+\; v^{2}\pdv{^{2}}{x'^{2}}. \end{aligned}

The two middle terms are equal because of Clairaut's theorem, since mixed partials of a twice-continuously-differentiable function commute. They are also equal because vv is a constant, so it passes through /x\partial/\partial x' without generating anything. Both facts are needed, and both are easy to use without noticing. Hence

2t2  =  2t2    2v2xt  +  v22x2. \pdv{^{2}}{t^{2}} \;=\; \pdv{^{2}}{t'^{2}} \;-\; 2v\,\frac{\partial^{2}}{\partial x'\,\partial t'} \;+\; v^{2}\,\pdv{^{2}}{x'^{2}}.

Assemble. Substitute both results into (2.1.14):

2ϕx2  =  1c2[2ϕt22v2ϕxt+v22ϕx2]. \pdv{^{2}\phi}{x'^{2}} \;=\; \frac{1}{c^{2}}\left[\pdv{^{2}\phi}{t'^{2}} - 2v\,\frac{\partial^{2}\phi}{\partial x'\,\partial t'} + v^{2}\,\pdv{^{2}\phi}{x'^{2}}\right].

Bring the v2v^{2} term to the left and collect the coefficient of 2ϕ/x2\partial^{2}\phi/\partial x'^{2}:

(1v2c2)2ϕx2  +  2vc22ϕxt  =  1c22ϕt2, \left(1 - \frac{v^{2}}{c^{2}}\right)\pdv{^{2}\phi}{x'^{2}} \;+\; \frac{2v}{c^{2}}\,\frac{\partial^{2}\phi}{\partial x'\,\partial t'} \;=\; \frac{1}{c^{2}}\,\pdv{^{2}\phi}{t'^{2}},

which is (2.1.18) below, rearranged.

Here is the result of that grind, written with everything on one side and with βv/c\beta \equiv v/c, as the convention of Part II requires:

(1β2)2ϕx2  +  2vc22ϕxt    1c22ϕt2  =  0. \big(1-\beta^{2}\big)\,\pdv{^{2}\phi}{x'^{2}} \;+\; \frac{2v}{c^{2}}\,\frac{\partial^{2}\phi}{\partial x'\,\partial t'} \;-\; \frac{1}{c^{2}}\,\pdv{^{2}\phi}{t'^{2}} \;=\; 0. (2.1.18)

Compare that with what we wanted, which was (2.1.14) with primes on it. Two things have gone wrong, and they are worth naming separately.

(i) A cross term has appeared. The middle term 2ϕ/xt\propto \partial^{2}\phi/\partial x'\,\partial t' has no counterpart in (2.1.14). It is odd in vv, so reversing the direction of motion changes its sign, and that is precisely how an equation encodes a preferred direction. A term that knows which way is "forwards" cannot appear in an isotropic law.

(ii) The leading coefficient has changed. 11 has become 1β21-\beta^{2}. Even if you could somehow dispose of the cross term, the coefficients no longer match.

The sharp statement

The wave equation changes form under a Galilean boost. If it holds in one inertial frame, it does not hold in any other.

This is not an approximation, not a small effect, and not a subtlety of interpretation. It is a three-line consequence of the chain rule. And since (2.1.14) is a consequence of Maxwell's equations, the same is true of them: Maxwell's equations are not Galilean invariant.

4.3 · What the transformed equation is actually saying

(2.1.18) is not wrong. It is the correct description of a wave in a medium, as seen by an observer moving through that medium. It is worth extracting exactly what it says, both to check the algebra and to see what is at stake.

So look for travelling-wave solutions in the new frame: ϕ=f(xut)\phi = f(x'-ut') for some speed uu and arbitrary shape ff, exactly as in Chapter 0.8 Problem 4. Each derivative brings down a factor, with x2ϕ=f\partial^{2}_{x'}\phi = f'', xtϕ=uf\partial_{x'}\partial_{t'}\phi = -u f'' and t2ϕ=u2f\partial^{2}_{t'}\phi = u^{2}f''. Substitute into (2.1.18) and cancel the common ff'', which is nonzero for any wave worth the name. What is left is a condition on uu alone:

(1β2)    2vuc2    u2c2  =  0u2+2vu(c2v2)=0. \big(1-\beta^{2}\big) \;-\; \frac{2vu}{c^{2}} \;-\; \frac{u^{2}}{c^{2}} \;=\; 0 \qquad\Longleftrightarrow\qquad u^{2} + 2vu - \big(c^{2}-v^{2}\big) = 0. (2.1.19)

That is a quadratic in uu. Solve it, either by completing the square or with the formula:

u  =  v±v2+c2v2  =  v±cu=cvoru=(c+v). \begin{aligned} u \;&=\; -v \pm \sqrt{v^{2} + c^{2} - v^{2}} \;=\; -v \pm c\\[4pt] &\Longrightarrow\qquad u = c-v \quad\text{or}\quad u = -(c+v). \end{aligned} (2.1.20)

Those are exactly the two speeds §3.1 predicted on physical grounds. In the moving frame, light chases off in the +x+x' direction at cvc-v and comes back the other way at c+vc+v. The algebra of §4.2 and the sound-in-a-wind picture of §3.1 are the same statement.

So (2.1.18) is perfectly consistent physics, for a wave in a medium. Sound obeys the analogous equation and nobody minds, because sound has a medium and its rest frame is not mysterious. The entire question is whether light is like that.

The fork

We now have a genuine contradiction between two things both believed in 1900, and there are exactly three ways out. They are mutually exclusive, and they are all uncomfortable.

(A) The principle of relativity fails for electromagnetism. Mechanics obeys it and optics does not. There is a preferred frame, the ether's, and Maxwell's equations hold only there. Every other frame gets (2.1.18). This was the majority view, it is entirely coherent, and it makes a prediction: you can measure your velocity through the ether by measuring the speed of light in different directions.

(B) Maxwell's equations are wrong, or at least are only the low-velocity limit of some Galilean-invariant electrodynamics still to be found. Several people tried, and the attempts (Hertz's, Ritz's) all made predictions that failed. They were also fighting uphill, since Maxwell's equations were passing every test anyone could devise.

(C) The Galilean transformation is wrong. Frames are related by something else, and (2.1.1) is only its small-vv approximation. Then the principle of relativity can hold for both mechanics and electromagnetism. The price is that Newton's laws, which are exactly Galilean invariant by §1.3, must be modified instead.

Notice that (C) requires giving up t=tt'=t. Nothing else in (2.1.1) has enough freedom in it. That is why (C) looked, in 1900, like the least attractive option available, and it is why the experiments came first.

In plain terms 2.1.4

Analogy has been carrying the argument, and analogy can always be resisted, so the case is now made by computation. Take the wave equation, strip it to one dimension, and rewrite it in coordinates sliding past at a steady rate, which is the chain rule and nothing else. The spatial derivative comes through untouched; the time derivative picks up an extra piece, because holding your position fixed in one frame means drifting backwards in the other.

Two things then go wrong at once. A cross term appears with no counterpart in the original, and it reverses sign when the motion does, which is precisely how an equation announces that it knows which way is forwards. The leading coefficient has also changed. So the equation holds in one frame and no other, and the demonstration is three lines long.

One entry in the calculation did all the damage: the statement that changing where you are does not change what time it is. The fork is therefore genuine, and there are exactly three ways through. Either the principle of relativity fails for electromagnetism and a preferred frame exists after all, or the equations of electromagnetism are wrong, or the transformation is wrong and the shared clock goes with it. The transformed equation is not nonsense; it correctly describes a wave in a medium seen by somebody moving through it. Everything turns on whether light is like that.

5 · Michelson–Morley

Branch (A) is the one that is directly testable, and testing it produced the most famous experiment in physics. The logic is short. The Earth orbits the Sun at about 30 kms130\ \mathrm{km\,s^{-1}}, so unless the ether happens to be dragged along perfectly, there is an ether wind blowing through the laboratory. Measure the speed of light along the wind and across it, and the two answers should differ.

The difficulty is that they differ by very little. From §4.3 the two one-way speeds are c±vc\pm v with v/c104v/c\approx10^{-4}. A direct timing measurement would therefore have to resolve one part in 10410^{4} of a quantity that was not itself known that well. The best absolute determination in 1887 was Michelson's own, from 1879, at 299910±50 kms1299\,910\pm50\ \mathrm{km\,s^{-1}}, which is good to about two parts in 10410^{4}. The difference he needed to see was smaller than the error bar on the whole. Michelson's answer was to stop measuring speed and start measuring interference, which compares two light paths against each other and can resolve a hundredth of a wavelength.

5.1 · The instrument

A source of monochromatic light shines on a half-silvered mirror set at 4545^{\circ}, called a beam splitter. Half the light passes through and continues to a mirror at distance L1L_{1}. Half reflects off at right angles and travels to a second mirror at distance L2L_{2}. Both beams return to the splitter, recombine, and go on to a detector. Whether they arrive in phase or out of phase depends on the difference in their travel times, so the recombined beam shows a pattern of interference fringes.

Two facts about this arrangement do the real work. First, the arms are perpendicular, so if there is an ether wind then one arm can lie along it and the other across it. Second, and this is the piece of experimental cunning without which the whole thing is impossible, you cannot know L1L_{1} and L2L_{2} to a fraction of a wavelength. The absolute fringe position is therefore uninterpretable. What you can do instead is rotate the whole apparatus by 9090^{\circ}, which exchanges the roles of the two arms without changing either length, and then watch the fringes move. §5.5 shows that the unknown mismatch L1L2L_{1}-L_{2} cancels out of that measurement exactly.

5.2 · The arm along the wind

Work in the ether's rest frame, where light travels at cc in all directions and the apparatus moves at vv in the +x+x direction. Working in the ether frame is the honest choice, because it is the frame where we are entitled to say light moves at cc, and that is the hypothesis under test.

Set t=0t=0 when a pulse leaves the splitter, and put the splitter at the origin at that instant. Thereafter the splitter is at vtvt and the far mirror at L1+vtL_{1}+vt, both sliding to the right at vv.

Outbound. The pulse moves right at cc, so it is at ctct, while the mirror is at L1+vtL_{1}+vt. They meet when

ctout  =  L1+vtouttout  =  L1cv. ct_{\text{out}} \;=\; L_{1} + v\,t_{\text{out}} \qquad\Longrightarrow\qquad t_{\text{out}} \;=\; \frac{L_{1}}{c-v}. (2.1.21)

The mirror is running away from the light, so the light closes the gap at only cvc-v. Notice that this is not a claim that the light moves at cvc-v. It moves at cc in the ether frame, as assumed. The cvc-v is a closing speed between two objects, and closing speeds may perfectly well exceed cc or fall below it without anything moving at that rate.

Return. Now the light travels left at cc from the mirror's position while the splitter advances to meet it. By the identical argument the gap closes at c+vc+v:

tback  =  L1c+v. t_{\text{back}} \;=\; \frac{L_{1}}{c+v}. (2.1.22)

Add them, put over a common denominator, and divide top and bottom by c2c^{2}:

t  =  L1cv+L1c+v  =  L1(c+v)+L1(cv)(cv)(c+v)  =  2L1cc2v2=  2L1/c1β2. \begin{aligned} t_{\parallel} \;=\; \frac{L_{1}}{c-v} + \frac{L_{1}}{c+v} \;&=\; \frac{L_{1}(c+v) + L_{1}(c-v)}{(c-v)(c+v)} \;=\; \frac{2L_{1}c}{c^{2}-v^{2}}\\[6pt] &=\; \frac{2L_{1}/c}{1-\beta^{2}}. \end{aligned} (2.1.23)

The round trip is slower than the 2L1/c2L_{1}/c it would take with no wind, by a factor 1/(1β2)1/(1-\beta^{2}). That is worth a sentence, because the intuition points the wrong way. The two legs seem as though they ought to average out, one gaining what the other loses. They do not. The reason is that the leg travelled at the lower speed lasts longer, so it gets more weight in the total. Losing time at cvc-v for a long while beats gaining it at c+vc+v briefly. The same asymmetry is why a round trip against and with a current always takes longer than the same trip in still water, and why your average speed on a there-and-back drive is the harmonic mean rather than the arithmetic one.

5.3 · The arm across the wind, done properly

This is where most treatments go soft, so we will be slow. The naive answer says the light goes straight across at cc, so that the round trip is 2L2/c2L_{2}/c and the transverse arm is unaffected. That answer is wrong, and it is wrong for a reason that becomes central in Chapter 2.2.

The mirror at the end of the transverse arm is moving sideways. During the time the light is in flight, that mirror slides a distance vtvt down the xx-axis. Light that left the splitter travelling straight along yy in the ether frame would arrive where the mirror used to be and miss it. To hit the mirror, the light must leave with a velocity that has an xx-component matching the apparatus's own motion.

Write that as an equation. Let the light's velocity in the ether frame be u=(ux,uy)\vv u = (u_{x},u_{y}) with, by hypothesis, u=c\abs{\vv u}=c. For the light to stay over the moving arm it needs ux=vu_{x}=v. Then

ux2+uy2=c2,ux=vuy=c2v2. u_{x}^{2} + u_{y}^{2} = c^{2}, \qquad u_{x}=v \qquad\Longrightarrow\qquad u_{y} = \sqrt{c^{2}-v^{2}}. (2.1.24)

The useful component of the light's velocity, the part that gets it across the arm, is not cc but c2v2\sqrt{c^{2}-v^{2}}. The rest of its speed budget is spent keeping up sideways. The arm is L2L_{2} long and the crossing is covered at uyu_{y}, so

t  =  2L2c2v2  =  2L2/c1β2, t_{\perp} \;=\; \frac{2L_{2}}{\sqrt{c^{2}-v^{2}}} \;=\; \frac{2L_{2}/c}{\sqrt{1-\beta^{2}}}, (2.1.25)

with the factor of 22 because the return trip is the mirror image of the outbound one and takes the same time.

The same result comes out of a right triangle, and that is the version to carry in your head. In time t1t_{1} the light travels a straight-line distance ct1ct_{1} in the ether frame, which is the hypotenuse. The apparatus has drifted vt1vt_{1}, which is the base. The arm is L2L_{2}, which is the height. Pythagoras:

(ct1)2  =  L22+(vt1)2t12(c2v2)=L22t1=L2c2v2, \begin{aligned} \big(ct_{1}\big)^{2} \;=\; L_{2}^{2} + \big(vt_{1}\big)^{2} \qquad&\Longrightarrow\qquad t_{1}^{2}\big(c^{2}-v^{2}\big) = L_{2}^{2}\\[5pt] &\Longrightarrow\qquad t_{1} = \frac{L_{2}}{\sqrt{c^{2}-v^{2}}}, \end{aligned} (2.1.26)

and t=2t1t_{\perp}=2t_{1} as before. Same answer, same square root. The square root has the same origin in both derivations: a fixed total speed cc has to be shared between "across" and "along", and Pythagoras charges you for the sharing. That triangle is going to reappear in Chapter 2.2 as the derivation of time dilation, with the arm replaced by a light clock, and there the factor 1/1β21/\sqrt{1-\beta^{2}} will be named γ\gamma.

⚠ Why this isn't obvious — the transverse arm is affected too

It is very natural to treat the transverse arm as a clean reference. The wind blows sideways across it, so surely it is untouched, and all the physics is in the parallel arm. That reading is wrong, and getting it wrong changes the answer by a factor of two.

Compare the two results at small β\beta, using the binomial series of Chapter 0.3:

t=2Lc(1+β2+),t=2Lc(1+12β2+). t_{\parallel} = \frac{2L}{c}\Big(1+\beta^{2}+\cdots\Big), \qquad t_{\perp} = \frac{2L}{c}\Big(1+\tfrac12\beta^{2}+\cdots\Big).

Both arms are slowed. The parallel arm is slowed by γ2\gamma^{2}, the transverse arm by γ\gamma, where γ=(1β2)1/2\gamma=(1-\beta^{2})^{-1/2}, so that γ2=(1β2)1\gamma^{2}=(1-\beta^{2})^{-1}. Those two exponents are exactly the ones sitting in (2.1.23) and (2.1.25). If you wrongly take t=2L/ct_{\perp}=2L/c you get a β2\beta^{2} coefficient of 11 instead of 12\tfrac12, and you predict twice the fringe shift.

And the difference between γ\gamma and γ2\gamma^{2} is not an incidental detail. It is the entire observable. The experiment cannot measure either round-trip time. It can only compare them. Everything Michelson and Morley could possibly have seen lives in the gap between those two exponents, and if the two arms had been affected identically the experiment would have had nothing to detect no matter how sensitive it was. Chapter 2.2's resolution has to reproduce exactly this structure, one power of γ\gamma along the motion and none across it. §6.2 shows that FitzGerald spotted the shape of the fix from precisely this observation, fifteen years before anyone knew why it was true.

5.4 · The difference

Take equal arms, L1=L2=LL_{1}=L_{2}=L, which is what the instrument is built to have and what §5.5 shows we do not actually need. Subtract (2.1.25) from (2.1.23):

Δt    tt  =  2Lc(11β2    11β2). \Delta t \;\equiv\; t_{\parallel} - t_{\perp} \;=\; \frac{2L}{c}\left(\frac{1}{1-\beta^{2}} \;-\; \frac{1}{\sqrt{1-\beta^{2}}}\right). (2.1.27)

That is exact. Now expand for small β\beta, since β104\beta\approx10^{-4} and what we want to know is how big the effect is. Chapter 0.3's binomial series gives, with s=β2s=\beta^{2},

11s=1+s+s2+,11s=1+12s+38s2+, \frac{1}{1-s} = 1 + s + s^{2} + \cdots, \qquad \frac{1}{\sqrt{1-s}} = 1 + \tfrac12 s + \tfrac38 s^{2} + \cdots, (2.1.28)

We want the difference of those two series, so subtract the second from the first term by term. The bracket in (2.1.27) becomes

(1+s+s2)(1+12s+38s2)+  =  12s+58s2+=  12β2(1+54β2+), \begin{aligned} \big(1+s+s^{2}\big) - \big(1+\tfrac12 s+\tfrac38 s^{2}\big) + \cdots \;&=\; \tfrac12 s + \tfrac58 s^{2} + \cdots\\[4pt] &=\; \tfrac12\beta^{2}\Big(1 + \tfrac54\beta^{2} + \cdots\Big), \end{aligned} (2.1.29)

Only the leading piece matters at β104\beta\approx10^{-4}, so keep the 12β2\tfrac12\beta^{2} term and drop everything after it. The round-trip difference is then

Δt    2Lcβ22  =  Lβ2c  =  Lv2c3. \Delta t \;\approx\; \frac{2L}{c}\cdot\frac{\beta^{2}}{2} \;=\; \frac{L\beta^{2}}{c} \;=\; \frac{L v^{2}}{c^{3}}. (2.1.30)

Look at what has dropped out. The two arms agree at order β0\beta^{0} and again at order β1\beta^{1}, and they first differ at order β2\beta^{2}. The effect is second order in v/cv/c, and that is the central practical fact about the entire subject. There is no first-order ether-wind effect in a round-trip measurement, because the delay on the way out is compensated to first order by the gain on the way back. Every experiment sensitive only to first order was therefore guaranteed a null result before it was built, and there had been several. Michelson's design was the first with second-order sensitivity, which is why it counts.

And that β2=v2/c2\beta^{2}=v^{2}/c^{2} is the same suppression factor that made the third law's failure so hard to notice in Chapter 1.1 §3.3. Both are the leading symptom of the same disease.

5.5 · What rotating the apparatus measures

Now the experimental point flagged in §5.1. You cannot measure Δt\Delta t directly. Doing so would require knowing L1L2L_{1}-L_{2} to a fraction of 590 nm590\ \mathrm{nm}, which nobody can build. So rotate instead.

Before rotation, with arm 1 along the wind and arm 2 across it, the time difference between the beams is, from (2.1.23) and (2.1.25),

Δtbefore  =  2L1/c1β2    2L2/c1β2. \Delta t_{\text{before}} \;=\; \frac{2L_{1}/c}{1-\beta^{2}} \;-\; \frac{2L_{2}/c}{\sqrt{1-\beta^{2}}}. (2.1.31)

Turn the table through 9090^{\circ}. Arm 1 is now across the wind and arm 2 along it, and neither length has changed. So

Δtafter  =  2L1/c1β2    2L2/c1β2. \Delta t_{\text{after}} \;=\; \frac{2L_{1}/c}{\sqrt{1-\beta^{2}}} \;-\; \frac{2L_{2}/c}{1-\beta^{2}}. (2.1.32)

Subtract. The terms regroup by denominator rather than by arm, and the arms enter only through their sum:

ΔtbeforeΔtafter  =  2(L1+L2)/c1β2    2(L1+L2)/c1β2=  2(L1+L2)c(11β211β2). \begin{aligned} \Delta t_{\text{before}} - \Delta t_{\text{after}} \;&=\; \frac{2\big(L_{1}+L_{2}\big)/c}{1-\beta^{2}} \;-\; \frac{2\big(L_{1}+L_{2}\big)/c}{\sqrt{1-\beta^{2}}}\\[6pt] &=\; \frac{2\big(L_{1}+L_{2}\big)}{c}\left(\frac{1}{1-\beta^{2}} - \frac{1}{\sqrt{1-\beta^{2}}}\right). \end{aligned} (2.1.33)

The unknown mismatch L1L2L_{1}-L_{2} has vanished identically. Only L1+L2L_{1}+L_{2} survives, and that is a quantity you can measure with a ruler. Set L1=L2=LL_{1}=L_{2}=L so that L1+L2=2LL_{1}+L_{2}=2L, and compare with (2.1.27). The change on rotation is exactly 2Δt2\Delta t.

Now turn that into something visible. A time difference τ\tau between the two beams corresponds to an optical path difference cτc\tau, and one whole fringe of movement corresponds to a path difference of one wavelength λ\lambda. So the number of fringes the pattern shifts as the apparatus turns is

ΔN  =  c2Δtλ  =  4Lλ(11β211β2)    2Lβ2λ. \Delta N \;=\; \frac{c\cdot 2\Delta t}{\lambda} \;=\; \frac{4L}{\lambda}\left(\frac{1}{1-\beta^{2}} - \frac{1}{\sqrt{1-\beta^{2}}}\right) \;\approx\; \frac{2L\beta^{2}}{\lambda}. (2.1.34)

Clean, and made of measurable things: an arm length, a wavelength, and the square of the Earth's speed in units of cc.

5.6 · The numbers, 1887

Michelson and Morley floated a sandstone slab on a bath of mercury so that it could be turned smoothly without flexing, and they folded each light path back and forth between multiple mirrors to lengthen it. ⚑ The effective one-way path was L11 mL\approx11\ \mathrm{m}, the source was a sodium lamp at λ590 nm\lambda\approx590\ \mathrm{nm}, and the assumed velocity was the Earth's orbital speed, v=30 kms1v=30\ \mathrm{km\,s^{-1}}. Then

β=3.0×1042.99792×108=1.0007×104,β2=1.0014×108, \beta = \frac{3.0\times10^{4}}{2.99792\times10^{8}} = 1.0007\times10^{-4}, \qquad \beta^{2} = 1.0014\times10^{-8}, (2.1.35)

Now feed that value of β2\beta^{2} into (2.1.34), along with the arm length and the wavelength above:

ΔN    2(11 m)(1.0014×108)5.90×107 m  =  0.373 fringes. \Delta N \;\approx\; \frac{2\big(11\ \mathrm{m}\big)\big(1.0014\times10^{-8}\big)}{5.90\times10^{-7}\ \mathrm{m}} \;=\; 0.373\ \text{fringes}. (2.1.36)

(The exact expression (2.1.34) gives 0.3733980.373398, and the expansion also gives 0.3733980.373398. At β=104\beta=10^{-4} the correction term 54β2\tfrac54\beta^{2} of (2.1.29) is one part in 10810^{8}, so the two agree to six figures, as they should.)

About four tenths of a fringe. That is not enormous. But Michelson's interferometer could see a shift of about a hundredth of a fringe, so the predicted effect was roughly forty times the detection threshold.

5.7 · The interactive

v = 30.00 km/s
L = 11.0 m
t∥ = 73.384101678 ns
t⊥ = 73.384101311 ns
Δt = 3.6743e-16 s = 367.4 as
ΔN = 0.3734 fringes
ΔN / bound = 37.3× — far above the threshold; crosses the bound at v = 4.91 km/s
What the ether wind should have done, against what the experiment could see. The upper panel is the interferometer: source, beam splitter, the arm along the wind (blue) and the arm across it (orange), and the detector where the beams recombine. The inset at the right is the triangle of (2.1.26), drawn hugely out of scale — at the real β104\beta\approx10^{-4} the slant would be invisible — showing why the transverse crossing takes L/c2v2L/\sqrt{c^{2}-v^{2}} and not L/cL/c.

The lower panel plots the predicted fringe shift against wind speed on log–log axes, where the β2\beta^{2} scaling of (2.1.34) shows up as a straight line of slope exactly 22 — double the wind speed, quadruple the shift. The shaded band is the 1887 experimental upper bound of about 0.010.01 fringes; the dashed vertical line is the Earth's orbital speed of 30 kms130\ \mathrm{km\,s^{-1}}, which is not adjustable by anybody.

Drag the wind speed down and watch where the curve crosses into the band. At the 1887 apparatus (L=11 mL=11\ \mathrm{m}) you have to get below about 4.9 kms14.9\ \mathrm{km\,s^{-1}} — a sixth of the Earth's orbital speed — before the prediction hides under the detection threshold. At the actual orbital speed the prediction sits a factor of 3737 above it. This is not a marginal call. Push the arm length up and it gets worse for the ether, since ΔN\Delta N is linear in LL; that is exactly why Michelson folded the path.

Every number is computed from the exact expressions (2.1.23), (2.1.25) and (2.1.34), never from the small-β\beta expansion. Because tt_{\parallel} and tt_{\perp} agree to nine significant figures at realistic β\beta, subtracting them directly in floating point would throw away most of the answer, so Δt\Delta t is evaluated from the algebraically identical but numerically stable rearrangement Δt=2Lcβ2/[(1β2)(1+1β2)]\Delta t = \frac{2L}{c}\,\beta^{2}\big/\big[(1-\beta^{2})(1+\sqrt{1-\beta^{2}})\big] — which you can check reduces to (2.1.30) as β0\beta\to0.

5.8 · The result

They found nothing.

⚑ Michelson and Morley reported in 1887 that the observed displacement was certainly less than a twentieth of the predicted 0.40.4 fringes, and probably less than a fortieth. That is an upper bound of about 0.010.01 fringes, which is to say the noise floor of the instrument. They repeated the measurement at different times of day and in different seasons, to catch the Earth's rotation and its orbital motion pointing the apparatus in every possible direction relative to any conceivable ether frame. Nothing.

This is not a small discrepancy

A great many experiments in physics disagree with prediction by twenty or thirty per cent and are eventually reconciled by a systematic effect somebody had missed. This is not one of them.

The prediction was 0.3730.373 fringes. The bound was 0.010.01. That is not a mismatch in a coefficient. It is the total absence of an effect that should have been unmissable, by a factor of nearly forty, in an apparatus specifically designed to have several times the required sensitivity. Nor can it be rescued by supposing the Earth happens to be at rest in the ether at the moment of measurement. Six months later the Earth's orbital velocity has reversed, so its speed relative to any fixed ether frame must then be at least 60 kms160\ \mathrm{km\,s^{-1}}, and the predicted shift is four times larger. Michelson and Morley looked. Nothing.

⚑ Modern versions replace the arms with cryogenic optical resonators and look for a directional dependence of the resonant frequency as the apparatus turns on a rotating table. The current bound on the fractional anisotropy of cc is below one part in 101710^{17}, some seven orders of magnitude tighter than what the 1887 apparatus could have detected. In a hundred and forty years of increasingly ferocious effort, the effect has not appeared.

In plain terms 2.1.5

An experiment that finds nothing is worth exactly as much as the effect it was built to find. Timing light directly was hopeless, so Michelson compared two perpendicular paths by interference, which resolves a hundredth of a wavelength, and turned the apparatus through a right angle so the arms exchanged roles. The unknown mismatch between their lengths cancels out of that comparison identically.

Both arms are slowed, and missing that is the standard error. The arm along the motion loses more time crawling against the flow than it regains running with it, because the slow leg lasts longer and so counts for more. The arm across is slowed too, by a smaller power, since part of the light's fixed speed budget goes on keeping up sideways. The whole observable is the gap between those two powers, and it is second order in the ratio of the Earth's speed to light's, which is why every earlier experiment sensitive only to first order was guaranteed a null result before it was built.

The predicted shift was about four tenths of a fringe, in an instrument that could see a hundredth. Nothing appeared, at any hour or season, and a hundred and forty years of increasingly ferocious repetition has not made it appear. That is not a discrepancy in a coefficient; it is the absence of something that should have been unmissable by a factor of nearly forty.

6 · The rescues, and why they failed

A null result does not by itself dispose of the ether. What it disposes of is the simplest ether: one at rest in some frame, through which the Earth moves freely and which does nothing to the apparatus. Several serious alternatives were proposed. It is worth being specific about them, because they were good physics, and because the last one was very nearly right.

6.1 · Ether drag

Here is the obvious first move. Suppose the Earth drags the local ether along with it, so that near the surface there is no wind at all. The interferometer then sits in still air, so to speak, and sees nothing. Two independent observations rule this out.

Stellar aberration. Point a telescope at a star directly overhead. The Earth is moving sideways at vv while the light travels down the tube, so the telescope must be tilted slightly into the direction of motion. Otherwise the eyepiece has moved out from under the light by the time it arrives. If the tube has length \ell, the light takes /c\ell/c to traverse it and the tube advances v/cv\ell/c in that time, so the required tilt is

tanθ  =  v/c  =  vc  =  βθβ. \tan\theta \;=\; \frac{v\ell/c}{\ell} \;=\; \frac{v}{c} \;=\; \beta \qquad\Longrightarrow\qquad \theta \approx \beta. (2.1.37)

With v=29.78 kms1v=29.78\ \mathrm{km\,s^{-1}} this comes to θ=9.933×105 rad\theta = 9.933\times10^{-5}\ \mathrm{rad}, and converting to arcseconds by multiplying by 206265206265,

θ  =  20.5. \theta \;=\; 20.5'' . (2.1.38)

⚑ Bradley discovered exactly this in 1728. Every star traces out a small ellipse over the course of a year, of angular radius 20.495520.4955'', and the effect is the single most reliable confirmation that the Earth moves. Now put that together with drag. If the ether were carried along with the Earth, the light would enter that co-moving ether at the top of the atmosphere and thereafter travel in the telescope's own rest frame. There would then be no aberration at all, or at most a fringe effect at the boundary. The aberration is observed, it has exactly the size v/cv/c predicted for undragged ether, and it is the same for all stars regardless of distance. Complete drag is dead.

Fizeau's water tube. In 1851 Fizeau sent light through a tube of water flowing at speed uu and measured how much the water carried the light along with it. Complete drag predicts that the light speed in the lab is c/n+uc/n + u, and no drag predicts c/nc/n. He measured neither. ⚑ The light was dragged by a fraction of the water's speed,

vlight  =  cn+fu,f  =  11n2, v_{\text{light}} \;=\; \frac{c}{n} + f\,u, \qquad f \;=\; 1 - \frac{1}{n^{2}}, (2.1.39)

which for water (n=1.333n=1.333) is f=0.437f=0.437. Fizeau confirmed that to within a few per cent, and Michelson's much more precise repetition in 1886 returned 0.434±0.020.434\pm0.02. A partial drag coefficient is a very strange thing for a mechanical medium to have. How does a substance know to transmit exactly 43.7%43.7\% of its own motion? No ether model ever gave a convincing account of it. (Chapter 2.2 will get (2.1.39) in one line, as the leading term of relativistic velocity addition, with no medium, no drag and no free parameter. It is one of the cleanest confirmations of the whole framework, and it is a nineteenth-century measurement.)

6.2 · The FitzGerald–Lorentz contraction

Now the interesting one. In 1889 FitzGerald, and independently Lorentz in 1892, made a proposal of startling economy: objects moving through the ether contract along their direction of motion by exactly the factor 1β2\sqrt{1-\beta^{2}}.

Watch what that does. The parallel arm lies along the motion, so its length in the ether frame is not LL but L1β2L\sqrt{1-\beta^{2}}. The transverse arm is perpendicular and is unaffected. Substitute into (2.1.23):

t  =  2(L1β2)/c1β2  =  2L/c1β2  =  t. t_{\parallel} \;=\; \frac{2\big(L\sqrt{1-\beta^{2}}\big)/c}{1-\beta^{2}} \;=\; \frac{2L/c}{\sqrt{1-\beta^{2}}} \;=\; t_{\perp}. (2.1.40)

The two round-trip times are equal. Not approximately equal, and not equal to order β2\beta^{2}. They are identically equal, for every β\beta, because 1β2/(1β2)=1/1β2\sqrt{1-\beta^{2}}/(1-\beta^{2}) = 1/\sqrt{1-\beta^{2}} is an algebraic identity. The predicted fringe shift is exactly zero to all orders, and Michelson–Morley is explained.

It is impossible to look at (2.1.40) and not feel that something is being got away with. The contraction factor was chosen for no reason except that it makes the answer come out. It was, in FitzGerald's own framing, a hypothesis with one job.

And yet it is not absurd. Lorentz gave it a physical rationale. Suppose the forces holding matter together are electromagnetic in origin, which by 1892 was a reasonable guess. Suppose also that electromagnetic fields are distorted by motion through the ether. Then the equilibrium spacing of the atoms in a solid might well change when the solid is set in motion. On that reading the contraction is a real, dynamical effect on matter, produced by the ether acting on the electrons in the rod. The rod is genuinely shorter. The ether is genuinely there. It just conspires to be undetectable.

⚠ Why this isn't obvious — a null result does not refute the ether

The tidy story is that Michelson–Morley disproved the ether and Einstein cleared away the wreckage. Both halves are false, and the second is the more misleading.

(2.1.40) is a proof that a null result cannot kill a sufficiently determined theory. Lorentz's ether, equipped with contraction and with the local time of §6.3, predicts exactly what Michelson and Morley saw. Not approximately. Exactly. By 1904 it reproduced every optical experiment then known. No measurement of the kind being discussed could have distinguished it from what replaced it, because by construction the two make the same predictions.

What an experiment refutes is a specific model, not a concept. It refuted the undragged, undistorting ether. It said nothing about ethers that contract their occupants.

So what did kill it? Not evidence. The ether stopped doing any work. Every quantity it was introduced to explain turned out to be calculable without it: the propagation speed, the failure of the wind to show up, the aberration, the Fizeau coefficient. And every attempt to detect it produced a null result that had to be patched by a new property whose only content was the null result it was patching. A frame of reference that no experiment can identify is not a frame of reference. It is a spare wheel bolted to the theory, and the moment you notice that removing it changes no prediction, it has been refuted in the only sense that matters.

That is a methodological point rather than an empirical one, and it deserves to be made without condescension. The nineteenth century was not being obtuse. They were doing exactly what you should do with a productive hypothesis under pressure, which is to modify it minimally and see whether it survives. It did survive, technically, and it was abandoned anyway. That tells you that "consistent with all the data" is not the only standard a theory is held to. Compare the situation honestly with the present. Chapter 7.9 will ask the same question about ideas in quantum gravity that are consistent with everything and predict nothing, the answer will be the same one, and it will be no more comfortable.

6.3 · Local time

Contraction alone is not quite enough, and Lorentz knew it. He wanted Maxwell's equations to come out with the same form in a moving frame, so that the theory would work at first order in β\beta for optical experiments generally and not only for Michelson–Morley. To get that, he found he had to introduce a further substitution, which in his 1895 form reads

t  =  t    vxc2. t' \;=\; t \;-\; \frac{vx}{c^{2}}. (2.1.41)

He called tt' the local time (Ortszeit), and he was explicit that he regarded it as a mathematical convenience with no physical meaning whatever. The real time was tt, the ether's time, universal and absolute in Newton's sense. The quantity tt' was an auxiliary variable, a change of integration variable that made the equations tractable, no more significant than substituting u=x2u=x^{2} in an integral.

By 1904 Lorentz had the full set. Written in modern notation, with γ=(1β2)1/2\gamma = (1-\beta^{2})^{-1/2}:

x=γ(xvt),y=y,z=z,t=γ(tvxc2). x' = \gamma\big(x - vt\big), \qquad y'=y, \qquad z'=z, \qquad t' = \gamma\left(t - \frac{vx}{c^{2}}\right). (2.1.42)

Those are the Lorentz transformations. Chapter 2.2 derives them from scratch and will not use this equation. It is here only so that the historical point can be made without ambiguity.

What Einstein actually contributed

Lorentz had (2.1.42) before Einstein. Poincaré had noticed that the transformations form a group, and in 1904 he had stated a principle of relativity covering electrodynamics. The algebra of special relativity was on paper, published, and being used.

So it is worth being exact about what 1905 added, because "Einstein derived the Lorentz transformation" is not it.

For Lorentz, xx' and tt' were auxiliary variables. There was a true frame, the ether's, with true coordinates x,tx,t. Moving rods were really shortened, by a dynamical action of the ether on the electrical forces inside matter. Moving clocks were not slowed at all, since tt' was not a time but a substitution. The transformation described what happens to matter when it moves through the ether. It was a theory of rods and clocks.

For Einstein, xx' and tt' are the coordinates a moving observer actually measures, obtained from two postulates about the symmetry of nature and nothing else. No ether appears in the derivation, because none is needed. There is no true frame, no true time, and nothing dynamical happens to the rod. The rod is shorter in that frame in the same sense that a pencil has a shorter shadow when you turn it, and the reason both observers can consistently say the other's rod is short is that they are disagreeing about which events are simultaneous. The transformation describes the structure of space and time. It is not a theory of rods and clocks. It is a statement about the arena they sit in, and it therefore applies to everything: nuclear forces, particle lifetimes, chemical clocks, and things nobody had thought of, rather than only to electromagnetically bound matter.

That difference is not philosophical decoration. It is what makes the theory predictive outside its domain of construction. Lorentz's contraction is a claim about electromagnetically bound rods, and there is no reason it should apply to a decaying muon. Einstein's is a claim about the coordinates themselves, so it applies to the muon whether or not anything holding it together is electromagnetic. Muons in the atmosphere do live longer by exactly γ\gamma, and the weak interaction that decays them was not discovered for another thirty years.

And it is why the ether then had nothing left to do. Lorentz's ether was necessary to his derivation. Without a medium, there is nothing for a rod to move through and no mechanism to contract it. Einstein's derivation never mentions it, so the ether becomes an object with no role in any calculation, detectable by no experiment, and postulated for no reason. It did not have to be refuted. It had to be noticed to be idle.

Grind box — how close was Lorentz, exactly?

Uncomfortably close, and the gap is instructive. Take Lorentz's own 1904 apparatus and ask what it gets right.

Right: the algebra. (2.1.42) is character for character what Chapter 2.2 will derive. The contraction factor, the local time, the group property, the invariance of Maxwell's equations under the substitution: all correct, and all published before 1905.

Right: the experimental predictions, for everything then measured. Michelson–Morley, Fizeau, aberration, the first-order optical experiments: Lorentz's theory reproduces them all.

Wrong: the status of tt'. Because it is a substitution rather than a time, Lorentz had no account of time dilation as a physical phenomenon, and no reason to expect a moving clock of any construction to run slow. Ask him about an unstable particle and the theory is silent.

Wrong: the asymmetry. In Lorentz's scheme the contraction is real for a rod moving through the ether and not for a rod at rest in it, so the two frames are not equivalent even though no experiment can tell them apart. Einstein's symmetry between the frames is exact, and it is the reason each observer finds the other's rods short. Lorentz's framework cannot make that statement, since on his account one of the two rods is really contracted and the other really is not.

Wrong: the generality. Lorentz's contraction was derived from a hypothesis about electromagnetic forces inside matter. Nothing in it covers gravitational binding, nuclear binding, or the decay rate of a particle. Every one of those obeys the same γ\gamma anyway, which on Lorentz's account would be an extraordinary coincidence and on Einstein's is a tautology.

The lesson generalises past this episode and is worth carrying: having the right equations is not the same as having the right theory. Two accounts can agree on every equation and every measured number and still differ in what they say the symbols mean. That difference shows up the moment you try to apply the theory somewhere it was not built for. Lorentz's electrodynamics and Einstein's relativity were empirically indistinguishable in 1905. They stopped being indistinguishable the moment anyone measured a particle lifetime.

In plain terms 2.1.6

FitzGerald proposed a repair of startling economy, and what makes it instructive is that it works. Suppose anything moving through the medium is shortened along its direction of motion by a particular factor, and the two round-trip times become equal, not approximately but identically, at every speed, because the shortening and the slowing are related by an algebraic identity. The predicted shift is then exactly zero.

It was not absurd either. If the forces holding matter together are electromagnetic, and fields are distorted by motion through the medium, atoms in a rod might settle at a different spacing once it is set moving, so the rod is genuinely shorter and the medium genuinely there. With that and one further substitution, the theory reproduced every optical experiment then known. A null result cannot kill a determined theory; what an experiment refutes is a model rather than a concept.

So the medium was not disproved. It stopped doing any work. Every quantity it had been introduced to explain turned out calculable without it, and every attempt to detect it produced a null result patched by a property whose only content was the null result it patched. A frame no experiment can pick out is not a frame of reference. The verdict is methodological rather than empirical, and deserves stating without condescension, since modifying a productive hypothesis minimally under pressure is what anybody should do.

7 · The two postulates

Everything above is the case for the prosecution. Here is what Einstein proposed instead, stated in the form Chapter 2.2 will use as input.

The postulates of special relativity

Postulate 1 (the principle of relativity). The laws of physics take the same form in all inertial frames. No experiment performed within a uniformly moving laboratory can reveal that laboratory's velocity. And this applies to all the laws, mechanical, electromagnetic, and any not yet discovered, rather than merely to the mechanical ones.

Postulate 2 (the invariance of cc). Light propagates in vacuum with a definite speed cc, the same in all inertial frames, independent of the state of motion of the emitting body.

Notice how little that is. There is no ether, no statement about what light is, no model of matter, no dynamics and no mechanism. Two sentences, one of which was universally accepted already. Chapter 2.2 shows that the pair of them determines the transformation between inertial frames essentially uniquely, and everything in Part II after that is consequences.

7.1 · Postulate 1 is not new

It is §1.4's principle with one word changed: all the laws, not just the mechanical ones. That single word is what rules out branch (A) of §4.3, where mechanics obeyed the principle and electromagnetism did not. We already have (2.1.18), in which Maxwell's equations demonstrably change form under (2.1.1). So asserting Postulate 1 forces you into branch (C), and the Galilean transformation must go.

7.2 · Postulate 2 is the radical one, and it is not a definition

This is where the strangeness lives, and it is worth being precise about why.

Consider a lamp on the front of a train moving at vv, and a lamp on the platform, flashing as they pass. Postulate 2 says both flashes travel at cc relative to the platform and at cc relative to the train. Not c+vc+v and cc. Not cc and cvc-v. The same cc, measured by both observers, for both flashes. (2.1.2), the rule that velocities add, is being denied for light, flatly.

Two objections need answering.

"Isn't this just a convention about how to set clocks?" No. There is a real convention lurking nearby, since synchronising distant clocks requires a rule and any rule involves an assumption about one-way light speed, and Chapter 2.3 will treat that carefully. But Postulate 2 has empirical content independent of any such convention, because the round-trip speed is measurable without synchronising anything: one clock, one mirror, one pulse. The claim that the two-way speed is the same in all frames and independent of the source's motion is a fact about the world, and it can be checked.

"Isn't the source-independence part obvious for a wave?" For a wave in a medium, yes. Sound from a moving whistle travels at 343 ms1343\ \mathrm{m\,s^{-1}} through the air regardless of the whistle, since once the disturbance is launched the medium takes over and forgets the source. That is precisely why the ether was attractive: it delivers source-independence for free. Postulate 2 asserts source-independence with no medium to enforce it, which is a much stranger claim. And it also asserts frame-independence, which no medium theory delivers at all.

So Postulate 2 is an empirical claim, and it must be tested. It has been, exhaustively. Worked example 2 gives the sharpest version: photons emitted by particles moving at 0.99975c0.99975c travel at cc, not at 1.99975c1.99975c, to within a part in 10410^{4}.

Handing off to Chapter 2.2

Two postulates. One transformation to find. The question Chapter 2.2 asks is exactly this: what transformation between inertial coordinates could possibly satisfy both? The answer, remarkably, is that there is essentially only one. Linearity is forced by the requirement that free particles move in straight lines in every frame. The constant in it is forced by demanding a speed that comes out the same in all frames. And what drops out is (2.1.42), with ttt'\neq t and everything that follows from it.

You already know one thing about the answer that Chapter 2.2 will have to reproduce. Whatever the transformation is, it must produce a contraction by 1β2\sqrt{1-\beta^{2}} along the direction of motion and nothing across it, because that is precisely the combination (2.1.40) needs to explain Michelson–Morley. The experiment has already told you the shape of the answer. What it could not tell you is why.

In plain terms 2.1.7

Notice how little is being asked for at the end of all that. Two sentences: the laws take the same form in every inertial frame, all of them and not merely the mechanical ones; and light travels in vacuum at one definite speed, the same for every observer and independent of the motion of its source.

The first sentence was already common property with one word altered, and that word forbids the branch on which mechanics obeys the principle while optics does not. The second is the radical one, and it is not a disguised convention about setting clocks: a round-trip speed can be measured with one clock, one mirror and one pulse, synchronising nothing, so the claim has empirical content and has been checked to absurdity. Photons from particles moving at nearly the limiting speed arrive at that speed, not at twice it.

What makes the second sentence strange is worth locating exactly. A medium delivers independence from the source for free, since once a disturbance is launched the medium takes over and forgets its origin; asserting that independence with no medium to enforce it is a far stronger claim, and asserting that the speed is also the same for every observer is one no medium delivers. The interferometer had already revealed the shape of the answer, one factor along the motion and none across it. What it could not reveal was why.

8 · Worked examples

Worked example 1 — the 1887 apparatus, all the way through

Take the Michelson–Morley interferometer as built: effective arm length L=11.0 mL = 11.0\ \mathrm{m} in each arm, sodium light at λ=590 nm\lambda = 590\ \mathrm{nm}, and the ether at rest relative to the Sun so that the wind speed is the Earth's orbital speed v=30.0 kms1v=30.0\ \mathrm{km\,s^{-1}}. Compute both round-trip times, their difference, and the fringe shift on a 9090^{\circ} rotation. Compare with the experimental bound.

Step 1 · the small parameter. With c=2.99792×108 ms1c = 2.99792\times10^{8}\ \mathrm{m\,s^{-1}},

β=3.00×1042.99792×108=1.0007×104,β2=1.0014×108. \beta = \frac{3.00\times10^{4}}{2.99792\times10^{8}} = 1.0007\times10^{-4}, \qquad \beta^{2} = 1.0014\times10^{-8}.

Already the shape of the problem is clear: we are chasing a part in 10810^{8}.

Step 2 · the two round-trip times. The no-wind time is 2L/c=22.0/2.99792458×108=7.33841009×108 s=73.38410094 ns2L/c = 22.0/2.99792458\times10^{8} = 7.33841009\times10^{-8}\ \mathrm{s} = 73.38410094\ \mathrm{ns}. Then from (2.1.23) and (2.1.25),

t=73.38410094 ns11.0014×108=73.384101678 ns,t=73.38410094 ns11.0014×108=73.384101311 ns. \begin{aligned} t_{\parallel} &= \frac{73.38410094\ \mathrm{ns}}{1-1.0014\times10^{-8}} = 73.384101678\ \mathrm{ns},\\[4pt] t_{\perp} &= \frac{73.38410094\ \mathrm{ns}}{\sqrt{1-1.0014\times10^{-8}}} = 73.384101311\ \mathrm{ns}. \end{aligned}

They agree to nine significant figures. This is why nobody attempted to time the two beams separately.

Step 3 · the difference. Subtracting nine matching digits is a bad idea numerically, so use the exact rearrangement

Δt=2Lcβ2(1β2)(1+1β2)=3.6743×1016 s, \Delta t = \frac{2L}{c}\cdot\frac{\beta^{2}}{\big(1-\beta^{2}\big)\big(1+\sqrt{1-\beta^{2}}\big)} = 3.6743\times10^{-16}\ \mathrm{s},

which is 367 attoseconds367\ \mathrm{attoseconds}. Cross-check that against the leading-order estimate (2.1.30): Lβ2/c=11×1.0014×108/2.99792×108=3.6743×1016 sL\beta^{2}/c = 11\times1.0014\times10^{-8}/2.99792\times10^{8} = 3.6743\times10^{-16}\ \mathrm{s}. Agreement to five figures. ✓

Step 4 · turn it into a distance. A time difference is invisible, but a path difference is not. Multiply by cc:

cΔt=1.1015×107 m=110.15 nm=0.1867λ. c\,\Delta t = 1.1015\times10^{-7}\ \mathrm{m} = 110.15\ \mathrm{nm} = 0.1867\,\lambda.

The two beams arrive about a fifth of a wavelength out of step. This is the entire physical content of the experiment: a fifth of a wavelength, from a 367367-attosecond delay, produced by the Earth moving at a ten-thousandth of the speed of light.

Step 5 · the fringe shift on rotation. Rotating exchanges the arms, so the path difference swings from +0.1867λ+0.1867\lambda to 0.1867λ-0.1867\lambda and the fringes move by twice that:

ΔN=2cΔtλ=2×0.1867=0.373 fringes. \Delta N = \frac{2c\,\Delta t}{\lambda} = 2\times 0.1867 = 0.373\ \text{fringes}.

Or directly from (2.1.34): ΔN2Lβ2/λ=2(11)(1.0014×108)/(5.90×107)=0.373\Delta N \approx 2L\beta^{2}/\lambda = 2(11)(1.0014\times10^{-8})/(5.90\times10^{-7}) = 0.373. ✓ (The exact expression gives 0.3733980.373398 and the expansion gives 0.3733980.373398, and they differ in the ninth decimal.)

Step 6 · compare. ⚑ Observed: less than 0.010.01 fringes. Predicted: 0.3730.373. The ratio is

0.3730.01=37.3. \frac{0.373}{0.01} = 37.3.

To hide the prediction under the bound you would need β2\beta^{2} smaller by a factor of 3737, which means vv smaller by 37=6.1\sqrt{37}=6.1, which means an ether wind below 4.9 kms14.9\ \mathrm{km\,s^{-1}}. That is a sixth of the Earth's orbital speed, and it is a speed the Earth cannot have relative to a fixed frame at all times of year.

A number worth having. The Solar System moves at about 220 kms1220\ \mathrm{km\,s^{-1}} around the galactic centre, and at about 370 kms1370\ \mathrm{km\,s^{-1}} relative to the frame in which the cosmic microwave background is isotropic. Had the ether been at rest in either, the 1887 apparatus would have shown 2020 fringes or 5757 fringes respectively, and the pattern would have swept past the crosshairs dozens of times as the table turned. There is no plausible ether frame for which this experiment is a close call.

Worked example 2 — what Galilean addition demands of light, and what is measured

Suppose light obeys (2.1.2) like everything else. Work out what a laboratory should measure for light from a moving source, and compare with experiment.

Step 1 · the prediction. If light leaves its source at cc relative to the source, and velocities add by (2.1.2), then a lab in which the source moves at vv directly towards the detector measures

clab=c+v, c_{\text{lab}} = c + v,

and cvc-v for a receding source. This is the "emission theory" or "ballistic" hypothesis, and it is what you get if you take the Galilean transformation seriously and drop the ether. Note that it is not the ether theory. The ether predicts cc relative to the ether regardless of the source. The two hypotheses are different, and both are testable.

Step 2 · the astronomical test. This is de Sitter's argument from 1913, and it costs nothing but arithmetic. Take a spectroscopic binary star at distance DD whose components orbit at speed vorbv_{\text{orb}}. Light emitted while a star approaches us travels, on this hypothesis, at c+vc+v. Half an orbit later, receding, it travels at cvc-v. The difference in arrival time after a journey DD is

Δtarr=DcvDc+v=2Dvc2v22Dvc2. \Delta t_{\text{arr}} = \frac{D}{c-v} - \frac{D}{c+v} = \frac{2Dv}{c^{2}-v^{2}} \approx \frac{2Dv}{c^{2}}.

Now put in numbers for a typical system, with D=100D = 100 light years =9.46×1017 m=9.46\times10^{17}\ \mathrm{m} and vorb=100 kms1v_{\text{orb}}=100\ \mathrm{km\,s^{-1}}:

Δtarr  =  2Dvc2=2×9.46×1017×1058.99×1016=2.11×106 s=24 days. \Delta t_{\text{arr}} \;=\; \frac{2Dv}{c^{2}} = \frac{2\times9.46\times10^{17}\times10^{5}}{8.99\times10^{16}} = 2.11\times10^{6}\ \mathrm{s} = 24\ \text{days}.

Compare that with an orbital period of, say, 55 days. The scrambling is five times the period. Fast-phase light would overtake slow-phase light, the star would appear at several places in its orbit at once, and the spectroscopic velocity curve would be unrecognisable, not merely distorted but multivalued. Nothing of the kind is seen.

Writing the hypothesis as clab=c+kvc_{\text{lab}}=c+kv multiplies the estimate by kk. Binary orbits are observed to fit Kepler's laws at the per-cent level, and holding the distortion inside that needs k0.01×(5/24)2×103k \lesssim 0.01\times(5/24) \approx 2\times10^{-3}. ⚑ De Sitter's published bound was of this order. ⚑ Brecher's 1977 analysis of X-ray binaries, where the timing is far sharper, tightened it to 2×1092\times10^{-9}.

Step 3 · the laboratory test, and it is brutal. ⚑ At CERN in 1964, Alväger and collaborators produced neutral pions in a proton beam and used the photons from π0γγ\pi^{0}\to\gamma\gamma. The pions were moving at βπ=0.99975\beta_{\pi} = 0.99975, which is as close to a light-speed source as anything ever built, and the photons' speed was measured by time of flight over a known baseline.

The Galilean prediction for photons emitted forwards is unambiguous:

clab=c+vπ=(1+0.99975)c=1.99975c=5.995×108 ms1. c_{\text{lab}} = c + v_{\pi} = \big(1 + 0.99975\big)c = 1.99975\,c = 5.995\times10^{8}\ \mathrm{m\,s^{-1}}.

Very nearly double the speed of light. What was measured was

clab=(2.9977±0.0004)×108 ms1, c_{\text{lab}} = \big(2.9977 \pm 0.0004\big)\times10^{8}\ \mathrm{m\,s^{-1}},

which is cc to within 1.31.3 parts in 10410^{4}. Writing the hypothesis as clab=c+kvπc_{\text{lab}} = c + k\,v_{\pi}, the measurement gives k=(0.7±1.3)×104k = (-0.7\pm1.3)\times10^{-4}. That is consistent with zero, and it excludes k=1k=1 by about seven thousand standard deviations.

What to take from this. Both experiments are ⚑ quoted results rather than derivations. The derivation in each case is the prediction they are compared against, which is the arithmetic above. And the logical structure of the pair is worth noting. The ether hypothesis is killed by Michelson–Morley plus aberration plus Fizeau. The emission hypothesis survives all three, and is killed instead by de Sitter and by π0\pi^{0} decay. Between them the two families of experiment eliminate every way of keeping Galilean velocity addition for light. Postulate 2 is what is left standing.

9 · Your turn

Problem 1 — FitzGerald's fix, and what it cannot fix

(a) Redo (2.1.33) assuming the arm along the wind contracts by 1β2\sqrt{1-\beta^{2}} while the transverse arm is unchanged, and keep L1L2L_{1}\neq L_{2}. Show that the fringe shift on rotation vanishes identically, for any β\beta and any arm lengths. (b) The cancellation in (2.1.40) is exact, so there is no residual at order β4\beta^{4} to find in the equal-arm case. Where, then, does an observable β\beta-dependence survive? Compute the total time difference between the two beams for unequal arms with contraction, and show that it depends on β\beta. (c) Estimate the size of that effect for L1L2=16 cmL_{1}-L_{2}=16\ \mathrm{cm} and λ=590 nm\lambda=590\ \mathrm{nm}, as the Earth's speed relative to a hypothetical ether changes from 3030 to 60 kms160\ \mathrm{km\,s^{-1}} over six months. (d) What extra ingredient removes even this?

Solution

(a) With contraction, the arm currently along the wind has ether-frame length Li1β2L_{i}\sqrt{1-\beta^{2}}. Its round-trip time is then (2Li/c)/1β2\big(2L_{i}/c\big)/\sqrt{1-\beta^{2}} by (2.1.40), which is the same functional form as the transverse arm. Hence

Δtbefore=2L1/c1β22L2/c1β2=2(L1L2)/c1β2, \Delta t_{\text{before}} = \frac{2L_{1}/c}{\sqrt{1-\beta^{2}}} - \frac{2L_{2}/c}{\sqrt{1-\beta^{2}}} = \frac{2\big(L_{1}-L_{2}\big)/c}{\sqrt{1-\beta^{2}}},

Now rotate. Arm 2 is along the wind and arm 1 across it, and the result is the same expression, because both arms now carry identical round-trip formulas and only the labels have swapped. The difference is zero identically. The fringe shift on rotation is zero for any β\beta, any L1L_{1} and any L2L_{2}. Michelson–Morley is fully explained, and not merely to leading order.

(b) Look at what survives. The beams still differ in arrival time, by

Δt=2(L1L2)c11β2=2(L1L2)cγ2(L1L2)c(1+12β2). \Delta t = \frac{2\big(L_{1}-L_{2}\big)}{c}\cdot\frac{1}{\sqrt{1-\beta^{2}}} = \frac{2\big(L_{1}-L_{2}\big)}{c}\,\gamma \approx \frac{2\big(L_{1}-L_{2}\big)}{c}\left(1+\tfrac12\beta^{2}\right).

This is not zero, and it depends on β\beta. Rotating the apparatus does not change it, so Michelson–Morley cannot see it. But the Earth's speed relative to any putative ether changes over the year as its orbital velocity swings around. The fringe pattern should therefore drift slowly with the seasons even if the apparatus is never moved. That is precisely the experiment Kennedy and Thorndike performed in 1932, with deliberately unequal arms, and it is the natural complement to Michelson–Morley. MM varies the direction at fixed speed, and KT varies the speed at fixed direction.

(c) The fringe count is N=cΔt/λ=2(L1L2)γ/λN = c\Delta t/\lambda = 2(L_{1}-L_{2})\gamma/\lambda, so the change between two speeds is

ΔN=2(L1L2)λ(γ2γ1)(L1L2)λ(β22β12). \Delta N = \frac{2\big(L_{1}-L_{2}\big)}{\lambda}\Big(\gamma_{2}-\gamma_{1}\Big) \approx \frac{\big(L_{1}-L_{2}\big)}{\lambda}\Big(\beta_{2}^{2}-\beta_{1}^{2}\Big).

With β1=1.0007×104\beta_{1}=1.0007\times10^{-4} and β2=2.0014×104\beta_{2}=2.0014\times10^{-4} we have β22β12=3.004×108\beta_{2}^{2}-\beta_{1}^{2} = 3.004\times10^{-8}, and therefore

ΔN=0.16×3.004×1085.90×107=8.1×103 fringes. \Delta N = \frac{0.16 \times 3.004\times10^{-8}}{5.90\times10^{-7}} = 8.1\times10^{-3}\ \text{fringes}.

That is small, under a hundredth of a fringe. But a stable interferometer watched over months can reach it, and ⚑ Kennedy and Thorndike saw nothing.

(d) Time dilation. The fringe count is a path difference divided by a wavelength, and the wavelength is set by the source, which is moving too. Suppose a moving clock runs slow by γ\gamma. Then light emitted by a moving atom has its period stretched by γ\gamma, and its wavelength stretched by γ\gamma as well. So λγλ0\lambda\to\gamma\lambda_{0} and

N=cΔtλ=2(L1L2)γγλ0=2(L1L2)λ0, N = \frac{c\,\Delta t}{\lambda} = \frac{2\big(L_{1}-L_{2}\big)\gamma}{\gamma\lambda_{0}} = \frac{2\big(L_{1}-L_{2}\big)}{\lambda_{0}},

independent of β\beta. Contraction alone explains Michelson–Morley. Contraction and time dilation together are what Kennedy–Thorndike needs. Historically this is the cleanest demonstration that length contraction is not a complete story. The two experiments together force both effects, and both drop out of Chapter 2.2 from the two postulates without either being assumed. Add the observed isotropy of cc from Michelson–Morley, and the three experiments between them pin the Lorentz transformation down uniquely. That is why this trio is sometimes called the experimental basis of special relativity.

Problem 2 — the transformed wave equation, and which term is the culprit

(a) Repeat §4 for a Galilean boost of speed vv along xx applied to the full three-dimensional wave equation 2ϕ=c2t2ϕ\nabla^{2}\phi = c^{-2}\partial_{t}^{2}\phi, and write the result. (b) Identify precisely which term breaks the symmetry, and give two independent arguments that it cannot appear in a law obeying the principle of relativity. (c) Show that if you keep only terms of order β0\beta^{0} you recover the original equation, then say why that is not a rescue.

Solution

(a) The boost affects only xx and tt, so y=y\partial_{y}=\partial_{y'} and z=z\partial_{z}=\partial_{z'} are untouched, and (2.1.17) applies verbatim to the rest. Substituting into 2ϕ=c2t2ϕ\nabla^{2}\phi = c^{-2}\partial_{t}^{2}\phi:

2ϕx2+2ϕy2+2ϕz2=1c2[2ϕt22v2ϕxt+v22ϕx2], \pdv{^{2}\phi}{x'^{2}} + \pdv{^{2}\phi}{y'^{2}} + \pdv{^{2}\phi}{z'^{2}} = \frac{1}{c^{2}}\left[\pdv{^{2}\phi}{t'^{2}} - 2v\,\frac{\partial^{2}\phi}{\partial x'\,\partial t'} + v^{2}\pdv{^{2}\phi}{x'^{2}}\right],

Bring the v2v^{2} term across and collect the coefficient of 2ϕ/x2\partial^{2}\phi/\partial x'^{2}:

(1β2)2ϕx2+2ϕy2+2ϕz2+2vc22ϕxt1c22ϕt2=0. \big(1-\beta^{2}\big)\pdv{^{2}\phi}{x'^{2}} + \pdv{^{2}\phi}{y'^{2}} + \pdv{^{2}\phi}{z'^{2}} + \frac{2v}{c^{2}}\frac{\partial^{2}\phi}{\partial x'\,\partial t'} - \frac{1}{c^{2}}\pdv{^{2}\phi}{t'^{2}} = 0.

Two symptoms are visible at once. The xx' derivative now carries a different coefficient from the yy' and zz' ones, so space is no longer isotropic. And there is a cross term.

(b) The culprit is 2vc22ϕxt\dfrac{2v}{c^{2}}\dfrac{\partial^{2}\phi}{\partial x'\,\partial t'}. Two arguments.

Parity. Under xxx'\to-x' every other term is unchanged, since each carries an even number of xx' derivatives. The cross term changes sign, because it carries exactly one. So the equation distinguishes +x+x' from x-x', which is to say it knows which way the wind blows. A law valid in a frame with no preferred direction cannot contain such a term.

Boost-parameter dependence. The coefficient is proportional to vv, and vv describes the relationship between two frames rather than anything in the physics. Any law whose coefficients depend on the observer's velocity is by definition not the same law for all observers, which is the negation of Postulate 1.

The modified coefficient 1β21-\beta^{2} is also fatal, but it is a weaker symptom. It is even in vv, so it survives parity, and you could imagine absorbing it by rescaling xx'. You cannot absorb the cross term that way. It is the one that cannot be rescaled away, and eliminating it is exactly what forces tt' to depend on xx in Chapter 2.2. That is worth noticing now. The term that breaks the symmetry is the term mixing space and time derivatives, and the repair will be a transformation that mixes space and time coordinates.

(c) Setting β=0\beta=0 in the boxed equation returns 2ϕ=c2t2ϕ\nabla^{2}\phi=c^{-2}\partial_{t}^{2}\phi exactly. To first order in β\beta only the cross term survives, at relative size 2v/c(ω/k)/c2β2v/c \cdot (\omega/k)/c \sim 2\beta. So for vcv\ll c the equation is nearly unchanged, and that is precisely why the discrepancy went unnoticed for forty years. Every terrestrial source moves at β104\beta\lesssim10^{-4}.

But it is not a rescue, for two reasons. First, "approximately invariant" is not a coherent notion for a symmetry principle. Either the laws are the same in all frames or they are not, and if they are not, then there exists in principle a measurement that identifies your frame, no matter how hard that measurement is. Second, and decisively, the experiments of §5 and Worked example 2 are designed to reach the order at which the discrepancy lives. Michelson and Morley reached β2\beta^{2}, and the π0\pi^{0} experiment reached β1\beta\approx1, where the "small" correction is a factor of two. At those sensitivities the effect is not small, and it is absent.

Problem 3 — stellar aberration and the drag hypothesis

(a) Derive the aberration angle for a star directly overhead from scratch, using the moving telescope picture, and evaluate it for the Earth's orbital speed. (b) Explain quantitatively why an observer sees the star's apparent position trace out a small ellipse over a year, and say what the ellipse's shape depends on. (c) Explain precisely why complete ether drag is incompatible with the observation, and why partial drag does not straightforwardly save it either. (d) A rain analogy is often used. State the analogy, then say exactly where it fails once Postulate 2 is adopted.

Solution

(a) Let the telescope tube have length \ell, and let the Earth move sideways at vv. Light enters the objective and takes /c\ell/c to reach the eyepiece. In that time the eyepiece has moved v/cv\ell/c in the direction of motion. For the light to land on the eyepiece, the tube must be tilted forwards by an angle θ\theta with

tanθ=v/c=vc. \tan\theta = \frac{v\ell/c}{\ell} = \frac{v}{c}.

The tube length cancels, which is the sign of a real effect rather than an instrumental one. With v=29.78 kms1v=29.78\ \mathrm{km\,s^{-1}} this gives θ=9.9335×105 rad\theta = 9.9335\times10^{-5}\ \mathrm{rad}, and multiplying by 206265rad1206265''\,\mathrm{rad^{-1}} turns that into θ=20.49\theta = 20.49''. ⚑ The measured constant of aberration is 20.495520.4955''.

(b) The tilt is always towards the instantaneous direction of the Earth's motion, and that direction rotates through 360360^{\circ} in a year as the Earth goes round its orbit. So the apparent position of any star swings around a closed curve of angular radius v/cv/c.

What shape that curve takes depends on where the star sits. A star at the pole of the ecliptic sees the Earth's velocity vector sweep out a full circle in the plane perpendicular to the line of sight, so its aberration ellipse is a circle of radius 20.520.5''. A star in the plane of the ecliptic sees only the component of the Earth's velocity perpendicular to the line of sight, which oscillates back and forth, so its ellipse degenerates to a line of half-length 20.520.5''. In between, the semi-minor axis is 20.5sinλ20.5''\sin\lambda, with λ\lambda the star's ecliptic latitude.

That pattern is the signature which identifies aberration: the same semi-major axis for every star, and a semi-minor axis depending only on ecliptic latitude and not at all on distance. It is also what distinguishes aberration from parallax, whose size does depend on distance and whose phase is shifted by a quarter of a year.

(c) The derivation in (a) assumed the light travels in a straight line at cc in a frame in which the telescope is moving. Suppose instead that the ether were completely dragged along by the Earth. Then within the dragged region light would propagate in the telescope's own rest frame, the eyepiece would not move relative to the medium during the transit, and there would be no tilt. You would see no annual aberration at all. It is observed, at full strength, so complete drag is excluded.

Partial drag does not rescue it either, and the reason is instructive. If the drag is partial, the aberration should depend on how much dragged medium the light passes through. It should differ for observations made through a long column of air against a short one. Better still, it should change if you fill the telescope tube with water, since water has a Fresnel drag coefficient of 0.440.44 by (2.1.39) and would drag the light substantially. ⚑ Airy performed exactly this experiment in 1871 with a water-filled telescope, and found the aberration unchanged, at the same 20.520.5''. A drag model has to explain why filling the instrument with a strongly dragging medium changes nothing, and the honest answer within ether theory is a conspiracy of cancellations. (Relativity gets Airy's null result immediately. Aberration is a property of the transformation between the source's frame and the observer's, and it has nothing to do with what the light passes through on the way in.)

(d) The analogy. Running through vertically falling rain, you must tilt your umbrella forwards, and the faster you run the more you tilt. The tilt angle satisfies tanθ=vyou/vrain\tan\theta = v_{\text{you}}/v_{\text{rain}}, which is (2.1.37) with the rain speed in place of cc.

Where it fails. The rain calculation is a Galilean velocity addition. In your frame the raindrops have velocity (vyou,vrain)(-v_{\text{you}}, -v_{\text{rain}}), and the tilt is the direction of that resultant. It also predicts that the drops arrive faster in your frame, at vyou2+vrain2\sqrt{v_{\text{you}}^{2}+v_{\text{rain}}^{2}}. For light, Postulate 2 forbids that. The photons arrive at cc, not at v2+c2\sqrt{v^{2}+c^{2}}. So the correct relativistic aberration formula cannot be the vector-addition one, and it is not. Chapter 2.2 derives cosθ=(cosθβ)/(1βcosθ)\cos\theta' = (\cos\theta-\beta)/(1-\beta\cos\theta), which agrees with (2.1.37) to first order in β\beta and differs at second order.

Since β104\beta\approx10^{-4} for the Earth, that correction enters at relative order β2=108\beta^{2}=10^{-8} and reaches at most β2/4=2.5×109 rad\beta^{2}/4 = 2.5\times10^{-9}\ \mathrm{rad}, which is 5×1045\times10^{-4} arcseconds. That is far below the precision of nineteenth-century astrometry. Hence the naive derivation gave the right answer for two centuries, and hence aberration could not have settled the question by itself. It is measurable now, and the relativistic formula is the one that matches.

Problem 4 — how fast would you have to go?

(a) Find the wind speed vv that would produce a fringe shift of exactly 1.0001.000 in the 1887 apparatus (L=11 mL=11\ \mathrm{m}, λ=590 nm\lambda=590\ \mathrm{nm}). Do it first from the leading-order formula, then check against the exact expression. (b) Comment on whether that speed is achievable or plausible. (c) Suppose instead you fix v=30 kms1v=30\ \mathrm{km\,s^{-1}} and ask how long the arms would have to be for a one-fringe shift. Is that achievable? (d) What does the comparison tell you about why the experiment was done in 1887 and not in 1830?

Solution

(a) From (2.1.34) at leading order, ΔN2Lβ2/λ\Delta N \approx 2L\beta^{2}/\lambda, so ΔN=1\Delta N=1 requires

β=λ2L=5.90×10722.0=2.682×108=1.6377×104, \beta = \sqrt{\frac{\lambda}{2L}} = \sqrt{\frac{5.90\times10^{-7}}{22.0}} = \sqrt{2.682\times10^{-8}} = 1.6377\times10^{-4}, v=βc=1.6377×104×2.99792×108=4.91×104 ms1=49.1 kms1. v = \beta c = 1.6377\times10^{-4}\times 2.99792\times10^{8} = 4.91\times10^{4}\ \mathrm{m\,s^{-1}} = 49.1\ \mathrm{km\,s^{-1}}.

Solving the exact expression (2.1.34) numerically gives β=1.637626×104\beta = 1.637626\times10^{-4} and v=49.0948 kms1v=49.0948\ \mathrm{km\,s^{-1}}, the same to six figures, since the correction is O(β2)=108O(\beta^{2})=10^{-8}.

(b) 49 kms149\ \mathrm{km\,s^{-1}} is only 1.61.6 times the Earth's orbital speed. It is laughably out of reach for a laboratory, since no apparatus has ever been moved at anything close to it. But it is utterly ordinary as an astronomical speed. That is the whole point of the experiment's design. It does not need you to move the apparatus, because the Earth is already moving. And it means the null result is not marginal. For the ether to hide, its rest frame would have to lie within about 5 kms15\ \mathrm{km\,s^{-1}} of the Earth at all times of year. That is impossible. The Earth's own velocity changes by 60 kms160\ \mathrm{km\,s^{-1}} between January and July, so any fixed frame is at least 30 kms130\ \mathrm{km\,s^{-1}} away from the Earth for half the year.

(c) Rearranging for LL at ΔN=1\Delta N=1 with β=1.0007×104\beta = 1.0007\times10^{-4}:

L=λ2β2=5.90×1072×1.0014×108=29.5 m. L = \frac{\lambda}{2\beta^{2}} = \frac{5.90\times10^{-7}}{2\times1.0014\times10^{-8}} = 29.5\ \mathrm{m}.

That is under thirty metres of optical path in each arm, held rigid to a fraction of a wavelength and rotatable. Michelson achieved 11 m11\ \mathrm{m} by folding the beam back and forth eight times across a 1.5 m1.5\ \mathrm{m} stone. Tripling that is difficult but not absurd, and later workers did reach tens of metres. So the two routes to sensitivity are not symmetric: you cannot change vv at all, and you can change LL by a factor of a few. The experiment lives or dies on LL. That is why every improvement in this line of work has been an improvement in effective path length, culminating in the modern optical resonators, where the light bounces 10510^{5} times and the effective path is kilometres.

(d) Because ΔN2Lβ2/λ\Delta N \approx 2L\beta^{2}/\lambda has β2=108\beta^{2}=10^{-8} in it, and an experiment sensitive to one part in 10810^{8} was not buildable earlier. Everything that mattered had to arrive first: sufficiently monochromatic sources, high-quality half-silvered mirrors, mechanical isolation by way of the mercury float, and above all the realisation that interferometry converts a hopeless timing measurement into a feasible displacement measurement. Michelson invented the instrument for this purpose.

There is a general pattern here worth noticing: the second-order effect is where the interesting physics is, and second-order effects only become visible when someone builds a null-comparison instrument. The same story runs through Cavendish's torsion balance, Eötvös's test of the equivalence principle (Chapter 3.1), and LIGO, whose optical layout is a Michelson interferometer with 4 km4\ \mathrm{km} arms looking for a displacement of 1018 m10^{-18}\ \mathrm{m}. The instrument that failed to find the ether became, a century later, the instrument that found gravitational waves.

The brick you just laid

You have the crisis, in a form you can state in three sentences and defend line by line.

One. Galilean relativity is exact for Newtonian mechanics. The transformation is x=xvtx'=x-vt, t=tt'=t. Velocities subtract. Accelerations are unchanged. Forces built out of relative positions and relative velocities are unchanged too, so F=ma\vv F=m\vv a has literally the same form in every inertial frame. And you have the separation that matters: the principle of relativity (P) and the transformation (G) that implements it are different claims, fused together for three centuries because t=tt'=t did not look like an assumption.

Two. Maxwell's equations produce a speed. Take the curl of Faraday, commute the derivatives, substitute Ampère–Maxwell, expand the left side with ×(×E)=(E)2E\nabla\times(\nabla\times\vv E)=\nabla(\nabla\cdot\vv E)-\nabla^{2}\vv E, and use E=0\nabla\cdot\vv E=0. Out comes the wave equation of Chapter 0.8, with c=1/μ0ϵ0=2.998×108 ms1c = 1/\sqrt{\mu_{0}\epsilon_{0}} = 2.998\times10^{8}\ \mathrm{m\,s^{-1}}, assembled from a capacitor and a pair of current-carrying wires. That is why light is an electromagnetic wave.

Three. The two are incompatible, and the proof is three lines of chain rule. Under a Galilean boost the wave equation acquires a cross term 2ϕ/xt\propto\partial^{2}\phi/\partial x'\partial t' and a modified coefficient 1β21-\beta^{2}. It changes form. So one of three things has to give. Either the principle of relativity fails for electromagnetism, or Maxwell is wrong, or the Galilean transformation is wrong. The first option makes a prediction, and Michelson and Morley went looking for it and did not find it, at a predicted 0.3730.373 fringes against a bound of 0.010.01.

You also have the anatomy of the effect, which matters more than the history. The parallel arm is slowed by γ2\gamma^{2} and the transverse arm by γ\gamma. The difference between them is second order in β\beta, and the whole observable lives in the gap between those two exponents. You know that FitzGerald's ad hoc contraction by 1β2\sqrt{1-\beta^{2}} cancels the effect exactly. You know that Lorentz had the transformation equations before Einstein did. And you know what 1905 actually contributed, which was not the algebra but the recognition that these are facts about space and time rather than about matter moving through a medium. That is why the ether then had nothing left to do.

Where this gets spent.

  • The two postulatesChapter 2.2, immediately. They are the entire input, and the Lorentz transformation is the unique output. Every strange consequence arrives there as a corollary rather than an assumption: the relativity of simultaneity, time dilation, length contraction.
  • The Pythagoras triangle of (2.1.26) → Chapter 2.2's light clock, where the identical triangle gives time dilation and the identical 1/1β21/\sqrt{1-\beta^{2}} is named γ\gamma.
  • ttt'\neq t, and t/x0\partial t'/\partial x\neq0 → Chapter 2.3, where the failure of simultaneity becomes geometry. Events acquire an invariant separation Δs2=c2Δt2Δx2\Delta s^{2}=c^{2}\Delta t^{2}-\Delta x^{2}, and the light cone is the set of events (2.1.10) can reach.
  • "Both sides transform the same way, so the equality survives"Chapter 2.4. That criterion, isolated in §1's grind box, is promoted there to the definition of a tensor equation, and it becomes the reason the rest of this book is written in indices.
  • The wave equation and the constant ccChapter 2.6, the resolution. Maxwell's equations turn out to be exactly Lorentz invariant, with no modification whatever. They were relativistic all along, forty years before anyone knew what that meant, and it was Newtonian mechanics that had to be changed. That chapter also pays Chapter 1.1's outstanding debt: the field momentum ε0E×B\varepsilon_{0}\vv E\times\vv B that repairs the third law's failure.
  • The whole argument's shape → Chapter 3.1, which runs it again. A principle believed on excellent grounds (relativity) meets a fact believed on excellent grounds (gravitational and inertial mass are equal), the two turn out to be incompatible, and the resolution is again that a piece of assumed structure has to go. This time the casualty is the flatness of spacetime. And Chapter 5.1 runs it a third time for quantum mechanics and relativity, where the resolution is that particle number cannot be conserved and fields become unavoidable.

One sentence to carry: the speed in a wave equation belongs to the frame in which the equation holds, and Maxwell's equations refused to name that frame. Everything in Part II is the consequence of taking that refusal at face value.