Part II · Special Relativity — Chapter 2.1
The Crisis of 1900
Three facts, each of them solid. Any two of them are comfortable together. All three cannot be true at once, and this chapter is the proof.
Part I rebuilt mechanics from an action principle, and it worked. One scalar function, varied over paths, gave back Newton's second law. It gave that law back in any coordinates. It handled constraints without ever naming a constraint force. And through Noether it turned every symmetry into a conservation law. The machinery is sound, and nothing in Part II takes any of it back.
Now we point it at light, and it breaks.
Here is the shape of the chapter, stated in advance so you can watch it close. There are three claims. (1) The laws of physics take the same form in every inertial frame, and the transformation relating those frames is , . That is Galilean relativity. It is older than Newton, and Newton's laws satisfy it exactly. (2) Maxwell's equations describe every electric and magnetic phenomenon anyone had measured by 1890, and they imply that electromagnetic disturbances propagate at a definite speed built out of two laboratory constants. (3) Experiment finds no trace of the frame that (1) and (2) together demand must exist.
Each claim is going to be established here rather than asserted. Claim (1) comes by direct calculation in §1. Claim (2) comes from deriving the wave equation out of Maxwell in §2. Claim (3) comes from working out what Michelson and Morley should have seen in §5 and comparing that with what they did see. Section 4 is the hinge. It shows by explicit chain rule that Maxwell's equations change form under the transformation in (1), and that is where the contradiction actually lives.
One thread is picked up from earlier. Chapter 1.1 §3.3 showed that Newton's third law fails for two charges in relative motion: an isolated pair of particles has , at a computable rate suppressed by . The only available repair was that the field carries the missing momentum, and that suspicious factor was left hanging. It is about to reappear as the size of the effect Michelson and Morley went looking for. It is the same , and the match is not a coincidence. Both are the leading signature of one fact: electromagnetism does not transform the way mechanics does.
What this chapter does not do. It does not derive the Lorentz transformation. That is Chapter 2.2's job, and doing it early would spoil the only honest reason to accept it. That reason is that by the end of §7 you will want some transformation that resolves the contradiction, and there turns out to be essentially one candidate. This chapter ends with two postulates and nothing else.
Tools you'll need — Chapter 0.3: the binomial series, and what "to leading order in " means quantitatively. Chapter 0.6 §5.2: the multivariable chain rule for a change of coordinates, which is all that §4 is, ground out. Chapter 0.7 §4 (curl), §3 (divergence), §7 (the Laplacian and the second-derivative identities): all four Maxwell equations are written in that language, and §2 manipulates them there. Chapter 0.8 §7.6: the wave equation, its solutions , and the fact that a wave equation's speed is a property of the medium. Chapter 1.1 §3.3: the third law failing for moving charges, and the loose end that failure left.
1 · Galilean relativity, stated precisely
Let's start by looking closely at an idea that has been central to physics for over four hundred years: Galilean relativity. It is that old, and it is correct. To understand the crisis that sparked modern physics, we first need to separate the core principle of relativity from the specific mathematical transformation we use to calculate it. For centuries these two concepts were treated as the exact same thing. Once we pry them apart, the path forward becomes much clearer. Chapter 2.2 is where they finally come apart for good.
1.1 · The transformation
Imagine two observers, each with their own coordinate system, or "frame", of reference. We will call the stationary frame and the moving frame . Let's give them both a measuring tape and a clock, and assume they perfectly agree on how to use them.
To keep the maths simple, we'll align their axes and have them start their stopwatches at exactly when they pass each other. After that moment, the frame slides away in the positive direction at a steady speed , as measured in . This setup is known as standard configuration, and it is the baseline for everything in Part II.
Now, suppose something happens — an "event" — and both observers write down where and when it occurred.
- The observer in frame records the position as and the time as .
- Meanwhile, the origin of the moving frame has shifted forward by a distance . Because of this shift, the moving observer will measure the event's position as .
We are also going to make a crucial assumption here, and we put it in explicitly so that we can watch it die later: both observers will record the exact same time . This leads us directly to the classical Galilean transformation:
Let's pause and look at that final equation, . It might look like simple bookkeeping, but it actually contains a massive physical claim: there is a single, universal clock ticking at the exact same rate for everyone in the universe, regardless of how fast they are moving, so that two events either are simultaneous or are not, with no reference to who is asking. Isaac Newton built his mechanics on this idea of absolute time, and he stated it in exactly those terms: absolute time, flowing equably, without relation to anything external. It feels so intuitive to our everyday experience that it is hard to realise it is merely an assumption.
As we dive deeper into relativity, the first three equations of (2.1.1) will survive into Chapter 2.2 with a few tweaks. This universal time equation will completely fall apart.
1.2 · Velocities add
Next, let's see how our observers view motion. Imagine a particle travelling along a path. The stationary observer tracks its position as . Using our transformation (2.1.1), the moving observer tracks its position as .
To find the particle's velocity, we just need to take the derivative. Since we assumed that time is the same for both observers (), differentiating with respect to is identical to differentiating with respect to :
This equation tells us something very familiar: velocities add or subtract depending on your point of view. If you throw a ball forward at inside a train moving at , a person standing on the platform sees the ball moving at , the speed of the throw plus the speed of the train. Nobody has ever been surprised by this, and for everyday objects it is incredibly accurate.
But notice the hidden mechanics here. The subtraction comes from the shift in space, while the ease of the derivative relies entirely on the assumption that . Chapter 2.2 keeps the first ingredient and destroys the second. When that assumption breaks later on, this rule for adding velocities will have to change too, and (2.1.2) is what changes.
1.3 · Accelerations do not change, so survives
Now differentiate (2.1.2) once more. The relative velocity is a constant, since that is what it means for one inertial frame to move uniformly relative to another, so its derivative vanishes:
In three dimensions the same computation on each component gives . Acceleration is a Galilean invariant. Two observers in uniform relative motion disagree about where a particle is and about how fast it is going, and they agree exactly about how it is accelerating.
That single fact is why Newtonian mechanics is compatible with the principle of relativity. Here is the argument. Newton's second law in reads , and we want to know what it looks like in , so transform it piece by piece. The mass is a property of the body and does not depend on who is watching. The acceleration is unchanged by (2.1.3). So the law in reads , and the only remaining question is whether the force is the same in both frames.
For the forces Newtonian physics actually uses, it is. Every one of them depends on the particles' relative positions, and possibly on their relative velocities:
The boost cancels in both differences. Gravity depends on , and so do Coulomb's law and Hooke's law. Even a drag force proportional to the relative velocity of body and fluid is untouched. So for each of them, and
The law has the same shape in the new frame. Not merely a shape that can be corrected into the old one by adding terms, but the identical shape, with the identical symbols meaning the identical things. Compare Chapter 1.1 §4.2, where writing in polar coordinates produced two extra terms out of nothing. Nothing like that happens here.
Grind box — the full Galilean group, and what "same form" is allowed to mean
(2.1.1) is one member of a larger family. The complete set of transformations that carry one inertial frame to another, in Newtonian physics, is
where is a constant rotation matrix ( parameters), a constant boost velocity (), a constant spatial displacement () and a constant time offset (). Ten parameters. They compose and invert among themselves, since the composition of two boosts is a boost with and the inverse of a rotation is a rotation. So they form a group, the Galilean group. Chapter 6.1 gives that word a definition. For now the content is just that "change of inertial frame" is a closed operation.
Differentiating twice kills and outright and leaves , so acceleration is not quite invariant under the full group. It rotates, as a vector must. But it rotates the same way does, so still maps to . That is the precise sense in which "the form is unchanged": both sides of the equation transform identically, so the equality is preserved. Hold onto that formulation. It is the exact criterion Chapter 2.4 turns into the definition of a tensor equation, and it is the reason index notation exists.
What is not allowed. Let the relative velocity depend on time, with . Then
and an extra term has appeared that is not any interaction between bodies. That is the fictitious force of Chapter 1.1 §1.1, and its presence is exactly why the first law is an existence claim about a restricted class of frames rather than a corollary of the second. The principle of relativity was never a claim about all frames. It is a claim about the inertial ones, and the boost velocity must be constant for the argument of §1.3 to go through.
One more piece of fine print, cheap now and expensive later. Nothing above proves that (2.1.1) is the only transformation with these properties. We derived from the Galilean transformation. We did not show that a theory respecting the principle of relativity has to use it. That gap is the whole of Chapter 2.2, and the reason nobody looked into it for three hundred years is that did not feel like an assumption.
1.4 · The principle itself
Now separate the two ideas that §1.1–1.3 ran together.
The laws of physics take the same form in all inertial frames. No experiment performed entirely within a uniformly moving laboratory can determine that laboratory's velocity.
This is not Einstein's. It is Galileo's, from 1632, and his statement of it is better than most modern ones. Shut yourself below decks on a large ship with some flies, a bowl of water, and a friend to throw you a ball. You will find that everything proceeds exactly as it did in port. The flies do not pile up at the stern, the ball needs no extra effort thrown forwards rather than backwards, and the drops fall straight into the vessel beneath. ⚑ (Paraphrased from the Dialogo. The reasoning is his, the words are not.) By 1900 nobody proposed to abandon it. It was, and is, one of the best-tested statements in physics.
The principle says something about the relationship between frames. It does not say what the transformation between them is. Those are separate questions, and the failure to notice they were separate is the single reason the crisis took forty years to resolve. Written out:
| Claim | Status in 1900 | |
|---|---|---|
| (P) | The laws take the same form in all inertial frames | Universally accepted; verified constantly |
| (G) | Inertial frames are related by (2.1.1), with | Assumed without comment by everyone |
Everybody held both. It had never occurred to anyone that they were different claims, because (G) was not perceived as a claim at all. It was perceived as what the words in (P) meant. The resolution of this chapter's crisis is that (P) is right and (G) is wrong. Keep the pair separate from here on, because the rest of the chapter is about the strain between them.
A principle and the transformation that implements it are separate claims, and for three centuries nobody noticed, because the second was never perceived as a claim at all but as what the words of the first meant. The principle is Galileo's, from 1632: shut yourself below decks with some flies and a bowl of water, and everything proceeds as it did in port, so no experiment inside a smoothly moving laboratory can reveal how fast it is going.
The implementation is the familiar one. Subtract the distance the other origin has slid, leave the two directions across the motion alone, and give both observers the same reading for when anything happened. That last instruction carries all the weight, because it asserts a single universal clock, so that two events either are simultaneous or are not, with no reference to who is asking. It is the assumption the previous part promised would have to go.
Mechanics does not mind in the slightest. Differentiate twice and the steady relative velocity vanishes, so both observers agree about acceleration; and every force Newtonian physics uses depends on separations between bodies, which the shift leaves untouched. The law reappears in the moving frame with the identical shape rather than one needing repair by extra terms, which is exactly what the turning basis directions of polar coordinates had inflicted on it.
2 · Maxwell's equations produce a speed
Now the second of the three claims. It has to be derived rather than quoted, because the number that falls out at the end is the whole point.
The four equations below are the empirical input of this chapter. Each one summarises a body of nineteenth-century laboratory work, from Coulomb, Gauss, Ørsted, Ampère and Faraday. The last term in the last equation is Maxwell's own addition, which Chapter 0.7 Problem 4 showed is forced by charge conservation and the identity . In vacuum (, ) and in SI units:
Nothing else about electromagnetism is assumed anywhere in this chapter. The two constants and are measured, and §2.3 says exactly how. Chapter 2.6 rewrites all four equations as a single manifestly Lorentz-covariant statement, and shows that they were never four equations at all. Chapter 6.3 derives them from a symmetry principle. Here they are input.
2.1 · One identity first
The derivation needs the curl of a curl, which Chapter 0.7 §7 did not have occasion to compute. It is worth deriving rather than looking up, because the pattern of the answer is not guessable.
Here means the Laplacian of Chapter 0.7 §7.5 applied to each Cartesian component separately. That definition is only available because the components are Cartesian, a caveat that becomes important in Chapter 3.3 and is harmless here.
Grind box — (2.1.6) by brute force, one component
Write , whose components are, from Chapter 0.7 §4,
Take the -component of and substitute:
Now the trick, and it is the only step with any content. Add and subtract the missing term . The first group then completes to a divergence and the second to a Laplacian:
Both regroupings used Clairaut's theorem (Chapter 0.6 §6.1) to swap for , and its hypothesis is continuous second partials. That is the same hypothesis that made work in Chapter 0.7 §7.2. The and components follow by cycling , under which every expression above is invariant. Hence (2.1.6). ✓
Why the answer looks like that. is a second-order operator, so it must be some combination of acting on . Only two such combinations produce a vector: the gradient of the divergence, and the Laplacian of the vector. The identity says the coefficients are and . In Chapter 3.5 this becomes the statement that the operator splits in exactly this way, and the minus sign is the one that makes the electromagnetic wave equation come out with the right relative sign between space and time.
2.2 · Take the curl of Faraday's law
We want an equation for alone. Faraday's law relates to , and Ampère–Maxwell relates back to . So apply one to the other and eliminate . The operation that lets Ampère–Maxwell in is the curl, because the curl of is precisely what that law tells us. Start from Faraday and hit both sides with :
The second step swapped a space derivative with a time derivative. That is Clairaut again. The curl is built from , none of which is , so as long as has continuous second partials the two operations commute. It is a small step and worth naming, because in Chapter 3.3 the analogous swap will fail, and the failure will be the curvature of spacetime.
Now substitute Ampère–Maxwell on the right, with :
Next expand the left-hand side with (2.1.6), using , which holds because we are in vacuum with no charge anywhere to source a divergence:
Set the two expressions equal and cancel the overall minus sign:
Look at what has happened. We started with four equations coupling two fields, performed three manipulations, and produced an equation for by itself with no in it. And the equation is not a new object. It is the wave equation of Chapter 0.8 §7.6, so let us set the two side by side:
with three spatial dimensions instead of one, and a vector where the string had a scalar. Now read off the coefficient:
Nobody put a speed in. Two constants measured with capacitors and current-carrying wires went in, and a propagation speed came out.
Grind box — the same for , and the plane-wave solution
The magnetic field obeys the identical equation. Run the argument the other way. Take the curl of Ampère–Maxwell:
using Faraday in the last step. Now expand the left with (2.1.6) and use , which holds always rather than merely in vacuum, since there are no magnetic charges. So
the same equation with the same speed. The two fields propagate together, which they must, since they are two aspects of one thing, as Chapter 2.6 makes literal.
A plane wave, and . Try a solution travelling along with pointing along , so that . Then and , so (2.1.10) requires
That is Chapter 0.8's relation between frequency, wavenumber and wave speed, now with the speed fixed by two laboratory constants. Notice that it holds for every . The speed does not depend on the wavelength, so vacuum is non-dispersive and a pulse of any shape travels undistorted. That is Chapter 0.8's , and it is why we see sharp images of distant stars rather than smeared ones.
One honest caveat. (2.1.10) is a consequence of Maxwell's equations, not an equivalent of them. Every solution of Maxwell solves it, and not every solution of it solves Maxwell. Feeding the plane wave back into the original four equations imposes two further conditions we did not need above. First, forces propagation direction, so the wave is transverse. Second, Faraday fixes , so is perpendicular to both and the direction of travel, in phase with , and smaller by a factor . None of that is needed for this chapter's argument, and all of it matters in Chapter 2.6.
2.3 · Put the numbers in
Now the moment. Here is where and come from, because the provenance is the point.
⚑ , the permittivity of free space, is fixed by measuring the capacitance of a parallel-plate capacitor. Put a known charge on two plates of known area and separation, measure the voltage, and falls out of . This is a tabletop electrostatics experiment involving no light, no motion and no waves. Its value is .
⚑ , the permeability of free space, is fixed by measuring the force between two parallel current-carrying wires. Run known currents through them, measure the attraction per unit length, and falls out of . This is a tabletop magnetostatics experiment involving no light and no waves either. Its value is .
Multiply them together and take the reciprocal square root:
The measured speed of light, from Fizeau's toothed wheel in 1849 and Foucault's rotating mirror in 1862, was ⚑ about .
A capacitor and a pair of wires told you the speed of light.
Not "a number of the same order". The speed of light, to the precision of the input data, out of two constants that were measured in experiments containing no light at all. There is no route by which optics could have leaked into either measurement. One is a static charge on metal plates, and the other a steady current in a wire. The only way the number can come out right is if light is the propagating disturbance (2.1.10) describes.
⚑ Maxwell drew exactly that conclusion in 1862, in one of the most consequential sentences in physics: we can scarcely avoid the inference that light consists in the transverse undulations of the same medium which is the cause of electric and magnetic phenomena. Optics was a subject two thousand years old, with its own laws, its own instruments and its own practitioners. It stopped being a separate science and became a chapter of electromagnetism. Hertz generated and detected the waves directly in 1887, and radio, X-rays and everything else followed from taking (2.1.12) seriously across the spectrum.
This is what a successful unification looks like from the inside. Not a philosophical reorganisation, but two numbers you measured for unrelated reasons combining into a third number you had already measured for a third reason.
A sharp reader will object. In the SI system in force for most of the twentieth century, was defined to be exactly , as part of the definition of the ampere. And since 1983 the metre has been defined by fixing exactly. If the constants are defined rather than measured, has anything been discovered?
Yes, and the objection is worth taking seriously, because it locates the physical content precisely. Unit conventions can move a numerical factor from one constant to another. They cannot create a relationship between independently measurable quantities. The invariant statement is this: the ratio of the electrostatic to the electromagnetic unit of charge has the dimensions of a speed, and that speed can be measured with nothing but charges, currents and a balance. Weber and Kohlrausch did precisely that in 1856, discharging a capacitor of measured electrostatic capacity through a galvanometer of measured electromagnetic sensitivity, and got ⚑ , six years before Maxwell's remark and with no reference to light whatever. That measurement is the physics. It cannot be conjured out of a choice of units, and it is the reason the coincidence was persuasive at the time.
In modern SI the bookkeeping runs the other way. Now is exact by definition, is a measured quantity differing from in the ninth decimal place, and (2.1.12) is the relation used to determine it. The equation is doing work in both directions. Only the label "defined" has moved.
The shape of the argument in this section is one you have run yourself, and having run it is what makes §2.3 feel inevitable rather than lucky.
Take a drug obeying the first-order kinetics of Chapter 0.1's callout. Two of its constants can be measured without ever watching a concentration fall. The volume of distribution comes from a single dilution: give a known dose, let it mix, measure the concentration once, and . The clearance comes from a steady state: infuse at a known rate until the concentration stops changing, and . Neither experiment contains a half-life, and neither requires waiting for anything to decay. Yet
gives a time, assembled out of a volume and a flow. That is (2.1.12) with different letters on it: a rate of propagation manufactured out of two constants measured in experiments containing no propagation.
The step that matters is the comparison, not the prediction. Computing is bookkeeping. Going out, measuring the terminal slope of a real concentration–time curve, and finding that it agrees is a test. And when it fails to agree, the disagreement is informative rather than embarrassing, because it says a second compartment is hiding, which is the case Chapter 0.8's callout works through in full. Maxwell's step is exactly this comparison, performed once. A capacitor and a current balance predicted , Fizeau's toothed wheel had already measured , and the two numbers had no business being the same.
One difference is worth naming, because it is why his conclusion was so much larger than yours. and are properties of a drug and a patient, so agreement validates a model of that pair and says nothing about anything else. and are properties of empty space, so agreement could not possibly be a fact about a particular apparatus, and the only object left available to be identified was light itself. A model check and a unification share their arithmetic and differ entirely in their consequences, and what separates them is what the constants belong to.
Nothing resembling a speed was put into the four equations of electromagnetism, and a speed came out. Apply the law tying a changing magnetic field to a circulating electric one to its partner, which ties a changing electric field to a circulating magnetic one, and the two fields uncouple, leaving one equation apiece. That equation is the one a chain of masses and springs produced back in the toolkit, and reading off its coefficient gives a speed built from two laboratory constants.
The provenance of those constants is the point. One is fixed by putting a known charge on two metal plates and measuring the voltage across them; the other by running known currents through parallel wires and measuring how hard they pull. Neither experiment contains any light, any motion, or any waves.
No channel exists by which optics could have leaked into a static charge on metal or a steady current in a wire, so the only way the number comes out right is if light is the disturbance those equations describe. Maxwell drew that conclusion in 1862, and optics, two thousand years old and with its own laws and its own instruments, stopped being a separate science. This is what a successful unification looks like from the inside: two numbers measured for unrelated reasons combining into a third that had already been measured for a third.
3 · The question nobody could answer: with respect to what?
Now the trouble starts, and it starts with a question that is embarrassing in its simplicity.
(2.1.12) gives a speed. A speed is a rate of change of position, and position is measured relative to something. So here is the question: relative to what is the light going at ?
3.1 · Every other wave answers this question easily
Chapter 0.8 §7.6 built the wave equation from a chain of masses and springs and got . Look at what that formula is made of: the tension and the linear density of the string. Both are properties of the medium, so the speed is a property of the medium. It is therefore the speed relative to the medium. If the string is being reeled in while the wave travels along it, an observer at the side of the room measures something else.
Every wave in nineteenth-century physics was like this, without exception:
| Wave | Speed | Relative to |
|---|---|---|
| Transverse wave on a string | the string | |
| Sound in air | the air | |
| Ripples on water | (deep water) | the water |
| Seismic -waves | the rock |
And the consequence is entirely familiar. Stand in a wind of speed and sound travelling downwind reaches you at , upwind at . That is (2.1.2) applied to a wave, and it is measured routinely. The speed in the formula is the speed in the rest frame of the medium, and in any other frame you add the medium's velocity.
3.2 · The wave equation itself makes the point
You do not even need the physical picture. The mathematics says it directly. The wave equation is written in some particular set of coordinates , and its solutions are and , which are disturbances moving at in those coordinates. Suppose the equation holds in one coordinate system with speed , and suppose coordinates transform by (2.1.1). Then in a coordinate system moving at the same disturbance moves at . The number in the equation is therefore attached to one preferred frame, namely the frame in which the equation takes that form.
Apply this to (2.1.10). Maxwell's equations, written as they always are written, hold in some frame. In that frame light goes at in every direction. In a frame moving at through it, light should go at in one direction and in the other. So Maxwell's equations appear to single out a preferred frame of reference. That flatly contradicts claim (P) of §1.4, and it does so for reasons that have nothing to do with any experiment.
3.3 · The ether, taken seriously
Given all of §3.1, there was one obvious move, and it was not a stupid one. If light is a wave, it is a wave in something. Name that something the luminiferous ether, and declare Maxwell's equations to hold in the ether's rest frame. The preferred frame of §3.2 is then no longer an embarrassment. It is only the rest frame of a medium, exactly like the air for sound. The principle of relativity survives untouched, because the ether is a physical object, and detecting motion relative to it is no more mysterious than feeling a breeze.
This was not a fudge. It was the only known way for a wave to work, and every other wave anyone had ever studied confirmed it. Refusing to posit a medium in 1880 would have been the strange move, not the sober one. And the hypothesis was productive. It made a sharp prediction: the Earth moves, so there must be an ether wind, so the speed of light must be anisotropic in the laboratory. That prediction is testable, and §5 tests it.
The ether was, however, a demanding object. Its properties were fixed by (2.1.10), and they do not sit comfortably together.
Grind box — what the ether had to be like, quantitatively
It had to be a solid. Light is transverse, since a plane wave has perpendicular to the direction of travel, as the previous grind box showed. A transverse wave is a shear disturbance, with neighbouring layers of the medium sliding past one another. Fluids do not resist shear, which is why sound in air is purely longitudinal and why the Earth's liquid outer core transmits no -waves. So the ether had to have rigidity. It had to be an elastic solid filling all of space, and the planets had to pass through it without measurable resistance.
And an extraordinarily stiff one. For a transverse wave in an elastic solid the speed is , with the shear modulus, which is the same square root as in Chapter 0.8. Setting fixes neither the stiffness nor the density but their ratio:
Compare steel: shear modulus , density , so . And indeed , which is the shear-wave speed in steel. The ether's stiffness-to-density ratio therefore had to exceed steel's by a factor
nearly ten billion. You may make the density as small as you like to keep the planets moving, but the ratio is not negotiable, and it is the ratio that has to be explained.
And it had to be undetectable in every other way. Transparent, frictionless, incompressible enough to fill the space between the stars, non-interacting with matter except through this one channel, and possessed of no measurable effect on anything except that light goes through it.
Nineteenth-century physicists were entirely aware of this list, and it bothered them. Kelvin called the ether one of the two clouds on the horizon of physics. But an awkward hypothesis that makes a testable prediction is a perfectly respectable scientific object, and the response was the right one: go and measure the wind.
Sound in air, ripples on a pond, a pulse running down a rope, tremors through rock: every one of them travels at a speed assembled out of properties of the stuff it travels in, and therefore travels at that speed relative to the stuff. Stand in a wind and sound reaches you faster downwind than up, by exactly the wind's speed. So asking what the new speed was measured against was not a foolish question but an obligatory one.
The mathematics says it more sharply than the analogy does. A wave equation is written in some particular set of coordinates and its solutions run at the stated speed in those, so to an observer drifting past at some rate the same disturbance runs at a different one. The equations of electromagnetism therefore single out one frame, contradicting the principle of relativity before any experiment has been performed.
Naming a medium was the sober response rather than a fudge, because it was the only known way for a wave to work. It cost something. The medium had to resist shearing in order to carry a transverse disturbance, so it had to be an elastic solid filling all of space, with a stiffness-to-density ratio ten billion times steel's, through which the planets passed without measurable drag. Nineteenth-century physicists knew that list and it bothered them. An awkward hypothesis making a sharp prediction is still a respectable object.
4 · Maxwell's equations are not Galilean invariant
Section 3 argued from physical analogy that Maxwell's equations must pick out a frame. This section proves it, by computation, with no analogy anywhere. It is the hinge of the chapter.
Strip the problem to its skeleton. Take one Cartesian component of and call it . Suppress and , so that the Laplacian is . Then (2.1.10) becomes the one-dimensional wave equation:
Here is the question. What does this equation look like in a frame moving at along ? Not what do its solutions look like, but what does the equation look like. If the principle of relativity holds with the Galilean transformation, the answer must be , character for character.
4.1 · The chain rule, set up carefully
We have new coordinates and , and a function that we may regard either as a function of or of . It is one physical field, described two ways. Chapter 0.6 §5.2 gives the rule for converting derivatives: each old derivative becomes a sum over the new ones, weighted by how each new coordinate responds to the old one.
The four partial derivatives of the new coordinates with respect to the old come straight from and :
Take a moment over . It says that changing where you are does not change what time it is. That is the mathematical form of absolute simultaneity. It is the only place enters this calculation, and it is the entry that Chapter 2.2 will make nonzero. Now substitute (2.1.16) into (2.1.15):
The spatial derivative is untouched, and the time derivative acquires an extra piece. That is the whole asymmetry, and everything below is bookkeeping. The physical reading of the second relation is worth having. Holding your position fixed in the old frame means drifting backwards at in the new one, so the rate of change you measure standing still picks up a term from the motion. It is the convective derivative of fluid mechanics, arrived at without meaning to.
4.2 · Second derivatives, and the term that ruins everything
Grind box — squaring the operators, every step
Space. Apply (2.1.17) twice:
Nothing to do.
Time. Here we must square a sum of two operators. The one thing to be careful about is that operators need not commute, so , and we may not jump to without checking. Write and :
The two middle terms are equal because of Clairaut's theorem, since mixed partials of a twice-continuously-differentiable function commute. They are also equal because is a constant, so it passes through without generating anything. Both facts are needed, and both are easy to use without noticing. Hence
Assemble. Substitute both results into (2.1.14):
Bring the term to the left and collect the coefficient of :
which is (2.1.18) below, rearranged.
Here is the result of that grind, written with everything on one side and with , as the convention of Part II requires:
Compare that with what we wanted, which was (2.1.14) with primes on it. Two things have gone wrong, and they are worth naming separately.
(i) A cross term has appeared. The middle term has no counterpart in (2.1.14). It is odd in , so reversing the direction of motion changes its sign, and that is precisely how an equation encodes a preferred direction. A term that knows which way is "forwards" cannot appear in an isotropic law.
(ii) The leading coefficient has changed. has become . Even if you could somehow dispose of the cross term, the coefficients no longer match.
The wave equation changes form under a Galilean boost. If it holds in one inertial frame, it does not hold in any other.
This is not an approximation, not a small effect, and not a subtlety of interpretation. It is a three-line consequence of the chain rule. And since (2.1.14) is a consequence of Maxwell's equations, the same is true of them: Maxwell's equations are not Galilean invariant.
4.3 · What the transformed equation is actually saying
(2.1.18) is not wrong. It is the correct description of a wave in a medium, as seen by an observer moving through that medium. It is worth extracting exactly what it says, both to check the algebra and to see what is at stake.
So look for travelling-wave solutions in the new frame: for some speed and arbitrary shape , exactly as in Chapter 0.8 Problem 4. Each derivative brings down a factor, with , and . Substitute into (2.1.18) and cancel the common , which is nonzero for any wave worth the name. What is left is a condition on alone:
That is a quadratic in . Solve it, either by completing the square or with the formula:
Those are exactly the two speeds §3.1 predicted on physical grounds. In the moving frame, light chases off in the direction at and comes back the other way at . The algebra of §4.2 and the sound-in-a-wind picture of §3.1 are the same statement.
So (2.1.18) is perfectly consistent physics, for a wave in a medium. Sound obeys the analogous equation and nobody minds, because sound has a medium and its rest frame is not mysterious. The entire question is whether light is like that.
We now have a genuine contradiction between two things both believed in 1900, and there are exactly three ways out. They are mutually exclusive, and they are all uncomfortable.
(A) The principle of relativity fails for electromagnetism. Mechanics obeys it and optics does not. There is a preferred frame, the ether's, and Maxwell's equations hold only there. Every other frame gets (2.1.18). This was the majority view, it is entirely coherent, and it makes a prediction: you can measure your velocity through the ether by measuring the speed of light in different directions.
(B) Maxwell's equations are wrong, or at least are only the low-velocity limit of some Galilean-invariant electrodynamics still to be found. Several people tried, and the attempts (Hertz's, Ritz's) all made predictions that failed. They were also fighting uphill, since Maxwell's equations were passing every test anyone could devise.
(C) The Galilean transformation is wrong. Frames are related by something else, and (2.1.1) is only its small- approximation. Then the principle of relativity can hold for both mechanics and electromagnetism. The price is that Newton's laws, which are exactly Galilean invariant by §1.3, must be modified instead.
Notice that (C) requires giving up . Nothing else in (2.1.1) has enough freedom in it. That is why (C) looked, in 1900, like the least attractive option available, and it is why the experiments came first.
Analogy has been carrying the argument, and analogy can always be resisted, so the case is now made by computation. Take the wave equation, strip it to one dimension, and rewrite it in coordinates sliding past at a steady rate, which is the chain rule and nothing else. The spatial derivative comes through untouched; the time derivative picks up an extra piece, because holding your position fixed in one frame means drifting backwards in the other.
Two things then go wrong at once. A cross term appears with no counterpart in the original, and it reverses sign when the motion does, which is precisely how an equation announces that it knows which way is forwards. The leading coefficient has also changed. So the equation holds in one frame and no other, and the demonstration is three lines long.
One entry in the calculation did all the damage: the statement that changing where you are does not change what time it is. The fork is therefore genuine, and there are exactly three ways through. Either the principle of relativity fails for electromagnetism and a preferred frame exists after all, or the equations of electromagnetism are wrong, or the transformation is wrong and the shared clock goes with it. The transformed equation is not nonsense; it correctly describes a wave in a medium seen by somebody moving through it. Everything turns on whether light is like that.
5 · Michelson–Morley
Branch (A) is the one that is directly testable, and testing it produced the most famous experiment in physics. The logic is short. The Earth orbits the Sun at about , so unless the ether happens to be dragged along perfectly, there is an ether wind blowing through the laboratory. Measure the speed of light along the wind and across it, and the two answers should differ.
The difficulty is that they differ by very little. From §4.3 the two one-way speeds are with . A direct timing measurement would therefore have to resolve one part in of a quantity that was not itself known that well. The best absolute determination in 1887 was Michelson's own, from 1879, at , which is good to about two parts in . The difference he needed to see was smaller than the error bar on the whole. Michelson's answer was to stop measuring speed and start measuring interference, which compares two light paths against each other and can resolve a hundredth of a wavelength.
5.1 · The instrument
A source of monochromatic light shines on a half-silvered mirror set at , called a beam splitter. Half the light passes through and continues to a mirror at distance . Half reflects off at right angles and travels to a second mirror at distance . Both beams return to the splitter, recombine, and go on to a detector. Whether they arrive in phase or out of phase depends on the difference in their travel times, so the recombined beam shows a pattern of interference fringes.
Two facts about this arrangement do the real work. First, the arms are perpendicular, so if there is an ether wind then one arm can lie along it and the other across it. Second, and this is the piece of experimental cunning without which the whole thing is impossible, you cannot know and to a fraction of a wavelength. The absolute fringe position is therefore uninterpretable. What you can do instead is rotate the whole apparatus by , which exchanges the roles of the two arms without changing either length, and then watch the fringes move. §5.5 shows that the unknown mismatch cancels out of that measurement exactly.
5.2 · The arm along the wind
Work in the ether's rest frame, where light travels at in all directions and the apparatus moves at in the direction. Working in the ether frame is the honest choice, because it is the frame where we are entitled to say light moves at , and that is the hypothesis under test.
Set when a pulse leaves the splitter, and put the splitter at the origin at that instant. Thereafter the splitter is at and the far mirror at , both sliding to the right at .
Outbound. The pulse moves right at , so it is at , while the mirror is at . They meet when
The mirror is running away from the light, so the light closes the gap at only . Notice that this is not a claim that the light moves at . It moves at in the ether frame, as assumed. The is a closing speed between two objects, and closing speeds may perfectly well exceed or fall below it without anything moving at that rate.
Return. Now the light travels left at from the mirror's position while the splitter advances to meet it. By the identical argument the gap closes at :
Add them, put over a common denominator, and divide top and bottom by :
The round trip is slower than the it would take with no wind, by a factor . That is worth a sentence, because the intuition points the wrong way. The two legs seem as though they ought to average out, one gaining what the other loses. They do not. The reason is that the leg travelled at the lower speed lasts longer, so it gets more weight in the total. Losing time at for a long while beats gaining it at briefly. The same asymmetry is why a round trip against and with a current always takes longer than the same trip in still water, and why your average speed on a there-and-back drive is the harmonic mean rather than the arithmetic one.
5.3 · The arm across the wind, done properly
This is where most treatments go soft, so we will be slow. The naive answer says the light goes straight across at , so that the round trip is and the transverse arm is unaffected. That answer is wrong, and it is wrong for a reason that becomes central in Chapter 2.2.
The mirror at the end of the transverse arm is moving sideways. During the time the light is in flight, that mirror slides a distance down the -axis. Light that left the splitter travelling straight along in the ether frame would arrive where the mirror used to be and miss it. To hit the mirror, the light must leave with a velocity that has an -component matching the apparatus's own motion.
Write that as an equation. Let the light's velocity in the ether frame be with, by hypothesis, . For the light to stay over the moving arm it needs . Then
The useful component of the light's velocity, the part that gets it across the arm, is not but . The rest of its speed budget is spent keeping up sideways. The arm is long and the crossing is covered at , so
with the factor of because the return trip is the mirror image of the outbound one and takes the same time.
The same result comes out of a right triangle, and that is the version to carry in your head. In time the light travels a straight-line distance in the ether frame, which is the hypotenuse. The apparatus has drifted , which is the base. The arm is , which is the height. Pythagoras:
and as before. Same answer, same square root. The square root has the same origin in both derivations: a fixed total speed has to be shared between "across" and "along", and Pythagoras charges you for the sharing. That triangle is going to reappear in Chapter 2.2 as the derivation of time dilation, with the arm replaced by a light clock, and there the factor will be named .
It is very natural to treat the transverse arm as a clean reference. The wind blows sideways across it, so surely it is untouched, and all the physics is in the parallel arm. That reading is wrong, and getting it wrong changes the answer by a factor of two.
Compare the two results at small , using the binomial series of Chapter 0.3:
Both arms are slowed. The parallel arm is slowed by , the transverse arm by , where , so that . Those two exponents are exactly the ones sitting in (2.1.23) and (2.1.25). If you wrongly take you get a coefficient of instead of , and you predict twice the fringe shift.
And the difference between and is not an incidental detail. It is the entire observable. The experiment cannot measure either round-trip time. It can only compare them. Everything Michelson and Morley could possibly have seen lives in the gap between those two exponents, and if the two arms had been affected identically the experiment would have had nothing to detect no matter how sensitive it was. Chapter 2.2's resolution has to reproduce exactly this structure, one power of along the motion and none across it. §6.2 shows that FitzGerald spotted the shape of the fix from precisely this observation, fifteen years before anyone knew why it was true.
5.4 · The difference
Take equal arms, , which is what the instrument is built to have and what §5.5 shows we do not actually need. Subtract (2.1.25) from (2.1.23):
That is exact. Now expand for small , since and what we want to know is how big the effect is. Chapter 0.3's binomial series gives, with ,
We want the difference of those two series, so subtract the second from the first term by term. The bracket in (2.1.27) becomes
Only the leading piece matters at , so keep the term and drop everything after it. The round-trip difference is then
Look at what has dropped out. The two arms agree at order and again at order , and they first differ at order . The effect is second order in , and that is the central practical fact about the entire subject. There is no first-order ether-wind effect in a round-trip measurement, because the delay on the way out is compensated to first order by the gain on the way back. Every experiment sensitive only to first order was therefore guaranteed a null result before it was built, and there had been several. Michelson's design was the first with second-order sensitivity, which is why it counts.
And that is the same suppression factor that made the third law's failure so hard to notice in Chapter 1.1 §3.3. Both are the leading symptom of the same disease.
5.5 · What rotating the apparatus measures
Now the experimental point flagged in §5.1. You cannot measure directly. Doing so would require knowing to a fraction of , which nobody can build. So rotate instead.
Before rotation, with arm 1 along the wind and arm 2 across it, the time difference between the beams is, from (2.1.23) and (2.1.25),
Turn the table through . Arm 1 is now across the wind and arm 2 along it, and neither length has changed. So
Subtract. The terms regroup by denominator rather than by arm, and the arms enter only through their sum:
The unknown mismatch has vanished identically. Only survives, and that is a quantity you can measure with a ruler. Set so that , and compare with (2.1.27). The change on rotation is exactly .
Now turn that into something visible. A time difference between the two beams corresponds to an optical path difference , and one whole fringe of movement corresponds to a path difference of one wavelength . So the number of fringes the pattern shifts as the apparatus turns is
Clean, and made of measurable things: an arm length, a wavelength, and the square of the Earth's speed in units of .
5.6 · The numbers, 1887
Michelson and Morley floated a sandstone slab on a bath of mercury so that it could be turned smoothly without flexing, and they folded each light path back and forth between multiple mirrors to lengthen it. ⚑ The effective one-way path was , the source was a sodium lamp at , and the assumed velocity was the Earth's orbital speed, . Then
Now feed that value of into (2.1.34), along with the arm length and the wavelength above:
(The exact expression (2.1.34) gives , and the expansion also gives . At the correction term of (2.1.29) is one part in , so the two agree to six figures, as they should.)
About four tenths of a fringe. That is not enormous. But Michelson's interferometer could see a shift of about a hundredth of a fringe, so the predicted effect was roughly forty times the detection threshold.
5.7 · The interactive
The lower panel plots the predicted fringe shift against wind speed on log–log axes, where the scaling of (2.1.34) shows up as a straight line of slope exactly — double the wind speed, quadruple the shift. The shaded band is the 1887 experimental upper bound of about fringes; the dashed vertical line is the Earth's orbital speed of , which is not adjustable by anybody.
Drag the wind speed down and watch where the curve crosses into the band. At the 1887 apparatus () you have to get below about — a sixth of the Earth's orbital speed — before the prediction hides under the detection threshold. At the actual orbital speed the prediction sits a factor of above it. This is not a marginal call. Push the arm length up and it gets worse for the ether, since is linear in ; that is exactly why Michelson folded the path.
Every number is computed from the exact expressions (2.1.23), (2.1.25) and (2.1.34), never from the small- expansion. Because and agree to nine significant figures at realistic , subtracting them directly in floating point would throw away most of the answer, so is evaluated from the algebraically identical but numerically stable rearrangement — which you can check reduces to (2.1.30) as .
5.8 · The result
They found nothing.
⚑ Michelson and Morley reported in 1887 that the observed displacement was certainly less than a twentieth of the predicted fringes, and probably less than a fortieth. That is an upper bound of about fringes, which is to say the noise floor of the instrument. They repeated the measurement at different times of day and in different seasons, to catch the Earth's rotation and its orbital motion pointing the apparatus in every possible direction relative to any conceivable ether frame. Nothing.
A great many experiments in physics disagree with prediction by twenty or thirty per cent and are eventually reconciled by a systematic effect somebody had missed. This is not one of them.
The prediction was fringes. The bound was . That is not a mismatch in a coefficient. It is the total absence of an effect that should have been unmissable, by a factor of nearly forty, in an apparatus specifically designed to have several times the required sensitivity. Nor can it be rescued by supposing the Earth happens to be at rest in the ether at the moment of measurement. Six months later the Earth's orbital velocity has reversed, so its speed relative to any fixed ether frame must then be at least , and the predicted shift is four times larger. Michelson and Morley looked. Nothing.
⚑ Modern versions replace the arms with cryogenic optical resonators and look for a directional dependence of the resonant frequency as the apparatus turns on a rotating table. The current bound on the fractional anisotropy of is below one part in , some seven orders of magnitude tighter than what the 1887 apparatus could have detected. In a hundred and forty years of increasingly ferocious effort, the effect has not appeared.
An experiment that finds nothing is worth exactly as much as the effect it was built to find. Timing light directly was hopeless, so Michelson compared two perpendicular paths by interference, which resolves a hundredth of a wavelength, and turned the apparatus through a right angle so the arms exchanged roles. The unknown mismatch between their lengths cancels out of that comparison identically.
Both arms are slowed, and missing that is the standard error. The arm along the motion loses more time crawling against the flow than it regains running with it, because the slow leg lasts longer and so counts for more. The arm across is slowed too, by a smaller power, since part of the light's fixed speed budget goes on keeping up sideways. The whole observable is the gap between those two powers, and it is second order in the ratio of the Earth's speed to light's, which is why every earlier experiment sensitive only to first order was guaranteed a null result before it was built.
The predicted shift was about four tenths of a fringe, in an instrument that could see a hundredth. Nothing appeared, at any hour or season, and a hundred and forty years of increasingly ferocious repetition has not made it appear. That is not a discrepancy in a coefficient; it is the absence of something that should have been unmissable by a factor of nearly forty.
6 · The rescues, and why they failed
A null result does not by itself dispose of the ether. What it disposes of is the simplest ether: one at rest in some frame, through which the Earth moves freely and which does nothing to the apparatus. Several serious alternatives were proposed. It is worth being specific about them, because they were good physics, and because the last one was very nearly right.
6.1 · Ether drag
Here is the obvious first move. Suppose the Earth drags the local ether along with it, so that near the surface there is no wind at all. The interferometer then sits in still air, so to speak, and sees nothing. Two independent observations rule this out.
Stellar aberration. Point a telescope at a star directly overhead. The Earth is moving sideways at while the light travels down the tube, so the telescope must be tilted slightly into the direction of motion. Otherwise the eyepiece has moved out from under the light by the time it arrives. If the tube has length , the light takes to traverse it and the tube advances in that time, so the required tilt is
With this comes to , and converting to arcseconds by multiplying by ,
⚑ Bradley discovered exactly this in 1728. Every star traces out a small ellipse over the course of a year, of angular radius , and the effect is the single most reliable confirmation that the Earth moves. Now put that together with drag. If the ether were carried along with the Earth, the light would enter that co-moving ether at the top of the atmosphere and thereafter travel in the telescope's own rest frame. There would then be no aberration at all, or at most a fringe effect at the boundary. The aberration is observed, it has exactly the size predicted for undragged ether, and it is the same for all stars regardless of distance. Complete drag is dead.
Fizeau's water tube. In 1851 Fizeau sent light through a tube of water flowing at speed and measured how much the water carried the light along with it. Complete drag predicts that the light speed in the lab is , and no drag predicts . He measured neither. ⚑ The light was dragged by a fraction of the water's speed,
which for water () is . Fizeau confirmed that to within a few per cent, and Michelson's much more precise repetition in 1886 returned . A partial drag coefficient is a very strange thing for a mechanical medium to have. How does a substance know to transmit exactly of its own motion? No ether model ever gave a convincing account of it. (Chapter 2.2 will get (2.1.39) in one line, as the leading term of relativistic velocity addition, with no medium, no drag and no free parameter. It is one of the cleanest confirmations of the whole framework, and it is a nineteenth-century measurement.)
6.2 · The FitzGerald–Lorentz contraction
Now the interesting one. In 1889 FitzGerald, and independently Lorentz in 1892, made a proposal of startling economy: objects moving through the ether contract along their direction of motion by exactly the factor .
Watch what that does. The parallel arm lies along the motion, so its length in the ether frame is not but . The transverse arm is perpendicular and is unaffected. Substitute into (2.1.23):
The two round-trip times are equal. Not approximately equal, and not equal to order . They are identically equal, for every , because is an algebraic identity. The predicted fringe shift is exactly zero to all orders, and Michelson–Morley is explained.
It is impossible to look at (2.1.40) and not feel that something is being got away with. The contraction factor was chosen for no reason except that it makes the answer come out. It was, in FitzGerald's own framing, a hypothesis with one job.
And yet it is not absurd. Lorentz gave it a physical rationale. Suppose the forces holding matter together are electromagnetic in origin, which by 1892 was a reasonable guess. Suppose also that electromagnetic fields are distorted by motion through the ether. Then the equilibrium spacing of the atoms in a solid might well change when the solid is set in motion. On that reading the contraction is a real, dynamical effect on matter, produced by the ether acting on the electrons in the rod. The rod is genuinely shorter. The ether is genuinely there. It just conspires to be undetectable.
The tidy story is that Michelson–Morley disproved the ether and Einstein cleared away the wreckage. Both halves are false, and the second is the more misleading.
(2.1.40) is a proof that a null result cannot kill a sufficiently determined theory. Lorentz's ether, equipped with contraction and with the local time of §6.3, predicts exactly what Michelson and Morley saw. Not approximately. Exactly. By 1904 it reproduced every optical experiment then known. No measurement of the kind being discussed could have distinguished it from what replaced it, because by construction the two make the same predictions.
What an experiment refutes is a specific model, not a concept. It refuted the undragged, undistorting ether. It said nothing about ethers that contract their occupants.
So what did kill it? Not evidence. The ether stopped doing any work. Every quantity it was introduced to explain turned out to be calculable without it: the propagation speed, the failure of the wind to show up, the aberration, the Fizeau coefficient. And every attempt to detect it produced a null result that had to be patched by a new property whose only content was the null result it was patching. A frame of reference that no experiment can identify is not a frame of reference. It is a spare wheel bolted to the theory, and the moment you notice that removing it changes no prediction, it has been refuted in the only sense that matters.
That is a methodological point rather than an empirical one, and it deserves to be made without condescension. The nineteenth century was not being obtuse. They were doing exactly what you should do with a productive hypothesis under pressure, which is to modify it minimally and see whether it survives. It did survive, technically, and it was abandoned anyway. That tells you that "consistent with all the data" is not the only standard a theory is held to. Compare the situation honestly with the present. Chapter 7.9 will ask the same question about ideas in quantum gravity that are consistent with everything and predict nothing, the answer will be the same one, and it will be no more comfortable.
6.3 · Local time
Contraction alone is not quite enough, and Lorentz knew it. He wanted Maxwell's equations to come out with the same form in a moving frame, so that the theory would work at first order in for optical experiments generally and not only for Michelson–Morley. To get that, he found he had to introduce a further substitution, which in his 1895 form reads
He called the local time (Ortszeit), and he was explicit that he regarded it as a mathematical convenience with no physical meaning whatever. The real time was , the ether's time, universal and absolute in Newton's sense. The quantity was an auxiliary variable, a change of integration variable that made the equations tractable, no more significant than substituting in an integral.
By 1904 Lorentz had the full set. Written in modern notation, with :
Those are the Lorentz transformations. Chapter 2.2 derives them from scratch and will not use this equation. It is here only so that the historical point can be made without ambiguity.
Lorentz had (2.1.42) before Einstein. Poincaré had noticed that the transformations form a group, and in 1904 he had stated a principle of relativity covering electrodynamics. The algebra of special relativity was on paper, published, and being used.
So it is worth being exact about what 1905 added, because "Einstein derived the Lorentz transformation" is not it.
For Lorentz, and were auxiliary variables. There was a true frame, the ether's, with true coordinates . Moving rods were really shortened, by a dynamical action of the ether on the electrical forces inside matter. Moving clocks were not slowed at all, since was not a time but a substitution. The transformation described what happens to matter when it moves through the ether. It was a theory of rods and clocks.
For Einstein, and are the coordinates a moving observer actually measures, obtained from two postulates about the symmetry of nature and nothing else. No ether appears in the derivation, because none is needed. There is no true frame, no true time, and nothing dynamical happens to the rod. The rod is shorter in that frame in the same sense that a pencil has a shorter shadow when you turn it, and the reason both observers can consistently say the other's rod is short is that they are disagreeing about which events are simultaneous. The transformation describes the structure of space and time. It is not a theory of rods and clocks. It is a statement about the arena they sit in, and it therefore applies to everything: nuclear forces, particle lifetimes, chemical clocks, and things nobody had thought of, rather than only to electromagnetically bound matter.
That difference is not philosophical decoration. It is what makes the theory predictive outside its domain of construction. Lorentz's contraction is a claim about electromagnetically bound rods, and there is no reason it should apply to a decaying muon. Einstein's is a claim about the coordinates themselves, so it applies to the muon whether or not anything holding it together is electromagnetic. Muons in the atmosphere do live longer by exactly , and the weak interaction that decays them was not discovered for another thirty years.
And it is why the ether then had nothing left to do. Lorentz's ether was necessary to his derivation. Without a medium, there is nothing for a rod to move through and no mechanism to contract it. Einstein's derivation never mentions it, so the ether becomes an object with no role in any calculation, detectable by no experiment, and postulated for no reason. It did not have to be refuted. It had to be noticed to be idle.
Grind box — how close was Lorentz, exactly?
Uncomfortably close, and the gap is instructive. Take Lorentz's own 1904 apparatus and ask what it gets right.
Right: the algebra. (2.1.42) is character for character what Chapter 2.2 will derive. The contraction factor, the local time, the group property, the invariance of Maxwell's equations under the substitution: all correct, and all published before 1905.
Right: the experimental predictions, for everything then measured. Michelson–Morley, Fizeau, aberration, the first-order optical experiments: Lorentz's theory reproduces them all.
Wrong: the status of . Because it is a substitution rather than a time, Lorentz had no account of time dilation as a physical phenomenon, and no reason to expect a moving clock of any construction to run slow. Ask him about an unstable particle and the theory is silent.
Wrong: the asymmetry. In Lorentz's scheme the contraction is real for a rod moving through the ether and not for a rod at rest in it, so the two frames are not equivalent even though no experiment can tell them apart. Einstein's symmetry between the frames is exact, and it is the reason each observer finds the other's rods short. Lorentz's framework cannot make that statement, since on his account one of the two rods is really contracted and the other really is not.
Wrong: the generality. Lorentz's contraction was derived from a hypothesis about electromagnetic forces inside matter. Nothing in it covers gravitational binding, nuclear binding, or the decay rate of a particle. Every one of those obeys the same anyway, which on Lorentz's account would be an extraordinary coincidence and on Einstein's is a tautology.
The lesson generalises past this episode and is worth carrying: having the right equations is not the same as having the right theory. Two accounts can agree on every equation and every measured number and still differ in what they say the symbols mean. That difference shows up the moment you try to apply the theory somewhere it was not built for. Lorentz's electrodynamics and Einstein's relativity were empirically indistinguishable in 1905. They stopped being indistinguishable the moment anyone measured a particle lifetime.
FitzGerald proposed a repair of startling economy, and what makes it instructive is that it works. Suppose anything moving through the medium is shortened along its direction of motion by a particular factor, and the two round-trip times become equal, not approximately but identically, at every speed, because the shortening and the slowing are related by an algebraic identity. The predicted shift is then exactly zero.
It was not absurd either. If the forces holding matter together are electromagnetic, and fields are distorted by motion through the medium, atoms in a rod might settle at a different spacing once it is set moving, so the rod is genuinely shorter and the medium genuinely there. With that and one further substitution, the theory reproduced every optical experiment then known. A null result cannot kill a determined theory; what an experiment refutes is a model rather than a concept.
So the medium was not disproved. It stopped doing any work. Every quantity it had been introduced to explain turned out calculable without it, and every attempt to detect it produced a null result patched by a property whose only content was the null result it patched. A frame no experiment can pick out is not a frame of reference. The verdict is methodological rather than empirical, and deserves stating without condescension, since modifying a productive hypothesis minimally under pressure is what anybody should do.
7 · The two postulates
Everything above is the case for the prosecution. Here is what Einstein proposed instead, stated in the form Chapter 2.2 will use as input.
Postulate 1 (the principle of relativity). The laws of physics take the same form in all inertial frames. No experiment performed within a uniformly moving laboratory can reveal that laboratory's velocity. And this applies to all the laws, mechanical, electromagnetic, and any not yet discovered, rather than merely to the mechanical ones.
Postulate 2 (the invariance of ). Light propagates in vacuum with a definite speed , the same in all inertial frames, independent of the state of motion of the emitting body.
Notice how little that is. There is no ether, no statement about what light is, no model of matter, no dynamics and no mechanism. Two sentences, one of which was universally accepted already. Chapter 2.2 shows that the pair of them determines the transformation between inertial frames essentially uniquely, and everything in Part II after that is consequences.
7.1 · Postulate 1 is not new
It is §1.4's principle with one word changed: all the laws, not just the mechanical ones. That single word is what rules out branch (A) of §4.3, where mechanics obeyed the principle and electromagnetism did not. We already have (2.1.18), in which Maxwell's equations demonstrably change form under (2.1.1). So asserting Postulate 1 forces you into branch (C), and the Galilean transformation must go.
7.2 · Postulate 2 is the radical one, and it is not a definition
This is where the strangeness lives, and it is worth being precise about why.
Consider a lamp on the front of a train moving at , and a lamp on the platform, flashing as they pass. Postulate 2 says both flashes travel at relative to the platform and at relative to the train. Not and . Not and . The same , measured by both observers, for both flashes. (2.1.2), the rule that velocities add, is being denied for light, flatly.
Two objections need answering.
"Isn't this just a convention about how to set clocks?" No. There is a real convention lurking nearby, since synchronising distant clocks requires a rule and any rule involves an assumption about one-way light speed, and Chapter 2.3 will treat that carefully. But Postulate 2 has empirical content independent of any such convention, because the round-trip speed is measurable without synchronising anything: one clock, one mirror, one pulse. The claim that the two-way speed is the same in all frames and independent of the source's motion is a fact about the world, and it can be checked.
"Isn't the source-independence part obvious for a wave?" For a wave in a medium, yes. Sound from a moving whistle travels at through the air regardless of the whistle, since once the disturbance is launched the medium takes over and forgets the source. That is precisely why the ether was attractive: it delivers source-independence for free. Postulate 2 asserts source-independence with no medium to enforce it, which is a much stranger claim. And it also asserts frame-independence, which no medium theory delivers at all.
So Postulate 2 is an empirical claim, and it must be tested. It has been, exhaustively. Worked example 2 gives the sharpest version: photons emitted by particles moving at travel at , not at , to within a part in .
Two postulates. One transformation to find. The question Chapter 2.2 asks is exactly this: what transformation between inertial coordinates could possibly satisfy both? The answer, remarkably, is that there is essentially only one. Linearity is forced by the requirement that free particles move in straight lines in every frame. The constant in it is forced by demanding a speed that comes out the same in all frames. And what drops out is (2.1.42), with and everything that follows from it.
You already know one thing about the answer that Chapter 2.2 will have to reproduce. Whatever the transformation is, it must produce a contraction by along the direction of motion and nothing across it, because that is precisely the combination (2.1.40) needs to explain Michelson–Morley. The experiment has already told you the shape of the answer. What it could not tell you is why.
Notice how little is being asked for at the end of all that. Two sentences: the laws take the same form in every inertial frame, all of them and not merely the mechanical ones; and light travels in vacuum at one definite speed, the same for every observer and independent of the motion of its source.
The first sentence was already common property with one word altered, and that word forbids the branch on which mechanics obeys the principle while optics does not. The second is the radical one, and it is not a disguised convention about setting clocks: a round-trip speed can be measured with one clock, one mirror and one pulse, synchronising nothing, so the claim has empirical content and has been checked to absurdity. Photons from particles moving at nearly the limiting speed arrive at that speed, not at twice it.
What makes the second sentence strange is worth locating exactly. A medium delivers independence from the source for free, since once a disturbance is launched the medium takes over and forgets its origin; asserting that independence with no medium to enforce it is a far stronger claim, and asserting that the speed is also the same for every observer is one no medium delivers. The interferometer had already revealed the shape of the answer, one factor along the motion and none across it. What it could not reveal was why.
8 · Worked examples
Take the Michelson–Morley interferometer as built: effective arm length in each arm, sodium light at , and the ether at rest relative to the Sun so that the wind speed is the Earth's orbital speed . Compute both round-trip times, their difference, and the fringe shift on a rotation. Compare with the experimental bound.
Step 1 · the small parameter. With ,
Already the shape of the problem is clear: we are chasing a part in .
Step 2 · the two round-trip times. The no-wind time is . Then from (2.1.23) and (2.1.25),
They agree to nine significant figures. This is why nobody attempted to time the two beams separately.
Step 3 · the difference. Subtracting nine matching digits is a bad idea numerically, so use the exact rearrangement
which is . Cross-check that against the leading-order estimate (2.1.30): . Agreement to five figures. ✓
Step 4 · turn it into a distance. A time difference is invisible, but a path difference is not. Multiply by :
The two beams arrive about a fifth of a wavelength out of step. This is the entire physical content of the experiment: a fifth of a wavelength, from a -attosecond delay, produced by the Earth moving at a ten-thousandth of the speed of light.
Step 5 · the fringe shift on rotation. Rotating exchanges the arms, so the path difference swings from to and the fringes move by twice that:
Or directly from (2.1.34): . ✓ (The exact expression gives and the expansion gives , and they differ in the ninth decimal.)
Step 6 · compare. ⚑ Observed: less than fringes. Predicted: . The ratio is
To hide the prediction under the bound you would need smaller by a factor of , which means smaller by , which means an ether wind below . That is a sixth of the Earth's orbital speed, and it is a speed the Earth cannot have relative to a fixed frame at all times of year.
A number worth having. The Solar System moves at about around the galactic centre, and at about relative to the frame in which the cosmic microwave background is isotropic. Had the ether been at rest in either, the 1887 apparatus would have shown fringes or fringes respectively, and the pattern would have swept past the crosshairs dozens of times as the table turned. There is no plausible ether frame for which this experiment is a close call.
Suppose light obeys (2.1.2) like everything else. Work out what a laboratory should measure for light from a moving source, and compare with experiment.
Step 1 · the prediction. If light leaves its source at relative to the source, and velocities add by (2.1.2), then a lab in which the source moves at directly towards the detector measures
and for a receding source. This is the "emission theory" or "ballistic" hypothesis, and it is what you get if you take the Galilean transformation seriously and drop the ether. Note that it is not the ether theory. The ether predicts relative to the ether regardless of the source. The two hypotheses are different, and both are testable.
Step 2 · the astronomical test. This is de Sitter's argument from 1913, and it costs nothing but arithmetic. Take a spectroscopic binary star at distance whose components orbit at speed . Light emitted while a star approaches us travels, on this hypothesis, at . Half an orbit later, receding, it travels at . The difference in arrival time after a journey is
Now put in numbers for a typical system, with light years and :
Compare that with an orbital period of, say, days. The scrambling is five times the period. Fast-phase light would overtake slow-phase light, the star would appear at several places in its orbit at once, and the spectroscopic velocity curve would be unrecognisable, not merely distorted but multivalued. Nothing of the kind is seen.
Writing the hypothesis as multiplies the estimate by . Binary orbits are observed to fit Kepler's laws at the per-cent level, and holding the distortion inside that needs . ⚑ De Sitter's published bound was of this order. ⚑ Brecher's 1977 analysis of X-ray binaries, where the timing is far sharper, tightened it to .
Step 3 · the laboratory test, and it is brutal. ⚑ At CERN in 1964, Alväger and collaborators produced neutral pions in a proton beam and used the photons from . The pions were moving at , which is as close to a light-speed source as anything ever built, and the photons' speed was measured by time of flight over a known baseline.
The Galilean prediction for photons emitted forwards is unambiguous:
Very nearly double the speed of light. What was measured was
which is to within parts in . Writing the hypothesis as , the measurement gives . That is consistent with zero, and it excludes by about seven thousand standard deviations.
What to take from this. Both experiments are ⚑ quoted results rather than derivations. The derivation in each case is the prediction they are compared against, which is the arithmetic above. And the logical structure of the pair is worth noting. The ether hypothesis is killed by Michelson–Morley plus aberration plus Fizeau. The emission hypothesis survives all three, and is killed instead by de Sitter and by decay. Between them the two families of experiment eliminate every way of keeping Galilean velocity addition for light. Postulate 2 is what is left standing.
9 · Your turn
Problem 1 — FitzGerald's fix, and what it cannot fix
(a) Redo (2.1.33) assuming the arm along the wind contracts by while the transverse arm is unchanged, and keep . Show that the fringe shift on rotation vanishes identically, for any and any arm lengths. (b) The cancellation in (2.1.40) is exact, so there is no residual at order to find in the equal-arm case. Where, then, does an observable -dependence survive? Compute the total time difference between the two beams for unequal arms with contraction, and show that it depends on . (c) Estimate the size of that effect for and , as the Earth's speed relative to a hypothetical ether changes from to over six months. (d) What extra ingredient removes even this?
Solution
(a) With contraction, the arm currently along the wind has ether-frame length . Its round-trip time is then by (2.1.40), which is the same functional form as the transverse arm. Hence
Now rotate. Arm 2 is along the wind and arm 1 across it, and the result is the same expression, because both arms now carry identical round-trip formulas and only the labels have swapped. The difference is zero identically. The fringe shift on rotation is zero for any , any and any . Michelson–Morley is fully explained, and not merely to leading order.
(b) Look at what survives. The beams still differ in arrival time, by
This is not zero, and it depends on . Rotating the apparatus does not change it, so Michelson–Morley cannot see it. But the Earth's speed relative to any putative ether changes over the year as its orbital velocity swings around. The fringe pattern should therefore drift slowly with the seasons even if the apparatus is never moved. That is precisely the experiment Kennedy and Thorndike performed in 1932, with deliberately unequal arms, and it is the natural complement to Michelson–Morley. MM varies the direction at fixed speed, and KT varies the speed at fixed direction.
(c) The fringe count is , so the change between two speeds is
With and we have , and therefore
That is small, under a hundredth of a fringe. But a stable interferometer watched over months can reach it, and ⚑ Kennedy and Thorndike saw nothing.
(d) Time dilation. The fringe count is a path difference divided by a wavelength, and the wavelength is set by the source, which is moving too. Suppose a moving clock runs slow by . Then light emitted by a moving atom has its period stretched by , and its wavelength stretched by as well. So and
independent of . Contraction alone explains Michelson–Morley. Contraction and time dilation together are what Kennedy–Thorndike needs. Historically this is the cleanest demonstration that length contraction is not a complete story. The two experiments together force both effects, and both drop out of Chapter 2.2 from the two postulates without either being assumed. Add the observed isotropy of from Michelson–Morley, and the three experiments between them pin the Lorentz transformation down uniquely. That is why this trio is sometimes called the experimental basis of special relativity.
Problem 2 — the transformed wave equation, and which term is the culprit
(a) Repeat §4 for a Galilean boost of speed along applied to the full three-dimensional wave equation , and write the result. (b) Identify precisely which term breaks the symmetry, and give two independent arguments that it cannot appear in a law obeying the principle of relativity. (c) Show that if you keep only terms of order you recover the original equation, then say why that is not a rescue.
Solution
(a) The boost affects only and , so and are untouched, and (2.1.17) applies verbatim to the rest. Substituting into :
Bring the term across and collect the coefficient of :
Two symptoms are visible at once. The derivative now carries a different coefficient from the and ones, so space is no longer isotropic. And there is a cross term.
(b) The culprit is . Two arguments.
Parity. Under every other term is unchanged, since each carries an even number of derivatives. The cross term changes sign, because it carries exactly one. So the equation distinguishes from , which is to say it knows which way the wind blows. A law valid in a frame with no preferred direction cannot contain such a term.
Boost-parameter dependence. The coefficient is proportional to , and describes the relationship between two frames rather than anything in the physics. Any law whose coefficients depend on the observer's velocity is by definition not the same law for all observers, which is the negation of Postulate 1.
The modified coefficient is also fatal, but it is a weaker symptom. It is even in , so it survives parity, and you could imagine absorbing it by rescaling . You cannot absorb the cross term that way. It is the one that cannot be rescaled away, and eliminating it is exactly what forces to depend on in Chapter 2.2. That is worth noticing now. The term that breaks the symmetry is the term mixing space and time derivatives, and the repair will be a transformation that mixes space and time coordinates.
(c) Setting in the boxed equation returns exactly. To first order in only the cross term survives, at relative size . So for the equation is nearly unchanged, and that is precisely why the discrepancy went unnoticed for forty years. Every terrestrial source moves at .
But it is not a rescue, for two reasons. First, "approximately invariant" is not a coherent notion for a symmetry principle. Either the laws are the same in all frames or they are not, and if they are not, then there exists in principle a measurement that identifies your frame, no matter how hard that measurement is. Second, and decisively, the experiments of §5 and Worked example 2 are designed to reach the order at which the discrepancy lives. Michelson and Morley reached , and the experiment reached , where the "small" correction is a factor of two. At those sensitivities the effect is not small, and it is absent.
Problem 3 — stellar aberration and the drag hypothesis
(a) Derive the aberration angle for a star directly overhead from scratch, using the moving telescope picture, and evaluate it for the Earth's orbital speed. (b) Explain quantitatively why an observer sees the star's apparent position trace out a small ellipse over a year, and say what the ellipse's shape depends on. (c) Explain precisely why complete ether drag is incompatible with the observation, and why partial drag does not straightforwardly save it either. (d) A rain analogy is often used. State the analogy, then say exactly where it fails once Postulate 2 is adopted.
Solution
(a) Let the telescope tube have length , and let the Earth move sideways at . Light enters the objective and takes to reach the eyepiece. In that time the eyepiece has moved in the direction of motion. For the light to land on the eyepiece, the tube must be tilted forwards by an angle with
The tube length cancels, which is the sign of a real effect rather than an instrumental one. With this gives , and multiplying by turns that into . ⚑ The measured constant of aberration is .
(b) The tilt is always towards the instantaneous direction of the Earth's motion, and that direction rotates through in a year as the Earth goes round its orbit. So the apparent position of any star swings around a closed curve of angular radius .
What shape that curve takes depends on where the star sits. A star at the pole of the ecliptic sees the Earth's velocity vector sweep out a full circle in the plane perpendicular to the line of sight, so its aberration ellipse is a circle of radius . A star in the plane of the ecliptic sees only the component of the Earth's velocity perpendicular to the line of sight, which oscillates back and forth, so its ellipse degenerates to a line of half-length . In between, the semi-minor axis is , with the star's ecliptic latitude.
That pattern is the signature which identifies aberration: the same semi-major axis for every star, and a semi-minor axis depending only on ecliptic latitude and not at all on distance. It is also what distinguishes aberration from parallax, whose size does depend on distance and whose phase is shifted by a quarter of a year.
(c) The derivation in (a) assumed the light travels in a straight line at in a frame in which the telescope is moving. Suppose instead that the ether were completely dragged along by the Earth. Then within the dragged region light would propagate in the telescope's own rest frame, the eyepiece would not move relative to the medium during the transit, and there would be no tilt. You would see no annual aberration at all. It is observed, at full strength, so complete drag is excluded.
Partial drag does not rescue it either, and the reason is instructive. If the drag is partial, the aberration should depend on how much dragged medium the light passes through. It should differ for observations made through a long column of air against a short one. Better still, it should change if you fill the telescope tube with water, since water has a Fresnel drag coefficient of by (2.1.39) and would drag the light substantially. ⚑ Airy performed exactly this experiment in 1871 with a water-filled telescope, and found the aberration unchanged, at the same . A drag model has to explain why filling the instrument with a strongly dragging medium changes nothing, and the honest answer within ether theory is a conspiracy of cancellations. (Relativity gets Airy's null result immediately. Aberration is a property of the transformation between the source's frame and the observer's, and it has nothing to do with what the light passes through on the way in.)
(d) The analogy. Running through vertically falling rain, you must tilt your umbrella forwards, and the faster you run the more you tilt. The tilt angle satisfies , which is (2.1.37) with the rain speed in place of .
Where it fails. The rain calculation is a Galilean velocity addition. In your frame the raindrops have velocity , and the tilt is the direction of that resultant. It also predicts that the drops arrive faster in your frame, at . For light, Postulate 2 forbids that. The photons arrive at , not at . So the correct relativistic aberration formula cannot be the vector-addition one, and it is not. Chapter 2.2 derives , which agrees with (2.1.37) to first order in and differs at second order.
Since for the Earth, that correction enters at relative order and reaches at most , which is arcseconds. That is far below the precision of nineteenth-century astrometry. Hence the naive derivation gave the right answer for two centuries, and hence aberration could not have settled the question by itself. It is measurable now, and the relativistic formula is the one that matches.
Problem 4 — how fast would you have to go?
(a) Find the wind speed that would produce a fringe shift of exactly in the 1887 apparatus (, ). Do it first from the leading-order formula, then check against the exact expression. (b) Comment on whether that speed is achievable or plausible. (c) Suppose instead you fix and ask how long the arms would have to be for a one-fringe shift. Is that achievable? (d) What does the comparison tell you about why the experiment was done in 1887 and not in 1830?
Solution
(a) From (2.1.34) at leading order, , so requires
Solving the exact expression (2.1.34) numerically gives and , the same to six figures, since the correction is .
(b) is only times the Earth's orbital speed. It is laughably out of reach for a laboratory, since no apparatus has ever been moved at anything close to it. But it is utterly ordinary as an astronomical speed. That is the whole point of the experiment's design. It does not need you to move the apparatus, because the Earth is already moving. And it means the null result is not marginal. For the ether to hide, its rest frame would have to lie within about of the Earth at all times of year. That is impossible. The Earth's own velocity changes by between January and July, so any fixed frame is at least away from the Earth for half the year.
(c) Rearranging for at with :
That is under thirty metres of optical path in each arm, held rigid to a fraction of a wavelength and rotatable. Michelson achieved by folding the beam back and forth eight times across a stone. Tripling that is difficult but not absurd, and later workers did reach tens of metres. So the two routes to sensitivity are not symmetric: you cannot change at all, and you can change by a factor of a few. The experiment lives or dies on . That is why every improvement in this line of work has been an improvement in effective path length, culminating in the modern optical resonators, where the light bounces times and the effective path is kilometres.
(d) Because has in it, and an experiment sensitive to one part in was not buildable earlier. Everything that mattered had to arrive first: sufficiently monochromatic sources, high-quality half-silvered mirrors, mechanical isolation by way of the mercury float, and above all the realisation that interferometry converts a hopeless timing measurement into a feasible displacement measurement. Michelson invented the instrument for this purpose.
There is a general pattern here worth noticing: the second-order effect is where the interesting physics is, and second-order effects only become visible when someone builds a null-comparison instrument. The same story runs through Cavendish's torsion balance, Eötvös's test of the equivalence principle (Chapter 3.1), and LIGO, whose optical layout is a Michelson interferometer with arms looking for a displacement of . The instrument that failed to find the ether became, a century later, the instrument that found gravitational waves.
You have the crisis, in a form you can state in three sentences and defend line by line.
One. Galilean relativity is exact for Newtonian mechanics. The transformation is , . Velocities subtract. Accelerations are unchanged. Forces built out of relative positions and relative velocities are unchanged too, so has literally the same form in every inertial frame. And you have the separation that matters: the principle of relativity (P) and the transformation (G) that implements it are different claims, fused together for three centuries because did not look like an assumption.
Two. Maxwell's equations produce a speed. Take the curl of Faraday, commute the derivatives, substitute Ampère–Maxwell, expand the left side with , and use . Out comes the wave equation of Chapter 0.8, with , assembled from a capacitor and a pair of current-carrying wires. That is why light is an electromagnetic wave.
Three. The two are incompatible, and the proof is three lines of chain rule. Under a Galilean boost the wave equation acquires a cross term and a modified coefficient . It changes form. So one of three things has to give. Either the principle of relativity fails for electromagnetism, or Maxwell is wrong, or the Galilean transformation is wrong. The first option makes a prediction, and Michelson and Morley went looking for it and did not find it, at a predicted fringes against a bound of .
You also have the anatomy of the effect, which matters more than the history. The parallel arm is slowed by and the transverse arm by . The difference between them is second order in , and the whole observable lives in the gap between those two exponents. You know that FitzGerald's ad hoc contraction by cancels the effect exactly. You know that Lorentz had the transformation equations before Einstein did. And you know what 1905 actually contributed, which was not the algebra but the recognition that these are facts about space and time rather than about matter moving through a medium. That is why the ether then had nothing left to do.
Where this gets spent.
- The two postulates → Chapter 2.2, immediately. They are the entire input, and the Lorentz transformation is the unique output. Every strange consequence arrives there as a corollary rather than an assumption: the relativity of simultaneity, time dilation, length contraction.
- The Pythagoras triangle of (2.1.26) → Chapter 2.2's light clock, where the identical triangle gives time dilation and the identical is named .
- , and → Chapter 2.3, where the failure of simultaneity becomes geometry. Events acquire an invariant separation , and the light cone is the set of events (2.1.10) can reach.
- "Both sides transform the same way, so the equality survives" → Chapter 2.4. That criterion, isolated in §1's grind box, is promoted there to the definition of a tensor equation, and it becomes the reason the rest of this book is written in indices.
- The wave equation and the constant → Chapter 2.6, the resolution. Maxwell's equations turn out to be exactly Lorentz invariant, with no modification whatever. They were relativistic all along, forty years before anyone knew what that meant, and it was Newtonian mechanics that had to be changed. That chapter also pays Chapter 1.1's outstanding debt: the field momentum that repairs the third law's failure.
- The whole argument's shape → Chapter 3.1, which runs it again. A principle believed on excellent grounds (relativity) meets a fact believed on excellent grounds (gravitational and inertial mass are equal), the two turn out to be incompatible, and the resolution is again that a piece of assumed structure has to go. This time the casualty is the flatness of spacetime. And Chapter 5.1 runs it a third time for quantum mechanics and relativity, where the resolution is that particle number cannot be conserved and fields become unavoidable.
One sentence to carry: the speed in a wave equation belongs to the frame in which the equation holds, and Maxwell's equations refused to name that frame. Everything in Part II is the consequence of taking that refusal at face value.