Part II · Special Relativity — Chapter 2.2

The Lorentz Transformation, Derived

Two sentences go in. One transformation comes out, and there is no second candidate. Everything else in this chapter is arithmetic.

Where we are

Chapter 2.1 ended holding two postulates and a wreck. The wreck is the Galilean transformation x=xvtx'=x-vt, t=tt'=t. Maxwell's equations change form under it. The ether frame it demands cannot be found. And every attempt to save it died in the laboratory, whether that attempt was a dragged ether, an emission theory, or contraction by fiat. The two postulates are what Einstein put in its place, and Chapter 2.1 deliberately stopped there.

This chapter asks the only question left: what transformation between inertial coordinates could satisfy both postulates? The answer is not "here is a clever guess that works." The answer is that the transformation is forced, once you also write down three structural assumptions that everyone makes silently. Those three are homogeneity, isotropy and reciprocity. The general form leaves one free function of vv, and the second postulate together with reciprocity pins that function down completely. Nobody is choosing anything.

Then we spend it. Simultaneity, time dilation and length contraction all arrive as corollaries, and we derive them in that order for a reason. The second and third are consequences of the first. The standard presentation does them backwards, and that is why so many people never recover. After that we look at how velocities combine, why they seem not to add, and the variable in which they do add. The chapter closes by taking the two famous paradoxes apart and showing that they are the same paradox wearing different clothes.

One warning about what this chapter is not. It is not geometry. We will produce an invariant combination c2t2x2c^{2}t^{2}-x^{2} almost by accident in §2, and then mostly leave it alone. Turning that invariant into a genuine geometry, meaning a spacetime with a notion of length, angle and straightness, is Chapter 2.3's job. That chapter is a better one for having this algebra already in hand.

Tools you'll need  — Chapter 2.1 §7: the two postulates, which are the entire input. Chapter 0.4: linear maps, matrices, matrix multiplication, and what it means for a set of maps to be closed under composition, which is what §3 is, applied. Chapter 0.3: Taylor and binomial expansions, used every time we ask "and what does this look like at ordinary speeds?" Chapter 0.1 §3: the derivative as the coefficient of a linear approximation. The linearity argument in §1 is that idea run in reverse, and §5 differentiates a coordinate transformation.

1 · What the postulates constrain

Let's start by writing down exactly what we are given. Then we write down exactly what we are quietly assuming on top of it, which is the part usually skipped.

1.1 · The input

The two postulates

P1 (principle of relativity). The laws of physics take the same form in all inertial frames. No experiment inside a uniformly moving laboratory can reveal its velocity.

P2 (invariance of cc). Light propagates in vacuum with speed cc, the same in all inertial frames, independent of the motion of the source.

Now fix the geometry once and use it for the whole chapter. There are two inertial frames SS and SS', with parallel axes. SS' moves in the +x+x direction of SS with speed vv, and the origins coincide at the moment both clocks read zero. This is the standard configuration, and it costs nothing, because any other arrangement is this one composed with a rotation and a shift of origin. As throughout Part II, we write the relative speed in units of cc:

β    vc. \beta \;\equiv\; \frac{v}{c}. (2.2.1)

Nothing is assumed yet about the range of β\beta. That β<1\abs{\beta}\lt1 will emerge in §2.3, where a larger value makes the transformation complex, and it will be proved unreachable in §5.

What we want are the four functions that take an event's coordinates (t,x,y,z)(t,x,y,z) in SS to its coordinates (t,x,y,z)(t',x',y',z') in SS'. An event is a point of spacetime: a flash, a collision, a clock tick. Both observers are describing the same one. Nothing physical changes when we relabel. The question is only what the labels are.

1.2 · The three assumptions nobody states

The postulates alone do not determine the transformation. They are constraints on a class of maps, and you have to say which class. Three assumptions do that work, and it is worth being blunt: these are physical assumptions, not logical necessities. Each is testable, each has been tested, and each could in principle have come out otherwise. This is where the real content of the chapter enters, and pretending otherwise is how the derivation acquires its undeserved reputation for being a magic trick.

(H) Homogeneity of space and time

No place is special and no moment is special. Concretely: the rule relating SS-coordinates to SS'-coordinates is the same rule everywhere and everywhen. If you and I are both inertial observers, the dictionary between our coordinates cannot depend on where in the room we set up, or on what year it is.

This forces the transformation to be linear. Here is the argument, and it is better than the hand-wave usually offered.

Let Φ\Phi be the map from SS-coordinates to SS'-coordinates, acting on events. Take any two events PP and QQ and let Δ=QP\Delta = Q-P be their coordinate separation in SS. Homogeneity says the separation the primed observer assigns, Φ(Q)Φ(P)\Phi(Q)-\Phi(P), can depend on Δ\Delta but not on where the pair sits. If it did depend on that, a rod's transformed length would tell you your absolute position, and space would have a marked spot. So there is a single function FF with

Φ(P+Δ)Φ(P)  =  F(Δ)for every event P. \Phi(P+\Delta) - \Phi(P) \;=\; F(\Delta) \qquad \text{for every event } P. (2.2.2)

We want to learn what kind of function FF can be, and the way to find out is to chain two separations together. Go from PP to P+Δ1P+\Delta_{1} to P+Δ1+Δ2P+\Delta_{1}+\Delta_{2}, and add the two pieces:

F(Δ1+Δ2)  =  Φ(P+Δ1+Δ2)Φ(P)  =  [Φ(P+Δ1+Δ2)Φ(P+Δ1)]+[Φ(P+Δ1)Φ(P)]  =  F(Δ2)+F(Δ1). \begin{aligned} F(\Delta_{1}+\Delta_{2}) &\;=\; \Phi(P+\Delta_{1}+\Delta_{2}) - \Phi(P)\\[3pt] &\;=\; \big[\Phi(P+\Delta_{1}+\Delta_{2}) - \Phi(P+\Delta_{1})\big] + \big[\Phi(P+\Delta_{1}) - \Phi(P)\big]\\[3pt] &\;=\; F(\Delta_{2}) + F(\Delta_{1}). \end{aligned} (2.2.3)

That is Cauchy's functional equation, and for a continuous FF its only solutions are the linear ones, F(Δ)=MΔF(\Delta)=M\Delta for a constant matrix MM. Since the origins coincide we also have Φ(0)=0\Phi(0)=0, and therefore Φ(Δ)=F(Δ)=MΔ\Phi(\Delta)=F(\Delta)=M\Delta. The transformation is a constant matrix acting on (ct,x,y,z)(ct,x,y,z). The entries of MM may depend on vv, and they had better, but they may not depend on the event.

Grind box — Cauchy's equation, and the loophole that "straight lines stay straight" leaves open

Cauchy. ⚑ The theorem is: if F:RnRmF:\R^{n}\to\R^{m} satisfies F(a+b)=F(a)+F(b)F(a+b)=F(a)+F(b) and is continuous at even one point, then FF is linear. Here is the sketch in one dimension. Additivity gives F(na)=nF(a)F(na)=nF(a) for integer nn, hence F(a)=F(na/n)=nF(a/n)F(a) = F(n\cdot a/n)=nF(a/n), so F(qa)=qF(a)F(qa)=qF(a) for every rational qq. Continuity then upgrades rational to real. (Without continuity the theorem is false, because there exist monstrous non-measurable additive functions, built with the axiom of choice. Nobody is proposing that spacetime is coordinatised by one.)

The other argument, and why it is not quite enough. The usual alternative runs like this. A free particle moves uniformly in a straight line in SS, by Newton's first law, and by P1 it must do so in SS' too. So the map takes straight worldlines to straight worldlines, and therefore it is linear. The conclusion is right, but the reasoning has a gap. The maps of Rn\R^{n} carrying straight lines to straight lines are not just the affine ones. They are the projective ones,

xμ  =  Aμνxν+bμcνxν+d, x'^{\mu} \;=\; \frac{A^{\mu}{}_{\nu}x^{\nu} + b^{\mu}}{c_{\nu}x^{\nu} + d},

which are perfectly good line-preserving maps with a nonconstant denominator. ⚑ (That every line-preserving bijection is of this form is the fundamental theorem of projective geometry, which we quote.) Such a map is not linear, and it does something revealing. It blows up on the hypersurface cνxν+d=0c_{\nu}x^{\nu}+d=0. That is a preferred set of events, a marked place and time in the universe. Homogeneity is exactly the assumption that kills the denominator, forcing cν=0c_{\nu}=0 and returning us to affine, hence linear once the origins coincide.

This is not an idle technicality. Some theories keep P1 and P2 but drop strict homogeneity, among them "de Sitter relativity" and the Fock–Lorentz transformations. They live precisely in that discarded denominator, and they are constrained by experiment rather than by logic. The assumption is doing work, and now you know which work.

(I) Isotropy of space

No direction is special either. Two consequences follow, and we need both.

First: y=yy'=y and z=zz'=z. Directions perpendicular to the boost are untouched. Suppose instead y=k(β)yy'=k(\beta)\,y for some factor kk. Boosting by v-v afterwards must return us to SS, so k(β)k(β)=1k(\beta)k(-\beta)=1. But isotropy says the transverse factor cannot know the sign of the boost. Reflect the xx-axis, and a +v+v boost becomes a v-v boost while yy is untouched, so k(β)=k(β)k(\beta)=k(-\beta). Put the two together and k2=1k^{2}=1, and since k(0)=1k(0)=1 and kk is continuous, k=1k=1.

Second: tt' and xx' cannot depend on yy or zz. By linearity, t=Dt+Ex+Gy+Hzt'=Dt+Ex+Gy+Hz. Rotate the whole apparatus by 180180^{\circ} about the xx-axis. This sends yyy\to-y, zzz\to-z and leaves the boost configuration exactly as it was, so it cannot change tt'. That forces G=H=0G=H=0. The same argument works for xx'.

Grind box — the transverse argument done physically, with rods and scratches

The algebraic argument above is airtight but feels like bookkeeping. Here is the physical version, which is the one to remember, because it shows that a transverse contraction would produce a flat contradiction rather than merely a strange world.

Two identical rods, each of proper length one metre, both aligned along yy, with their lower ends riding along the xx-axis. One is at rest in SS, the other at rest in SS', and they slide past each other. Fix a sharp scriber at the top of each rod.

Suppose SS finds the moving rod short, k<1k \lt 1. Then as they pass, the SS-rod's scriber passes above the tip of the SS'-rod, and the SS'-rod's scriber scratches the SS-rod's side below its tip. Whether a scratch appears on a rod is a local coincidence of matter with matter. It is a single event, and every observer agrees it happened.

Now look from SS'. By P1, SS''s situation is the mirror of SS's. SS' sees SS moving at v-v, and by isotropy the physics cannot distinguish +v+v from v-v here. So SS' must equally find the other rod short, and predicts the scratches the other way round. Both frames cannot be right about which rod carries a scratch, because that is one event and not two. The only escape is k=1k=1. The tips meet exactly, nobody scratches anybody, and both frames agree.

Note carefully what did not happen: no such contradiction arises along xx. There, "which rod is shorter" is settled by measurements at two separate places, and §4.3 will show that two separate places is precisely where frames stop agreeing about what "at the same time" means. Transverse: one event, no room to disagree. Longitudinal: two events, all the room in the world.

(R) Reciprocity

If SS' moves at +v+v as measured in SS, then SS moves at v-v as measured in SS'. Equivalently: the transformation from SS' back to SS is the same transformation with vvv\to-v.

This sounds like a tautology and it is not. It is the statement that the two frames stand in a symmetric relation. Neither is privileged in a way that would let SS measure SS' receding at one rate while SS' measures SS receding at another. It follows from P1 together with isotropy, and we could derive it. We will instead assume it and flag it, because assuming it makes visible exactly how much is being asked.

⚠ Why this isn't obvious

Notice what has and has not been assumed. We have not assumed t=tt'=t. We have not assumed that lengths are invariant, that simultaneity is absolute, that velocities add, or that there is a universal "now". Every one of those was baked into the Galilean transformation without comment for three hundred years, and every one of them is about to fail.

What we have assumed is that spacetime is uniform and directionless, and that the relation between inertial observers is symmetric. Those three, plus the two postulates, are the complete input. If the answer surprises you, the surprise has to be located in one of those five statements. There is nowhere else for it to hide. That is the point of listing them.

In plain terms 2.2.1

Two postulates cannot determine anything on their own, because a constraint is always a constraint on some class of candidates and nobody has yet said which class. Three further assumptions do that work, and the derivation's reputation for being a conjuring trick comes entirely from leaving them unspoken. Each is a physical claim, each is testable, and each could have come out otherwise.

The first is that no place and no moment is special, so the dictionary between two observers' coordinates is the same wherever and whenever they set up. That forces the dictionary to be linear, by an argument better than the usual one. The familiar route, that free particles travel in straight lines for everybody so straight lines must go to straight lines, leaves a loophole, because the maps preserving straightness include ones carrying a denominator, and a denominator that vanishes somewhere marks out a special place and time in the universe.

The second is that no direction is special, which leaves the two transverse directions alone and keeps the sideways coordinates out of the other two equations. The third is that the relation between the observers is symmetric, each receding from the other at one rate. Those three, with the two postulates, are the complete input; if what comes out is surprising, the surprise is living in one of five statements and has nowhere else to hide.

2 · The derivation, in full

By §1 the transformation is linear, the transverse coordinates are untouched, and nothing depends on yy or zz. So the whole problem lives in the (t,x)(t,x) plane, and it reads

x  =  Ax+Bt,t  =  Dt+Ex,y=y,z=z, x' \;=\; A\,x + B\,t, \qquad t' \;=\; D\,t + E\,x, \qquad y'=y, \qquad z'=z, (2.2.4)

where A,B,D,EA,B,D,E are constants, meaning functions of vv alone. That is four unknowns, and we have three conditions to impose. Three conditions for four unknowns sounds like one short, but it is not: condition (ii) will hand us two equations rather than one, because a light pulse can be sent in either direction along xx and both directions have to work. Let's impose them one at a time.

2.1 · Condition (i): the origin of SS' moves at vv

This is what "SS' moves at vv" means, made operational. The spatial origin of SS' is the locus x=0x'=0, and by construction it travels along x=vtx=vt in SS. We want the relation between AA and BB that this forces, so substitute x=0x'=0 and x=vtx=vt into (2.2.4):

0  =  A(vt)+Btfor all tB=Av. 0 \;=\; A(vt) + Bt \quad\text{for all } t \qquad\Longrightarrow\qquad B = -Av. (2.2.5)

Feed that value of BB back into the first line of (2.2.4), and the spatial part of the transformation is already forced into the familiar shape

x  =  A(xvt). x' \;=\; A\,(x - vt). (2.2.6)

Notice how little that cost. And notice what it did not do. It says nothing whatsoever about tt'. The Galilean transformation is the special case A=1A=1, D=1D=1, E=0E=0. Every departure from Galileo is now compressed into three numbers.

2.2 · Condition (ii): a light pulse goes at cc in both frames

This is Postulate 2, and it is the crucial input. It is the only place where the physics of light enters. At the moment the origins coincide (t=t=0t=t'=0, x=x=0x=x'=0) a flash is emitted from the common origin. Consider the part of it travelling in the +x+x direction. In SS it obeys x=ctx=ct. In SS', by P2, it must obey x=ctx'=ct', with the same cc and not cvc-v.

Our aim is to turn that demand into equations relating AA, DD and EE. So put x=ctx=ct into (2.2.6) and into t=Dt+Ext'=Dt+Ex, and see what each gives:

x=A(ctvt)=At(cv),t=Dt+E(ct)=t(D+Ec), x' = A\,(ct - vt) = At\,(c-v), \qquad t' = Dt + E(ct) = t\,(D + Ec), (2.2.7)

Both of those are proportional to tt, so imposing x=ctx'=ct' will divide the time variable out and leave a relation among the three constants alone. Do it:

A(cv)  =  c(D+Ec)  =  cD+Ec2. A\,(c-v) \;=\; c\,(D + Ec) \;=\; cD + Ec^{2}. (2.2.8)

That is one equation for two unknowns, so we need a second, and the flash gives us one for free. Now do the same for the part of it going in the x-x direction, x=ctx=-ct, which must obey x=ctx'=-ct':

A(c+v)  =  c(DEc)A(c+v)  =  cDEc2. -A\,(c+v) \;=\; -c\,(D - Ec) \qquad\Longrightarrow\qquad A\,(c+v) \;=\; cD - Ec^{2}. (2.2.9)

Two linear equations, two unknowns DD and EE in terms of AA. To isolate DD we want the Ec2Ec^{2} terms gone, so add (2.2.8) and (2.2.9). They cancel, and

2Ac  =  2cDD  =  A. 2Ac \;=\; 2cD \qquad\Longrightarrow\qquad D \;=\; A. (2.2.10)

That is one of the two. For the other we want the cDcD terms gone instead, so subtract the same pair of equations rather than adding them:

2Av  =  2Ec2E  =  Avc2. -2Av \;=\; 2Ec^{2} \qquad\Longrightarrow\qquad E \;=\; -\frac{Av}{c^{2}}. (2.2.11)

Three of the four unknowns are gone, and everything is now expressed through AA:

x  =  A(xvt),t  =  A(tvxc2). x' \;=\; A\left(x - vt\right), \qquad t' \;=\; A\left(t - \frac{vx}{c^{2}}\right). (2.2.12)

Let's pause on (2.2.12), because two things have already happened that no amount of staring at Galileo would have suggested.

The time equation has picked up a position term. tt' depends on xx. Two events at the same tt but different xx have different tt'. Simultaneity has already broken, before we know what AA is, and it broke because of the second postulate and nothing else. Everything in §4 is contained in the term Avx/c2-Avx/c^{2}.

The same coefficient AA appears in both. That was not imposed. It fell out of adding the two light conditions. The transformation is symmetric between ctct and xx, and the symmetry is easier to see if we measure time in metres, which means using x0=ctx^{0}=ct as the coordinate rather than tt, as Part II does throughout. So multiply the second equation by cc:

ct  =  A(ctβx),x  =  A(xβct). ct' \;=\; A\,(ct - \beta x), \qquad x' \;=\; A\,(x - \beta\,ct). (2.2.13)

Now ctct and xx enter on exactly the same footing. That symmetry is the first hint that time is about to become a coordinate like the others, which is the whole thesis of Chapter 2.3.

2.3 · Condition (iii): reciprocity fixes AA

One unknown left. Reciprocity says the inverse of the boost by +v+v is the boost by v-v. So apply the boost with parameter β\beta and then the boost with parameter β-\beta, and demand that we get back what we started with. Writing A+=A(β)A_{+}=A(\beta) and A=A(β)A_{-}=A(-\beta):

x  =  A(x+βct)  =  A[A+(xβct)+βA+(ctβx)]  =  AA+[xβct+βctβ2x]  =  AA+(1β2)x. \begin{aligned} x'' &\;=\; A_{-}\big(x' + \beta\,ct'\big)\\[3pt] &\;=\; A_{-}\Big[A_{+}(x-\beta ct) + \beta A_{+}(ct-\beta x)\Big]\\[3pt] &\;=\; A_{-}A_{+}\Big[x - \beta ct + \beta ct - \beta^{2}x\Big] \;=\; A_{-}A_{+}\big(1-\beta^{2}\big)\,x. \end{aligned} (2.2.14)

We started with xx and we must end with xx, whatever xx is, so the whole prefactor in front of it has to be one:

A(β)A(β)(1β2)  =  1. A(\beta)\,A(-\beta)\,\big(1-\beta^{2}\big) \;=\; 1. (2.2.15)

(The same manipulation on ctct'' gives the same condition, and the grind box does it.) That is one equation in two unknowns, A(β)A(\beta) and A(β)A(-\beta), unless AA happens to be an even function of β\beta. Isotropy delivers exactly that, and here is how. Reflect the xx-axis by setting X=xX=-x and X=xX'=-x'. Then (2.2.13) becomes

ct=A(β)(ct(β)X),X=A(β)(X(β)ct), ct' = A(\beta)\big(ct - (-\beta)X\big), \qquad X' = A(\beta)\big(X - (-\beta)\,ct\big), (2.2.16)

which is a boost with parameter β-\beta and coefficient A(β)A(\beta). Space has no handedness, so the reflected description is as good as the original, and the coefficient for parameter β-\beta must be what it always is: A(β)=A(β)A(-\beta) = A(\beta). Feed that into (2.2.15):

A2(1β2)=1A=±11β2. A^{2}\big(1-\beta^{2}\big) = 1 \qquad\Longrightarrow\qquad A = \pm\frac{1}{\sqrt{1-\beta^{2}}}. (2.2.17)

Take the positive root, and here is why we may. A(0)=+1A(0)=+1, because no boost is no change. AA is continuous. And AA never vanishes, so it cannot cross zero to reach the negative branch. The surviving root is the one quantity this whole chapter turns on, so give it a name:

γ    11β2  =  (1β2)1/2. \gamma \;\equiv\; \frac{1}{\sqrt{1-\beta^{2}}} \;=\; \big(1-\beta^{2}\big)^{-1/2}. (2.2.18)

Substitute that name back into the two lines of (2.2.13), and nothing is left undetermined. The whole quarry of the chapter is now on the page.

  ct=γ(ctβx),x=γ(xβct),y=y,z=z.   \boxed{\; ct' = \gamma\big(ct - \beta x\big), \qquad x' = \gamma\big(x - \beta\,ct\big), \qquad y'=y, \qquad z'=z. \;} (2.2.19)

That is the Lorentz transformation for a boost along xx, in the form this book uses everywhere. Three things are worth noting straight away. First, γ1\gamma\geq1 always. Second, γ\gamma\to\infty as β1\beta\to1. Third, β>1\beta \gt 1 would make γ\gamma imaginary, which is a first cheap indication that cc is a barrier. §5 makes that indication respectable.

Grind box — the inverse, checked both ways, and the time equation of the reciprocity step

The time half of (2.2.14). Same manoeuvre:

ct=A(ct+βx)=AA+[(ctβx)+β(xβct)]=AA+(1β2)ct, \begin{aligned} ct'' &= A_{-}\big(ct' + \beta x'\big) = A_{-}A_{+}\Big[(ct-\beta x) + \beta(x - \beta ct)\Big]\\[3pt] &= A_{-}A_{+}\big(1-\beta^{2}\big)\,ct, \end{aligned}

which gives condition (2.2.15) again. Consistent, as it must be: two equations, one condition.

Inverting (2.2.19) directly. This time do not appeal to reciprocity at all. Just solve. Multiply the first equation by γ\gamma and the second by γβ\gamma\beta, and add:

γct+γβx=γ2[(ctβx)+β(xβct)]=γ2(1β2)ct=ct, \gamma\,ct' + \gamma\beta\,x' = \gamma^{2}\Big[(ct-\beta x) + \beta(x-\beta ct)\Big] = \gamma^{2}\big(1-\beta^{2}\big)\,ct = ct,

since γ2(1β2)=1\gamma^{2}(1-\beta^{2})=1 is (2.2.18) squared and rearranged. The same trick with the roles swapped gives xx. Hence

ct=γ(ct+βx),x=γ(x+βct), ct = \gamma\big(ct' + \beta x'\big), \qquad x = \gamma\big(x' + \beta\,ct'\big),

which is (2.2.19) with ββ\beta\to-\beta and primes swapped. So reciprocity, which we assumed, is also a theorem of the result. That is a genuine consistency check rather than a circularity, because the algebra above never used it.

Numbers worth memorising. γ\gamma is stubbornly close to 11 until β\beta is large:

β\beta0.010.010.10.10.50.50.80.80.90.90.990.990.9990.999
γ\gamma1.000051.000051.0051.0051.1551.1551.6671.6672.2942.2947.0897.08922.3722.37

Half the speed of light buys you a 15% effect. This is exactly why nineteenth-century physics could be built without noticing. At β=104\beta=10^{-4}, the Earth's orbital speed, γ1=5×109\gamma-1 = 5\times10^{-9}.

2.4 · A consistency check we did not pay for

Postulate 2 was imposed only along the xx-axis, one pulse forward and one back. But light goes in all directions, and a flash from the origin makes an expanding sphere, x2+y2+z2=c2t2x^{2}+y^{2}+z^{2}=c^{2}t^{2} in SS. If the derivation is any good, that sphere had better still be a sphere in SS', expanding at the same cc. We never asked for that. Let's check it.

The quantity to watch is the one that measures the radius of the sphere against the time it has had to expand. So compute the combination c2t2x2c^{2}t'^{2}-x'^{2} from (2.2.19):

c2t2x2=γ2(ctβx)2γ2(xβct)2=γ2[c2t22βctx+β2x2x2+2βctxβ2c2t2]=γ2(1β2)(c2t2x2)  =  c2t2x2. \begin{aligned} c^{2}t'^{2}-x'^{2} &= \gamma^{2}\big(ct-\beta x\big)^{2} - \gamma^{2}\big(x-\beta ct\big)^{2}\\[3pt] &= \gamma^{2}\Big[c^{2}t^{2} - 2\beta\,ctx + \beta^{2}x^{2} - x^{2} + 2\beta\,ctx - \beta^{2}c^{2}t^{2}\Big]\\[3pt] &= \gamma^{2}\big(1-\beta^{2}\big)\big(c^{2}t^{2}-x^{2}\big) \;=\; c^{2}t^{2}-x^{2}. \end{aligned} (2.2.20)

The cross terms cancel identically and γ2(1β2)=1\gamma^{2}(1-\beta^{2})=1 finishes it. Since y=yy'=y and z=zz'=z, those two coordinates come along untouched, and we have proved something considerably stronger than we asked for:

c2t2x2y2z2  =  c2t2x2y2z2. c^{2}t'^{2} - x'^{2} - y'^{2} - z'^{2} \;=\; c^{2}t^{2} - x^{2} - y^{2} - z^{2}. (2.2.21)

So the set c2t2x2y2z2=0c^{2}t^{2}-x^{2}-y^{2}-z^{2}=0, which is the expanding light sphere, maps to itself. Light emitted from the common origin is a sphere of radius ctct in SS and a sphere of radius ctct' in SS', each centred on that observer's own origin, and the two origins are moving apart. Both are right. If that sounds impossible, hold the thought until §4.1. The two observers disagree about which points of the sphere are "simultaneous", and that disagreement is exactly large enough.

The quantity that survived

(2.2.21) was not an input. We demanded invariance of the light speed, which is a statement about one particular family of events, the ones with c2t2x2y2z2=0c^{2}t^{2}-x^{2}-y^{2}-z^{2}=0. What came back is invariance of the whole quantity, for every pair of events, whatever its value. That is an enormous strengthening, and it is the single most important line in this chapter.

The combination Δs2=c2Δt2Δx2Δy2Δz2\Delta s^{2}=c^{2}\Delta t^{2}-\Delta x^{2}-\Delta y^{2}-\Delta z^{2} is called the interval. Everybody computes the same number for it. It plays the role that Δx2+Δy2+Δz2\Delta x^{2}+\Delta y^{2}+\Delta z^{2} plays for rotations in ordinary space, the thing that does not care how you turned your head. Chapter 2.3 takes that analogy seriously and builds a geometry out of it, and Chapter 2.4 turns the invariance into the algebraic condition ημνΛμρΛνσ=ηρσ\eta_{\mu\nu}\Lambda^{\mu}{}_{\rho}\Lambda^{\nu}{}_{\sigma}=\eta_{\rho\sigma} that defines the Lorentz group. Everything downstream of here rests on one minus sign.

2.5 · The low-speed limit: Galileo was not wrong, he was first-order

A new theory that contradicted the old one everywhere would be a theory about a different universe. The correct relationship is subtler. Relativity must reduce to Galileo where Galileo was tested. So let's expand and see.

What we need first is γ\gamma for small β\beta, and the binomial series (Chapter 0.3) supplies it:

γ=(1β2)1/2=1+12β2+38β4+ \gamma = \big(1-\beta^{2}\big)^{-1/2} = 1 + \tfrac12\beta^{2} + \tfrac38\beta^{4} + \cdots (2.2.22)

so γ=1+O(β2)\gamma = 1 + O(\beta^{2}). Substituting that into (2.2.19) and keeping terms through first order in β\beta:

x=xvt+O(β2),t=tvxc2+O(β2). x' = x - vt + O(\beta^{2}), \qquad t' = t - \frac{vx}{c^{2}} + O(\beta^{2}). (2.2.23)

Now let cc\to\infty at fixed vv, xx and tt, which is what "ordinary speeds" means. The second term of tt' vanishes too, and what is left is x=xvtx'=x-vt and t=tt'=t. That is the Galilean transformation. Not approximately, and not in spirit. Exactly, as a limit.

That is the honest statement of the relationship, and it is worth saying what it means. Galilean relativity is not a mistake that got corrected. It is the leading term of an expansion, and the reason it looked like the whole truth for three centuries is that nobody could measure the next term. The same pattern will recur. Newtonian gravity is the leading term of general relativity (Chapter 3.6), and classical mechanics is a limit of quantum mechanics (Chapter 4.10). Old theories almost never die. They get demoted to leading order.

But look closely at (2.2.23), because the two corrections do not enter at the same order, and that asymmetry explains the entire history of the subject:

  • Lengths and time rates are wrong by a factor γ\gamma, that is by 12β2\tfrac12\beta^{2}, which is second order. At β=104\beta=10^{-4} that is 5×1095\times10^{-9}, a part in two hundred million.
  • Simultaneity is wrong by vx/c2-vx/c^{2}, which is first order in β\beta, but multiplied by a distance. At v=30 kms1v=30\ \mathrm{km\,s^{-1}} and x=1 mx=1\ \mathrm{m} this is 3.3×1013 s3.3\times10^{-13}\ \mathrm{s}, which is small. But scale xx up to a satellite orbit, 2×107 m2\times10^{7}\ \mathrm{m}, and it becomes 6.7 μs6.7\ \mathrm{\mu s}. That is 2.0 km2.0\ \mathrm{km} of light travel, and it would wreck a navigation system in an afternoon.

So the largest relativistic effect at low speed is the one about clocks at different places, and it is the one nobody notices, because noticing it requires two clocks far apart and a way to compare them. That is exactly the experiment nineteenth-century physics could not do and twenty-first-century engineering does continuously.

In plain terms 2.2.2

Everything in the derivation is bookkeeping except one step, and that step is the only place where any physics of light enters. Demanding that the moving origin move at the stated rate fixes the spatial equation up to one unknown factor, and costs almost nothing, since the old transformation is the case where that factor is one.

The light condition then goes in, once for a pulse each way along the axis. Adding the two resulting equations forces the same unknown factor into the time equation as sits in the spatial one; subtracting them forces the time equation to acquire a term proportional to position. That term is the end of simultaneity, and it arrives before the remaining factor is known. Symmetry between the observers pins that factor down, and no second candidate appears.

The bonus is larger than the result. Invariance was demanded only of the events light connects, and what comes back is a combination of the time and space separations that is the same for every pair, whatever its value. Quantities invariant for reasons nobody demanded are worth building a subject on. The old transformation survives as the leading term of an expansion, and its two corrections do not enter at the same order: lengths and rates are wrong at second order, while simultaneity is wrong at first order multiplied by a distance, so the largest low-speed effect concerns distant clocks.

3 · Matrix form and the group property

Chapter 0.4 taught you to see a linear map as a matrix, and composition as matrix multiplication. Let's do that here. Write the four coordinates as xμ=(ct,x,y,z)x^{\mu}=(ct,x,y,z), with time measured in metres so that all four carry the same units. Then (2.2.19) is

(ctxyz)=Λ(ctxyz),Λμν=(γγβ00γβγ0000100001). \begin{pmatrix} ct'\\ x'\\ y'\\ z'\end{pmatrix} = \Lambda\begin{pmatrix} ct\\ x\\ y\\ z\end{pmatrix}, \qquad \Lambda^{\mu}{}_{\nu} = \begin{pmatrix} \gamma & -\gamma\beta & 0 & 0\\ -\gamma\beta & \gamma & 0 & 0\\ 0 & 0 & 1 & 0\\ 0 & 0 & 0 & 1\end{pmatrix}. (2.2.24)

The index placement Λμν\Lambda^{\mu}{}_{\nu} puts the upper index on the row and the lower one on the column, and it is not decoration. Chapter 2.4 explains exactly what promise it encodes, and derives this same matrix as a Jacobian xμ/xν\partial x'^{\mu}/\partial x^{\nu}. For now it is a matrix.

Two features are worth naming immediately. It is symmetric, which we will use in Problem 3. And its determinant is γ2γ2β2=γ2(1β2)=1\gamma^{2}-\gamma^{2}\beta^{2}=\gamma^{2}(1-\beta^{2})=1, so boosts preserve four-dimensional volume in spacetime. That fact matters in statistical mechanics and in Chapter 5.8.

3.1 · Two boosts make a boost

Now the question that decides whether we have a theory or a formula. Boost from SS to SS' at β1\beta_{1}, then boost from SS' to SS'' at β2\beta_{2}, along the same axis. Is the composite a boost? To find out we multiply the two matrices, and only the upper-left 2×22\times2 block does anything:

Λ(β2)Λ(β1)=γ2(1β2β21)γ1(1β1β11)=γ1γ2(1+β1β2(β1+β2)(β1+β2)1+β1β2). \begin{aligned} \Lambda(\beta_{2})\Lambda(\beta_{1}) &= \gamma_{2}\begin{pmatrix}1 & -\beta_{2}\\ -\beta_{2} & 1\end{pmatrix} \gamma_{1}\begin{pmatrix}1 & -\beta_{1}\\ -\beta_{1} & 1\end{pmatrix}\\[6pt] &= \gamma_{1}\gamma_{2}\begin{pmatrix} 1+\beta_{1}\beta_{2} & -(\beta_{1}+\beta_{2})\\ -(\beta_{1}+\beta_{2}) & 1+\beta_{1}\beta_{2}\end{pmatrix}. \end{aligned} (2.2.25)

We want that product to look like a single boost matrix, which has ones on the diagonal. So pull out the common factor 1+β1β21+\beta_{1}\beta_{2}, and give the ratio that survives on the off-diagonal a name of its own:

β3    β1+β21+β1β2. \beta_{3} \;\equiv\; \frac{\beta_{1}+\beta_{2}}{1+\beta_{1}\beta_{2}}. (2.2.26)

Then the product is γ1γ2(1+β1β2)\gamma_{1}\gamma_{2}(1+\beta_{1}\beta_{2}) times the matrix (1β3β31)\left(\begin{smallmatrix}1 & -\beta_{3}\\ -\beta_{3} & 1\end{smallmatrix}\right), and it is a boost provided that prefactor equals γ3=(1β32)1/2\gamma_{3}=(1-\beta_{3}^{2})^{-1/2}. It does equal that, and the grind box does the algebra, so

Λ(β2)Λ(β1)  =  Λ ⁣(β1+β21+β1β2). \Lambda(\beta_{2})\,\Lambda(\beta_{1}) \;=\; \Lambda\!\left(\frac{\beta_{1}+\beta_{2}}{1+\beta_{1}\beta_{2}}\right). (2.2.27)
Grind box — the prefactor really is γ3\gamma_{3}, and one identity that keeps reappearing

Compute 1β321-\beta_{3}^{2} from (2.2.26):

1β32=(1+β1β2)2(β1+β2)2(1+β1β2)2. 1-\beta_{3}^{2} = \frac{(1+\beta_{1}\beta_{2})^{2}-(\beta_{1}+\beta_{2})^{2}}{(1+\beta_{1}\beta_{2})^{2}}.

Expand the numerator and watch the cross terms die:

(1+ab)2(a+b)2=1+2ab+a2b2a22abb2=1a2b2+a2b2  =  (1a2)(1b2). \begin{aligned} (1+ab)^{2}-(a+b)^{2} &= 1 + 2ab + a^{2}b^{2} - a^{2} - 2ab - b^{2}\\[3pt] &= 1 - a^{2} - b^{2} + a^{2}b^{2} \;=\; \big(1-a^{2}\big)\big(1-b^{2}\big). \end{aligned}

So

1β32=1β121β221+β1β2=1γ1γ2(1+β1β2), \sqrt{1-\beta_{3}^{2}} = \frac{\sqrt{1-\beta_{1}^{2}}\,\sqrt{1-\beta_{2}^{2}}}{1+\beta_{1}\beta_{2}} = \frac{1}{\gamma_{1}\gamma_{2}\big(1+\beta_{1}\beta_{2}\big)},

hence γ3=γ1γ2(1+β1β2)\gamma_{3}=\gamma_{1}\gamma_{2}(1+\beta_{1}\beta_{2}), which is the prefactor. ✓

Keep that identity. The factorisation (1+ab)2(a+b)2=(1a2)(1b2)(1+ab)^{2}-(a+b)^{2}=(1-a^{2})(1-b^{2}), along with its sibling (1ab)2(ab)2=(1a2)(1b2)(1-ab)^{2}-(a-b)^{2}=(1-a^{2})(1-b^{2}), will do the work again in §5, where it proves that you cannot reach cc by combining sub-cc speeds. It appears once more in §6, where it turns out to be the hyperbolic addition formula in disguise. Three appearances of one factorisation is usually a sign that something structural is going on, and it is.

The other group axioms. Λ(0)=I\Lambda(0)=I is the identity. Λ(β)\Lambda(-\beta) is the inverse, since (2.2.26) with β2=β1\beta_{2}=-\beta_{1} gives β3=0\beta_{3}=0. Associativity is free, because matrix multiplication is associative. So the boosts along a fixed axis form a one-parameter group, and §6 finds the parameter that makes it look like one.

This matters more than it looks. Closure is not automatic. It is a statement that the set of transformations is a self-consistent family, that "the frames of physics" is a coherent notion, and that no chain of boosts can ever take you outside the theory. Had the composition of two boosts failed to be a boost, the whole scheme would have been incoherent.

But look hard at (2.2.26). The composed parameter is not β1+β2\beta_{1}+\beta_{2}. Boost by 0.5c0.5c and then by 0.5c0.5c again and you get

β3=0.5+0.51+0.25=11.25=0.8, \beta_{3} = \frac{0.5+0.5}{1+0.25} = \frac{1}{1.25} = 0.8, (2.2.28)

not 11. Two half-light-speed boosts give eight tenths, and no finite chain of sub-cc boosts will ever reach cc. §5 proves that properly and §6 explains it. Velocity is the wrong variable to be adding.

⚠ One caveat, flagged now and paid in Problem 3

Everything above was for boosts along one fixed axis. Boosts along different axes do not compose into a pure boost. The product picks up a spatial rotation, called the Wigner rotation. Problem 3 derives that the effect must exist, in two lines, from the fact that (2.2.24) is a symmetric matrix. The physical consequence is called Thomas precession. It is a measurable contribution to atomic fine structure, and it is a genuinely relativistic effect with no Newtonian ancestor at all.

In plain terms 2.2.3

Whether a family of transformations closes on itself decides whether you have a theory or a formula. Compose two boosts along one axis and the result is another boost, with a parameter emphatically not the sum of the two, since half the speed of light followed by half the speed of light delivers eight tenths. Closure is neither automatic nor decorative, because had it failed there would be no coherent notion of the frames of physics, and some chain of changes of viewpoint would have carried you outside the theory.

Two smaller facts fall out along the way. The transformation has determinant one, so it preserves four-dimensional volume, which is the determinant doing the job the toolkit gave it. And the whole apparatus is now a matrix acting on four numbers, so composing changes of frame has become multiplying matrices and nothing more.

One caveat is worth flagging because it has no ancestor in older physics. Boosts along a single axis behave; boosts along different axes do not compose into a boost at all, and the product leaves a spatial rotation behind. That rotation is measurable, contributing to the fine structure of atomic spectra, and nothing in Newtonian mechanics predicts anything of the kind. The composition rule also says plainly that velocity is the wrong quantity to be adding.

4 · The three consequences

Now we cash it in. Three results, and every one of them is a two-line substitution into (2.2.19). The difficulty in this material is never the algebra. It is knowing which two events you are transforming. So each subsection begins by stating the operational question before any symbol moves: who measures what, with which apparatus, and using which pair of events.

4.1 · Relativity of simultaneity

The operational question. Two firecrackers explode. An observer in SS, using clocks that have been synchronised throughout SS, records both explosions at the same reading. What do the clocks of SS' record, given that they have been synchronised throughout SS'?

Take two events with Δt=0\Delta t=0 and Δx0\Delta x\neq0 in SS. We want their time separation in SS', so apply (2.2.19) to the separation itself. That is legitimate because the transformation is linear, which is what §1 bought us:

cΔt=γ(cΔtβΔx)=γβΔxΔt=γβΔxc    0. c\,\Delta t' = \gamma\big(c\,\Delta t - \beta\,\Delta x\big) = -\gamma\beta\,\Delta x \qquad\Longrightarrow\qquad \Delta t' = -\frac{\gamma\beta\,\Delta x}{c} \;\neq\; 0. (2.2.29)

Events simultaneous in SS are not simultaneous in SS', unless they happen at the same place. The sign matters and is worth reading off. Suppose Δx>0\Delta x \gt 0, so that the second event is further along the direction of SS''s motion. Then Δt<0\Delta t' \lt 0, so in SS' the leading event happened first. The frame that is moving forward sees the forward event as earlier.

A concrete case

A train of proper length 300 m300\ \mathrm{m} passes a platform at β=0.6\beta=0.6, so γ=1.25\gamma=1.25. Lightning strikes both ends of the train, simultaneously as judged on the platform.

In the platform frame SS the train is contracted (§4.3) to 300/1.25=240 m300/1.25=240\ \mathrm{m}, so the two strike events have Δx=240 m\Delta x=240\ \mathrm{m} and Δt=0\Delta t=0. Putting those numbers into (2.2.29) gives, on the train,

Δt=γβΔxc=1.25×0.6×2402.998×108 s=6.00×107 s=600 ns, \Delta t' = -\frac{\gamma\beta\Delta x}{c} = -\frac{1.25\times0.6\times240}{2.998\times10^{8}}\ \mathrm{s} = -6.00\times10^{-7}\ \mathrm{s} = -600\ \mathrm{ns}, (2.2.30)

and Δx=γΔx=1.25×240=300 m\Delta x' = \gamma\Delta x = 1.25\times240 = 300\ \mathrm{m} ✓, which is the train's own length, as it must be. So on the train, the front is struck 600 ns600\ \mathrm{ns} before the rear. Not "appears to be". It is, in every sense that the train's own clocks can express.

Check it without the transformation, as a sanity test. Both flashes travel at cc in the platform frame and meet at the midpoint of the two strike locations. Meanwhile the train's midpoint is moving toward the front strike, so it runs into the front flash first. Now argue in the train frame. The train observer sits midway between the two scorch marks on the train, both flashes travel to him at cc over equal distances, and the front flash arrives first. Equal distances, equal speeds, different arrival times, and therefore different emission times. The front strike happened first. Two independent routes, one answer. ✓

Why this one goes first

Time dilation and length contraction are not independent phenomena sitting alongside the failure of simultaneity. They are consequences of it. Every apparent paradox in the rest of this chapter, and in the rest of the subject, is resolved by finding the place where somebody assumed two distant events had a frame-independent time order.

Readers who meet dilation and contraction first almost never recover, because they acquire a picture in which clocks and rulers are mysteriously defective while "now" remains universal. It does not. There is no universal "now". There is a family of them, one per state of motion, and they cut through spacetime at different angles. Hold that and nothing that follows can confuse you.

4.2 · Time dilation

The operational question. A single clock rides along with SS' and ticks twice. The elapsed time it reads between the ticks is Δτ\Delta\tau. That quantity is called the clock's proper time, and it is what that particular clock actually displays. What is the elapsed time between those same two ticks according to SS, where the reading must be taken off two different synchronised clocks, one at each tick's location?

The clock is at rest in SS', so the two ticks have Δx=0\Delta x'=0 and Δt=Δτ\Delta t'=\Delta\tau. We want Δt\Delta t, so use the inverse transformation from the grind box in §2.3:

cΔt=γ(cΔt+βΔx)=γcΔτ  Δt=γΔτ   c\,\Delta t = \gamma\big(c\,\Delta t' + \beta\,\Delta x'\big) = \gamma\,c\,\Delta\tau \qquad\Longrightarrow\qquad \boxed{\;\Delta t = \gamma\,\Delta\tau\;} (2.2.31)

Since γ1\gamma\geq1, more time elapses in SS than the moving clock records. A moving clock runs slow, by exactly γ\gamma. Note the asymmetry in the apparatus, which is where all the content is: one clock on one side, two clocks on the other.

The symmetric version, and why it is not a contradiction

Run the identical argument the other way. A clock at rest in SS ticks twice, Δx=0\Delta x=0, Δt=ΔτS\Delta t=\Delta\tau_{S}. Then (2.2.19) gives directly cΔt=γ(cΔτS0)c\Delta t'=\gamma(c\Delta\tau_{S}-0), so Δt=γΔτS\Delta t'=\gamma\Delta\tau_{S}. Each observer finds the other's clock slow, by the same factor. Both statements are true. Both are derived from the same transformation. There is no contradiction, and here is the accounting that shows it.

Set β=0.6\beta=0.6, γ=1.25\gamma=1.25. In SS there are two clocks, C1C_{1} at x=0x=0 and C2C_{2} at x=12x=12 light-seconds, synchronised in SS. A traveller carrying clock CC' flies from C1C_{1} to C2C_{2}.

Told by SS: the trip takes Δt=12/0.6=20 s\Delta t = 12/0.6 = 20\ \mathrm{s}. The traveller's clock, moving, runs slow, and reads 20/1.25=16 s20/1.25 = 16\ \mathrm{s} on arrival. C2C_{2} reads 20 s20\ \mathrm{s}.

Told by SS': the traveller is at rest and the whole SS apparatus streams past at 0.6c0.6c. The distance from C1C_{1} to C2C_{2} is contracted to 12/1.25=9.612/1.25=9.6 light-seconds, so it takes 9.6/0.6=16 s9.6/0.6 = 16\ \mathrm{s}. That matches the traveller's own clock, since it is his own clock. During those 1616 seconds the SS clocks, being the moving ones now, tick only 16/1.25=12.8 s16/1.25 = 12.8\ \mathrm{s}.

Contradiction? C2C_{2} must read 20 s20\ \mathrm{s} when the traveller arrives. That is a local coincidence, two objects at the same place, and no one can disagree about it. Yet SS' says the SS clocks advanced only 12.8 s12.8\ \mathrm{s}. The missing 7.2 s7.2\ \mathrm{s} is exactly the simultaneity offset. At the moment the traveller passes C1C_{1}, which we may call t=0t'=0, the SS clocks that SS' regards as simultaneous satisfy t=βx/ct = \beta x/c. So C2C_{2}, sitting at x=12x=12 light-seconds, already reads

tC2=βxc=0.6×12=7.2 s. t_{C_{2}} = \frac{\beta x}{c} = 0.6\times 12 = 7.2\ \mathrm{s}. (2.2.32)

In the traveller's reckoning, C2C_{2} was never synchronised with C1C_{1} in the first place. It was 7.2 s7.2\ \mathrm{s} ahead before the trip began. Add that head start to what the clock ticked during the trip, and the books balance:

7.2head start  +  12.8ticked during the trip  =  20.0what C2 reads \underbrace{7.2}_{\text{head start}} \;+\; \underbrace{12.8}_{\text{ticked during the trip}} \;=\; \underbrace{20.0}_{\text{what }C_{2}\text{ reads}} \quad\checkmark (2.2.33)

Both frames agree on every local coincidence, meaning on which clock reads what when it is next to which other clock. They disagree only about the synchronisation of distant clocks, which is not an observable but a convention imposed frame by frame. The symmetry of time dilation is consistent precisely because simultaneity fails, and the failure is exactly the right size. It could not have been any other size, because (2.2.29) and (2.2.31) came out of the same matrix.

⚠ Why this isn't obvious

"Each sees the other's clock run slow" sounds like a<ba \lt b and b<ab \lt a. It is not, because the two statements are not comparisons of the same things. Each is a comparison of one clock against a pair of clocks synchronised in the other frame, and the two frames use different pairs, synchronised differently. Strip out the words and write down which events are being compared, and the appearance of contradiction evaporates every single time.

Here is the test to apply whenever you feel the vertigo. Ask whether the disputed claim is about a local coincidence. "Clock CC' reads 1616 while sitting next to C2C_{2}, which reads 2020" is a coincidence, two objects at one place making one event, and every frame agrees on it. "Clock CC' reads 1616 while C2C_{2}, twelve light-seconds away, reads 2020" is not a coincidence. It invokes a simultaneity convention, and frames are entitled to differ. All the paradoxes are manufactured by smuggling the second kind of statement in wearing the clothes of the first.

4.3 · Length contraction

The operational question, and it is the whole difficulty. A rod lies at rest in SS' along the xx'-axis. Its ends are at xLx'_{L} and xRx'_{R}, always, so its length in SS' is unambiguous: L0=xRxLL_{0}=x'_{R}-x'_{L}, called the proper length. Now ask SS for the rod's length. In SS the rod is moving, so "where its ends are" is a function of time, and to get a length you must locate both ends at the same moment. At the same moment according to whom? According to SS, because that is what it means for SS to measure a length.

So a length measurement is a pair of events (locate left end, locate right end) constrained by Δt=0\Delta t=0 in the measuring frame. The moment you write that down, the answer is forced, because the constraint is a simultaneity constraint and §4.1 already told you those are frame-dependent.

Apply (2.2.19) to the pair. In SS we have Δx=L\Delta x = L, which is what we want, and Δt=0\Delta t=0, which is the definition of measuring. In SS' we have Δx=L0\Delta x' = L_{0}, since the rod is at rest there and it does not matter when you look. Then

Δx=γ(ΔxβcΔt)=γΔxL0=γL, \Delta x' = \gamma\big(\Delta x - \beta\,c\,\Delta t\big) = \gamma\,\Delta x \qquad\Longrightarrow\qquad L_{0} = \gamma L, (2.2.34)

which is the answer, but written the wrong way round: it gives the proper length in terms of the measured one. Divide through by γ\gamma to put the measured length on the left, where we want it:

  L=L0γ=L01β2    L0.   \boxed{\;L = \frac{L_{0}}{\gamma} = L_{0}\sqrt{1-\beta^{2}} \;\leq\; L_{0}.\;} (2.2.35)

A moving rod is short, by γ\gamma, along its direction of motion only, since y=yy'=y and z=zz'=z mean nothing happens across it. And Chapter 2.1's loose end is now tied. This is precisely the 1β2\sqrt{1-\beta^{2}} that FitzGerald and Lorentz had to postulate to explain Michelson–Morley. Here it is a theorem.

It follows from simultaneity — watch

The derivation used Δt=0\Delta t=0, which by (2.2.29) means the two measurement events are not simultaneous in SS':

Δt=γβΔxc=γβLc=βL0c. \Delta t' = -\frac{\gamma\beta\,\Delta x}{c} = -\frac{\gamma\beta L}{c} = -\frac{\beta L_{0}}{c}. (2.2.36)

From the rod's own frame, then, SS located the right end first and the left end βL0/c\beta L_{0}/c later. Since the rod is not going anywhere in SS', that sloppiness costs SS nothing. The ends are where they always are, which is why SS still gets a definite answer. But it is exactly why the answer is not L0L_{0}. To see that the choice of event pair is doing all the work, let's redo the calculation with a pair that is simultaneous in SS' instead:

Δt=0    cΔt=βΔx  Δx=γ(Δxβ2Δx)=Δxγ    Δx=γΔx. \begin{aligned} \Delta t' = 0 \;&\Longrightarrow\; c\,\Delta t = \beta\,\Delta x\\[4pt] &\Longrightarrow\; \Delta x' = \gamma\big(\Delta x - \beta^{2}\,\Delta x\big) = \frac{\Delta x}{\gamma} \;\Longrightarrow\; \Delta x = \gamma\,\Delta x'. \end{aligned} (2.2.37)

Same rod, same frames, different pair of events, and the answer flips from L0/γL_{0}/\gamma to γL0\gamma L_{0}. The contraction is not a property of the rod. It is a property of the measurement, and specifically of whose simultaneity slices the measurement uses. Nothing was squeezed. Two observers sliced the same four-dimensional object at different angles and got different cross-sections, exactly as two people slicing the same sausage at different angles get ellipses of different lengths. Chapter 2.3 makes that picture precise. Here the algebra already contains it.

⚠ Why this isn't obvious — contraction is not "looking squashed"

Length contraction is a statement about a simultaneous measurement. It is not a statement about what a camera records, and the two differ, because light from different parts of an object takes different times to reach the lens. A photograph does not capture a simultaneity slice. It captures the events whose light arrives together.

Let's work the difference out for a cube of side LL flying past at β\beta, seen from far away and edge-on. Light from the rear face must set out earlier than light from the front face, earlier by L/cL/c, which is the extra distance, in order to arrive at the camera together. During that L/cL/c the cube moves a distance βL\beta L, so the rear face is displaced sideways in the image by βL\beta L, and you can see round the back of the cube. Meanwhile the front face, measured simultaneously, is contracted to L1β2L\sqrt{1-\beta^{2}}. Now compare with a cube that is not moving at all but has been rotated by an angle θ\theta. Its front face projects to LcosθL\cos\theta, and one side face becomes visible with projected width LsinθL\sin\theta. Match the two:

sinθ=β,cosθ=1β2 \sin\theta = \beta, \qquad \cos\theta = \sqrt{1-\beta^{2}} \quad\checkmark

The photograph of a rapidly moving cube is the photograph of a rotated cube, turned by θ=arcsinβ\theta=\arcsin\beta. ⚑ The general theorem is the Terrell–Penrose result of 1959: a small object of any shape appears rotated rather than contracted, and a sphere always photographs as a perfectly circular disc. The cube calculation above is the whole idea, and the general case is bookkeeping over solid angle.

So contraction is real and it is not visual. If someone shows you a picture of a squashed spaceship, the picture is wrong. The spaceship is shorter when measured. It does not look shorter when photographed.

β = 0.300 γ = 1.04828 φ = 0.30952 (light lines fixed at 45°)
S: Δt = +1.2000 yr (A first) S′: Δt′ = +0.2516 yr (A first)
twin worldlines: off
The whole chapter on one diagram. The coordinates are those of the unprimed frame SS, with xx across and ctct up. Inside this figure only, units are chosen with c=1c=1, meaning years for time and light-years for distance, so light travels one unit of distance per unit of time and the light lines sit at exactly 4545^{\circ}. Everything else in the book keeps cc explicit. (i) The orange lines are the light cone through the origin. Drag β\beta and watch them: they do not move. That is Postulate 2, drawn. (ii) The blue lines are the primed axes. The ctct' axis is the worldline of the SS' origin (x=βctx=\beta ct), and the xx' axis is the set of events SS' calls simultaneous with the origin (ct=βxct=\beta x). They scissor symmetrically toward the light line, which is the geometric content of (2.2.19): the same β\beta appears in both equations. The faint dashed lines parallel to xx' are SS''s other simultaneity slices, its "nows". (iii) Two spacelike-separated events, AA and BB, with AA earlier in SS by 1.2 yr1.2\ \mathrm{yr}. The dashed purple lines drop each event onto the ctct' axis along SS''s simultaneity slices. Push β\beta past 0.3750.375 and the order reverses. Δt\Delta t' changes sign, and BB, which happened later for SS, happens earlier for SS'. Nothing has been signalled, and no cause has overtaken an effect. These two events cannot influence each other at all, which is exactly the licence that lets their order be frame-dependent. (iv) The twin toggle draws two worldlines from (0,0)(0,0) to (0,4 yr)(0,4\ \mathrm{yr}): the straight one at x=0x=0, and the bent one that flies out at the current β\beta for two years and returns. Proper time along each is computed numerically by summing Δt2Δx2/c2\sqrt{\Delta t^{2}-\Delta x^{2}/c^{2}} over 8000 steps along the path, with no formula used at all, and the bent path always accumulates less. The green dashed lines are the traveller's simultaneity slices immediately before and immediately after the turn, and where they cross x=0x=0 they mark the jump in the stay-at-home twin's clock that §7.1 is about.
Familiar ground — time dilation is an accelerated failure time model

Muon decay is first-order. A population at rest obeys the survival function S0(τ)=eτ/τ0S_{0}(\tau)=\ee^{-\tau/\tau_{0}} with τ0=2.197 μs\tau_{0}=2.197\ \mathrm{\mu s}, which is Chapter 0.1's clearance curve with a different label on the constant. Nothing in (2.2.31) touches that law. What it touches is the argument fed into it.

Watched from the ground, a beam of muons moving at βc\beta c has survival

S(t)  =  S0 ⁣(tγ)  =  exp ⁣(tγτ0), S(t) \;=\; S_{0}\!\left(\frac{t}{\gamma}\right) \;=\; \exp\!\left(-\frac{t}{\gamma\tau_{0}}\right),

and a survival function of the form S(t)=S0(t/φ)S(t)=S_{0}(t/\varphi) is, by definition, an accelerated failure time model with acceleration factor φ\varphi. Special relativity is an AFT model whose covariate is speed and whose factor is φ=γ\varphi=\gamma. That is not a resemblance, it is the same equation, and the reason relativity picks AFT rather than proportional hazards is the physical content of this whole chapter. Proportional hazards would say that motion alters the muon's internal decay rate, which is a claim about muons. AFT says the muon's clock has been rescaled and the decay rate per tick of that clock is untouched, which is a claim about geometry. It is therefore a claim that every particle obeys identically, whatever it is made of.

For an exponential the two models coincide, since scaling the time axis and scaling the hazard are the same operation there. That is exactly why the cosmic-ray calculation of §8 is easy, and exactly why it is easy to take the wrong moral from it. Had the muon decayed with a Weibull hazard the two readings would have parted company, and only the AFT one would have survived the experiment.

What breaks. φ\varphi is not fitted. It is γ=(1β2)1/2\gamma=(1-\beta^{2})^{-1/2}, forced by §2 with no free parameter and no residual heterogeneity whatever. Two muons at the same speed have exactly the same acceleration factor, so there is no frailty term and nothing left over to model. And the covariate is not a property of the subject at all but of the observer. That is why §8's second telling returns the identical count of survivors at the detector, even though it is done in the muon's own frame, where φ=1\varphi=1 and the atmosphere is contracted instead. It has to return the same count. Whether a given muon reaches the ground is one event, and the two frames are describing it, not deciding it.

In plain terms 2.2.4

Three consequences come out of one substitution, and the order in which they are met decides whether the subject ever becomes intelligible. The root fact is that two events at the same time but different places for one observer are not at the same time for another, and the two famous effects follow from it rather than accompany it. Readers who meet them the other way round acquire a picture in which clocks and rulers are mysteriously defective while the word now stays universal, and they rarely recover.

It does not stay universal. Each state of motion carries its own family of nows, cutting through spacetime at different angles, which is what the local conservation laws of the last part were built to survive. It is also why each observer finding the other's clock slow is no contradiction: one clock is compared against a pair synchronised in the other frame, the two observers use different pairs, and the offset between them is exactly the size that balances the books.

Contraction turns out to be a property of the measurement rather than of the rod. Locating both ends at one moment is a demand about simultaneity, so different observers use different pairs of events and get different answers, and a pair the rod's own frame calls simultaneous flips the answer the other way. Nothing was squeezed, and nothing was done to the rod.

5 · Velocity addition

An object moves at velocity uu along xx in frame SS. What velocity uu' does SS' assign it? Galileo says u=uvu'=u-v, which is the velocity-composition rule of Chapter 2.1 §1.2, and it is exactly what Postulate 2 refuses to accept for light. So let's derive the replacement.

Velocity is a ratio of displacements rather than of coordinates, so what we want is the transformation applied to small displacements. Take differentials of (2.2.19). That is legitimate because the coefficients are constants, so the differentials obey the same linear relations as the coordinates do:

dx=γ(dxvdt),dt=γ(dtvdxc2). \dd x' = \gamma\big(\dd x - v\,\dd t\big), \qquad \dd t' = \gamma\left(\dd t - \frac{v\,\dd x}{c^{2}}\right). (2.2.38)

The velocity in SS' is dx/dt\dd x'/\dd t', so divide the first of these by the second. The factors of γ\gamma cancel, which is a small mercy, and a hint that the result depends on β\beta less violently than one might fear:

u=dxdt=dxvdtdtvdx/c2=dxdtv1vc2dxdt, u' = \frac{\dd x'}{\dd t'} = \frac{\dd x - v\,\dd t}{\dd t - v\,\dd x/c^{2}} = \frac{\dfrac{\dd x}{\dd t} - v}{1 - \dfrac{v}{c^{2}}\dfrac{\dd x}{\dd t}}, (2.2.39)

and the last step is only relabelling. Write u=dx/dtu=\dd x/\dd t for the velocity in SS, and the rule stands in the form this book quotes everywhere:

  u=uv1uvc2  equivalentlyβu=βuβv1βuβv. \boxed{\; u' = \frac{u-v}{1 - \dfrac{uv}{c^{2}}} \;} \qquad\text{equivalently}\qquad \beta_{u'} = \frac{\beta_{u}-\beta_{v}}{1-\beta_{u}\beta_{v}}. (2.2.40)

The dimensionless form is (2.2.26) with a sign flipped, as it should be: composing two boosts and transforming a velocity are the same operation seen from two sides.

5.1 · The postulate reproduces itself

The first thing to ask of a new rule is whether it respects the postulate that produced it. So set u=cu=c and turn the crank:

u=cv1cvc2=cv1vc=c(cv)cv=c. u' = \frac{c-v}{1-\dfrac{cv}{c^{2}}} = \frac{c-v}{1-\dfrac{v}{c}} = \frac{c\,(c-v)}{c-v} = c. (2.2.41)

That holds for any vv with v<c\abs{v}\lt c. This is a real check, not a tautology dressed up. We imposed the invariance of cc only for a pulse emitted from the origin at t=0t=0 along ±x\pm x. The transformation that resulted now certifies that any object moving at cc in SS moves at cc in every frame, anywhere, at any time, in either direction. A theory that failed this would have been inconsistent with its own founding assumption.

5.2 · You cannot get to cc by adding speeds

Claim. If βu<1\abs{\beta_{u}}\lt1 and βv<1\abs{\beta_{v}}\lt1 then βu<1\abs{\beta_{u'}}\lt1.

Proof. First, the denominator never vanishes, since βuβv<1\abs{\beta_{u}\beta_{v}}\lt1 gives 1βuβv>01-\beta_{u}\beta_{v}\gt0. What we want is the sign of 1βu21-\beta_{u'}^{2}, so compute it, using the identity from the §3 grind box in its second form, (1ab)2(ab)2=(1a2)(1b2)(1-ab)^{2}-(a-b)^{2}=(1-a^{2})(1-b^{2}):

1βu2=(1βuβv)2(βuβv)2(1βuβv)2=(1βu2)(1βv2)(1βuβv)2. 1-\beta_{u'}^{2} = \frac{\big(1-\beta_{u}\beta_{v}\big)^{2}-\big(\beta_{u}-\beta_{v}\big)^{2}}{\big(1-\beta_{u}\beta_{v}\big)^{2}} = \frac{\big(1-\beta_{u}^{2}\big)\big(1-\beta_{v}^{2}\big)}{\big(1-\beta_{u}\beta_{v}\big)^{2}}. (2.2.42)

Every factor on the right is strictly positive, so 1βu2>01-\beta_{u'}^{2}\gt0, so βu<1\abs{\beta_{u'}}\lt1. \blacksquare

Read (2.2.42) once more, because it says more than the claim. It says the "distance from cc", measured by 1β21-\beta^{2}, is multiplicative under composition. Combine two sub-light speeds and the product of their deficits is what survives. You can multiply small positive numbers together forever and never reach zero. The light barrier is not a wall someone installed. It is the statement that a product of positives is positive.

5.3 · And at ordinary speeds, Galileo

To see what the new rule costs at everyday speeds, take uv/c21uv/c^{2}\ll1 and expand the denominator (Chapter 0.3):

u=(uv)(1+uvc2+)=uv+uv(uv)c2+ u' = (u-v)\left(1+\frac{uv}{c^{2}}+\cdots\right) = u - v + \frac{uv(u-v)}{c^{2}} + \cdots (2.2.43)

Now put numbers in. Two cars approach each other at 30 ms130\ \mathrm{m\,s^{-1}} each, so u=+30u=+30 and v=30v=-30. The Galilean answer is 60 ms160\ \mathrm{m\,s^{-1}} and the correction is

uv(uv)c2=(30)(30)(60)8.99×1016=6.0×1013 ms1, \frac{uv(u-v)}{c^{2}} = \frac{(30)(-30)(60)}{8.99\times10^{16}} = -6.0\times10^{-13}\ \mathrm{m\,s^{-1}}, (2.2.44)

a relative error of one part in 101410^{14}. This is why the Galilean rule survived so long unchallenged: on a motorway it is wrong in the fourteenth decimal place.

⚠ Why this isn't obvious — some speeds may exceed cc, and it is fine

Nothing above forbids coordinate speeds greater than cc. What is forbidden is the transport of matter, energy or information faster than cc. Two examples, both computed.

A spotlight on a distant wall. Rotate a laser at angular rate ω\omega. The spot on a wall at distance RR moves at vspot=ωRv_{\text{spot}}=\omega R, with no upper bound. Point a laser at the Moon (R=3.84×108 mR=3.84\times10^{8}\ \mathrm{m}) and spin it at one revolution per second, ω=2π s1\omega=2\pi\ \mathrm{s^{-1}}:

vspot=2π×3.84×108=2.41×109 ms1=8.05c. v_{\text{spot}} = 2\pi\times3.84\times10^{8} = 2.41\times10^{9}\ \mathrm{m\,s^{-1}} = 8.05\,c.

No rule is broken, because the spot is not a thing. The photons landing at one place and the photons landing at the next place travelled independently from the laser, and nothing went sideways. You cannot send a message along the wall this way, because each point of the wall learns only about the laser, never about its neighbours.

A pair of scissors. Two straight edges cross at a small angle θ\theta, one sliding toward the other at speed vv. The moving edge is y=xtanθ+vty = x\tan\theta + vt, so the intersection with y=0y=0 sits at x=vt/tanθx = -vt/\tan\theta and travels at v/tanθv/\tan\theta, which is unbounded as θ0\theta\to0. With v=100 ms1v=100\ \mathrm{m\,s^{-1}} and θ=107 rad\theta = 10^{-7}\ \mathrm{rad} the crossing point moves at 3.3c3.3\,c. Again nothing is transported. The intersection is a geometric coincidence rather than an object. And a real pair of scissors would in any case not be rigid, since closing the handles sends a wave down the blade at the speed of sound in steel.

The rule to carry: ask what is being transported. If the answer is "nothing", any speed is allowed.

In plain terms 2.2.5

The rule replacing the addition of velocities earns its place by passing a test nobody arranged for it. Invariance of the light speed was imposed for one pulse leaving one origin in one of two directions, and the rule that resulted now certifies that anything travelling at that speed in one frame travels at it in every frame, anywhere, at any time, in any direction.

The barrier that follows is not a wall anybody installed. Rate a speed by the quantity measuring its distance from the limit, and combining two speeds multiplies the two ratings together, up to a positive factor. Small positive numbers can be multiplied together forever without reaching zero, so no finite chain of ordinary boosts arrives at the limit, and the whole of the light barrier is the observation that a product of positive numbers is positive.

None of this forbids coordinate speeds above the limit, and it is worth being clear which is which. Sweep a laser across the face of the Moon and the spot crosses it many times faster than light, because the photons landing at one place and those landing next door travelled independently and nothing went sideways. Close a pair of shears at a shallow enough angle and the crossing point outruns light for the same reason. Ask what is being transported, and where the answer is nothing, any speed is permitted.

6 · Rapidity — the parameter that behaves

Velocities do not add. Something must, because §3 showed that boosts along an axis form a one-parameter group, and a one-parameter group has an additive parameter almost by definition. Composing two elements should add their labels, the way rotating by 3030^{\circ} and then 4040^{\circ} gives 7070^{\circ}. We chose a bad label, that is all. Let's find the good one.

The place to look is the composition rule (2.2.26) itself,

β3=β1+β21+β1β2, \beta_{3} = \frac{\beta_{1}+\beta_{2}}{1+\beta_{1}\beta_{2}}, (2.2.45)

which is not a formula anyone would invent. But it is a formula everyone has seen. It is the addition theorem for the hyperbolic tangent. So let's define the rapidity ϕ\phi to be the thing whose hyperbolic tangent the velocity is:

β  =  tanhϕ,equivalentlyϕ=artanhβ=12ln ⁣1+β1β. \beta \;=\; \tanh\phi, \qquad\text{equivalently}\qquad \phi = \operatorname{artanh}\beta = \tfrac12\ln\!\frac{1+\beta}{1-\beta}. (2.2.46)

This is a legitimate change of variable. tanh\tanh maps R\R one-to-one onto (1,1)(-1,1), which is exactly the allowed range of β\beta, and it is smooth with a smooth inverse. Every subluminal velocity has exactly one rapidity, and every real number is somebody's rapidity.

6.1 · What γ\gamma becomes

Before we can rewrite the boost we need γ\gamma in the new variable. Divide the identity cosh2ϕsinh2ϕ=1\cosh^{2}\phi-\sinh^{2}\phi=1 by cosh2ϕ\cosh^{2}\phi to get 1tanh2ϕ=1/cosh2ϕ1-\tanh^{2}\phi=1/\cosh^{2}\phi. Hence

γ=11β2=11tanh2ϕ=coshϕ,γβ=coshϕtanhϕ=sinhϕ. \gamma = \frac{1}{\sqrt{1-\beta^{2}}} = \frac{1}{\sqrt{1-\tanh^{2}\phi}} = \cosh\phi, \qquad \gamma\beta = \cosh\phi\tanh\phi = \sinh\phi. (2.2.47)

Both entries of the boost matrix are now single hyperbolic functions, so substitute them into (2.2.24) and keep only the block that does anything:

(ctx)=(coshϕsinhϕsinhϕcoshϕ)(ctx), \begin{pmatrix} ct'\\ x'\end{pmatrix} = \begin{pmatrix} \cosh\phi & -\sinh\phi\\ -\sinh\phi & \cosh\phi\end{pmatrix}\begin{pmatrix} ct\\ x\end{pmatrix}, (2.2.48)

and the transformation reads ct=ctcoshϕxsinhϕct'=ct\cosh\phi - x\sinh\phi, x=xcoshϕctsinhϕx'=x\cosh\phi-ct\sinh\phi.

6.2 · Rapidities add

The claim to test is that composing two boosts adds their rapidities, so multiply two of these matrices together and see what comes out:

(coshϕ2sinhϕ2sinhϕ2coshϕ2)(coshϕ1sinhϕ1sinhϕ1coshϕ1)=  (c1c2+s1s2(s1c2+c1s2)(c1s2+s1c2)s1s2+c1c2), \begin{aligned} &\begin{pmatrix} \cosh\phi_{2} & -\sinh\phi_{2}\\ -\sinh\phi_{2} & \cosh\phi_{2}\end{pmatrix} \begin{pmatrix} \cosh\phi_{1} & -\sinh\phi_{1}\\ -\sinh\phi_{1} & \cosh\phi_{1}\end{pmatrix}\\[6pt] &\qquad=\; \begin{pmatrix} c_{1}c_{2}+s_{1}s_{2} & -(s_{1}c_{2}+c_{1}s_{2})\\ -(c_{1}s_{2}+s_{1}c_{2}) & s_{1}s_{2}+c_{1}c_{2}\end{pmatrix}, \end{aligned} (2.2.49)

writing ci=coshϕic_{i}=\cosh\phi_{i} and si=sinhϕis_{i}=\sinh\phi_{i}. The hyperbolic addition formulas are cosh(a+b)=coshacoshb+sinhasinhb\cosh(a+b)=\cosh a\cosh b+\sinh a\sinh b and sinh(a+b)=sinhacoshb+coshasinhb\sinh(a+b)=\sinh a\cosh b+\cosh a\sinh b, and they turn that matrix into

  Λ(ϕ2)Λ(ϕ1)  =  Λ(ϕ1+ϕ2).   \boxed{\;\Lambda(\phi_{2})\,\Lambda(\phi_{1}) \;=\; \Lambda(\phi_{1}+\phi_{2}).\;} (2.2.50)

Rapidities add. Plainly, exactly, with no correction term. And taking tanh\tanh of both sides recovers (2.2.26), since tanh(ϕ1+ϕ2)=(tanhϕ1+tanhϕ2)/(1+tanhϕ1tanhϕ2)\tanh(\phi_{1}+\phi_{2})=(\tanh\phi_{1}+\tanh\phi_{2})/(1+\tanh\phi_{1}\tanh\phi_{2}). So §5 and §6 are the same statement written in two variables.

Check the arithmetic of (2.2.28) in the new currency: artanh(0.5)=0.549306\operatorname{artanh}(0.5)=0.549306, twice that is 1.098612=ln31.098612=\ln3, and tanh(ln3)=(313)/(3+13)=0.8\tanh(\ln3)= (3-\tfrac13)/(3+\tfrac13)=0.8 ✓. Two boosts of 0.5c0.5c do not give cc. They give rapidity 2×0.54932\times0.5493, which is velocity 0.8c0.8c.

The moral

Velocities "fail to add" for the same reason angles-as-slopes fail to add. Compose two rotations and their angles add. Their slopes combine by (m1+m2)/(1m1m2)(m_{1}+m_{2})/(1-m_{1}m_{2}), which nobody finds shocking, because nobody expected slopes to be the natural parameter. Velocity is the slope of a worldline. Rapidity is its angle. We spent three hundred years adding the slopes.

And the change of variable makes the light barrier trivial rather than mysterious. Rapidity runs over the whole of R\R and adds without limit, so you can boost forever. But β=tanhϕ\beta=\tanh\phi saturates. ϕ=1\phi=1 gives β=0.762\beta=0.762, ϕ=3\phi=3 gives 0.9950.995, and ϕ=10\phi=10 gives 0.99999999590.9999999959. The barrier is not a barrier at all. It is the horizontal asymptote of tanh\tanh, seen from inside a badly chosen coordinate.

Two forward pointers, flagged ⚑

(a) This is a Lie group in embryo. The boosts along one axis are a family Λ(ϕ)\Lambda(\phi) with Λ(0)=I\Lambda(0)=I, Λ(ϕ2)Λ(ϕ1)=Λ(ϕ1+ϕ2)\Lambda(\phi_{2})\Lambda(\phi_{1})=\Lambda(\phi_{1}+\phi_{2}) and Λ\Lambda smooth in ϕ\phi. That is a smooth one-parameter group. ⚑ Chapter 6.1 shows that any such family is Λ(ϕ)=exp(ϕK)\Lambda(\phi)=\exp(\phi K) for a fixed matrix KK, its generator, obtained by differentiating at the identity. Here K=dΛdϕ0=(0110)K = \dv{\Lambda}{\phi}\big|_{0} = \left(\begin{smallmatrix}0 & -1\\ -1 & 0\end{smallmatrix}\right), and indeed K2=IK^{2}=I, so exp(ϕK)=Icoshϕ+Ksinhϕ\exp(\phi K)=I\cosh\phi + K\sinh\phi, which is (2.2.48) exactly. The full Lorentz group has six such parameters, three boosts and three rotations, and is called SO(1,3)\mathrm{SO}(1,3).

(b) The hyperbolic functions are the metric's minus sign, surfacing. A rotation in the xyxy plane is (cosθsinθsinθcosθ)\left(\begin{smallmatrix}\cos\theta & -\sin\theta\\ \sin\theta & \cos\theta\end{smallmatrix}\right) and preserves x2+y2x^{2}+y^{2}. A boost in the (ct,x)(ct,x) plane is (2.2.48) and preserves c2t2x2c^{2}t^{2}-x^{2}, per (2.2.20). Circular functions preserve a sum of squares. Hyperbolic functions preserve a difference. The two matrices differ in exactly one respect: the rotation's off-diagonal entries carry opposite signs while the boost's carry the same sign, so the boost matrix is symmetric and the rotation's off-diagonal part is not. That single sign is the - in ημν=diag(1,1,1,1)\eta_{\mu\nu}=\mathrm{diag}(1,-1,-1,-1). Chapter 2.3 makes this geometric: a boost is a rotation through an imaginary angle, or better, an honest rotation in a geometry whose circles are hyperbolas.

Grind box — hyperbolic functions from scratch, in case they are rusty

Definitions, and everything else follows:

coshϕ=eϕ+eϕ2,sinhϕ=eϕeϕ2,tanhϕ=sinhϕcoshϕ. \cosh\phi = \frac{\ee^{\phi}+\ee^{-\phi}}{2}, \qquad \sinh\phi = \frac{\ee^{\phi}-\ee^{-\phi}}{2}, \qquad \tanh\phi=\frac{\sinh\phi}{\cosh\phi}.

The Pythagorean identity. Multiply out:

cosh2ϕsinh2ϕ=(eϕ+eϕ)2(eϕeϕ)24=4eϕeϕ4=1. \cosh^{2}\phi-\sinh^{2}\phi = \frac{\big(\ee^{\phi}+\ee^{-\phi}\big)^{2}-\big(\ee^{\phi}-\ee^{-\phi}\big)^{2}}{4} = \frac{4\,\ee^{\phi}\ee^{-\phi}}{4} = 1.

Compare cos2+sin2=1\cos^{2}+\sin^{2}=1. The point (cosθ,sinθ)(\cos\theta,\sin\theta) traces a circle, and the point (coshϕ,sinhϕ)(\cosh\phi,\sinh\phi) traces a hyperbola. That hyperbola is Chapter 2.3's central object.

Addition. Multiply exponentials:

coshacoshb+sinhasinhb=14[(ea+ea)(eb+eb)+(eaea)(ebeb)]=14[2ea+b+2eab]  =  cosh(a+b). \begin{aligned} \cosh a\cosh b + \sinh a \sinh b &= \tfrac14\Big[\big(\ee^{a}+\ee^{-a}\big)\big(\ee^{b}+\ee^{-b}\big) + \big(\ee^{a}-\ee^{-a}\big)\big(\ee^{b}-\ee^{-b}\big)\Big]\\[3pt] &= \tfrac14\Big[2\ee^{a+b}+2\ee^{-a-b}\Big] \;=\; \cosh(a+b). \end{aligned}

The sinh\sinh formula goes the same way. Dividing them gives the tanh\tanh addition theorem, which is (2.2.45).

The inverse. Solve β=tanhϕ\beta=\tanh\phi for ϕ\phi. Writing w=e2ϕw=\ee^{2\phi},

β=eϕeϕeϕ+eϕ=w1w+1    w=1+β1β    ϕ=12ln1+β1β. \beta = \frac{\ee^{\phi}-\ee^{-\phi}}{\ee^{\phi}+\ee^{-\phi}} = \frac{w-1}{w+1} \;\Longrightarrow\; w = \frac{1+\beta}{1-\beta} \;\Longrightarrow\; \phi = \tfrac12\ln\frac{1+\beta}{1-\beta}.

Useful values: artanh(0.5)=0.5493\operatorname{artanh}(0.5)=0.5493, artanh(0.8)=1.0986=ln3\operatorname{artanh}(0.8)=1.0986=\ln3, artanh(0.9)=1.4722\operatorname{artanh}(0.9)=1.4722, artanh(0.99)=2.6467\operatorname{artanh}(0.99)=2.6467. Note how slowly rapidity grows once β\beta is close to 11. Going from 0.9c0.9c to 0.99c0.99c costs about 1.171.17 of rapidity, roughly the same as going from rest to 0.8c0.8c. In Chapter 2.5 that observation acquires teeth, because the energy needed is what grows without bound.

Derivatives. cosh=sinh\cosh'=\sinh, sinh=cosh\sinh'=\cosh, tanh=1/cosh2=1tanh2\tanh'=1/\cosh^{2}=1-\tanh^{2}, all immediate from the exponential definitions. Problem 4 uses the last one.

In plain terms 2.2.6

Compose two rotations of a plane and their angles add, as anybody expects of a family labelled by one number. Boosts along an axis are such a family, so something of theirs adds, and it is not velocity. Take the composition rule seriously and it is a formula everybody has seen, the addition theorem for the hyperbolic tangent, so what adds is the angle whose hyperbolic tangent the velocity is.

Call that angle the rapidity. Rapidities add exactly, with no correction term, and their matrices multiply the way two turns of a plane do. This collects the pairing between a generator and the family it builds, promised when a self-partnered map was first exponentiated and again when conserved quantities were made to push things around: one fixed matrix, exponentiated by the parameter, generates the family.

Why velocities looked defective is now visible, and it is not a fact about nature. Velocity is the slope of a worldline and rapidity is its angle, and slopes have never added: two turns of a plane combine their angles cleanly while their slopes combine by an ugly quotient that shocks nobody, since nobody expected slopes to be the natural label. Three centuries were spent adding the slopes. The barrier stops being mysterious in the same stroke, since the angle runs over every real number and adds without limit while its hyperbolic tangent creeps towards one and never arrives.

7 · The paradoxes, dismantled

There are perhaps a dozen famous special-relativity paradoxes and they are all the same paradox. Here are the two that matter, worked to the bottom, and then the general statement.

7.1 · The twin paradox

The setup. Two twins, both aged 2020. One stays on Earth. The other flies to a star 44 light-years away at β=0.8\beta=0.8 (γ=5/3\gamma=5/3), turns round, and comes back. Here is the Earth-frame arithmetic. The outbound leg takes 4/0.8=54/0.8=5 years, so the round trip takes 1010 years. The traveller's clock is the moving one, so by (2.2.31) it records 10/γ=10×0.6=610/\gamma = 10\times0.6 = 6 years. The traveller returns aged 2626, and the stay-at-home is 3030.

The apparent paradox, stated as sharply as it deserves. Motion is relative. From the traveller's point of view, Earth receded at 0.8c0.8c and came back. By §4.2, each observer finds the other's clock slow, symmetrically and by the same γ\gamma. So the traveller should equally conclude that Earth's clock ran slow throughout. His own 66 years, divided by γ=5/3\gamma=5/3, gives 3.63.6 years for the Earth twin, who should therefore be the younger one on his return. Both cannot be true. When the two stand next to each other at the end, comparing wrinkles is a local coincidence, and there is one fact of the matter.

Where the symmetry fails. It fails at the turnaround, and it fails physically, not verbally. The stay-at-home twin occupies a single inertial frame for the whole story. The traveller does not. He is at rest in one inertial frame going out, and in a different one coming back. The switch is detectable from inside the ship without looking out of the window. An accelerometer registers it, coffee spills, and the traveller is pressed into his seat. There is no symmetry to appeal to, because the two histories are physically different, and no amount of "motion is relative" makes them the same. One worldline is straight. The other has a kink.

That is the resolution in one sentence, and people rightly find it unsatisfying. It explains why the twins may differ without showing where the missing 6.46.4 years went, that being the gap between the traveller's naive 3.63.6 and the actual 1010. So let's track them.

Following the traveller's own accounting, year by year

Work in Earth-frame coordinates (ct,x)(ct,x) measured in light-years and years, with the departure at the origin. The turnaround is the event T=(ct,x)=(5,4)T=(ct,x)=(5,4).

Outbound leg. The traveller is at rest in the frame SS' moving at β=+0.8\beta=+0.8. His own elapsed time is, from (2.2.19) applied to the event TT,

ctT=γ(ctTβxT)=53(50.8×4)=53×1.8=3 years. ct'_{T} = \gamma\big(ct_{T}-\beta x_{T}\big) = \tfrac53\big(5-0.8\times4\big) = \tfrac53\times1.8 = 3\ \text{years}. (2.2.51)

The next thing we need is his verdict on Earth. Which of Earth's events does he call simultaneous with his turnaround? Those with the same tt', that is those with ctβx=1.8ct-\beta x = 1.8. On Earth x=0x=0, so

ctEarth=1.8 years. ct_{\text{Earth}} = 1.8\ \text{years}. (2.2.52)

Consistent with time dilation as the traveller sees it: his 33 years, divided by γ=5/3\gamma=5/3, gives 1.81.8 years on the Earth clock. From out there, Earth is the moving one and Earth is running slow. ✓

Inbound leg. After the turn he is at rest in SS'' moving at β=0.8\beta=-0.8. Its simultaneity slices are ct+βx=constct+\beta x = \text{const}, so run the same question through the same turnaround event T=(5,4)T=(5,4) and see which Earth event he now calls "now":

ct+0.8x=5+3.2=8.2ctEarth=8.2 years. ct + 0.8x = 5 + 3.2 = 8.2 \qquad\Longrightarrow\qquad ct_{\text{Earth}} = 8.2\ \text{years}. (2.2.53)

The jump. An instant before the turn, the traveller's "now on Earth" is year 1.81.8. An instant after, it is year 8.28.2. Nothing happened to Earth. The traveller changed which slice of spacetime he calls "now", and the slice swung through Earth's worldline:

Δtjump=8.21.8=6.4 years  =  2βLc  =  2×0.8×4. \Delta t_{\text{jump}} = 8.2 - 1.8 = 6.4\ \text{years} \;=\; \frac{2\beta L}{c} \;=\; 2\times0.8\times4. (2.2.54)

That jump is the piece the naive argument left out, so add the three contributions to the Earth twin's clock and watch the books balance exactly:

1.8outbound  +  6.4turnaround jump  +  1.8inbound  =  10.0 years \underbrace{1.8}_{\text{outbound}} \;+\; \underbrace{6.4}_{\text{turnaround jump}} \;+\; \underbrace{1.8}_{\text{inbound}} \;=\; 10.0\ \text{years} \quad\checkmark (2.2.55)

while the traveller's own clock reads 3+3=63+3=6 years. Every year is accounted for. The traveller's claim "Earth's clock ran slow" was correct on each leg. What he cannot do is stitch the two legs' notions of "now" together and pretend they were one. The naive symmetric argument gave Earth 1.8+1.8=3.61.8+1.8=3.6 years, and the 6.46.4 it was missing is exactly the span his simultaneity slice swept across Earth's worldline during the turn. That turn was the one moment at which he was not an inertial observer, and therefore the one moment at which his slices were not entitled to line up.

The version to actually remember

Forget clocks for a moment. Proper time is the length of a worldline. Between two events it is Δτ=Δt2Δx2/c2\Delta\tau=\sqrt{\Delta t^{2}-\Delta x^{2}/c^{2}} along each straight piece, and you add the pieces. Two different paths between the same two endpoints have different lengths. Nobody is scandalised that a detour through the next town adds mileage.

Check the twin numbers this way. The straight path gives Δτ=1020=10\Delta\tau=\sqrt{10^{2}-0}=10 years. The bent path has two legs, each 5242=3\sqrt{5^{2}-4^{2}}=3, so 66 in total. Same endpoints, different lengths, no paradox. And there is no need to mention acceleration at all, except to say that a kink is what makes a path longer or shorter than a straight one.

The one genuinely strange thing is the sign. Because of the minus in Δt2Δx2/c2\Delta t^{2}-\Delta x^{2}/c^{2}, the straight path is the one with the most proper time, not the least. Detours in spacetime save you time rather than costing it. That inverted extremum is not a curiosity. It is the principle from which Chapter 3.1 derives the motion of freely falling bodies, and it is why a satellite's orbit is the path that maximises the proper time of the clock aboard it.

Turn on the twin toggle in the figure above and watch this happen. The traveller's two simultaneity slices are the green dashed lines through the turnaround, and where they cross x=0x=0 they bracket exactly the jump. Drag β\beta and the numbers move, but the three pieces always sum to the stay-at-home twin's total.

7.2 · The pole and the barn

The setup. A runner carries a pole of proper length 15 m15\ \mathrm{m} through a barn of proper length 10 m10\ \mathrm{m} with a door at each end, at β=0.8\beta=0.8 so γ=5/3\gamma=5/3. In the barn's frame the pole is contracted to 15×0.6=9 m15\times0.6=9\ \mathrm{m}, so it fits, with a metre to spare. The farmer shuts both doors simultaneously at the moment it is inside, and the pole is briefly wholly enclosed. That last clause is the claim to be tested.

The apparent paradox. In the runner's frame the pole is 15 m15\ \mathrm{m} and the barn is contracted to 10×0.6=6 m10\times0.6=6\ \mathrm{m}. The pole is two and a half times the barn's length. It cannot possibly be enclosed. But "the pole was inside the barn with both doors shut" ought to be a fact, not an opinion.

The resolution. It is not a fact, because it is not a local coincidence. It is a statement that two events at different places, front door shuts and back door shuts, are simultaneous. By §4.1 that is frame-dependent, and here the disagreement is enormous. With Δx=10 m\Delta x = 10\ \mathrm{m},

cΔt=γβΔx=53×0.8×10=403=13.33 m, c\,\Delta t' = -\gamma\beta\,\Delta x = -\tfrac53\times0.8\times10 = -\tfrac{40}{3} = -13.33\ \mathrm{m}, (2.2.56)

so in the runner's frame the two door-closings are 13.33 m13.33\ \mathrm{m} of light travel apart, which is 44.5 ns44.5\ \mathrm{ns}, with the far door going first. They do not shut together at all. The far door shuts and reopens long before the pole's nose gets there. Much later, the near door shuts behind the pole's tail. The pole passes through a barn that is opening and closing around it like a slow shutter, and it is never enclosed for an instant.

Both frames are right, and here is the decisive point: they agree on everything that is not a convention. Was the pole ever hit by a door? No, in both frames. Did the nose emerge from the far door before the far door shut? Problem 2 tabulates every event in both frames, and the answer is the same in both. What they disagree on is only "did the two doors shut at the same time", which is not an observable but a labelling.

7.3 · The general lesson

Every special-relativity paradox is the same paradox

Somewhere in the statement, someone has assumed simultaneity is absolute. Find that assumption and the paradox dissolves. There is no second mechanism.

The reliable procedure, which will get you through any of them:

  1. List the events. Not objects, and not "the rod" or "the ship", but events, each with a definite (t,x)(t,x) in one frame. Anything that is not an event is not yet a physics question.
  2. Ask which claims are local coincidences. Two things at the same place at the same time: a door hitting a pole, two clocks side by side, a twin shaking a hand. These are frame-independent and every frame must agree on them, full stop.
  3. Transform. Apply (2.2.19) to each event and tabulate.
  4. Locate the smuggled assumption. It will be a claim of the form "and at that moment, over there, ...", which is a statement about distant simultaneity dressed as a fact.

Apply that to the twins, where the smuggled claim is "while the traveller aged 3 years, Earth aged 1.8, and that was still true after the turn". Apply it to the pole, where the smuggled claim is "both doors shut at the same time" taken as an absolute. Apply it to any of the others: the ladder and the trapdoor, the rigid rod pushed at one end, the submarine that both sinks and floats. Same disease, same cure, every time.

In plain terms 2.2.7

One twin comes back younger than the other, and the asymmetry permitting it is physical rather than verbal. The one who stayed occupied a single inertial frame throughout; the one who travelled was at rest in one frame going out and in a different one coming back, and the switch is detectable from inside the ship, with no window needed. One history is bent and the other is not.

Saying so explains why the two may differ without showing where the missing years went, and they can be tracked. The traveller's claim that the other clock was running slow is correct on each leg taken separately. What he may not do is stitch the two legs' notions of the present together, because at the turn his family of nows swings through an enormous span of the other worldline, and that span is precisely the amount the naive accounting was missing.

Every paradox in the subject is this one in different clothes. Somewhere in the statement, someone has treated the simultaneity of two distant events as a fact rather than a labelling, and the cure never varies. List the events; ask which claims are local coincidences of one thing with another at one place, since everybody agrees on those; transform the rest; and the smuggled assumption will be standing in plain view. What this pile of results is waiting for is the idea that organises it.

8 · Worked examples

Worked example 1 — cosmic-ray muons, done in both frames

Muons are produced by cosmic-ray collisions about 15 km15\ \mathrm{km} up in the atmosphere. ⚑ Their mean lifetime at rest is τ0=2.2 μs\tau_{0}=2.2\ \mathrm{\mu s}, measured in the laboratory, and they decay exponentially, so that the fraction surviving after a proper time τ\tau is eτ/τ0\ee^{-\tau/\tau_{0}} (Chapter 0.1 §5, the same first-order kinetics as drug clearance). Take β=0.998\beta=0.998. What fraction reaches sea level, according to (a) Newtonian expectations and (b) relativity? Then redo (b) entirely in the muon's frame, where the atmosphere is contracted rather than the clock dilated, and check the two accounts agree.

Step 0: the Lorentz factor.

γ=110.9982=13.996×103=15.82. \gamma = \frac{1}{\sqrt{1-0.998^{2}}} = \frac{1}{\sqrt{3.996\times10^{-3}}} = 15.82.

Step 1: Newtonian. No dilation, no contraction, so the muon crosses 15 km15\ \mathrm{km} at 0.998c0.998c and its internal clock is the same as ours.

t=150000.998×2.998×108=5.013×105 s=50.13 μs, t = \frac{15\,000}{0.998\times2.998\times10^{8}} = 5.013\times10^{-5}\ \mathrm{s} = 50.13\ \mathrm{\mu s},

which is 50.13/2.2=22.7950.13/2.2 = 22.79 lifetimes. Surviving fraction:

e22.79=1.27×1010. \ee^{-22.79} = 1.27\times10^{-10}.

About one muon in eight billion. The predicted flux at sea level is, to any instrument ever built, zero. ⚑ What is actually observed is a copious flux, of order one muon per square centimetre per minute at sea level. The Newtonian prediction is not slightly wrong.

Step 2: relativity, told from the ground. The muon's internal clock is the moving one, so by (2.2.31) the proper time it experiences during our 50.13 μs50.13\ \mathrm{\mu s} is

Δτ=Δtγ=50.1315.82=3.169 μs=1.441 lifetimes, \Delta\tau = \frac{\Delta t}{\gamma} = \frac{50.13}{15.82} = 3.169\ \mathrm{\mu s} = 1.441\ \text{lifetimes}, e1.441=0.2368. \ee^{-1.441} = 0.2368.

Nearly a quarter of them make it. The two predictions differ by a factor of 1.9×1091.9\times10^{9}. That is why this is not a subtle test but a demonstration you can do with a scintillator on a bench, and why ⚑ Rossi and Hall's 1941 mountain-versus-sea-level muon counts were decisive.

Step 3: the same physics, told from the muon. This is the half that gets skipped, and it is the half that matters, because in the muon's own frame nothing dilates its clock. Its clock is at rest. It lives 2.2 μs2.2\ \mathrm{\mu s} on average, exactly as advertised. And it had better still reach the ground, because whether a given muon arrives is a local coincidence, and frames are not allowed to disagree about it.

What is different in the muon's frame is the atmosphere, which is now the moving object. The 15 km15\ \mathrm{km} column of air is a proper length, since it is at rest relative to the ground, so by (2.2.35) it is contracted to

L=L0γ=1500015.82=948.2 m. L = \frac{L_{0}}{\gamma} = \frac{15\,000}{15.82} = 948.2\ \mathrm{m}.

That column rushes past the muon at 0.998c0.998c, taking

Δtmuon=948.20.998×2.998×108=3.169 μs, \Delta t_{\text{muon}} = \frac{948.2}{0.998\times2.998\times10^{8}} = 3.169\ \mathrm{\mu s},

and the surviving fraction is e3.169/2.2=e1.441=0.2368\ee^{-3.169/2.2}=\ee^{-1.441}=0.2368. Identical.

What this teaches. The two frames disagree about why the muon survives. One says the clock ran slow, the other says the trip was short. They agree exactly on the observable, which is the number of muons per second landing on the detector. That is the correct relationship between frames in every problem in this chapter: mechanisms are frame-dependent, observables are not. Note also that neither account works without the other's ingredient. Use dilation and contraction together and you would double-count by a factor of γ\gamma. Use neither and you get 101010^{-10}. There is exactly one factor of γ\gamma in the problem, and each frame places it somewhere different.

Worked example 2 — closing speed versus relative speed

In a laboratory, two particles fly directly toward each other, each at 0.9c0.9c as measured in the lab. (a) At what rate does the gap between them close, in lab coordinates? (b) What is the speed of one as measured in the rest frame of the other? (c) Reconcile.

(a) The closing rate. In lab coordinates, particle 1 is at x1=d/2+0.9ctx_{1}=-d/2+0.9ct and particle 2 at x2=d/20.9ctx_{2}=d/2-0.9ct. The gap is x2x1=d1.8ctx_{2}-x_{1}=d-1.8ct, so

ddt(x2x1)=1.8c. -\dv{}{t}\big(x_{2}-x_{1}\big) = 1.8\,c.

The gap shrinks at 1.8c1.8c. This is correct and no rule is violated. It is not the speed of anything. It is the rate of change of a difference of two coordinates, both measured in one frame. Nothing is transported at 1.8c1.8c. As with §5's spotlight, ask what is being carried, and the answer is nothing.

(b) The relative speed. This is a different question: what does particle 2's own frame measure for particle 1? Take SS' to be particle 2's rest frame, so βv=0.9\beta_{v}=-0.9 (it moves in the x-x direction in the lab), and the object of interest has βu=+0.9\beta_{u}=+0.9. Then by (2.2.40),

βu=βuβv1βuβv=0.9(0.9)1(0.9)(0.9)=1.81.81=0.994475. \beta_{u'} = \frac{\beta_{u}-\beta_{v}}{1-\beta_{u}\beta_{v}} = \frac{0.9-(-0.9)}{1-(0.9)(-0.9)} = \frac{1.8}{1.81} = 0.994475.

So u=0.9945cu' = 0.9945\,c. Not 1.8c1.8c, and not even cc. Each particle sees the other approaching at 99.45%99.45\% of light speed.

Cross-check with rapidity. artanh(0.9)=1.472219\operatorname{artanh}(0.9)=1.472219, and the relative rapidity is the sum, 2×1.472219=2.9444392\times1.472219=2.944439. Then tanh(2.944439)=0.994475\tanh(2.944439)=0.994475 ✓. That is the same number, obtained by addition rather than by a formula with a denominator in it. This is §6 earning its keep.

(c) Reconciling. The two answers are answers to two different questions, and the language of everyday physics does not distinguish them because at low speed they coincide. For β1\beta\ll1, βuβuβv\beta_{u'}\approx\beta_{u}-\beta_{v}, which is the closing rate. At high speed they separate dramatically. Here is the rule. A closing rate is a statement about one frame's coordinates and can be up to 2c2c. A relative velocity is a statement about one object's rest frame and is always less than cc. Only the second one is a velocity in the sense that (2.2.40) and Chapter 2.5's dynamics care about.

Why the difference has teeth. In a collider the quantity that determines what you can produce is the energy in the centre-of-mass frame, which depends on the relative velocity, not the closing rate. Two beams at 0.9c0.9c head-on do not collide at 1.8c1.8c in any useful sense. They collide at a relative γ\gamma of cosh(2.944)=9.53\cosh(2.944)=9.53, which is γrel=γ1γ2(1+β1β2)=2.2942×1.81=9.53\gamma_{\text{rel}} = \gamma_{1}\gamma_{2}(1+\beta_{1}\beta_{2}) = 2.294^{2}\times1.81 = 9.53 ✓. That is the composition law of §3, doing an accelerator physicist's arithmetic.

9 · Your turn

Problem 1 · the transverse velocity, which changes even though y=yy'=y

An object has velocity components (ux,uy)(u_{x},u_{y}) in SS. Derive uxu'_{x} and uyu'_{y} in SS'. Explain why uyuyu'_{y}\neq u_{y} even though the yy coordinate itself is untouched. Then check the formula on light. Take a pulse moving in the +y+y direction at speed cc in SS, verify that SS' also measures speed cc, and find the angle by which the pulse's direction is tilted.

Solution

Derivation. Take differentials of (2.2.19) including the transverse equation dy=dy\dd y'=\dd y:

dx=γ(dxvdt),dy=dy,dt=γ(dtvdxc2). \dd x' = \gamma\big(\dd x - v\dd t\big), \qquad \dd y' = \dd y, \qquad \dd t' = \gamma\left(\dd t - \frac{v\,\dd x}{c^{2}}\right).

Divide dy\dd y' by dt\dd t' and then divide numerator and denominator by dt\dd t:

uy=dyγ(dtvdx/c2)=uyγ(1uxvc2),ux=uxv1uxvc2. u'_{y} = \frac{\dd y}{\gamma\big(\dd t - v\dd x/c^{2}\big)} = \frac{u_{y}}{\gamma\left(1-\dfrac{u_{x}v}{c^{2}}\right)}, \qquad u'_{x} = \frac{u_{x}-v}{1-\dfrac{u_{x}v}{c^{2}}}.

Why it changes. A velocity is a ratio, Δy/Δt\Delta y/\Delta t. The numerator is indeed untouched. The denominator is not. Time intervals transform, so a displacement that is unchanged, divided by a duration that is not, gives a changed rate. The lesson generalises far beyond this problem: an invariant numerator does not make an invariant quotient, which is exactly why Chapter 2.5 will insist on differentiating with respect to proper time rather than coordinate time when it builds the four-velocity.

The special case ux=0u_{x}=0. Then ux=vu'_{x}=-v and uy=uy/γu'_{y}=u_{y}/\gamma, so the transverse velocity is reduced by exactly γ\gamma, since the transverse displacement is unchanged while the time to make it is dilated.

Check on light. Put ux=0u_{x}=0, uy=cu_{y}=c. Then ux=vu'_{x}=-v and uy=c/γu'_{y}=c/\gamma, so

u2=v2+c2γ2=v2+c2(1β2)=v2+c2v2=c2. \abs{u'}^{2} = v^{2} + \frac{c^{2}}{\gamma^{2}} = v^{2} + c^{2}\big(1-\beta^{2}\big) = v^{2}+c^{2}-v^{2} = c^{2}.

Speed cc ✓. Note how the two pieces conspire: the v2v^{2} gained along xx is precisely the c2β2c^{2}\beta^{2} lost along yy. Postulate 2 again, checking itself in a direction we never imposed it.

The tilt. The direction makes an angle θ\theta' with the yy'-axis given by

tanθ=uxuy=vc/γ=γβ,equivalentlysinθ=β. \tan\theta' = \frac{\abs{u'_{x}}}{u'_{y}} = \frac{v}{c/\gamma} = \gamma\beta, \qquad\text{equivalently}\qquad \sin\theta' = \beta.

This is stellar aberration. The apparent position of a star shifts by θβ\theta'\approx\beta as the Earth's velocity changes around its orbit, giving an annual ellipse of angular radius β=9.94×105 rad\beta = 9.94\times10^{-5}\ \mathrm{rad}, which in arcseconds (×206265\times\,206265) is 20.520.5''. Bradley measured it in 1727, and it was one of the constraints Chapter 2.1 §6.1 used to kill the fully dragged ether. Chapter 2.1's Problem 3 obtained θv/c\theta\approx v/c from a first-order argument. Here sinθ=β\sin\theta'=\beta is the exact statement, and the two agree to one part in 10810^{8} at the Earth's orbital speed. Note also that sinθ=β\sin\theta'=\beta is the same relation as the Terrell–Penrose rotation angle in §4.3. That is not a coincidence, since both are the statement that a boost tilts light rays.

Problem 2 · the pole and the barn, with every event tabulated

Pole of proper length 15 m15\ \mathrm{m}, barn of proper length 10 m10\ \mathrm{m} with doors at x=0x=0 (entrance) and x=10 mx=10\ \mathrm{m} (exit), runner at β=0.8\beta=0.8 in the +x+x direction. Use coordinates (ct,x)(ct,x) with both in metres. Let event AA be the pole's tail reaching the entrance, and make it the origin of both frames. (a) Find the barn-frame times at which the pole is fully inside, and choose the farmer's simultaneous door-shutting to be at the midpoint of that window. (b) Tabulate the four events AA, BB (nose reaches the exit door), FF (entrance door shuts) and RR (exit door shuts) in both frames. (c) In the runner's frame, show explicitly that neither door ever touches the pole.

Solution

(a) The window. γ=5/3\gamma=5/3, so the pole is contracted to 15/γ=9 m15/\gamma=9\ \mathrm{m} in the barn frame. At ct=0ct=0 the tail is at x=0x=0 and the nose at x=9x=9. The nose reaches x=10x=10 after the pole advances 1 m1\ \mathrm{m}, which takes ct=1/0.8=1.25 mct = 1/0.8 = 1.25\ \mathrm{m}. So the pole is wholly inside for 0<ct<1.25 m0 \lt ct \lt 1.25\ \mathrm{m} (that is, for 4.17 ns4.17\ \mathrm{ns}), and the farmer shuts both doors at ct=0.625 mct = 0.625\ \mathrm{m}.

(b) The table. Transform with ct=γ(ctβx)ct'=\gamma(ct-\beta x), x=γ(xβct)x'=\gamma(x-\beta ct) and γ=5/3\gamma=5/3, β=0.8\beta=0.8:

Eventctct (m)xx (m)ctct' (m)xx' (m)
AA   tail reaches entrance00000000
BB   nose reaches exit1.251.25101011.25-11.251515
FF   entrance door shuts0.6250.62500+1.0417+1.04170.8333-0.8333
RR   exit door shuts0.6250.625101012.2917-12.2917+15.8333+15.8333

Read the fourth column. In the barn frame FF and RR are simultaneous. In the runner's frame RR happens at ct=12.29ct'=-12.29 and FF at ct=+1.04ct'=+1.04, so the exit door shuts 13.33 m/c=44.5 ns13.33\ \mathrm{m}/c = 44.5\ \mathrm{ns} earlier, matching (2.2.56). Note also xB=15 mx'_{B}=15\ \mathrm{m}, the pole's proper length ✓. And in the runner's frame the nose reaches the exit at ct=11.25ct'=-11.25, which is well before the tail reaches the entrance at ct=0ct'=0. The pole is sticking out of both ends at once, which is what a 15 m15\ \mathrm{m} pole does in a 6 m6\ \mathrm{m} barn.

(c) Nobody gets hit. In the runner's frame the pole is at rest occupying 0x150\leq x'\leq15. Track the doors. The exit door is the worldline x=10x=10; substituting ct=ct/γ+β10ct = ct'/\gamma + \beta\cdot 10 into x=γ(xβct)x'=\gamma(x-\beta ct) gives

xexit(ct)=γ(1β2)10βct=60.8ct. x'_{\text{exit}}(ct') = \gamma\big(1-\beta^{2}\big)10 - \beta\,ct' = 6 - 0.8\,ct'.

It reaches the nose, x=15x'=15, when 60.8ct=156-0.8ct'=15, i.e. ct=11.25ct'=-11.25 ✓ (event BB). It shuts at ct=12.29ct'=-12.29, which is earlier, and at that moment it sits at x=6+0.8(12.2917)=15.83 mx'=6+0.8(12.2917)=15.83\ \mathrm{m}, which is beyond the nose. The door shuts in empty space in front of the pole, reopens, and the pole then passes through.

The entrance door is x=0x=0, worldline xent(ct)=0.8ctx'_{\text{ent}}(ct')=-0.8\,ct'. It shuts at ct=+1.0417ct'=+1.0417, when it sits at x=0.833 mx'=-0.833\ \mathrm{m}, which is behind the tail at x=0x'=0. It too shuts in empty space.

The point. Every frame-independent question gets the same answer in both frames. Did a door hit the pole? Did the nose pass the exit before or after the exit door shut? Only "were the doors shut simultaneously?" differs, and that was never a fact about the world. If you insist on a barn that genuinely traps the pole, you must shut the doors and keep them shut. Then the runner's frame will report that the nose crumples against the closed exit door while the tail is still outside, which is a real, invariant and rather expensive event that both frames agree on.

Problem 3 · two boosts in different directions leave a rotation behind

Work in the (ct,x,y)(ct,x,y) subspace. (a) Show that a pure boost matrix, in any direction, is symmetric. (b) Compute M=Λy(ϕ2)Λx(ϕ1)M=\Lambda_{y}(\phi_{2})\Lambda_{x}(\phi_{1}) explicitly and show it is not symmetric unless one rapidity vanishes. Conclude that it cannot be a pure boost, so boosts do not form a group. (c) Show that the leftover is a rotation, and estimate its size for small rapidities. Quote (⚑) the exact angle and check your estimate against it.

Solution

(a) Pure boosts are symmetric. A boost along xx is (2.2.48) padded with identity rows, manifestly symmetric. A boost in a general direction n^\hat n is that one conjugated by a rotation RR that takes n^\hat n to x^\hat x: Λn^=RTΛxR\Lambda_{\hat n} = R^{\mathsf T}\Lambda_{x}R (with RR acting trivially on ctct). Then Λn^T=RTΛxTR=RTΛxR=Λn^\Lambda_{\hat n}^{\mathsf T} = R^{\mathsf T}\Lambda_{x}^{\mathsf T}R = R^{\mathsf T}\Lambda_{x}R = \Lambda_{\hat n} ✓. So symmetry is the algebraic signature of "pure boost, no rotation".

(b) The product. With ci=coshϕic_{i}=\cosh\phi_{i}, si=sinhϕis_{i}=\sinh\phi_{i} and rows/columns ordered (ct,x,y)(ct,x,y):

Λx(ϕ1)=(c1s10s1c10001),Λy(ϕ2)=(c20s2010s20c2), \Lambda_{x}(\phi_{1}) = \begin{pmatrix} c_{1} & -s_{1} & 0\\ -s_{1} & c_{1} & 0\\ 0&0&1\end{pmatrix}, \qquad \Lambda_{y}(\phi_{2}) = \begin{pmatrix} c_{2} & 0 & -s_{2}\\ 0&1&0\\ -s_{2} & 0 & c_{2}\end{pmatrix}, M=Λy(ϕ2)Λx(ϕ1)=(c1c2s1c2s2s1c10s2c1s1s2c2). M = \Lambda_{y}(\phi_{2})\,\Lambda_{x}(\phi_{1}) = \begin{pmatrix} c_{1}c_{2} & -s_{1}c_{2} & -s_{2}\\ -s_{1} & c_{1} & 0\\ -s_{2}c_{1} & s_{1}s_{2} & c_{2}\end{pmatrix}.

Compare entries across the diagonal. M01=s1c2M_{01}=-s_{1}c_{2} sits opposite M10=s1M_{10}=-s_{1}, and these agree only if c2=1c_{2}=1, that is ϕ2=0\phi_{2}=0. Likewise M12=0M_{12}=0 sits opposite M21=s1s2M_{21}=s_{1}s_{2}. So MM is not symmetric, hence by (a) not a pure boost. Two boosts in different directions do not compose to a boost. The set of pure boosts is therefore not closed, and is not a group. Only the boosts along a fixed axis are, which is what §3 actually proved.

(c) What is left over. ⚑ Every proper orthochronous Lorentz transformation factors uniquely as M=BRM = B\,R with BB a pure boost and RR a rotation (the polar decomposition; we quote it). Since MM is not symmetric, RIR\neq I: there is a residual rotation, the Wigner rotation, in the xyxy plane.

Its size for small rapidities follows from non-commutativity. Write Λx(ϕ)=exp(ϕKx)\Lambda_{x}(\phi)=\exp(\phi K_{x}) and Λy(ϕ)=exp(ϕKy)\Lambda_{y}(\phi)=\exp(\phi K_{y}) with generators

Kx=(010100000),Ky=(001000100). K_{x} = \begin{pmatrix}0&-1&0\\-1&0&0\\0&0&0\end{pmatrix}, \qquad K_{y}=\begin{pmatrix}0&0&-1\\0&0&0\\-1&0&0\end{pmatrix}.

Multiplying out, [Ky,Kx]=KyKxKxKy[K_{y},K_{x}] = K_{y}K_{x}-K_{x}K_{y} has entries only in the xyxy block:

[Ky,Kx]=(000001010)    Jz, [K_{y},K_{x}] = \begin{pmatrix}0&0&0\\0&0&-1\\0&1&0\end{pmatrix} \;\equiv\; J_{z},

the generator of rotations about zz. ⚑ By the Baker–Campbell–Hausdorff formula (Chapter 6.1), eAeB=exp ⁣(A+B+12[A,B]+)\ee^{A}\ee^{B}=\exp\!\big(A+B+\tfrac12[A,B]+\cdots\big), so

M=exp ⁣(ϕ2Ky+ϕ1Kx+12ϕ1ϕ2Jz+), M = \exp\!\Big(\phi_{2}K_{y}+\phi_{1}K_{x} + \tfrac12\phi_{1}\phi_{2}J_{z}+\cdots\Big),

giving a rotation angle Ω12ϕ1ϕ212β1β2\Omega\approx\tfrac12\phi_{1}\phi_{2}\approx\tfrac12\beta_{1}\beta_{2} at low speed. It is second order in the velocities, which is why it has no Newtonian counterpart and why nobody stumbled on it before 1926.

⚑ The exact angle, quoted, for two perpendicular boosts:

tanΩ=γ1γ2β1β2γ1+γ2. \tan\Omega = \frac{\gamma_{1}\gamma_{2}\,\beta_{1}\beta_{2}}{\gamma_{1}+\gamma_{2}}.

Check the low-speed limit. With γi1\gamma_{i}\to1 and tanΩΩ\tan\Omega\to\Omega we get Ωβ1β2/2\Omega\to\beta_{1}\beta_{2}/2 ✓, matching the commutator estimate. Numerically, for β1=β2=0.6\beta_{1}=\beta_{2}=0.6 the exact formula gives Ω=0.2213 rad=12.7\Omega=0.2213\ \mathrm{rad}=12.7^{\circ}, and the two leading-order estimates bracket it. 12β1β2=0.1800\tfrac12\beta_{1}\beta_{2}=0.1800 comes in from below and 12ϕ1ϕ2=0.2402\tfrac12\phi_{1}\phi_{2}=0.2402 from above, since β<ϕ\beta\lt\phi for β>0\beta\gt0. They are the same estimate to leading order and differ only at the next one. For β1=β2=0.01\beta_{1}=\beta_{2}=0.01 the exact angle is 5.00025×1055.00025\times10^{-5}, against 5.00000×1055.00000\times10^{-5} and 5.00033×1055.00033\times10^{-5}, which is agreement to five figures.

Why it matters. An electron in an atom is continuously boosted in changing directions by the nuclear Coulomb field. The accumulated Wigner rotations make its spin precess, an effect called Thomas precession, and this contributes a factor of 12\tfrac12 to the spin–orbit coupling in atomic fine structure. Without it, the predicted fine-structure splitting is wrong by a factor of two, and the discrepancy stood unexplained for two years. A pure kinematic effect, with no force involved at all, showing up in a spectrometer.

Problem 4 · a rocket at constant proper acceleration

A rocket accelerates so that its crew feel a constant aa. That is, its acceleration measured in its own instantaneous rest frame is always aa. (a) Show that its rapidity grows linearly in proper time, ϕ=aτ/c\phi = a\tau/c, and hence β=tanh(aτ/c)\beta=\tanh(a\tau/c). (b) Find the lab-frame time and distance as functions of τ\tau. (c) With a=g=9.81 ms2a=g=9.81\ \mathrm{m\,s^{-2}}, compute β\beta, the lab time and the distance after 11 year of proper time, and then the proper time needed to reach the galactic centre, 2600026\,000 light-years away. Comment.

Solution

(a) The rapidity grows linearly. This is where §6 pays for itself. In a proper-time interval dτ\dd\tau, the rocket's instantaneous rest frame sees it acquire a velocity dv=adτ\dd v = a\,\dd\tau, hence a rapidity increment

dϕ=artanh ⁣(adτc)=adτc+O ⁣(dτ3). \dd\phi = \operatorname{artanh}\!\left(\frac{a\,\dd\tau}{c}\right) = \frac{a\,\dd\tau}{c} + O\!\big(\dd\tau^{3}\big).

Because rapidities add ((2.2.50)), these increments accumulate directly. No composition formula is needed, and needing one is exactly what would have made the velocity version painful. Integrating from rest,

ϕ(τ)=aτc,β(τ)=tanh ⁣aτc,γ(τ)=cosh ⁣aτc. \phi(\tau) = \frac{a\tau}{c}, \qquad \beta(\tau) = \tanh\!\frac{a\tau}{c}, \qquad \gamma(\tau)=\cosh\!\frac{a\tau}{c}.

(b) Lab time and distance. By (2.2.31), dt=γdτ=cosh(aτ/c)dτ\dd t = \gamma\,\dd\tau = \cosh(a\tau/c)\,\dd\tau, so

t(τ)=casinh ⁣aτc. t(\tau) = \frac{c}{a}\sinh\!\frac{a\tau}{c}.

And dx=vdt=ctanh(aτ/c)cosh(aτ/c)dτ=csinh(aτ/c)dτ\dd x = v\,\dd t = c\tanh(a\tau/c)\cdot\cosh(a\tau/c)\,\dd\tau = c\sinh(a\tau/c)\,\dd\tau, so

x(τ)=c2a[cosh ⁣aτc1]. x(\tau) = \frac{c^{2}}{a}\left[\cosh\!\frac{a\tau}{c} - 1\right].

Eliminating τ\tau with cosh2sinh2=1\cosh^{2}-\sinh^{2}=1 gives (x+c2/a)2c2t2=(c2/a)2\left(x+c^{2}/a\right)^{2}-c^{2}t^{2}=\left(c^{2}/a\right)^{2}, so the worldline is a hyperbola. Chapter 2.3 will show that this is the spacetime analogue of a circle, the curve of constant "curvature". Chapter 3.1 will note that its crew experience something remarkably like a gravitational field, complete with a horizon behind them.

(c) Numbers. The natural scale is

cg=2.998×1089.81=3.06×107 s=0.969 yr,c2g=0.969 ly. \frac{c}{g} = \frac{2.998\times10^{8}}{9.81} = 3.06\times10^{7}\ \mathrm{s} = 0.969\ \mathrm{yr}, \qquad \frac{c^{2}}{g} = 0.969\ \mathrm{ly}.

A pleasant accident of our units: one year at one gee is very nearly one rapidity unit.

proper time τ\tauϕ=gτ/c\phi = g\tau/cβ\betaγ\gammalab time ttdistance xx
1 yr1\ \mathrm{yr}1.0321.0320.77480.77481.5821.5821.187 yr1.187\ \mathrm{yr}0.564 ly0.564\ \mathrm{ly}
2 yr2\ \mathrm{yr}2.0652.0650.96830.96834.0054.0053.756 yr3.756\ \mathrm{yr}2.911 ly2.911\ \mathrm{ly}
5 yr5\ \mathrm{yr}5.1615.1610.9999340.99993487.287.284.5 yr84.5\ \mathrm{yr}83.5 ly83.5\ \mathrm{ly}
10 yr10\ \mathrm{yr}10.3210.320.99999999800.99999999801.52×1041.52\times10^{4}1.47×104 yr1.47\times10^{4}\ \mathrm{yr}1.47×104 ly1.47\times10^{4}\ \mathrm{ly}

For the galactic centre, invert the distance formula: coshϕ=1+xa/c2=1+26000/0.969=2.684×104\cosh\phi = 1 + xa/c^{2} = 1+26\,000/0.969 = 2.684\times10^{4}, so ϕ=10.89\phi = 10.89 and

τ=cϕg=0.9687×10.891=10.55 years of crew time. \tau = \frac{c\phi}{g} = 0.9687\times10.891 = 10.55\ \text{years of crew time}.

(Lab time: 2600126\,001 years. The Galaxy does not get a discount.)

Comment. Two things are worth extracting. First, the crew never detect anything odd. Their accelerometer reads a constant gg forever, their rapidity climbs without limit at 1.031.03 per year, and at no point do they approach a barrier. It is only the velocity that saturates, and velocity is the badly chosen coordinate. This is (2.2.46) made visceral. Second, the crew-time cost of a journey grows only logarithmically with distance once you are relativistic, since ϕln(2ax/c2)\phi\approx\ln(2ax/c^{2}) for large xx. The Andromeda galaxy, 2.52.5 million light-years away, costs 14.9714.97 years of crew time against 10.5510.55 for the galactic centre. The Universe is traversable within a lifetime and simultaneously unreachable, because you can never come home. Chapter 2.5 supplies the reason nobody does it: the fuel.

The brick you just laid

You derived (2.2.19), which says ct=γ(ctβx)ct'=\gamma(ct-\beta x), x=γ(xβct)x'=\gamma(x-\beta ct), y=yy'=y and z=zz'=z, from the two postulates plus homogeneity, isotropy and reciprocity, with every step forced and every assumption named. You showed it reduces to Galileo as cc\to\infty, that boosts along an axis compose into boosts, that the composition is addition in rapidity and not in velocity, and that the quantity c2t2x2y2z2c^{2}t^{2}-x^{2}-y^{2}-z^{2} is invariant even though you only ever demanded it for light. Simultaneity, dilation and contraction came out as corollaries in that order, and the two classical paradoxes turned out to be one paradox with two costumes.

Where this gets spent. Chapter 2.3 takes the invariant interval (2.2.21) and the hyperbolic form (2.2.48) and shows that together they define a geometry. That geometry is spacetime, in which a boost is a rotation and proper time is arc length. Chapter 2.4 takes the matrix (2.2.24), calls it Λμν\Lambda^{\mu}{}_{\nu}, and defines a tensor as anything that transforms with it. The rapidity form and the group property of §3 and §6 are quoted there verbatim. Chapter 2.5 applies the same transformation to momentum and energy instead of position and time, and E=mc2E=mc^{2} falls out of it. The γ\gamma you derived here is the γ\gamma in E=γmc2E=\gamma mc^{2}. Chapter 2.6 shows that the electric and magnetic fields transform into each other under exactly this boost, which retroactively explains Chapter 2.1's entire crisis. Chapter 3.1 keeps the transformation but only locally, which is the whole idea of general relativity. And Chapter 6.1 returns to the rapidity ϕ\phi and names what it has been all along: the parameter of a one-parameter Lie group, with the boost generator KK as its Lie-algebra element.