Part II · Special Relativity — Chapter 2.2
The Lorentz Transformation, Derived
Two sentences go in. One transformation comes out, and there is no second candidate. Everything else in this chapter is arithmetic.
Chapter 2.1 ended holding two postulates and a wreck. The wreck is the Galilean transformation , . Maxwell's equations change form under it. The ether frame it demands cannot be found. And every attempt to save it died in the laboratory, whether that attempt was a dragged ether, an emission theory, or contraction by fiat. The two postulates are what Einstein put in its place, and Chapter 2.1 deliberately stopped there.
This chapter asks the only question left: what transformation between inertial coordinates could satisfy both postulates? The answer is not "here is a clever guess that works." The answer is that the transformation is forced, once you also write down three structural assumptions that everyone makes silently. Those three are homogeneity, isotropy and reciprocity. The general form leaves one free function of , and the second postulate together with reciprocity pins that function down completely. Nobody is choosing anything.
Then we spend it. Simultaneity, time dilation and length contraction all arrive as corollaries, and we derive them in that order for a reason. The second and third are consequences of the first. The standard presentation does them backwards, and that is why so many people never recover. After that we look at how velocities combine, why they seem not to add, and the variable in which they do add. The chapter closes by taking the two famous paradoxes apart and showing that they are the same paradox wearing different clothes.
One warning about what this chapter is not. It is not geometry. We will produce an invariant combination almost by accident in §2, and then mostly leave it alone. Turning that invariant into a genuine geometry, meaning a spacetime with a notion of length, angle and straightness, is Chapter 2.3's job. That chapter is a better one for having this algebra already in hand.
Tools you'll need — Chapter 2.1 §7: the two postulates, which are the entire input. Chapter 0.4: linear maps, matrices, matrix multiplication, and what it means for a set of maps to be closed under composition, which is what §3 is, applied. Chapter 0.3: Taylor and binomial expansions, used every time we ask "and what does this look like at ordinary speeds?" Chapter 0.1 §3: the derivative as the coefficient of a linear approximation. The linearity argument in §1 is that idea run in reverse, and §5 differentiates a coordinate transformation.
1 · What the postulates constrain
Let's start by writing down exactly what we are given. Then we write down exactly what we are quietly assuming on top of it, which is the part usually skipped.
1.1 · The input
P1 (principle of relativity). The laws of physics take the same form in all inertial frames. No experiment inside a uniformly moving laboratory can reveal its velocity.
P2 (invariance of ). Light propagates in vacuum with speed , the same in all inertial frames, independent of the motion of the source.
Now fix the geometry once and use it for the whole chapter. There are two inertial frames and , with parallel axes. moves in the direction of with speed , and the origins coincide at the moment both clocks read zero. This is the standard configuration, and it costs nothing, because any other arrangement is this one composed with a rotation and a shift of origin. As throughout Part II, we write the relative speed in units of :
Nothing is assumed yet about the range of . That will emerge in §2.3, where a larger value makes the transformation complex, and it will be proved unreachable in §5.
What we want are the four functions that take an event's coordinates in to its coordinates in . An event is a point of spacetime: a flash, a collision, a clock tick. Both observers are describing the same one. Nothing physical changes when we relabel. The question is only what the labels are.
1.2 · The three assumptions nobody states
The postulates alone do not determine the transformation. They are constraints on a class of maps, and you have to say which class. Three assumptions do that work, and it is worth being blunt: these are physical assumptions, not logical necessities. Each is testable, each has been tested, and each could in principle have come out otherwise. This is where the real content of the chapter enters, and pretending otherwise is how the derivation acquires its undeserved reputation for being a magic trick.
(H) Homogeneity of space and time
No place is special and no moment is special. Concretely: the rule relating -coordinates to -coordinates is the same rule everywhere and everywhen. If you and I are both inertial observers, the dictionary between our coordinates cannot depend on where in the room we set up, or on what year it is.
This forces the transformation to be linear. Here is the argument, and it is better than the hand-wave usually offered.
Let be the map from -coordinates to -coordinates, acting on events. Take any two events and and let be their coordinate separation in . Homogeneity says the separation the primed observer assigns, , can depend on but not on where the pair sits. If it did depend on that, a rod's transformed length would tell you your absolute position, and space would have a marked spot. So there is a single function with
We want to learn what kind of function can be, and the way to find out is to chain two separations together. Go from to to , and add the two pieces:
That is Cauchy's functional equation, and for a continuous its only solutions are the linear ones, for a constant matrix . Since the origins coincide we also have , and therefore . The transformation is a constant matrix acting on . The entries of may depend on , and they had better, but they may not depend on the event.
Grind box — Cauchy's equation, and the loophole that "straight lines stay straight" leaves open
Cauchy. ⚑ The theorem is: if satisfies and is continuous at even one point, then is linear. Here is the sketch in one dimension. Additivity gives for integer , hence , so for every rational . Continuity then upgrades rational to real. (Without continuity the theorem is false, because there exist monstrous non-measurable additive functions, built with the axiom of choice. Nobody is proposing that spacetime is coordinatised by one.)
The other argument, and why it is not quite enough. The usual alternative runs like this. A free particle moves uniformly in a straight line in , by Newton's first law, and by P1 it must do so in too. So the map takes straight worldlines to straight worldlines, and therefore it is linear. The conclusion is right, but the reasoning has a gap. The maps of carrying straight lines to straight lines are not just the affine ones. They are the projective ones,
which are perfectly good line-preserving maps with a nonconstant denominator. ⚑ (That every line-preserving bijection is of this form is the fundamental theorem of projective geometry, which we quote.) Such a map is not linear, and it does something revealing. It blows up on the hypersurface . That is a preferred set of events, a marked place and time in the universe. Homogeneity is exactly the assumption that kills the denominator, forcing and returning us to affine, hence linear once the origins coincide.
This is not an idle technicality. Some theories keep P1 and P2 but drop strict homogeneity, among them "de Sitter relativity" and the Fock–Lorentz transformations. They live precisely in that discarded denominator, and they are constrained by experiment rather than by logic. The assumption is doing work, and now you know which work.
(I) Isotropy of space
No direction is special either. Two consequences follow, and we need both.
First: and . Directions perpendicular to the boost are untouched. Suppose instead for some factor . Boosting by afterwards must return us to , so . But isotropy says the transverse factor cannot know the sign of the boost. Reflect the -axis, and a boost becomes a boost while is untouched, so . Put the two together and , and since and is continuous, .
Second: and cannot depend on or . By linearity, . Rotate the whole apparatus by about the -axis. This sends , and leaves the boost configuration exactly as it was, so it cannot change . That forces . The same argument works for .
Grind box — the transverse argument done physically, with rods and scratches
The algebraic argument above is airtight but feels like bookkeeping. Here is the physical version, which is the one to remember, because it shows that a transverse contraction would produce a flat contradiction rather than merely a strange world.
Two identical rods, each of proper length one metre, both aligned along , with their lower ends riding along the -axis. One is at rest in , the other at rest in , and they slide past each other. Fix a sharp scriber at the top of each rod.
Suppose finds the moving rod short, . Then as they pass, the -rod's scriber passes above the tip of the -rod, and the -rod's scriber scratches the -rod's side below its tip. Whether a scratch appears on a rod is a local coincidence of matter with matter. It is a single event, and every observer agrees it happened.
Now look from . By P1, 's situation is the mirror of 's. sees moving at , and by isotropy the physics cannot distinguish from here. So must equally find the other rod short, and predicts the scratches the other way round. Both frames cannot be right about which rod carries a scratch, because that is one event and not two. The only escape is . The tips meet exactly, nobody scratches anybody, and both frames agree.
Note carefully what did not happen: no such contradiction arises along . There, "which rod is shorter" is settled by measurements at two separate places, and §4.3 will show that two separate places is precisely where frames stop agreeing about what "at the same time" means. Transverse: one event, no room to disagree. Longitudinal: two events, all the room in the world.
(R) Reciprocity
If moves at as measured in , then moves at as measured in . Equivalently: the transformation from back to is the same transformation with .
This sounds like a tautology and it is not. It is the statement that the two frames stand in a symmetric relation. Neither is privileged in a way that would let measure receding at one rate while measures receding at another. It follows from P1 together with isotropy, and we could derive it. We will instead assume it and flag it, because assuming it makes visible exactly how much is being asked.
Notice what has and has not been assumed. We have not assumed . We have not assumed that lengths are invariant, that simultaneity is absolute, that velocities add, or that there is a universal "now". Every one of those was baked into the Galilean transformation without comment for three hundred years, and every one of them is about to fail.
What we have assumed is that spacetime is uniform and directionless, and that the relation between inertial observers is symmetric. Those three, plus the two postulates, are the complete input. If the answer surprises you, the surprise has to be located in one of those five statements. There is nowhere else for it to hide. That is the point of listing them.
Two postulates cannot determine anything on their own, because a constraint is always a constraint on some class of candidates and nobody has yet said which class. Three further assumptions do that work, and the derivation's reputation for being a conjuring trick comes entirely from leaving them unspoken. Each is a physical claim, each is testable, and each could have come out otherwise.
The first is that no place and no moment is special, so the dictionary between two observers' coordinates is the same wherever and whenever they set up. That forces the dictionary to be linear, by an argument better than the usual one. The familiar route, that free particles travel in straight lines for everybody so straight lines must go to straight lines, leaves a loophole, because the maps preserving straightness include ones carrying a denominator, and a denominator that vanishes somewhere marks out a special place and time in the universe.
The second is that no direction is special, which leaves the two transverse directions alone and keeps the sideways coordinates out of the other two equations. The third is that the relation between the observers is symmetric, each receding from the other at one rate. Those three, with the two postulates, are the complete input; if what comes out is surprising, the surprise is living in one of five statements and has nowhere else to hide.
2 · The derivation, in full
By §1 the transformation is linear, the transverse coordinates are untouched, and nothing depends on or . So the whole problem lives in the plane, and it reads
where are constants, meaning functions of alone. That is four unknowns, and we have three conditions to impose. Three conditions for four unknowns sounds like one short, but it is not: condition (ii) will hand us two equations rather than one, because a light pulse can be sent in either direction along and both directions have to work. Let's impose them one at a time.
2.1 · Condition (i): the origin of moves at
This is what " moves at " means, made operational. The spatial origin of is the locus , and by construction it travels along in . We want the relation between and that this forces, so substitute and into (2.2.4):
Feed that value of back into the first line of (2.2.4), and the spatial part of the transformation is already forced into the familiar shape
Notice how little that cost. And notice what it did not do. It says nothing whatsoever about . The Galilean transformation is the special case , , . Every departure from Galileo is now compressed into three numbers.
2.2 · Condition (ii): a light pulse goes at in both frames
This is Postulate 2, and it is the crucial input. It is the only place where the physics of light enters. At the moment the origins coincide (, ) a flash is emitted from the common origin. Consider the part of it travelling in the direction. In it obeys . In , by P2, it must obey , with the same and not .
Our aim is to turn that demand into equations relating , and . So put into (2.2.6) and into , and see what each gives:
Both of those are proportional to , so imposing will divide the time variable out and leave a relation among the three constants alone. Do it:
That is one equation for two unknowns, so we need a second, and the flash gives us one for free. Now do the same for the part of it going in the direction, , which must obey :
Two linear equations, two unknowns and in terms of . To isolate we want the terms gone, so add (2.2.8) and (2.2.9). They cancel, and
That is one of the two. For the other we want the terms gone instead, so subtract the same pair of equations rather than adding them:
Three of the four unknowns are gone, and everything is now expressed through :
Let's pause on (2.2.12), because two things have already happened that no amount of staring at Galileo would have suggested.
The time equation has picked up a position term. depends on . Two events at the same but different have different . Simultaneity has already broken, before we know what is, and it broke because of the second postulate and nothing else. Everything in §4 is contained in the term .
The same coefficient appears in both. That was not imposed. It fell out of adding the two light conditions. The transformation is symmetric between and , and the symmetry is easier to see if we measure time in metres, which means using as the coordinate rather than , as Part II does throughout. So multiply the second equation by :
Now and enter on exactly the same footing. That symmetry is the first hint that time is about to become a coordinate like the others, which is the whole thesis of Chapter 2.3.
2.3 · Condition (iii): reciprocity fixes
One unknown left. Reciprocity says the inverse of the boost by is the boost by . So apply the boost with parameter and then the boost with parameter , and demand that we get back what we started with. Writing and :
We started with and we must end with , whatever is, so the whole prefactor in front of it has to be one:
(The same manipulation on gives the same condition, and the grind box does it.) That is one equation in two unknowns, and , unless happens to be an even function of . Isotropy delivers exactly that, and here is how. Reflect the -axis by setting and . Then (2.2.13) becomes
which is a boost with parameter and coefficient . Space has no handedness, so the reflected description is as good as the original, and the coefficient for parameter must be what it always is: . Feed that into (2.2.15):
Take the positive root, and here is why we may. , because no boost is no change. is continuous. And never vanishes, so it cannot cross zero to reach the negative branch. The surviving root is the one quantity this whole chapter turns on, so give it a name:
Substitute that name back into the two lines of (2.2.13), and nothing is left undetermined. The whole quarry of the chapter is now on the page.
That is the Lorentz transformation for a boost along , in the form this book uses everywhere. Three things are worth noting straight away. First, always. Second, as . Third, would make imaginary, which is a first cheap indication that is a barrier. §5 makes that indication respectable.
Grind box — the inverse, checked both ways, and the time equation of the reciprocity step
The time half of (2.2.14). Same manoeuvre:
which gives condition (2.2.15) again. Consistent, as it must be: two equations, one condition.
Inverting (2.2.19) directly. This time do not appeal to reciprocity at all. Just solve. Multiply the first equation by and the second by , and add:
since is (2.2.18) squared and rearranged. The same trick with the roles swapped gives . Hence
which is (2.2.19) with and primes swapped. So reciprocity, which we assumed, is also a theorem of the result. That is a genuine consistency check rather than a circularity, because the algebra above never used it.
Numbers worth memorising. is stubbornly close to until is large:
Half the speed of light buys you a 15% effect. This is exactly why nineteenth-century physics could be built without noticing. At , the Earth's orbital speed, .
2.4 · A consistency check we did not pay for
Postulate 2 was imposed only along the -axis, one pulse forward and one back. But light goes in all directions, and a flash from the origin makes an expanding sphere, in . If the derivation is any good, that sphere had better still be a sphere in , expanding at the same . We never asked for that. Let's check it.
The quantity to watch is the one that measures the radius of the sphere against the time it has had to expand. So compute the combination from (2.2.19):
The cross terms cancel identically and finishes it. Since and , those two coordinates come along untouched, and we have proved something considerably stronger than we asked for:
So the set , which is the expanding light sphere, maps to itself. Light emitted from the common origin is a sphere of radius in and a sphere of radius in , each centred on that observer's own origin, and the two origins are moving apart. Both are right. If that sounds impossible, hold the thought until §4.1. The two observers disagree about which points of the sphere are "simultaneous", and that disagreement is exactly large enough.
(2.2.21) was not an input. We demanded invariance of the light speed, which is a statement about one particular family of events, the ones with . What came back is invariance of the whole quantity, for every pair of events, whatever its value. That is an enormous strengthening, and it is the single most important line in this chapter.
The combination is called the interval. Everybody computes the same number for it. It plays the role that plays for rotations in ordinary space, the thing that does not care how you turned your head. Chapter 2.3 takes that analogy seriously and builds a geometry out of it, and Chapter 2.4 turns the invariance into the algebraic condition that defines the Lorentz group. Everything downstream of here rests on one minus sign.
2.5 · The low-speed limit: Galileo was not wrong, he was first-order
A new theory that contradicted the old one everywhere would be a theory about a different universe. The correct relationship is subtler. Relativity must reduce to Galileo where Galileo was tested. So let's expand and see.
What we need first is for small , and the binomial series (Chapter 0.3) supplies it:
so . Substituting that into (2.2.19) and keeping terms through first order in :
Now let at fixed , and , which is what "ordinary speeds" means. The second term of vanishes too, and what is left is and . That is the Galilean transformation. Not approximately, and not in spirit. Exactly, as a limit.
That is the honest statement of the relationship, and it is worth saying what it means. Galilean relativity is not a mistake that got corrected. It is the leading term of an expansion, and the reason it looked like the whole truth for three centuries is that nobody could measure the next term. The same pattern will recur. Newtonian gravity is the leading term of general relativity (Chapter 3.6), and classical mechanics is a limit of quantum mechanics (Chapter 4.10). Old theories almost never die. They get demoted to leading order.
But look closely at (2.2.23), because the two corrections do not enter at the same order, and that asymmetry explains the entire history of the subject:
- Lengths and time rates are wrong by a factor , that is by , which is second order. At that is , a part in two hundred million.
- Simultaneity is wrong by , which is first order in , but multiplied by a distance. At and this is , which is small. But scale up to a satellite orbit, , and it becomes . That is of light travel, and it would wreck a navigation system in an afternoon.
So the largest relativistic effect at low speed is the one about clocks at different places, and it is the one nobody notices, because noticing it requires two clocks far apart and a way to compare them. That is exactly the experiment nineteenth-century physics could not do and twenty-first-century engineering does continuously.
Everything in the derivation is bookkeeping except one step, and that step is the only place where any physics of light enters. Demanding that the moving origin move at the stated rate fixes the spatial equation up to one unknown factor, and costs almost nothing, since the old transformation is the case where that factor is one.
The light condition then goes in, once for a pulse each way along the axis. Adding the two resulting equations forces the same unknown factor into the time equation as sits in the spatial one; subtracting them forces the time equation to acquire a term proportional to position. That term is the end of simultaneity, and it arrives before the remaining factor is known. Symmetry between the observers pins that factor down, and no second candidate appears.
The bonus is larger than the result. Invariance was demanded only of the events light connects, and what comes back is a combination of the time and space separations that is the same for every pair, whatever its value. Quantities invariant for reasons nobody demanded are worth building a subject on. The old transformation survives as the leading term of an expansion, and its two corrections do not enter at the same order: lengths and rates are wrong at second order, while simultaneity is wrong at first order multiplied by a distance, so the largest low-speed effect concerns distant clocks.
3 · Matrix form and the group property
Chapter 0.4 taught you to see a linear map as a matrix, and composition as matrix multiplication. Let's do that here. Write the four coordinates as , with time measured in metres so that all four carry the same units. Then (2.2.19) is
The index placement puts the upper index on the row and the lower one on the column, and it is not decoration. Chapter 2.4 explains exactly what promise it encodes, and derives this same matrix as a Jacobian . For now it is a matrix.
Two features are worth naming immediately. It is symmetric, which we will use in Problem 3. And its determinant is , so boosts preserve four-dimensional volume in spacetime. That fact matters in statistical mechanics and in Chapter 5.8.
3.1 · Two boosts make a boost
Now the question that decides whether we have a theory or a formula. Boost from to at , then boost from to at , along the same axis. Is the composite a boost? To find out we multiply the two matrices, and only the upper-left block does anything:
We want that product to look like a single boost matrix, which has ones on the diagonal. So pull out the common factor , and give the ratio that survives on the off-diagonal a name of its own:
Then the product is times the matrix , and it is a boost provided that prefactor equals . It does equal that, and the grind box does the algebra, so
Grind box — the prefactor really is , and one identity that keeps reappearing
Compute from (2.2.26):
Expand the numerator and watch the cross terms die:
So
hence , which is the prefactor. ✓
Keep that identity. The factorisation , along with its sibling , will do the work again in §5, where it proves that you cannot reach by combining sub- speeds. It appears once more in §6, where it turns out to be the hyperbolic addition formula in disguise. Three appearances of one factorisation is usually a sign that something structural is going on, and it is.
The other group axioms. is the identity. is the inverse, since (2.2.26) with gives . Associativity is free, because matrix multiplication is associative. So the boosts along a fixed axis form a one-parameter group, and §6 finds the parameter that makes it look like one.
This matters more than it looks. Closure is not automatic. It is a statement that the set of transformations is a self-consistent family, that "the frames of physics" is a coherent notion, and that no chain of boosts can ever take you outside the theory. Had the composition of two boosts failed to be a boost, the whole scheme would have been incoherent.
But look hard at (2.2.26). The composed parameter is not . Boost by and then by again and you get
not . Two half-light-speed boosts give eight tenths, and no finite chain of sub- boosts will ever reach . §5 proves that properly and §6 explains it. Velocity is the wrong variable to be adding.
Everything above was for boosts along one fixed axis. Boosts along different axes do not compose into a pure boost. The product picks up a spatial rotation, called the Wigner rotation. Problem 3 derives that the effect must exist, in two lines, from the fact that (2.2.24) is a symmetric matrix. The physical consequence is called Thomas precession. It is a measurable contribution to atomic fine structure, and it is a genuinely relativistic effect with no Newtonian ancestor at all.
Whether a family of transformations closes on itself decides whether you have a theory or a formula. Compose two boosts along one axis and the result is another boost, with a parameter emphatically not the sum of the two, since half the speed of light followed by half the speed of light delivers eight tenths. Closure is neither automatic nor decorative, because had it failed there would be no coherent notion of the frames of physics, and some chain of changes of viewpoint would have carried you outside the theory.
Two smaller facts fall out along the way. The transformation has determinant one, so it preserves four-dimensional volume, which is the determinant doing the job the toolkit gave it. And the whole apparatus is now a matrix acting on four numbers, so composing changes of frame has become multiplying matrices and nothing more.
One caveat is worth flagging because it has no ancestor in older physics. Boosts along a single axis behave; boosts along different axes do not compose into a boost at all, and the product leaves a spatial rotation behind. That rotation is measurable, contributing to the fine structure of atomic spectra, and nothing in Newtonian mechanics predicts anything of the kind. The composition rule also says plainly that velocity is the wrong quantity to be adding.
4 · The three consequences
Now we cash it in. Three results, and every one of them is a two-line substitution into (2.2.19). The difficulty in this material is never the algebra. It is knowing which two events you are transforming. So each subsection begins by stating the operational question before any symbol moves: who measures what, with which apparatus, and using which pair of events.
4.1 · Relativity of simultaneity
The operational question. Two firecrackers explode. An observer in , using clocks that have been synchronised throughout , records both explosions at the same reading. What do the clocks of record, given that they have been synchronised throughout ?
Take two events with and in . We want their time separation in , so apply (2.2.19) to the separation itself. That is legitimate because the transformation is linear, which is what §1 bought us:
Events simultaneous in are not simultaneous in , unless they happen at the same place. The sign matters and is worth reading off. Suppose , so that the second event is further along the direction of 's motion. Then , so in the leading event happened first. The frame that is moving forward sees the forward event as earlier.
A concrete case
A train of proper length passes a platform at , so . Lightning strikes both ends of the train, simultaneously as judged on the platform.
In the platform frame the train is contracted (§4.3) to , so the two strike events have and . Putting those numbers into (2.2.29) gives, on the train,
and ✓, which is the train's own length, as it must be. So on the train, the front is struck before the rear. Not "appears to be". It is, in every sense that the train's own clocks can express.
Check it without the transformation, as a sanity test. Both flashes travel at in the platform frame and meet at the midpoint of the two strike locations. Meanwhile the train's midpoint is moving toward the front strike, so it runs into the front flash first. Now argue in the train frame. The train observer sits midway between the two scorch marks on the train, both flashes travel to him at over equal distances, and the front flash arrives first. Equal distances, equal speeds, different arrival times, and therefore different emission times. The front strike happened first. Two independent routes, one answer. ✓
Time dilation and length contraction are not independent phenomena sitting alongside the failure of simultaneity. They are consequences of it. Every apparent paradox in the rest of this chapter, and in the rest of the subject, is resolved by finding the place where somebody assumed two distant events had a frame-independent time order.
Readers who meet dilation and contraction first almost never recover, because they acquire a picture in which clocks and rulers are mysteriously defective while "now" remains universal. It does not. There is no universal "now". There is a family of them, one per state of motion, and they cut through spacetime at different angles. Hold that and nothing that follows can confuse you.
4.2 · Time dilation
The operational question. A single clock rides along with and ticks twice. The elapsed time it reads between the ticks is . That quantity is called the clock's proper time, and it is what that particular clock actually displays. What is the elapsed time between those same two ticks according to , where the reading must be taken off two different synchronised clocks, one at each tick's location?
The clock is at rest in , so the two ticks have and . We want , so use the inverse transformation from the grind box in §2.3:
Since , more time elapses in than the moving clock records. A moving clock runs slow, by exactly . Note the asymmetry in the apparatus, which is where all the content is: one clock on one side, two clocks on the other.
The symmetric version, and why it is not a contradiction
Run the identical argument the other way. A clock at rest in ticks twice, , . Then (2.2.19) gives directly , so . Each observer finds the other's clock slow, by the same factor. Both statements are true. Both are derived from the same transformation. There is no contradiction, and here is the accounting that shows it.
Set , . In there are two clocks, at and at light-seconds, synchronised in . A traveller carrying clock flies from to .
Told by : the trip takes . The traveller's clock, moving, runs slow, and reads on arrival. reads .
Told by : the traveller is at rest and the whole apparatus streams past at . The distance from to is contracted to light-seconds, so it takes . That matches the traveller's own clock, since it is his own clock. During those seconds the clocks, being the moving ones now, tick only .
Contradiction? must read when the traveller arrives. That is a local coincidence, two objects at the same place, and no one can disagree about it. Yet says the clocks advanced only . The missing is exactly the simultaneity offset. At the moment the traveller passes , which we may call , the clocks that regards as simultaneous satisfy . So , sitting at light-seconds, already reads
In the traveller's reckoning, was never synchronised with in the first place. It was ahead before the trip began. Add that head start to what the clock ticked during the trip, and the books balance:
Both frames agree on every local coincidence, meaning on which clock reads what when it is next to which other clock. They disagree only about the synchronisation of distant clocks, which is not an observable but a convention imposed frame by frame. The symmetry of time dilation is consistent precisely because simultaneity fails, and the failure is exactly the right size. It could not have been any other size, because (2.2.29) and (2.2.31) came out of the same matrix.
"Each sees the other's clock run slow" sounds like and . It is not, because the two statements are not comparisons of the same things. Each is a comparison of one clock against a pair of clocks synchronised in the other frame, and the two frames use different pairs, synchronised differently. Strip out the words and write down which events are being compared, and the appearance of contradiction evaporates every single time.
Here is the test to apply whenever you feel the vertigo. Ask whether the disputed claim is about a local coincidence. "Clock reads while sitting next to , which reads " is a coincidence, two objects at one place making one event, and every frame agrees on it. "Clock reads while , twelve light-seconds away, reads " is not a coincidence. It invokes a simultaneity convention, and frames are entitled to differ. All the paradoxes are manufactured by smuggling the second kind of statement in wearing the clothes of the first.
4.3 · Length contraction
The operational question, and it is the whole difficulty. A rod lies at rest in along the -axis. Its ends are at and , always, so its length in is unambiguous: , called the proper length. Now ask for the rod's length. In the rod is moving, so "where its ends are" is a function of time, and to get a length you must locate both ends at the same moment. At the same moment according to whom? According to , because that is what it means for to measure a length.
So a length measurement is a pair of events (locate left end, locate right end) constrained by in the measuring frame. The moment you write that down, the answer is forced, because the constraint is a simultaneity constraint and §4.1 already told you those are frame-dependent.
Apply (2.2.19) to the pair. In we have , which is what we want, and , which is the definition of measuring. In we have , since the rod is at rest there and it does not matter when you look. Then
which is the answer, but written the wrong way round: it gives the proper length in terms of the measured one. Divide through by to put the measured length on the left, where we want it:
A moving rod is short, by , along its direction of motion only, since and mean nothing happens across it. And Chapter 2.1's loose end is now tied. This is precisely the that FitzGerald and Lorentz had to postulate to explain Michelson–Morley. Here it is a theorem.
It follows from simultaneity — watch
The derivation used , which by (2.2.29) means the two measurement events are not simultaneous in :
From the rod's own frame, then, located the right end first and the left end later. Since the rod is not going anywhere in , that sloppiness costs nothing. The ends are where they always are, which is why still gets a definite answer. But it is exactly why the answer is not . To see that the choice of event pair is doing all the work, let's redo the calculation with a pair that is simultaneous in instead:
Same rod, same frames, different pair of events, and the answer flips from to . The contraction is not a property of the rod. It is a property of the measurement, and specifically of whose simultaneity slices the measurement uses. Nothing was squeezed. Two observers sliced the same four-dimensional object at different angles and got different cross-sections, exactly as two people slicing the same sausage at different angles get ellipses of different lengths. Chapter 2.3 makes that picture precise. Here the algebra already contains it.
Length contraction is a statement about a simultaneous measurement. It is not a statement about what a camera records, and the two differ, because light from different parts of an object takes different times to reach the lens. A photograph does not capture a simultaneity slice. It captures the events whose light arrives together.
Let's work the difference out for a cube of side flying past at , seen from far away and edge-on. Light from the rear face must set out earlier than light from the front face, earlier by , which is the extra distance, in order to arrive at the camera together. During that the cube moves a distance , so the rear face is displaced sideways in the image by , and you can see round the back of the cube. Meanwhile the front face, measured simultaneously, is contracted to . Now compare with a cube that is not moving at all but has been rotated by an angle . Its front face projects to , and one side face becomes visible with projected width . Match the two:
The photograph of a rapidly moving cube is the photograph of a rotated cube, turned by . ⚑ The general theorem is the Terrell–Penrose result of 1959: a small object of any shape appears rotated rather than contracted, and a sphere always photographs as a perfectly circular disc. The cube calculation above is the whole idea, and the general case is bookkeeping over solid angle.
So contraction is real and it is not visual. If someone shows you a picture of a squashed spaceship, the picture is wrong. The spaceship is shorter when measured. It does not look shorter when photographed.
Muon decay is first-order. A population at rest obeys the survival function with , which is Chapter 0.1's clearance curve with a different label on the constant. Nothing in (2.2.31) touches that law. What it touches is the argument fed into it.
Watched from the ground, a beam of muons moving at has survival
and a survival function of the form is, by definition, an accelerated failure time model with acceleration factor . Special relativity is an AFT model whose covariate is speed and whose factor is . That is not a resemblance, it is the same equation, and the reason relativity picks AFT rather than proportional hazards is the physical content of this whole chapter. Proportional hazards would say that motion alters the muon's internal decay rate, which is a claim about muons. AFT says the muon's clock has been rescaled and the decay rate per tick of that clock is untouched, which is a claim about geometry. It is therefore a claim that every particle obeys identically, whatever it is made of.
For an exponential the two models coincide, since scaling the time axis and scaling the hazard are the same operation there. That is exactly why the cosmic-ray calculation of §8 is easy, and exactly why it is easy to take the wrong moral from it. Had the muon decayed with a Weibull hazard the two readings would have parted company, and only the AFT one would have survived the experiment.
What breaks. is not fitted. It is , forced by §2 with no free parameter and no residual heterogeneity whatever. Two muons at the same speed have exactly the same acceleration factor, so there is no frailty term and nothing left over to model. And the covariate is not a property of the subject at all but of the observer. That is why §8's second telling returns the identical count of survivors at the detector, even though it is done in the muon's own frame, where and the atmosphere is contracted instead. It has to return the same count. Whether a given muon reaches the ground is one event, and the two frames are describing it, not deciding it.
Three consequences come out of one substitution, and the order in which they are met decides whether the subject ever becomes intelligible. The root fact is that two events at the same time but different places for one observer are not at the same time for another, and the two famous effects follow from it rather than accompany it. Readers who meet them the other way round acquire a picture in which clocks and rulers are mysteriously defective while the word now stays universal, and they rarely recover.
It does not stay universal. Each state of motion carries its own family of nows, cutting through spacetime at different angles, which is what the local conservation laws of the last part were built to survive. It is also why each observer finding the other's clock slow is no contradiction: one clock is compared against a pair synchronised in the other frame, the two observers use different pairs, and the offset between them is exactly the size that balances the books.
Contraction turns out to be a property of the measurement rather than of the rod. Locating both ends at one moment is a demand about simultaneity, so different observers use different pairs of events and get different answers, and a pair the rod's own frame calls simultaneous flips the answer the other way. Nothing was squeezed, and nothing was done to the rod.
5 · Velocity addition
An object moves at velocity along in frame . What velocity does assign it? Galileo says , which is the velocity-composition rule of Chapter 2.1 §1.2, and it is exactly what Postulate 2 refuses to accept for light. So let's derive the replacement.
Velocity is a ratio of displacements rather than of coordinates, so what we want is the transformation applied to small displacements. Take differentials of (2.2.19). That is legitimate because the coefficients are constants, so the differentials obey the same linear relations as the coordinates do:
The velocity in is , so divide the first of these by the second. The factors of cancel, which is a small mercy, and a hint that the result depends on less violently than one might fear:
and the last step is only relabelling. Write for the velocity in , and the rule stands in the form this book quotes everywhere:
The dimensionless form is (2.2.26) with a sign flipped, as it should be: composing two boosts and transforming a velocity are the same operation seen from two sides.
5.1 · The postulate reproduces itself
The first thing to ask of a new rule is whether it respects the postulate that produced it. So set and turn the crank:
That holds for any with . This is a real check, not a tautology dressed up. We imposed the invariance of only for a pulse emitted from the origin at along . The transformation that resulted now certifies that any object moving at in moves at in every frame, anywhere, at any time, in either direction. A theory that failed this would have been inconsistent with its own founding assumption.
5.2 · You cannot get to by adding speeds
Claim. If and then .
Proof. First, the denominator never vanishes, since gives . What we want is the sign of , so compute it, using the identity from the §3 grind box in its second form, :
Every factor on the right is strictly positive, so , so .
Read (2.2.42) once more, because it says more than the claim. It says the "distance from ", measured by , is multiplicative under composition. Combine two sub-light speeds and the product of their deficits is what survives. You can multiply small positive numbers together forever and never reach zero. The light barrier is not a wall someone installed. It is the statement that a product of positives is positive.
5.3 · And at ordinary speeds, Galileo
To see what the new rule costs at everyday speeds, take and expand the denominator (Chapter 0.3):
Now put numbers in. Two cars approach each other at each, so and . The Galilean answer is and the correction is
a relative error of one part in . This is why the Galilean rule survived so long unchallenged: on a motorway it is wrong in the fourteenth decimal place.
Nothing above forbids coordinate speeds greater than . What is forbidden is the transport of matter, energy or information faster than . Two examples, both computed.
A spotlight on a distant wall. Rotate a laser at angular rate . The spot on a wall at distance moves at , with no upper bound. Point a laser at the Moon () and spin it at one revolution per second, :
No rule is broken, because the spot is not a thing. The photons landing at one place and the photons landing at the next place travelled independently from the laser, and nothing went sideways. You cannot send a message along the wall this way, because each point of the wall learns only about the laser, never about its neighbours.
A pair of scissors. Two straight edges cross at a small angle , one sliding toward the other at speed . The moving edge is , so the intersection with sits at and travels at , which is unbounded as . With and the crossing point moves at . Again nothing is transported. The intersection is a geometric coincidence rather than an object. And a real pair of scissors would in any case not be rigid, since closing the handles sends a wave down the blade at the speed of sound in steel.
The rule to carry: ask what is being transported. If the answer is "nothing", any speed is allowed.
The rule replacing the addition of velocities earns its place by passing a test nobody arranged for it. Invariance of the light speed was imposed for one pulse leaving one origin in one of two directions, and the rule that resulted now certifies that anything travelling at that speed in one frame travels at it in every frame, anywhere, at any time, in any direction.
The barrier that follows is not a wall anybody installed. Rate a speed by the quantity measuring its distance from the limit, and combining two speeds multiplies the two ratings together, up to a positive factor. Small positive numbers can be multiplied together forever without reaching zero, so no finite chain of ordinary boosts arrives at the limit, and the whole of the light barrier is the observation that a product of positive numbers is positive.
None of this forbids coordinate speeds above the limit, and it is worth being clear which is which. Sweep a laser across the face of the Moon and the spot crosses it many times faster than light, because the photons landing at one place and those landing next door travelled independently and nothing went sideways. Close a pair of shears at a shallow enough angle and the crossing point outruns light for the same reason. Ask what is being transported, and where the answer is nothing, any speed is permitted.
6 · Rapidity — the parameter that behaves
Velocities do not add. Something must, because §3 showed that boosts along an axis form a one-parameter group, and a one-parameter group has an additive parameter almost by definition. Composing two elements should add their labels, the way rotating by and then gives . We chose a bad label, that is all. Let's find the good one.
The place to look is the composition rule (2.2.26) itself,
which is not a formula anyone would invent. But it is a formula everyone has seen. It is the addition theorem for the hyperbolic tangent. So let's define the rapidity to be the thing whose hyperbolic tangent the velocity is:
This is a legitimate change of variable. maps one-to-one onto , which is exactly the allowed range of , and it is smooth with a smooth inverse. Every subluminal velocity has exactly one rapidity, and every real number is somebody's rapidity.
6.1 · What becomes
Before we can rewrite the boost we need in the new variable. Divide the identity by to get . Hence
Both entries of the boost matrix are now single hyperbolic functions, so substitute them into (2.2.24) and keep only the block that does anything:
and the transformation reads , .
6.2 · Rapidities add
The claim to test is that composing two boosts adds their rapidities, so multiply two of these matrices together and see what comes out:
writing and . The hyperbolic addition formulas are and , and they turn that matrix into
Rapidities add. Plainly, exactly, with no correction term. And taking of both sides recovers (2.2.26), since . So §5 and §6 are the same statement written in two variables.
Check the arithmetic of (2.2.28) in the new currency: , twice that is , and ✓. Two boosts of do not give . They give rapidity , which is velocity .
Velocities "fail to add" for the same reason angles-as-slopes fail to add. Compose two rotations and their angles add. Their slopes combine by , which nobody finds shocking, because nobody expected slopes to be the natural parameter. Velocity is the slope of a worldline. Rapidity is its angle. We spent three hundred years adding the slopes.
And the change of variable makes the light barrier trivial rather than mysterious. Rapidity runs over the whole of and adds without limit, so you can boost forever. But saturates. gives , gives , and gives . The barrier is not a barrier at all. It is the horizontal asymptote of , seen from inside a badly chosen coordinate.
(a) This is a Lie group in embryo. The boosts along one axis are a family with , and smooth in . That is a smooth one-parameter group. ⚑ Chapter 6.1 shows that any such family is for a fixed matrix , its generator, obtained by differentiating at the identity. Here , and indeed , so , which is (2.2.48) exactly. The full Lorentz group has six such parameters, three boosts and three rotations, and is called .
(b) The hyperbolic functions are the metric's minus sign, surfacing. A rotation in the plane is and preserves . A boost in the plane is (2.2.48) and preserves , per (2.2.20). Circular functions preserve a sum of squares. Hyperbolic functions preserve a difference. The two matrices differ in exactly one respect: the rotation's off-diagonal entries carry opposite signs while the boost's carry the same sign, so the boost matrix is symmetric and the rotation's off-diagonal part is not. That single sign is the in . Chapter 2.3 makes this geometric: a boost is a rotation through an imaginary angle, or better, an honest rotation in a geometry whose circles are hyperbolas.
Grind box — hyperbolic functions from scratch, in case they are rusty
Definitions, and everything else follows:
The Pythagorean identity. Multiply out:
Compare . The point traces a circle, and the point traces a hyperbola. That hyperbola is Chapter 2.3's central object.
Addition. Multiply exponentials:
The formula goes the same way. Dividing them gives the addition theorem, which is (2.2.45).
The inverse. Solve for . Writing ,
Useful values: , , , . Note how slowly rapidity grows once is close to . Going from to costs about of rapidity, roughly the same as going from rest to . In Chapter 2.5 that observation acquires teeth, because the energy needed is what grows without bound.
Derivatives. , , , all immediate from the exponential definitions. Problem 4 uses the last one.
Compose two rotations of a plane and their angles add, as anybody expects of a family labelled by one number. Boosts along an axis are such a family, so something of theirs adds, and it is not velocity. Take the composition rule seriously and it is a formula everybody has seen, the addition theorem for the hyperbolic tangent, so what adds is the angle whose hyperbolic tangent the velocity is.
Call that angle the rapidity. Rapidities add exactly, with no correction term, and their matrices multiply the way two turns of a plane do. This collects the pairing between a generator and the family it builds, promised when a self-partnered map was first exponentiated and again when conserved quantities were made to push things around: one fixed matrix, exponentiated by the parameter, generates the family.
Why velocities looked defective is now visible, and it is not a fact about nature. Velocity is the slope of a worldline and rapidity is its angle, and slopes have never added: two turns of a plane combine their angles cleanly while their slopes combine by an ugly quotient that shocks nobody, since nobody expected slopes to be the natural label. Three centuries were spent adding the slopes. The barrier stops being mysterious in the same stroke, since the angle runs over every real number and adds without limit while its hyperbolic tangent creeps towards one and never arrives.
7 · The paradoxes, dismantled
There are perhaps a dozen famous special-relativity paradoxes and they are all the same paradox. Here are the two that matter, worked to the bottom, and then the general statement.
7.1 · The twin paradox
The setup. Two twins, both aged . One stays on Earth. The other flies to a star light-years away at (), turns round, and comes back. Here is the Earth-frame arithmetic. The outbound leg takes years, so the round trip takes years. The traveller's clock is the moving one, so by (2.2.31) it records years. The traveller returns aged , and the stay-at-home is .
The apparent paradox, stated as sharply as it deserves. Motion is relative. From the traveller's point of view, Earth receded at and came back. By §4.2, each observer finds the other's clock slow, symmetrically and by the same . So the traveller should equally conclude that Earth's clock ran slow throughout. His own years, divided by , gives years for the Earth twin, who should therefore be the younger one on his return. Both cannot be true. When the two stand next to each other at the end, comparing wrinkles is a local coincidence, and there is one fact of the matter.
Where the symmetry fails. It fails at the turnaround, and it fails physically, not verbally. The stay-at-home twin occupies a single inertial frame for the whole story. The traveller does not. He is at rest in one inertial frame going out, and in a different one coming back. The switch is detectable from inside the ship without looking out of the window. An accelerometer registers it, coffee spills, and the traveller is pressed into his seat. There is no symmetry to appeal to, because the two histories are physically different, and no amount of "motion is relative" makes them the same. One worldline is straight. The other has a kink.
That is the resolution in one sentence, and people rightly find it unsatisfying. It explains why the twins may differ without showing where the missing years went, that being the gap between the traveller's naive and the actual . So let's track them.
Following the traveller's own accounting, year by year
Work in Earth-frame coordinates measured in light-years and years, with the departure at the origin. The turnaround is the event .
Outbound leg. The traveller is at rest in the frame moving at . His own elapsed time is, from (2.2.19) applied to the event ,
The next thing we need is his verdict on Earth. Which of Earth's events does he call simultaneous with his turnaround? Those with the same , that is those with . On Earth , so
Consistent with time dilation as the traveller sees it: his years, divided by , gives years on the Earth clock. From out there, Earth is the moving one and Earth is running slow. ✓
Inbound leg. After the turn he is at rest in moving at . Its simultaneity slices are , so run the same question through the same turnaround event and see which Earth event he now calls "now":
The jump. An instant before the turn, the traveller's "now on Earth" is year . An instant after, it is year . Nothing happened to Earth. The traveller changed which slice of spacetime he calls "now", and the slice swung through Earth's worldline:
That jump is the piece the naive argument left out, so add the three contributions to the Earth twin's clock and watch the books balance exactly:
while the traveller's own clock reads years. Every year is accounted for. The traveller's claim "Earth's clock ran slow" was correct on each leg. What he cannot do is stitch the two legs' notions of "now" together and pretend they were one. The naive symmetric argument gave Earth years, and the it was missing is exactly the span his simultaneity slice swept across Earth's worldline during the turn. That turn was the one moment at which he was not an inertial observer, and therefore the one moment at which his slices were not entitled to line up.
Forget clocks for a moment. Proper time is the length of a worldline. Between two events it is along each straight piece, and you add the pieces. Two different paths between the same two endpoints have different lengths. Nobody is scandalised that a detour through the next town adds mileage.
Check the twin numbers this way. The straight path gives years. The bent path has two legs, each , so in total. Same endpoints, different lengths, no paradox. And there is no need to mention acceleration at all, except to say that a kink is what makes a path longer or shorter than a straight one.
The one genuinely strange thing is the sign. Because of the minus in , the straight path is the one with the most proper time, not the least. Detours in spacetime save you time rather than costing it. That inverted extremum is not a curiosity. It is the principle from which Chapter 3.1 derives the motion of freely falling bodies, and it is why a satellite's orbit is the path that maximises the proper time of the clock aboard it.
Turn on the twin toggle in the figure above and watch this happen. The traveller's two simultaneity slices are the green dashed lines through the turnaround, and where they cross they bracket exactly the jump. Drag and the numbers move, but the three pieces always sum to the stay-at-home twin's total.
7.2 · The pole and the barn
The setup. A runner carries a pole of proper length through a barn of proper length with a door at each end, at so . In the barn's frame the pole is contracted to , so it fits, with a metre to spare. The farmer shuts both doors simultaneously at the moment it is inside, and the pole is briefly wholly enclosed. That last clause is the claim to be tested.
The apparent paradox. In the runner's frame the pole is and the barn is contracted to . The pole is two and a half times the barn's length. It cannot possibly be enclosed. But "the pole was inside the barn with both doors shut" ought to be a fact, not an opinion.
The resolution. It is not a fact, because it is not a local coincidence. It is a statement that two events at different places, front door shuts and back door shuts, are simultaneous. By §4.1 that is frame-dependent, and here the disagreement is enormous. With ,
so in the runner's frame the two door-closings are of light travel apart, which is , with the far door going first. They do not shut together at all. The far door shuts and reopens long before the pole's nose gets there. Much later, the near door shuts behind the pole's tail. The pole passes through a barn that is opening and closing around it like a slow shutter, and it is never enclosed for an instant.
Both frames are right, and here is the decisive point: they agree on everything that is not a convention. Was the pole ever hit by a door? No, in both frames. Did the nose emerge from the far door before the far door shut? Problem 2 tabulates every event in both frames, and the answer is the same in both. What they disagree on is only "did the two doors shut at the same time", which is not an observable but a labelling.
7.3 · The general lesson
Somewhere in the statement, someone has assumed simultaneity is absolute. Find that assumption and the paradox dissolves. There is no second mechanism.
The reliable procedure, which will get you through any of them:
- List the events. Not objects, and not "the rod" or "the ship", but events, each with a definite in one frame. Anything that is not an event is not yet a physics question.
- Ask which claims are local coincidences. Two things at the same place at the same time: a door hitting a pole, two clocks side by side, a twin shaking a hand. These are frame-independent and every frame must agree on them, full stop.
- Transform. Apply (2.2.19) to each event and tabulate.
- Locate the smuggled assumption. It will be a claim of the form "and at that moment, over there, ...", which is a statement about distant simultaneity dressed as a fact.
Apply that to the twins, where the smuggled claim is "while the traveller aged 3 years, Earth aged 1.8, and that was still true after the turn". Apply it to the pole, where the smuggled claim is "both doors shut at the same time" taken as an absolute. Apply it to any of the others: the ladder and the trapdoor, the rigid rod pushed at one end, the submarine that both sinks and floats. Same disease, same cure, every time.
One twin comes back younger than the other, and the asymmetry permitting it is physical rather than verbal. The one who stayed occupied a single inertial frame throughout; the one who travelled was at rest in one frame going out and in a different one coming back, and the switch is detectable from inside the ship, with no window needed. One history is bent and the other is not.
Saying so explains why the two may differ without showing where the missing years went, and they can be tracked. The traveller's claim that the other clock was running slow is correct on each leg taken separately. What he may not do is stitch the two legs' notions of the present together, because at the turn his family of nows swings through an enormous span of the other worldline, and that span is precisely the amount the naive accounting was missing.
Every paradox in the subject is this one in different clothes. Somewhere in the statement, someone has treated the simultaneity of two distant events as a fact rather than a labelling, and the cure never varies. List the events; ask which claims are local coincidences of one thing with another at one place, since everybody agrees on those; transform the rest; and the smuggled assumption will be standing in plain view. What this pile of results is waiting for is the idea that organises it.
8 · Worked examples
Muons are produced by cosmic-ray collisions about up in the atmosphere. ⚑ Their mean lifetime at rest is , measured in the laboratory, and they decay exponentially, so that the fraction surviving after a proper time is (Chapter 0.1 §5, the same first-order kinetics as drug clearance). Take . What fraction reaches sea level, according to (a) Newtonian expectations and (b) relativity? Then redo (b) entirely in the muon's frame, where the atmosphere is contracted rather than the clock dilated, and check the two accounts agree.
Step 0: the Lorentz factor.
Step 1: Newtonian. No dilation, no contraction, so the muon crosses at and its internal clock is the same as ours.
which is lifetimes. Surviving fraction:
About one muon in eight billion. The predicted flux at sea level is, to any instrument ever built, zero. ⚑ What is actually observed is a copious flux, of order one muon per square centimetre per minute at sea level. The Newtonian prediction is not slightly wrong.
Step 2: relativity, told from the ground. The muon's internal clock is the moving one, so by (2.2.31) the proper time it experiences during our is
Nearly a quarter of them make it. The two predictions differ by a factor of . That is why this is not a subtle test but a demonstration you can do with a scintillator on a bench, and why ⚑ Rossi and Hall's 1941 mountain-versus-sea-level muon counts were decisive.
Step 3: the same physics, told from the muon. This is the half that gets skipped, and it is the half that matters, because in the muon's own frame nothing dilates its clock. Its clock is at rest. It lives on average, exactly as advertised. And it had better still reach the ground, because whether a given muon arrives is a local coincidence, and frames are not allowed to disagree about it.
What is different in the muon's frame is the atmosphere, which is now the moving object. The column of air is a proper length, since it is at rest relative to the ground, so by (2.2.35) it is contracted to
That column rushes past the muon at , taking
and the surviving fraction is . Identical. ✓
What this teaches. The two frames disagree about why the muon survives. One says the clock ran slow, the other says the trip was short. They agree exactly on the observable, which is the number of muons per second landing on the detector. That is the correct relationship between frames in every problem in this chapter: mechanisms are frame-dependent, observables are not. Note also that neither account works without the other's ingredient. Use dilation and contraction together and you would double-count by a factor of . Use neither and you get . There is exactly one factor of in the problem, and each frame places it somewhere different.
In a laboratory, two particles fly directly toward each other, each at as measured in the lab. (a) At what rate does the gap between them close, in lab coordinates? (b) What is the speed of one as measured in the rest frame of the other? (c) Reconcile.
(a) The closing rate. In lab coordinates, particle 1 is at and particle 2 at . The gap is , so
The gap shrinks at . This is correct and no rule is violated. It is not the speed of anything. It is the rate of change of a difference of two coordinates, both measured in one frame. Nothing is transported at . As with §5's spotlight, ask what is being carried, and the answer is nothing.
(b) The relative speed. This is a different question: what does particle 2's own frame measure for particle 1? Take to be particle 2's rest frame, so (it moves in the direction in the lab), and the object of interest has . Then by (2.2.40),
So . Not , and not even . Each particle sees the other approaching at of light speed.
Cross-check with rapidity. , and the relative rapidity is the sum, . Then ✓. That is the same number, obtained by addition rather than by a formula with a denominator in it. This is §6 earning its keep.
(c) Reconciling. The two answers are answers to two different questions, and the language of everyday physics does not distinguish them because at low speed they coincide. For , , which is the closing rate. At high speed they separate dramatically. Here is the rule. A closing rate is a statement about one frame's coordinates and can be up to . A relative velocity is a statement about one object's rest frame and is always less than . Only the second one is a velocity in the sense that (2.2.40) and Chapter 2.5's dynamics care about.
Why the difference has teeth. In a collider the quantity that determines what you can produce is the energy in the centre-of-mass frame, which depends on the relative velocity, not the closing rate. Two beams at head-on do not collide at in any useful sense. They collide at a relative of , which is ✓. That is the composition law of §3, doing an accelerator physicist's arithmetic.
9 · Your turn
Problem 1 · the transverse velocity, which changes even though
An object has velocity components in . Derive and in . Explain why even though the coordinate itself is untouched. Then check the formula on light. Take a pulse moving in the direction at speed in , verify that also measures speed , and find the angle by which the pulse's direction is tilted.
Solution
Derivation. Take differentials of (2.2.19) including the transverse equation :
Divide by and then divide numerator and denominator by :
Why it changes. A velocity is a ratio, . The numerator is indeed untouched. The denominator is not. Time intervals transform, so a displacement that is unchanged, divided by a duration that is not, gives a changed rate. The lesson generalises far beyond this problem: an invariant numerator does not make an invariant quotient, which is exactly why Chapter 2.5 will insist on differentiating with respect to proper time rather than coordinate time when it builds the four-velocity.
The special case . Then and , so the transverse velocity is reduced by exactly , since the transverse displacement is unchanged while the time to make it is dilated.
Check on light. Put , . Then and , so
Speed ✓. Note how the two pieces conspire: the gained along is precisely the lost along . Postulate 2 again, checking itself in a direction we never imposed it.
The tilt. The direction makes an angle with the -axis given by
This is stellar aberration. The apparent position of a star shifts by as the Earth's velocity changes around its orbit, giving an annual ellipse of angular radius , which in arcseconds () is . Bradley measured it in 1727, and it was one of the constraints Chapter 2.1 §6.1 used to kill the fully dragged ether. Chapter 2.1's Problem 3 obtained from a first-order argument. Here is the exact statement, and the two agree to one part in at the Earth's orbital speed. Note also that is the same relation as the Terrell–Penrose rotation angle in §4.3. That is not a coincidence, since both are the statement that a boost tilts light rays.
Problem 2 · the pole and the barn, with every event tabulated
Pole of proper length , barn of proper length with doors at (entrance) and (exit), runner at in the direction. Use coordinates with both in metres. Let event be the pole's tail reaching the entrance, and make it the origin of both frames. (a) Find the barn-frame times at which the pole is fully inside, and choose the farmer's simultaneous door-shutting to be at the midpoint of that window. (b) Tabulate the four events , (nose reaches the exit door), (entrance door shuts) and (exit door shuts) in both frames. (c) In the runner's frame, show explicitly that neither door ever touches the pole.
Solution
(a) The window. , so the pole is contracted to in the barn frame. At the tail is at and the nose at . The nose reaches after the pole advances , which takes . So the pole is wholly inside for (that is, for ), and the farmer shuts both doors at .
(b) The table. Transform with , and , :
| Event | (m) | (m) | (m) | (m) |
|---|---|---|---|---|
| tail reaches entrance | ||||
| nose reaches exit | ||||
| entrance door shuts | ||||
| exit door shuts |
Read the fourth column. In the barn frame and are simultaneous. In the runner's frame happens at and at , so the exit door shuts earlier, matching (2.2.56). Note also , the pole's proper length ✓. And in the runner's frame the nose reaches the exit at , which is well before the tail reaches the entrance at . The pole is sticking out of both ends at once, which is what a pole does in a barn.
(c) Nobody gets hit. In the runner's frame the pole is at rest occupying . Track the doors. The exit door is the worldline ; substituting into gives
It reaches the nose, , when , i.e. ✓ (event ). It shuts at , which is earlier, and at that moment it sits at , which is beyond the nose. The door shuts in empty space in front of the pole, reopens, and the pole then passes through.
The entrance door is , worldline . It shuts at , when it sits at , which is behind the tail at . It too shuts in empty space.
The point. Every frame-independent question gets the same answer in both frames. Did a door hit the pole? Did the nose pass the exit before or after the exit door shut? Only "were the doors shut simultaneously?" differs, and that was never a fact about the world. If you insist on a barn that genuinely traps the pole, you must shut the doors and keep them shut. Then the runner's frame will report that the nose crumples against the closed exit door while the tail is still outside, which is a real, invariant and rather expensive event that both frames agree on.
Problem 3 · two boosts in different directions leave a rotation behind
Work in the subspace. (a) Show that a pure boost matrix, in any direction, is symmetric. (b) Compute explicitly and show it is not symmetric unless one rapidity vanishes. Conclude that it cannot be a pure boost, so boosts do not form a group. (c) Show that the leftover is a rotation, and estimate its size for small rapidities. Quote (⚑) the exact angle and check your estimate against it.
Solution
(a) Pure boosts are symmetric. A boost along is (2.2.48) padded with identity rows, manifestly symmetric. A boost in a general direction is that one conjugated by a rotation that takes to : (with acting trivially on ). Then ✓. So symmetry is the algebraic signature of "pure boost, no rotation".
(b) The product. With , and rows/columns ordered :
Compare entries across the diagonal. sits opposite , and these agree only if , that is . Likewise sits opposite . So is not symmetric, hence by (a) not a pure boost. Two boosts in different directions do not compose to a boost. The set of pure boosts is therefore not closed, and is not a group. Only the boosts along a fixed axis are, which is what §3 actually proved.
(c) What is left over. ⚑ Every proper orthochronous Lorentz transformation factors uniquely as with a pure boost and a rotation (the polar decomposition; we quote it). Since is not symmetric, : there is a residual rotation, the Wigner rotation, in the plane.
Its size for small rapidities follows from non-commutativity. Write and with generators
Multiplying out, has entries only in the block:
the generator of rotations about . ⚑ By the Baker–Campbell–Hausdorff formula (Chapter 6.1), , so
giving a rotation angle at low speed. It is second order in the velocities, which is why it has no Newtonian counterpart and why nobody stumbled on it before 1926.
⚑ The exact angle, quoted, for two perpendicular boosts:
Check the low-speed limit. With and we get ✓, matching the commutator estimate. Numerically, for the exact formula gives , and the two leading-order estimates bracket it. comes in from below and from above, since for . They are the same estimate to leading order and differ only at the next one. For the exact angle is , against and , which is agreement to five figures.
Why it matters. An electron in an atom is continuously boosted in changing directions by the nuclear Coulomb field. The accumulated Wigner rotations make its spin precess, an effect called Thomas precession, and this contributes a factor of to the spin–orbit coupling in atomic fine structure. Without it, the predicted fine-structure splitting is wrong by a factor of two, and the discrepancy stood unexplained for two years. A pure kinematic effect, with no force involved at all, showing up in a spectrometer.
Problem 4 · a rocket at constant proper acceleration
A rocket accelerates so that its crew feel a constant . That is, its acceleration measured in its own instantaneous rest frame is always . (a) Show that its rapidity grows linearly in proper time, , and hence . (b) Find the lab-frame time and distance as functions of . (c) With , compute , the lab time and the distance after year of proper time, and then the proper time needed to reach the galactic centre, light-years away. Comment.
Solution
(a) The rapidity grows linearly. This is where §6 pays for itself. In a proper-time interval , the rocket's instantaneous rest frame sees it acquire a velocity , hence a rapidity increment
Because rapidities add ((2.2.50)), these increments accumulate directly. No composition formula is needed, and needing one is exactly what would have made the velocity version painful. Integrating from rest,
(b) Lab time and distance. By (2.2.31), , so
And , so
Eliminating with gives , so the worldline is a hyperbola. Chapter 2.3 will show that this is the spacetime analogue of a circle, the curve of constant "curvature". Chapter 3.1 will note that its crew experience something remarkably like a gravitational field, complete with a horizon behind them.
(c) Numbers. The natural scale is
A pleasant accident of our units: one year at one gee is very nearly one rapidity unit.
| proper time | lab time | distance | |||
|---|---|---|---|---|---|
For the galactic centre, invert the distance formula: , so and
(Lab time: years. The Galaxy does not get a discount.)
Comment. Two things are worth extracting. First, the crew never detect anything odd. Their accelerometer reads a constant forever, their rapidity climbs without limit at per year, and at no point do they approach a barrier. It is only the velocity that saturates, and velocity is the badly chosen coordinate. This is (2.2.46) made visceral. Second, the crew-time cost of a journey grows only logarithmically with distance once you are relativistic, since for large . The Andromeda galaxy, million light-years away, costs years of crew time against for the galactic centre. The Universe is traversable within a lifetime and simultaneously unreachable, because you can never come home. Chapter 2.5 supplies the reason nobody does it: the fuel.
You derived (2.2.19), which says , , and , from the two postulates plus homogeneity, isotropy and reciprocity, with every step forced and every assumption named. You showed it reduces to Galileo as , that boosts along an axis compose into boosts, that the composition is addition in rapidity and not in velocity, and that the quantity is invariant even though you only ever demanded it for light. Simultaneity, dilation and contraction came out as corollaries in that order, and the two classical paradoxes turned out to be one paradox with two costumes.
Where this gets spent. Chapter 2.3 takes the invariant interval (2.2.21) and the hyperbolic form (2.2.48) and shows that together they define a geometry. That geometry is spacetime, in which a boost is a rotation and proper time is arc length. Chapter 2.4 takes the matrix (2.2.24), calls it , and defines a tensor as anything that transforms with it. The rapidity form and the group property of §3 and §6 are quoted there verbatim. Chapter 2.5 applies the same transformation to momentum and energy instead of position and time, and falls out of it. The you derived here is the in . Chapter 2.6 shows that the electric and magnetic fields transform into each other under exactly this boost, which retroactively explains Chapter 2.1's entire crisis. Chapter 3.1 keeps the transformation but only locally, which is the whole idea of general relativity. And Chapter 6.1 returns to the rapidity and names what it has been all along: the parameter of a one-parameter Lie group, with the boost generator as its Lie-algebra element.