Part II · Special Relativity — Chapter 2.3
Minkowski Geometry
Special relativity is Euclidean geometry with one sign flipped. Everything strange about it is that sign, working.
Chapter 2.2 produced a transformation and then a pile of consequences. Simultaneity fails. Moving clocks run slow. Moving rods are short. Velocities do not add, and rapidities do. Each of those was derived, each is correct, and taken together they are a list. Lists are how a subject looks before anyone has found its organising idea.
This chapter finds the organising idea. It is already sitting in 2.2 §2.4, noticed in passing and then left alone: every one of those transformations preserves the combination
.
One number, agreed on by everybody. We are going to take that seriously and see what it forces. What it forces is a geometry. By that we mean a four-dimensional space with a notion of distance, in which the Lorentz transformations are the rigid motions, boosts are rotations, worldlines have lengths, and the length of a worldline is the time a clock carried along it reads. After that, the list stops being a list. Time dilation is a statement about lengths. The twin paradox is a statement about triangles. The light barrier is a statement about which directions exist.
Here is the thesis in one sentence, and it is worth reading twice: special relativity is Euclidean geometry with one sign flipped. Nearly every difference between the two subjects traces back to a single minus sign in the distance formula. Why boosts are hyperbolic rather than circular. Why the straight path is the longest and not the shortest. Why some pairs of distinct points are at zero distance. Why there is a speed limit at all. All of those come from that one sign. We will flip it, and then follow it everywhere it goes.
One thing this chapter does not do honestly. It writes , and , and uses index notation informally, without defining what an index means. That debt is real, it is flagged where it is incurred, and Chapter 2.4 pays it in full.
Tools you'll need — Chapter 2.2: the boost , , the matrix form, and the rapidity with . Nothing from 2.2 is re-derived here. Chapter 0.4 §3 and §5: linear maps as matrices, composition as matrix multiplication, and the determinant. Section 1 asks which linear maps preserve a given quadratic expression, which is exactly a 0.4 question. Chapter 0.5 §1: the inner-product axioms, and the observation that once you have a bilinear form you have lengths, angles and orthogonality. This chapter removes one of those axioms on purpose and watches which conclusions survive. Chapter 1.2 §9, worked example 2: extremising gives straight lines in the plane and great circles on a sphere. Section 6 runs that machine one more time, in spacetime.
1 · The invariant interval
Everything in this chapter is about two events. Not two objects, and not two particles, but two events, each one a point of spacetime with coordinates in some inertial frame. The single combination of their coordinates that the whole subject turns on is this one:
Here and so on. The quantity is called the interval between the two events, or sometimes the spacetime separation.
Three warnings come with it, and all three are paid off before the chapter ends.
- The notation is a single symbol. It is not the square of something called , and it can come out negative.
- It is not a distance.
- The minus signs are not a typographical accident. They are the entire subject.
1.1 · It is the same in every frame
The claim to be proved is that two observers in relative motion, handed the same two events, compute the same . The proof is direct substitution and it is short. Start from the boost of Chapter 2.2,
and apply it to a separation rather than to a single event. That step is legitimate for exactly one reason, and 2.2 §1.2 worked hard to earn it. The transformation is linear, so it acts on differences the same way it acts on coordinates. Our goal is to see what happens to the two-dimensional part , so substitute the boost into it and grind:
The last line used , which is the definition of rearranged.
Look at what became of the cross terms. They cancelled identically, and they had to. They enter the two brackets with the same coefficient and with opposite overall sign, so nothing else was available to them.
The other two directions cost us nothing, because and . We may therefore subtract those two squares from each side without disturbing the equality, and that gives the full four-dimensional statement:
Every inertial observer, handed the same two events, computes the same . They will disagree about . They will disagree about . They will disagree about which event came first, sometimes. They agree about this.
Chapter 2.2 reached this same equation in its §2.4 and treated it as a consistency check. We had demanded that light go at , and the check confirmed that a spherical light pulse stays spherical. That reading undersells the result badly, and it is worth being precise about why.
Postulate 2 constrains only the pairs of events with , which is one three-dimensional family out of all the pairs there are. What came back out is a statement about every pair of events, at every value of . The theory handed us more invariance than we asked for. A quantity that turns out to be invariant for no reason you demanded is exactly the kind of quantity worth building a subject on.
1.2 · The reframe: this is what "length" means
Now comes the move that makes the rest of the book easier. Let's ask what a rotation is.
You probably answer "a rigid turn about an axis". That is a picture rather than a definition. The definition used in mathematics is algebraic, and it is the one Chapter 0.4 §5 was quietly preparing you for:
That is the whole content. You do not need to mention turning, or axes, or angles. Demand that a linear map leave unchanged for every vector, and rotations together with reflections are what you get. They form the group , and requiring the determinant to be throws away the reflections.
Notice which way the logic runs. It runs from the quadratic form to the transformations, and not the other way round.
Now notice what the preserved quantity is. It is what we call length. That is not a separate fact about the world sitting alongside the definition. It is the definition of length: the thing on which all observers with differently-oriented axes agree.
Read (2.3.4) again with that in mind. It says:
Lorentz transformations are the linear maps that preserve .
So plays the role that squared length plays in Euclidean geometry, and the Lorentz transformations play the role that rotations play. Change the quadratic form from to , and the group of "rigid motions" changes from rotations to boosts-and-rotations. Nothing else about the logic changes at all.
That is not an analogy. It is the same definition with a different quadratic form in it.
Two remarks on how much that box is claiming.
The first is that a converse is involved. What we proved above is that Lorentz transformations preserve . The statement that the linear maps preserving are exactly the Lorentz transformations runs the other way, and converses need proof. The grind box below proves it in the case that matters, which is one space dimension. It also finds three extra maps hiding in the answer, which turn out to be parity, time reversal, and the two applied together.
The second is that the word "geometry" is a promise. Calling this a geometry commits us to lengths, angles, straight lines and triangles behaving like a geometry. Sections 3 to 6 make good on each of those. Section 6 finds that one of them behaves like a geometry with the inequality reversed. That is the sign flip, arriving where it hurts.
Grind box — the converse: every linear map preserving , classified
Work in the plane, which is where all the content lives. Write the unknown map as a general matrix and let the invariance condition tell us what its entries have to be. So let
and demand for all . The word "all" is what makes this usable, because two quadratics that agree everywhere must agree coefficient by coefficient. So expand the left side and see what the coefficients are:
Matching coefficients of the three independent monomials gives three equations:
Solve the first. We have , so , and any number of modulus at least one can be written as a . Put with and . Then , and by allowing to take either sign we can write without loss of generality.
Solve the third. The same argument applies, since likewise, so and with . Notice that we have had to allow a second angle here. The middle equation is what will force it to equal the first.
Use the second. Substituting the four entries into ,
so , since vanishes only at zero. Therefore
Read the factorisation on the right. The second factor is exactly Chapter 2.2's boost in rapidity form. The first factor is one of four sign choices: is the identity, is parity , is time reversal , and is both. Note .
The conclusion. The linear maps preserving are precisely the boosts, possibly composed with parity and/or time reversal. So the converse holds, up to those discrete factors. That is the same structure as in Euclidean geometry, where preserving gives rotations possibly composed with a reflection.
The group therefore falls into four disconnected pieces. The piece containing the identity is the one with , the pure boosts, and it is the piece physics uses without comment. It has a name, the proper orthochronous Lorentz group. Proper means , and orthochronous means it does not reverse the direction of time. ⚑ That the same four-component structure survives in the full four-dimensional group, with rotations included, is quoted here and unpacked in Chapter 6.1.
A detail worth noticing. Nothing in this derivation used a postulate, a light ray, or a physical assumption. Given the quadratic form, the transformations are pure algebra. All the physics went into choosing rather than , and Chapter 2.2 is where that choice was forced.
1.3 · The sign convention, and the one half the world uses instead
Nothing above depended on whether we write or . The two differ by an overall factor of , they are preserved by exactly the same transformations, and every physical statement can be phrased in either. Books nevertheless choose, and then never mention it again, which is how readers get ambushed.
This book uses the timelike-positive convention, often written and sometimes called the "mostly minus" or "particle physics" signature:
The reason is Parts V to VII. In particle physics the object you write down forty times a day is the four-momentum, and its invariant square in this convention is . That is mass squared: positive, with no sign to remember. In the other convention it is , and every mass-shell condition in the book would carry a minus sign for no benefit. Since Parts V–VII are much longer than Part III, we optimise for them.
The other convention, or "mostly plus", is standard in general relativity, where the object you write forty times a day is a spatial line element and you would rather it came out positive. If you open a relativity textbook and find , nothing is wrong. Multiply by and carry on.
Here is what the choice does and does not touch. What changes: the sign of classifying timelike versus spacelike, the sign in , and the sign of itself. What does not change: a single physical prediction. It is a units-style choice, like measuring angles in degrees, and it carries the same amount of physics.
Out of a demand made about one narrow family of events came an equality holding for every pair of them, and that overshoot is what makes a geometry available. The combination of time separation and space separation that all observers compute alike has a name, the interval, and the way to see what it is doing is to ask first what a rotation is.
The answer taught first is a rigid turn about an axis, which is a picture rather than a definition. The definition is algebraic: a rotation is a linear map leaving the sum of the squared coordinates alone. Nothing about turning is needed, and the logic runs from the quadratic expression to the transformations rather than the other way. The quantity those transformations preserve is what the word length means, not a separate fact about the world but the name for whatever observers with differently oriented axes agree on.
Read the invariance again with that in mind, and the whole chapter is one substitution. Flip the sign of the three spatial terms in the quadratic expression, and the rigid motions change from rotations into the transformations of relativity. That is not an analogy; it is the same definition with a different expression inside it. The converse holds too, since the maps preserving the flipped expression are exactly the boosts, up to a reflection of space or of time.
2 · Coordinates, the metric, and one line of index notation
We are about to need a compact way to write (2.3.1) that does not degenerate into four-fold bookkeeping every time. The notation that does this is one of the great labour-saving devices in physics. We introduce it here informally, and label it honestly as informal, because §3 onward reads a great deal better in it.
2.1 · One symbol for four coordinates
Group the four coordinates of an event into a single object with an index:
Greek indices run over ; Latin indices over the spatial only. The time coordinate is and not , so that all four entries have the dimensions of length. That is not cosmetic: Chapter 2.2 §2.2 already showed that and enter the boost symmetrically, and writing instead would hide the symmetry behind a units conversion.
One warning about the superscript. It is an index, not a power. So is the -coordinate. Where we mean the square of something, we will write it so that the context is unmissable. This collision of notation is genuinely annoying, and it is also universal, and everyone lives with it.
2.2 · The metric, and the summation convention
Now define a array of numbers
called the Minkowski metric. It carries the minus signs of (2.3.1) and does nothing else. With it, the interval is
The second form uses the Einstein summation convention: an index appearing twice in a single term, once up and once down, is summed over its whole range, and the summation sign is not written. It is a convention about typography and nothing more. The saving is nonetheless substantial. The double sum has sixteen terms, of which twelve vanish because is diagonal, and you never want to write that out twice.
Before leaning on the shorthand, let's check that (2.3.9) really does say what it should. Since is diagonal, the only surviving terms are the ones with , so write those four out:
That is ✓. So the metric is doing exactly one job at this stage: supplying a sign for each term.
In ordinary three-dimensional space the corresponding array is , which supplies all plus signs. That is why nobody ever writes it, and why the metric is invisible in first-year physics. The whole of the difference between Euclidean geometry and Minkowski geometry is the difference between and .
Grind box — the summation convention, written out once in full, and the rules for using it
The unabbreviated version of (2.3.9), all sixteen terms, with to avoid symbol collisions:
Three rules, all of which will be justified properly in Chapter 2.4 and all of which you should obey immediately.
(1) A repeated index is dead. In the labels and have been summed away, so the result carries no index and does not depend on what you called them. Renaming throughout changes nothing. Such indices are called dummy indices, and they behave exactly like the in .
(2) An unrepeated index is alive, and must match on both sides. is a legal equation: is summed, is free and appears once on each side, so the equation is four equations. By contrast is not an equation at all, because it does not say which component equals which.
(3) Never use the same letter three times. is ill-formed: the convention cannot tell which pair you meant to sum. If you need two separate sums, use two separate letters. This rule catches more algebra errors than any other.
What is not being claimed here. We have written some indices up () and some down (), and given no reason. There is a reason, it is important, and it is not that one is a row and the other a column. Chapter 2.4 §3 gives it. Until then, treat the placement as a spelling rule you are obeying on trust, and note that the trust is being tracked: every expression in this chapter has each summed index appearing once up and once down, and you may check that it does.
2.3 · What a four-vector is, provisionally
We need a working notion of "vector" in spacetime, and there is one object whose behaviour under a boost we already know completely. The displacement between two nearby events, , transforms in exactly the way (2.3.2) says it does. We take that as our template:
A four-vector is a set of four quantities , one per frame, that transform between frames the same way does:
with the same matrix that transforms coordinate displacements.
This is deliberately informal. It says "these ones, like that one", which is a definition by example, and it leaves at least three questions unanswered. Why are some indices up and others down? What is the object , as opposed to what its components do? Does the definition depend on the transformation being a Lorentz transformation? Chapter 2.4 answers all three and turns this paragraph into a theorem. Nothing in this chapter needs the answers. Everything in this chapter needs the template.
Why bother now? Because the payoff arrives immediately and repeatedly. If and are both four-vectors, then the combination
is a single number that every observer agrees on. The reason is exactly the reason is agreed on, since is the special case .
What we have built is a machine for manufacturing invariants, and we will keep feeding it: in this chapter, the four-velocity in §7, the four-momentum in Chapter 2.5 (where turns out to be the whole of relativistic kinematics), and the four-current in Chapter 2.6.
2.4 · The condition on
One more line before we start using any of this, and it is the line Chapter 2.4 and Chapter 6.1 both build on. What we want is the condition a matrix has to satisfy in order to be a Lorentz transformation at all. So write the transformation in matrix form, as Chapter 2.2 §3 did, and demand invariance of the interval. In index notation:
Since the separation is arbitrary and both sides are symmetric in , the coefficients must match term by term:
Let's look at what that boxed line is saying. Compare with the condition defining a rotation matrix, , that is, . Same equation, one symbol different.
⚑ (2.3.13) is the defining equation of the Lorentz group, written , where the notation records that the metric has one plus and three minuses. Chapter 6.1 will take it as the starting point rather than as the conclusion, differentiate it near the identity, and read off the six generators. Problem 4 does the counting by hand and gets six.
Two consequences drop out of that condition without any further work. Here is the first. We want to know what the condition says about the overall scale of , so take determinants of (2.3.13) and use from Chapter 0.4 §5:
So Lorentz transformations preserve four-dimensional volume in spacetime, up to sign. Chapter 2.2 §3 checked for the boost explicitly.
The second consequence is the group property. The set of matrices satisfying (2.3.13) is closed under multiplication and inversion. Chapter 2.2 §3.1 verified that by brute force for boosts along one axis, and here it takes one line: .
Grind box — checking on the boost, entry by entry
Only the upper-left block does anything, so drop and and take with
The boost matrix is symmetric, so . Take the triple product in two stages and compute first. Multiplying on the left by flips the sign of the second row:
Now multiply that result on the left by , which is itself:
The diagonal entries are the hyperbolic Pythagorean identity. The off-diagonal ones cancel because the matrix is symmetric. Note which identity did the work: . In the Euclidean case the same computation with and runs on instead. One sign, two geometries, and that is what §3 is about.
You already compute quadratic forms with a matrix in the middle, and you compute them for exactly the reason (2.3.9) exists. Given a panel of measurements with covariance , the Mahalanobis distance of an observation from the mean is
written on the right in this section's notation. Its entire purpose is to be a number that does not depend on the units the individual features were recorded in, nor on how strongly they happen to be correlated: one agreed answer, assembled from components every laboratory writes down differently. That is the paragraph above (2.3.9), word for word, with standing where stands.
The parallel runs past the formula. Chapter 0.5 §6 guarantees that , being real and symmetric, is diagonalised by an orthogonal change of basis. That change of basis is the whitening transform, after which is a plain sum of squares and the maps preserving it are ordinary rotations. Here the same manoeuvre stops one step short. (2.3.8) is already diagonal, and no real change of basis can turn its diagonal into four plus signs, because the number of minus signs is an invariant of the form and not an artefact of the basis.
What that one sign buys is the whole of §4. is positive definite, so always, with equality only for an observation sitting exactly at the mean. There is nothing there to classify and no structure to find. is not positive definite, so comes out positive, negative or exactly zero, its zero set is a whole cone rather than a single point, and which of the three a given pair of events yields is the causal structure of the world.
The same fact can be stated through the transformations. Those preserving form a compact family, so no amount of whitened rotating takes you far. Those preserving do not, which is why has no upper bound while a correlation coefficient has two.
Two honest differences. is estimated from a sample and is a property of a population, whereas is neither estimated nor a property of anything anybody can vary. And a Mahalanobis distance runs from a point to a distribution, while the interval runs between two events, with no distribution anywhere in it.
Notation deserves attention when it makes a structure visible rather than merely shorter, and two decisions do that here. The first is to measure time in the same units as distance, so that two coordinates entering the transformation on an equal footing are seen to do so instead of being held apart by a conversion factor. The second is to collect the four coordinates of an event under one symbol carrying a label.
The signs then need somewhere to live, and they live in a small square array with one plus and three minuses down its diagonal, called the metric. Its entire job at this stage is to supply a sign to each term. In ordinary space the corresponding array is all plus signs, which is why nobody ever writes it and why the metric is invisible in first-year physics. The whole difference between the two geometries is the difference between those two arrays.
What the arrangement buys is a machine for manufacturing agreement. Take any two collections of four numbers that transform between frames the way a displacement between events does, combine them through the metric, and the number resulting is one every observer computes alike. The interval is the special case where both collections are the same displacement, and the machine will be fed repeatedly: a velocity shortly, a momentum in the next chapter, a current after that.
3 · Boosts are hyperbolic rotations
This is the section the chapter exists for. Everything before it was setup and everything after it is consequence.
3.1 · Two matrices, side by side
Put the Euclidean rotation and the Lorentz boost next to each other and read across.
A rotation in the plane preserves and is built from a pair of functions satisfying :
Now the boost, in the form Chapter 2.2 §6.1 put it into. A boost in the plane preserves , and it is built from a pair of functions satisfying :
The parallel is not loose, so let's set out the three points of contact. Each matrix is built from the two functions whose squares obey the identity matching its own quadratic form. Each preserves that form. Each is parametrised by one real number. And in both cases the parameter is additive:
The first of those follows from the circular addition formulas and the second from the hyperbolic ones, and Chapter 2.2 §6.2 multiplied the matrices out to check. That additivity deserves a name.
The rapidity of Chapter 2.2 is not merely "the variable in which velocities add". It is the hyperbolic angle of the rotation, and it adds for exactly the reason Euclidean angles add: composing two rotations of the same kind about the same axis adds their angles, because that is what a one-parameter group of rotations does.
Chapter 2.2 arrived at by noticing that is the addition theorem and taking the hint. Here the hint is explained. Velocity is , and of an angle is a slope. Slopes never add. Angles do. Three centuries of physics added the slopes.
Two differences are worth flagging before we exploit the parallel, because they are the places where the two geometries part company.
The signs in the off-diagonal entries. The rotation matrix has above the diagonal and below. The boost has in both places. So the boost matrix is symmetric and the rotation is not. That is not cosmetic. Chapter 2.2's Problem 3 used precisely this fact to prove that two boosts in different directions leave a rotation behind.
The range of the parameter. The angle is periodic. Rotate by and you are back where you started, and the orbit of a point closes. The rapidity is not periodic and runs over all of , so you may boost forever and never return. In the same way and are bounded, while and are not.
That last contrast has a name. The rotation group is compact and the boosts are not, and this is the deepest structural difference between them. It is responsible for a great deal in Part VI. It is also the reason has no upper bound while a rotation never stretches anything.
3.2 · The orbits: circles become hyperbolae
Now ask the question that turns algebra into a picture. Take one point and apply every transformation in the family. What curve does it sweep out?
Euclidean. Start at and rotate through every . Since is preserved, every image point satisfies with fixed. The orbit is a circle. And every point of that circle is reached, since the angle covers the full range. So: the orbits of the rotation group are the circles centred on the origin, and a circle is precisely a locus of constant distance from the origin. This is so familiar that it takes an effort to notice it is a theorem.
Minkowski. Start at and boost through every . Since is preserved, every image satisfies
That is the equation of a hyperbola, a rectangular one with asymptotes . So much for the curve the image stays on. We also want to know how the point moves along it, so write the starting point in the form for and apply (2.3.16) to it:
So all the boost does is slide the hyperbolic angle: . Compare the Euclidean statement that a rotation slides the polar angle, . The structure is identical, and the parametrisation is doing exactly the job that does.
The curves are the loci of events at fixed interval from the origin. They are what "a circle of radius " means when the distance function has a minus sign in it, and the three families they fall into are the whole of §4's causal structure, drawn:
- : two branches opening upward and downward, entirely inside the light cone. The upper branch is the set of events at proper time into the future of the origin.
- : two branches opening left and right, entirely outside the light cone, at proper distance from the origin.
- : the hyperbola degenerates into its own asymptotes, the two lines , which are the light cone. This is the one case with no Euclidean counterpart, because has only the single point as a solution while has a whole pair of lines.
A boost moves points along these curves and never across them. Which is to say: a boost can change an event's time coordinate by a factor of a thousand and its position coordinate likewise, and cannot change its interval from the origin by one part in . The figure below is built to make you watch exactly that happen.
One consequence deserves saying out loud, because it kills a persistent confusion. The light cone is an orbit too. A boost slides a null point along the ray , multiplying both coordinates by . (Put into (2.3.19), or just note that .) It can shrink the coordinates toward zero or blow them up without bound, and it can never move the point off the line.
That is Postulate 2 in geometric dress. The light cone is not merely a surface that happens to be invariant. It is the degenerate member of the family of invariant hyperbolae, and no continuous boost can carry a point across it.
3.3 · Why the boosted axes look wrong, and why the ticks are wrong too
Chapter 2.2's figure showed the primed axes scissoring symmetrically toward the light line, and readers invariably ask two questions about it. First: why are the and axes not perpendicular, if is a perfectly good frame with perpendicular axes? Second: where along those axes is "one second", and why is it not where I would put it? Both have the same answer, and we can now give it.
Perpendicularity. In Euclidean geometry, two directions are perpendicular when their dot product vanishes. In this geometry the dot product is (2.3.11), with the minus signs in it. The right word here is orthogonal, and it means .
So let's test the two primed axes against that condition. The axis is the set of events with , so it points along . The axis is the set with , so it points along . Their Minkowski dot product is
They are orthogonal, exactly, at every . They do not look it, because your eye is applying the Euclidean dot product of the page, in which their product is . The axes are perpendicular in the geometry and not on the paper, and it is the paper that is wrong.
There is a pleasing detail here. The two axes are mirror images in the line , since swapping maps one to the other. That reflection is what "scissoring symmetrically about the light line" means, and it is forced by the fact that light itself is orthogonal to itself. A null vector has .
Calibration. Where is the event "one second of time, at 's origin"? By definition it is at , , so its interval from the origin is . Intervals are invariant, so in the unprimed frame it is the point where the axis meets the invariant hyperbola . From the parametrisation, that point is
So the invariant hyperbolae calibrate the axes. To mark off unit ticks along any observer's time axis, intersect that axis with the hyperbolae . To mark off unit ticks along the space axis, intersect it with . The recipe works for every observer, because the hyperbolae belong to no observer.
How far off is the eye? The unit tick on the unprimed axis sits at Euclidean distance from the origin on the page. The unit tick on the axis sits at page-distance
Put into that, where exactly, which is a small gift. The page-distance comes out as . So the primed observer's second is drawn 46% longer than yours. It is not longer. The page is lying, uniformly and predictably, by the factor :
| page stretch |
Once you know that, length contraction and time dilation become visible on a diagram rather than mysterious. Chapter 2.2 §4.3 said that two observers "slice the same four-dimensional object at different angles". The slices are the lines parallel to each observer's axis. The calibration is why the same physical rod, sliced two ways, is assigned two different lengths without anything being squeezed.
This is the single most common source of confusion when reading spacetime diagrams, so it is worth stating in its bluntest form: on a spacetime diagram, longer on the page means shorter in proper time. Your visual system measures distances with because that is the geometry of the paper. The diagram's content is . These do not merely differ by a scale factor. They run in opposite directions.
Numbers make it concrete. Take Chapter 2.2's twins, in years and light-years, from to :
| Worldline | Length on the page | Proper time |
|---|---|---|
| straight, at rest | yr | |
| out to at and back | yr | |
| constant proper acceleration, same turning point | yr |
Perfectly monotone, and backwards. The path that looks longest is the one along which least time elapses. Every drawn segment that tilts away from vertical is buying page-length with proper time, at a rate set by the minus sign. So train yourself to stop trusting the picture for lengths. Keep trusting it for everything else, meaning incidence, ordering, and which side of a light ray something is on, because those it gets right.
3.4 · Seeing it
Grind box — why deserves to be called an angle: it is twice a sector area, exactly as is
Calling an "angle" is so far only an analogy of algebraic form. Here is the fact that makes it more than that, and the computation is one line in each geometry.
Chapter 0.7 §5 derived Green's theorem, one corollary of which is that the area swept by the radius vector from the origin along a curve is
Circle. Parametrise the unit circle as , . Then
Hyperbola. Parametrise the unit hyperbola as , . Then
The same integrand, the same value, the same conclusion: the parameter is twice the area of the sector it sweeps out, in both geometries. That is the honest sense in which is an angle. And it explains additivity without any trigonometry: applying two transformations sweeps two adjacent sectors, and areas add. Angles add because areas add, in the circular case and the hyperbolic case alike.
It also explains why the analogy could not have been avoided. "Angle" was never fundamentally about turning. It was about the invariant measure along the orbit of a one-parameter group, and both groups have one. Chapter 6.1 will call this measure the group parameter and will not care which geometry it came from.
A numerical check, since the claim is easy to test. Take . The sector bounded by the -axis, the hyperbola, and the ray to has area , which is ✓. (The integral evaluates to between the limits.)
Take one event and apply every transformation in the family to it, then ask what curve the image traces. In ordinary space the answer is a circle, since the sum of squares is preserved and the angle sweeps through everything, and a circle is a locus of constant distance from the centre. Flip the sign and the same question returns a hyperbola, with the light rays through the starting point as its asymptotes.
So hyperbolas are the circles of this geometry, the loci of events at fixed interval from a given one. Boosting slides a point along the hyperbola it started on and never carries it to a neighbouring one, which shows up as large changes in both coordinates alongside no change in the interval. The light rays are the degenerate member, where the curve collapses onto its own asymptotes, and that member has no Euclidean counterpart, because a sum of squares vanishes at one point while a difference vanishes along two lines.
This disposes of the complaint that boosted axes look wrong. They are genuinely at right angles in this geometry and fail to look it because the eye applies the geometry of the paper; the invariant hyperbolas supply the tick marks, since they belong to no observer. Longer on the paper means less elapsed time, so the two measures do not merely differ by a scale factor, they run in opposite directions.
4 · Causal structure
In Euclidean space, is positive unless both points coincide. Every pair of distinct points is at a positive distance, and there is nothing more to say. Flip the sign and a new phenomenon appears immediately: can be positive, negative, or zero, and which it is turns out to be the most important thing about a pair of events.
4.1 · Three kinds of separation
For two events with interval :
- Timelike if : . Light has time to spare. A material particle can be present at both.
- Null (or lightlike) if : . Only light connects them, and only just.
- Spacelike if : . Not even light can get from one to the other.
The classification is Lorentz-invariant, and the proof is two words: is invariant, so the sign of is invariant. That is the entire argument, and its brevity is the point. Everything else in this section is unpacking what that one-line theorem means.
It is worth being clear about how strong that is. Whether two events are timelike separated is not a matter of who is looking, how fast they are moving, or which coordinates they chose. It is a fact about the pair, of exactly the same standing as "these two points are three metres apart" in Euclidean geometry. It is a frame-independent property, and it is the first one we have met that is not just a number but a structure.
4.2 · The light cone
Fix an event and ask which events are null-separated from it. With at the origin, the condition reads
At each time this is a sphere of radius , which is the expanding flash of light emitted at . Stack those spheres up the time axis and you get a cone. Do the same for and you get a second cone opening downward, made of the light that could have converged on . Together they are the light cone of .
The name comes from what it looks like drawn. Suppress and and it is the pair of lines in the figure of §3.4. Suppress only and it is an honest cone.
The cone partitions spacetime into three regions, and every one of them is invariant, because the cone is invariant and the sign of inside it is invariant (§4.3):
| Region | Condition | Meaning |
|---|---|---|
| Absolute future | and | Events can influence. Every observer agrees they happen after . |
| Absolute past | and | Events that can influence . Every observer agrees they happen before . |
| Elsewhere | No causal contact in either direction. Time order is frame-dependent. |
The word absolute is doing real work there and is not decoration. It marks the parts of "before" and "after" that survive the loss of absolute simultaneity. Chapter 2.2 destroyed the universal "now", and it is easy to conclude from that wreckage that time order is entirely conventional. It is not. Precisely the causally connectable part of it is absolute. That is the minimum required for physics to make sense, and, as we are about to see, it is also exactly the maximum available.
"Elsewhere" is the genuinely new region, and it has no Newtonian counterpart. In Galilean physics every event is either before, after, or simultaneous with , and the three cases exhaust everything. Here the third of those has swollen from a three-dimensional slice into a four-dimensional region, and it contains, at this instant, everything in the universe more than a few light-nanoseconds away from you. The Andromeda galaxy has no "now". It has a two-million-year-thick band of events about which your question "what is happening there right now?" has no frame-independent answer.
4.3 · The causality theorem
Claim. If two events are timelike or null separated, every inertial observer agrees on which one happened first.
Proof. Suppose . The first thing to extract from that hypothesis is a bound on how far apart in space the two events are, relative to how far apart in time, so write the condition out and then drop the two transverse terms:
The division there is legitimate because . (If then , with equality only for coincident events.) Now choose axes so that the boost is along . That costs nothing, since a rotation changes neither nor the classification. What we want is a comparison between and , so take the time equation of (2.3.2) and factor out of it:
The sign of now rests entirely on the sign of that bracket, because is positive and carries the sign we are testing. So bound the bracket below, using the inequality we have just derived:
The last step used , which is not an extra assumption but a property of every Lorentz transformation. Since and the bracket is strictly positive, has the same sign as in every frame.
Notice exactly where the proof lives: in the collision of two inequalities, from the separation being timelike, and from the observer being an observer. Neither alone is enough. If observers could exceed , or if influences could be spacelike, the bracket could reach zero and change sign. Causality is protected by the product of two things being less than one.
A sharper version comes free. We would like a lower bound on the size of and not just on its sign, so rearrange the invariance relation into and take square roots:
with equality exactly in the frame where . So for a timelike pair, does not merely fail to vanish. It is bounded away from zero by a frame-independent amount, and no observer can even make the two events nearly simultaneous. Worked example 1 puts numbers on this.
4.4 · Spacelike: the order is genuinely up for grabs
Now the other case, and it is not a near miss. Claim. If two events are spacelike separated, there is a frame in which they are simultaneous, a frame in which precedes , and a frame in which precedes .
Proof. Take and rotate the axes so that the spatial separation lies along . Then and , so and we may divide by it:
Now go looking for the frame in which the two events are simultaneous. Reading (2.3.25), the bracket there vanishes, and with it , at the single velocity
By (2.3.28) that value satisfies , so it is an allowed velocity and the frame exists. Now for the two orders. The quantity is a continuous, strictly monotonic function of that passes through zero at , its derivative in never vanishing for spacelike separations. So it is positive on one side of and negative on the other. Boosts with slightly less than give one order, and boosts with slightly more give the other.
In the simultaneous frame the invariance of the interval gives , so the two events are separated by
That quantity is the proper distance. It is the spacelike counterpart of proper time, and it is the distance measured by the one observer for whom the two events are "at the same moment". The two cases mirror each other exactly. Just as is the smallest any observer can report for a timelike pair, is the smallest any observer can report for a spacelike pair, by the identical algebra.
Two events being reversible in order sounds alarming until you notice what it costs to arrange, and what it does not permit. Reversal requires , and that is possible exactly when the separation is spacelike. The condition for the reordering to be achievable and the condition for the events to be causally disconnected are the same condition. That is not luck. It is the theorem below.
4.5 · Why is a causal speed limit and not merely light's speed
Assemble the two results. Time order is absolute for causally connectable pairs, and reversible for causally disconnected ones. Nothing so far forbids a faster-than-light influence. We have just not allowed one yet. Let's allow one and watch what happens.
Suppose some influence travels from event to event at speed in some frame. It could be a signal, a particle, or a "quantum of whatever". Then , so and are spacelike separated. By §4.4 there is then an inertial frame in which precedes . But was the cause and the effect. So:
If any influence could travel faster than , then some inertial observer would see the effect happen before the cause.
That does not mean "would see it that way as an optical illusion". It means the observer would find, in coordinates constructed with their own synchronised clocks, that the effect is recorded at an earlier time than the cause. And by Postulate 1 that observer's description is as valid as anybody's.
An effect preceding its cause is uncomfortable, but it is not yet a contradiction. The contradiction takes one more step, and that step is constructive. With two superluminal signals you can send a message into your own past, and here is how.
The construction
Work with , in seconds and light-seconds, and take the idealised case of a signal that is instantaneous in the frame of whoever sends it. (The finite- version is in the grind box. It needs and is otherwise identical.)
Step 1. Alice is at rest in frame at . At she sends Bob a signal that is instantaneous in , meaning that it travels along the line . Bob sits at light-second, so the reception event is
Step 2. Bob is not at rest in . He moves at , so he is at rest in . We are going to need the reception event in his coordinates, so transform with the boost:
Step 3. On receiving the message, Bob immediately replies with a signal that is instantaneous in his own frame, so it travels along the line . The question is where that line crosses Alice's worldline . On we have , so
Alice receives Bob's reply at : six tenths of a second before she sent the original message. In general the return time is , so the effect is available at will and can be made as large as you like by increasing .
Step 4, the contradiction. Alice now writes a program: if a reply arrives before , do not send the message. The premise of the loop is that the message was sent. The conclusion is that it was not. This is not a paradox of interpretation, a puzzle about free will, or a subtlety about what "cause" means. It is a straightforward logical contradiction, and it is produced by three ingredients: the existence of superluminal signalling, the relativity of simultaneity, and Postulate 1, which grants Bob the same right as Alice to call his own frame's instantaneous signal instantaneous. Since the last two are established, the first must go.
This is why the constant is a causal speed limit and not merely "the speed of light". Nothing in the argument above mentioned light. The limit is a property of the causal structure of spacetime, of the light cone as the boundary between "can influence" and "cannot". Light travels at because light is massless, and not the other way round.
The distinction is not pedantry, and it pays off twice later. Suppose the photon turned out to have a tiny mass. Light would then travel slightly slower than , and not one line of this chapter would change. The constant would remain the limit, and light would merely stop saturating it.
The second payoff is in general relativity, where there is no global speed to speak of. What survives there is exactly this: the light cone at each point, the local causal structure, which Part III keeps while discarding almost everything else. That is why Chapter 3.1 can throw away flat spacetime and still have a theory.
Grind box — the loop with a finite superluminal speed, and the exact threshold
The instantaneous case above is clean, but it looks like a limiting trick. It is not one. Any works, provided the relative velocity of the two parties is large enough, and here is that condition derived.
Units . Alice at sends at a signal of speed in toward , and Bob receives it at
Bob is at rest in with velocity , so transform that reception event into his coordinates:
Bob replies with a signal of speed in his own frame, travelling in the direction: . Alice's worldline is, in , the line . Setting the two equal and solving,
Substitute the two transformed coordinates and factor out :
Alice's own clock reads on her worldline, so the reply arrives at
Since , the sign is the sign of the bracket, and the reply beats the message () exactly when
Sanity checks. The threshold is less than for every , since ✓, so an allowed always exists. As the threshold goes to : with truly instantaneous signals any relative motion suffices, matching the main text. As the threshold goes to : at exactly light speed you would need an observer at exactly light speed, of which there are none, and causality is safe. The barrier is not a convention. It is precisely where the construction stops working.
Numbers. Take , so the threshold is . With light-second:
| reply arrives at (s) |
At threshold the reply arrives exactly as the message departs; beyond it, before.
One honest caveat. What the argument rules out is superluminal signalling, meaning the transmission of information or influence, which is inconsistent with relativity plus logic. It does not forbid superluminal coordinate speeds, which Chapter 2.2 §5 already showed are harmless (the spotlight, the scissors), and it does not by itself forbid a particle that always moves faster than and cannot be used to signal. ⚑ Whether such a particle is consistent quantum-mechanically is a separate question with a well-known answer, which is that it is not, and the argument there is Chapter 5.1's rather than ours.
Chapter 0.5 §1 laid down the axioms an inner product must satisfy, and one of them was positive-definiteness: for every . Everything downstream in that chapter leaned on it. The Cauchy–Schwarz inequality needs it. So does the triangle inequality, the definition of as a real number, the interpretation of as a cosine, and the whole apparatus of angles.
does not satisfy it, and the violation is deliberate. The vector is not zero, and . So we have a nonzero vector of zero length, and vectors of negative length squared, and no way to define as a real number for all . Calling a "metric" is standard and is a slight abuse. The precise term is a non-degenerate symmetric bilinear form of signature , and the word non-degenerate marks the axiom we did keep.
What survives: bilinearity, symmetry, non-degeneracy (if for every then ), the invariance of under the transformations, the notion of orthogonality , and, most important of all, the ability to classify vectors into three invariant types. That is enough to do geometry with, and §5 and §6 go on to do exactly that.
What does not: positivity, the triangle inequality (§6 reverses it), Cauchy–Schwarz (§6 reverses that too, for timelike vectors), and the idea that "distance zero" means "same point". Vectors orthogonal to themselves exist, and they are exactly the null vectors. That sentence would be flatly false in Chapter 0.5, and here it is the defining feature of the light cone.
So the honest summary is: we removed one axiom from the definition of an inner product, and in exchange the geometry acquired a causal structure. That trade is what the rest of physics is built on.
In ordinary space the distance between two distinct points is positive and there is nothing further to say. Reverse one sign and the same quantity comes out positive, negative or exactly zero, and which of the three it is turns out to be the most important thing about the pair. The proof that the classification belongs to the pair rather than the observer is two words long: the quantity is invariant, so its sign is.
What has appeared is not another agreed number but an agreed structure: what survives of before and after once the universal present is gone. Events light or matter could reach keep their order in every frame, while events too far apart for anything to cross between have an order different observers reverse without anybody being wrong, since nothing was ever at stake in it. Time order is absolute where it could matter and negotiable where it could not.
The speed limit follows, and it belongs to the geometry rather than to light. Let some influence outrun it and the two events it joins are of the negotiable kind, so some observer records the effect before the cause. Two such influences are worse: one person signals another, the second replies with an influence instantaneous in his own frame, the reply arrives before the first was sent, and she declines to send it. Light travels at the limit because light is massless.
5 · Proper time is the length of a worldline
So far the interval has connected two events by a straight line. Now let's bend it.
A particle's history in spacetime is a curve, called its worldline. Along that curve, take two neighbouring events separated by . Suppose the particle moves slower than light, which by §4 is exactly the statement that its worldline is timelike at every point. Then , and we may define
What we want is expressed through the coordinate time and the particle's ordinary speed, since those are the things an observer measures. So divide the right-hand side by and pull back out in front:
Here is the particle's instantaneous speed in units of . It is a function of , because the particle may do whatever it likes along the way. Now add the contributions up by integrating along the path:
This quantity is the proper time along the worldline. Three things about it, in increasing order of importance.
It is invariant. Each is invariant, and a sum of invariants is invariant. Two observers computing (2.3.36) for the same path will use different , different , and different intermediate values, and will get the same number.
It is a functional of the path, not a function of the endpoints. This is the language of Chapter 1.2. Here takes an entire curve as input and returns one number, in precisely the way that the arc-length functional of 1.2 §1 does. Set the two of them side by side:
There is the sign again, in the only place it could possibly be. Proper time is arc length. It is the length of the worldline, measured with the geometry that spacetime actually has. Everything §6 does follows from taking that sentence at face value.
It is what a clock reads. This is the part that keeps proper time from being an abstraction. Take the interval between two neighbouring events on a clock's own worldline. In the clock's instantaneous rest frame it is not moving, so and . That is to say, the interval is the time elapsed on that clock. Since is invariant, the same holds computed in any frame at all.
So is the accumulated reading of a clock carried along the path. And (2.3.36) says that a moving clock accumulates less of it than the coordinate time, by exactly at each instant. That is Chapter 2.2's time dilation, now derived as a statement about the length of a curve rather than about clocks being defective.
Grind box — the assumption hidden in "a clock measures proper time"
The argument above computed in the clock's instantaneous rest frame, and then added the pieces up. There is an assumption buried in "instantaneous", and it deserves naming because it is a physical input rather than a mathematical one.
The clock hypothesis. An ideal clock's rate depends only on its instantaneous velocity, not on its acceleration. Equivalently: an accelerating clock reads the same as a momentarily co-moving inertial clock, at every moment.
This is not a theorem of special relativity. It is an extra assumption, and it is easy to see that it could fail: a pendulum clock in an accelerating rocket manifestly does not obey it, and neither does any device whose mechanism is disturbed by being shaken. What the hypothesis really asserts is that some physical processes are ideal in this sense, so that "proper time" is realisable and not merely definable.
⚑ Experimentally it holds to an astonishing degree. Muons circulating in a storage ring at experience a proper acceleration of order , and their decay rate is slowed by exactly the factor predicted from their speed alone, with no acceleration-dependent correction at the level. Atomic clocks flown on aircraft and the clocks aboard GPS satellites agree with (2.3.36) integrated along their actual trajectories. We quote those results. Deriving them would require the internal dynamics of the clock.
Why it matters twice. First, without it, §6's twin calculation would be incomplete, because you could always object that the turnaround does something unmodelled to the traveller's clock. Second, in general relativity the clock hypothesis is promoted to a postulate. There, proper time along a worldline is defined to be what a clock carried along it reads, and that definition is the bridge between the geometry and any measurement at all. Chapter 3.3 leans on it in its first paragraph.
Straight separations have carried everything so far, and no real history is straight. Bend the path, chop it into pieces short enough to be straight, take the interval along each and add them, and what accumulates is the proper time. It is a functional of the whole path rather than a function of its endpoints, in the sense the action was, and it is the arc length of the worldline in the geometry spacetime actually has.
Set it beside the arc length of a curve in the plane and the two expressions differ in one place only, the sign under the root. That single difference does all the work of the argument to come. The quantity is invariant besides, because each piece is, so two observers who disagree about every intermediate number agree about the total.
It is also what a clock reads, which keeps it from being an abstraction: in a clock's own momentary rest frame the interval between neighbouring events on its path is the time it displays. Dilation stops being a statement about defective clocks and becomes one about the length of a curve. One physical assumption hides in the word momentary and deserves naming: an ideal clock's rate depends on its speed and not on its acceleration. That is no theorem, a pendulum in a launching rocket violates it, and it holds experimentally to a degree hard to credit.
6 · The twin paradox is a theorem about triangles
Now the payoff. In the plane, the straight line is the shortest path between two points, and everyone has known that since childhood. In spacetime the corresponding theorem is true with the inequality turned around, and the reversal is the minus sign, doing its most consequential work.
6.1 · The theorem
Among all timelike worldlines connecting two fixed timelike-separated events, the straight one has the greatest proper time. Every other path takes strictly less, and the more it deviates, the less it takes.
Proof. Let the two events be and , timelike separated. The whole proof rests on picking the right frame first, so let's pick it. By §4.3 and the invariance of the interval there is a frame in which the two events occur at the same place. Take , which satisfies precisely because the separation is timelike, exactly mirroring the construction in §4.4.
Work in that frame from here on. There , and the coordinate time between the events is , which by (2.3.27) is .
The straight worldline joining them is the one at rest in this frame. Along it , so (2.3.36) gives .
Now take any timelike path between the same two events. It has the same endpoints, so it runs over the same coordinate-time interval, and that lets us compare the two proper times integral by integral:
The inequality holds because for every , with equality only at . So every path loses, and it loses strictly unless it is at rest for the entire journey, which is to say unless it is the straight one.
That is the whole proof. It is four lines because we chose the right frame first, and we were allowed to choose it because the answer is frame-independent. The integrand is a penalty for moving, charged continuously, and the straight path is the one that never pays it.
6.2 · The same result from Chapter 1.2's machinery
The direct argument is airtight but it uses a special frame. The variational argument does not, and it is the one that survives into Part III, so it is worth doing even though we already have the answer.
Extremise the functional (2.3.36). In one space dimension, with ,
This is a Chapter 1.2 problem of exactly the type solved in its worked example 2. The Euler–Lagrange equation is
Now notice that contains no at all, which is the statement that spacetime is homogeneous and has no preferred place in it. So the second term vanishes, and what is left says that the first bracket is a constant of the motion:
Solve for exactly as 1.2 did for the plane geodesic: square, clear the denominator, and collect,
Constant velocity, which is to say , a straight worldline. We assumed nothing about the answer, only stationarity.
There is a free bonus sitting in that last expression. Whatever real constant you pick, the resulting comes out strictly less than . The extremal paths are automatically timelike, and the variational principle cannot produce a superluminal answer even if you ask it to.
The variational calculation identifies the straight line as stationary. Which kind of stationary point it is has to come from somewhere else, and §6.1 supplies the answer: a maximum. Chapter 1.2 §5.1 made the same distinction for the action and warned that stationary need not mean minimal. Here is the cleanest possible example of that warning being necessary.
Look at what the calculation actually needed. It needed a functional of the form with , and the Euler–Lagrange machinery of Chapter 1.2. It did not need flatness, it did not need to be constant, and it did not need Lorentz transformations.
Chapter 3.3 will keep this variational principle word for word and make exactly one change. It replaces the constant array by a position-dependent one, , which is ten functions of where you are. The same functional, the same Euler–Lagrange equation, the same request for a stationary path. The resulting equation is called the geodesic equation, and its solutions are called free fall.
That is the whole of general relativity's kinematics, and this section is where it starts to be earned. Chapter 1.2 already showed you the pattern once: the same gave a straight line in the plane and a great circle on a sphere, and the only thing that changed was . Gravity is that substitution performed on spacetime. A planet orbits the Sun for the same reason a great-circle route from London to Tokyo goes over the Arctic. It is going as straight as the geometry allows.
6.3 · The reverse triangle inequality
The theorem of §6.1 has a clean algebraic form worth stating on its own, because it is the exact mirror image of the most familiar inequality in geometry.
First the vocabulary. Call a four-vector timelike if , and future-pointing if . For such vectors write , which is a real number. With those two words in place the theorem reads:
with equality only when and are parallel. Compare the Euclidean statement , and notice that the inequality has flipped. In the plane, going via a third point makes the journey longer. In spacetime it makes the elapsed time shorter, and the two-leg twin path is precisely , with the outbound leg and the inbound one. Problem 1 asks you to prove (2.3.43). The proof turns on a reversed Cauchy–Schwarz inequality , which is also flipped, and flipped for the same reason.
6.4 · Closing the loop with Chapter 2.2
Chapter 2.2 §7.1 worked the twin paradox in detail. The traveller goes light-years at , turns, and returns. Earth ages years and the traveller ages . The resolution offered there was that the traveller changes inertial frames while the stay-at-home does not, followed by a careful accounting of the years that the traveller's simultaneity slice sweeps across Earth's worldline during the turnaround. All correct, and all now unnecessary.
Here is the same result in one sentence. The travelling twin ages less because a bent timelike path is shorter, and shorter means less proper time.
Two paths, same endpoints, different lengths. Nobody is scandalised that a detour adds mileage. The only unfamiliar thing is that in this geometry a detour subtracts time.
Notice what the geometric statement does not require. It never mentions acceleration, it never uses a non-inertial frame, and it never needs the traveller's point of view at all. Acceleration enters only as the answer to a different question, which is why is one path bent? And a kink in a worldline is detectable from the inside, by an accelerometer, which is why there was never a symmetry between the twins to appeal to.
The often-heard slogan "the twin paradox is really about acceleration" is therefore half wrong. Acceleration is what makes a path bent. Bentness is what costs proper time. And the amount of time lost is not a function of the acceleration but of the whole shape of the path. Worked example 2 makes that concrete by comparing three paths between the same two events, one of them smoothly accelerated throughout with no kink anywhere.
Grind box — the second variation, and why the extremum is a maximum rather than a minimum
Section 6.1 proved the maximum by a global argument. Here is the local version, which is what Chapter 1.2 §5.1's second-variation machinery would say, and it exposes the sign in a different place.
Work in the frame where the endpoints coincide spatially, and perturb the straight path by a small with . Then and
by the binomial series of Chapter 0.3. Read off the two variations:
The first variation vanishes for every perturbation, confirming stationarity without solving anything. The second variation is negative definite, meaning strictly negative unless , which is to say unless the perturbation is nothing at all. Hence a maximum, and a genuine one rather than a saddle.
Where the sign came from. Compare the Euclidean calculation for a straight line in the plane, perturbed the same way: , giving : a minimum. The two calculations are identical except for the sign inside the square root, which is the sign in the metric, which flips the sign of the second variation, which flips minimum into maximum. One symbol, all the way through.
And a warning that will matter in Part III. This is a local statement. It says the straight path beats all nearby paths. In flat spacetime the global statement is true as well, and that is what §6.1 proved. In curved spacetime it can fail. There can be several geodesics between the same two events, and only one of them maximises. That is the conjugate-point phenomenon of Chapter 1.2 §5.1, and in Chapter 3.8 the several geodesics become the several images of a gravitationally lensed quasar.
The most familiar sentence in geometry is that a straight line is the shortest route between two points, and here it fails in the most interesting available way. Among all the paths a material object could take between two given events, the unaccelerated one accumulates the greatest elapsed time, and every other loses strictly.
The proof is four lines because the frame may be chosen first, and it may be chosen because the answer does not depend on the choice. Work where the two events happen at the same place. The path that stays put accumulates the full coordinate time, while every other pays a penalty at each instant for moving, a penalty never zero and never negative. So the travelling twin ages less for the same reason a detour through the next town adds mileage, except that in this geometry the detour subtracts, and the triangle inequality has turned around.
Two debts are settled by that. The warning that a stationary path need not be a minimum, issued when the action principle arrived, gets its cleanest illustration, since this stationary path is a maximum. And the calculation needed nothing but a functional of that shape and the machinery of the last part, not flatness and not constancy of the array of signs. Replace that array by one whose entries vary from place to place and the identical calculation returns the paths of free fall.
7 · Four-velocity, briefly
One construction before we stop. It is short, it is forced, and Chapter 2.5 is built entirely on it.
Ordinary velocity is , and Chapter 2.2's Problem 1 exposed its defect. The numerator is a piece of a four-vector, but the denominator is a coordinate, so the quotient transforms in a mess. Both the top and the bottom change when you change frames, and they do not change compatibly.
The repair is to divide by something invariant instead, and §5 has just supplied the only natural candidate. Define the four-velocity
Since transforms as a four-vector and is a number every observer agrees on, is a four-vector. The numerator picks up a factor of and the denominator picks up nothing. That is the entire reason for the definition.
We will want its components in terms of ordinary velocity, and they follow at once from together with the chain rule:
And now the property that makes this the right object to have defined. Feed into the invariant-manufacturing machine of §2.3 and compute its own square:
That holds identically, for every particle, at every moment, in every frame, whatever it is doing. Check it against the components: ✓.
The four-velocity is a four-vector of constant Minkowski length . It can point in any timelike future direction, and that is all the freedom it has. It lives on the invariant hyperbola of §3, which in four dimensions is a three-dimensional hyperboloid, and boosting a particle slides its four-velocity along that surface exactly as the figure in §3.4 slides an event along a hyperbola.
This is why and not is the right generalisation of velocity. The constraint is frame-independent, whereas "speed less than " is a statement about components. Chapter 2.5 multiplies by the rest mass to get , whose invariant square is then with no work at all, and that single equation contains , , and the massless case as its three readings.
You will hear it said that everything moves through spacetime at speed , and that time dilation is what happens when motion through space is "bought" out of motion through time. As a mnemonic it is not bad. As a derivation it is not one, and it is worth separating the two.
True: for every particle, by (2.3.47). That is a real theorem and it is what the slogan is gesturing at.
Careful: the components of are , and these do not trade off the way a budget does. Writing the invariant out,
So as spatial motion grows, the "motion through time" grows too. It does not shrink. There is no fixed pot of being divided up. The two pieces combine with a minus sign, which is precisely why the total can stay at while both pieces run off to infinity. A picture built on "you only have so much speed to share out" will give you the wrong sign on the first question that tests it.
The honest version of the slogan: a particle's four-velocity is a unit timelike vector (scaled by ), and what encodes the velocity is its direction rather than its magnitude. Time dilation is the statement that a unit vector tilted away from your time axis has a smaller projection onto it, and that projection is . That is a statement about a fixed-length vector being tilted, which is exactly §3's picture, and it is derivable. The budget metaphor is not.
One construction remains, and it is forced rather than chosen. Ordinary velocity has a defect that surfaces once frames matter: its numerator and denominator belong to different worlds, the displacement being part of a four-dimensional object while the elapsed time is one observer's coordinate. The repair is to divide by something nobody disagrees about, and the worldline's length is the one candidate.
What comes out has a fixed length, the same for every particle at every moment whatever it does, so all its freedom is in its direction. Dilation then reads as a fixed-length object tilted from your time axis having a smaller projection along it. The slogan about everything moving through spacetime at one speed gestures at this and is a mnemonic rather than a derivation, since the two pieces combine with a minus sign and both run to infinity rather than trading against each other.
That is as far as three chapters carry it. The speed limit is built into the geometry rather than into any material; observers slice one fabric at different angles and agree on the interval; the straightest history carries the most time. But the word used throughout for a collection of four numbers has been a definition by resemblance, and the up-and-down placement of its labels has been a spelling rule obeyed on trust. Saying what makes such a collection an object rather than a list comes next.
8 · Worked examples
Four events are given in a frame , in units where : time in years, distance in light-years, so coordinates are written with throughout.
| Event | ||||
|---|---|---|---|---|
(a) Classify all six pairs. (b) For the spacelike pair , find the frame in which they are simultaneous and a frame in which their order is reversed, and verify the interval is unchanged. (c) Show that no boost reverses the timelike pair , and find the frame that comes closest.
(a) The classification. Compute for each pair. Six pairs, six subtractions:
| Pair | Type | Invariant meaning | |||
|---|---|---|---|---|---|
| timelike | yr | ||||
| spacelike | ly | ||||
| null | on the light cone | ||||
| spacelike | ly | ||||
| spacelike | ly | ||||
| timelike | yr |
Now read the structure off. and both lie in the absolute future of 's light cone or on it, with strictly inside and exactly on it, which means a light signal emitted at arrives precisely at . is in 's elsewhere, so nothing that happens at can affect , and nothing at can affect . Yet can affect (), and so can . Causal influence is not a transitive-and-total ordering. It is a partial order, and the light cones are its structure.
Note also that the classification survives translation: only differences entered, so moving the origin changes nothing. That is Chapter 2.2's homogeneity, still holding.
(b) The spacelike pair . In , happens years before ( going from to ) and light-years to the right.
Simultaneous frame. By (2.3.29),
which has modulus less than ✓, as guaranteed for a spacelike pair. That frame moves in the direction at two thirds of light speed. In it, and
and ✓, which is the proper distance, exactly as (2.3.30) promised. The two events are light-years apart and simultaneous, for that observer.
Reversed frame. Any beyond in the same direction does it; take , :
The sign of has flipped. In it was that came first, and in this frame does. And the interval is untouched:
Both frames are correct. Neither event caused the other, and neither could, being spacelike separated, so nothing whatever is at stake in the disagreement.
(c) The timelike pair , which no boost can reorder. Here , . From (2.3.25),
Since , the bracket is positive for every admissible . At worst gives , still positive, while . So always, and in fact can be made arbitrarily large but never small. Tabulate:
| (yr) | |||||
| (ly) | |||||
The last row is the point of the whole exercise: two coordinates swinging over an order of magnitude and changing sign, and one number sitting perfectly still.
The closest approach. The minimum of is at , the frame in which the two events happen at the same place, and there
confirming (2.3.27). A clock can be present at both events, since they are timelike separated. It needs to travel light-year in years, which is . Such a clock reads years between them, and that is the proper time of the pair. No observer can shrink the gap below it, and none can make the two events simultaneous, let alone reversed.
The moral. Time order is not "relative". It is absolute exactly where it could matter and relative exactly where it could not, and the boundary between the two regimes is the light cone. That is a far more disciplined statement than "everything is relative", and it is the one the theory actually makes.
Two events: departure and reunion , again in years and light-years with . Compute the proper time along three timelike paths joining them: (a) the straight one; (b) Chapter 2.2's two-leg trip out to at and back; (c) a smooth worldline of constant proper acceleration with the same turning point. Verify that the straight one wins.
(a) Straight. The path has , so (2.3.36) gives
Equivalently, straight from the interval: .
(b) Two straight legs. Outbound from to , inbound from to . Each leg is straight, so its proper time is the interval along it:
Speed on each leg: , so and ✓ by the other route. This is Chapter 2.2 §7.1's traveller, and the deficit is years.
(c) Constant proper acceleration. Chapter 2.2's Problem 4 showed that a worldline of constant proper acceleration is a hyperbola , which §3 now identifies as an invariant hyperbola, translated. Take the branch that leaves , turns around, and comes back, symmetric about :
where was fixed by demanding , and then follows by symmetry. Choose so that the turning point matches path (b)'s, namely :
Is it timelike everywhere? Write . Then , whose modulus is less than for every ✓. The fastest it ever goes is at the endpoints, .
Proper time. Along the path,
so , which integrates to an inverse hyperbolic sine:
Since , we get (the square root comes out to exactly , so the argument is exactly ), hence
(Checked by direct numerical summation of over two million steps along the path: , against the closed form .)
The comparison.
| Path | Proper time | Deficit | Length on the page |
|---|---|---|---|
| (a) straight | yr | — | |
| (b) two legs, | yr | yr | |
| (c) constant proper acceleration | yr | yr |
The straight path wins, as §6.1 guarantees it must. And the ordering of the last column is exactly the reverse of the second, with more page-length going together with less proper time, monotonically. That is the ⚠ callout of §3.3 in numerical form.
Two things worth extracting.
Acceleration is not the mechanism. Path (c) has no kink anywhere. It is smooth, its proper acceleration is constant, and it never changes inertial frame abruptly. It nonetheless loses more proper time than path (b), which does have a kink. So "the twin who accelerates ages less" is not a law. The amount lost depends on the whole shape of the path, not on where or how hard it accelerated. The correct statement remains the geometric one: bent is shorter, and (c) is bent more.
The numbers are not exotic. The proper acceleration of path (c) is , and Chapter 2.2's Problem 4 gives light-years, so
A crew accelerating at , slightly less than standing on Earth, makes a round trip that takes ten years by Earth's clocks and four years eleven months by theirs, reaching light-years out and of light speed at the extremes. This is not a thought experiment about impossible machines. It is a comfortable ride, and the only obstacle is fuel, which is Chapter 2.5's problem.
9 · Your turn
Problem 1 — the reverse triangle inequality
Let and be future-pointing timelike four-vectors: , , , . Write . (a) Prove the reversed Cauchy–Schwarz inequality , with equality only when and are parallel. (b) Deduce . (c) Explain in one sentence which Euclidean fact each of these is the mirror image of, and where the sign flip entered.
Solution
(a) Reversed Cauchy–Schwarz. Both sides of the claimed inequality are invariant, so we may evaluate them in whatever frame is convenient. That is the standard and enormously useful move here, and it is licensed by (2.3.4). Choose the rest frame of . Since is timelike and future-pointing, §6.1's construction supplies a boost making its spatial part vanish, so
because then and . In that frame write . The dot product is easy:
Now use 's own invariant, , to write . We take the positive root, since is future-pointing, and by the causality theorem of §4.3 that property is frame-independent. Hence
with equality if and only if , i.e. is at rest in 's frame, i.e. .
The whole reversal happened in one place: enters with a plus sign after the minus in the metric has been moved to the other side. In Euclidean space the corresponding step subtracts and you get .
(b) The triangle inequality. is future-pointing (its time component is a sum of positives) and timelike (by (a): ), so is defined. Expand and apply (a):
Take square roots, which is legitimate because both sides are positive, and the result is , with equality only for parallel vectors.
(c) The mirrors. (a) mirrors of Chapter 0.5 §1, and (b) mirrors . Both flip because the metric is not positive-definite: in Chapter 0.5 the proof of Cauchy–Schwarz went by demanding for all , and that step is exactly the axiom §4's ⚠ callout removed.
What it says physically. Take and to be the two legs of the twin's journey, as four-vectors from departure to turnaround and turnaround to reunion. Then is the straight path, (stay-at-home's proper time) and (traveller's). The inequality is the twin paradox, and it is a two-line theorem about vectors. Numerically, with and in the units of §6.4: , , ✓, and ✓, with plenty to spare, because the two legs are far from parallel.
Problem 2 — simultaneity and reversal for a spacelike pair
Two events in a frame , in years and light-years with : and . (a) Classify the pair. (b) Find the frame in which they are simultaneous, and the distance between them in that frame. (c) Find a frame in which precedes , and verify the interval. (d) Show that no frame can make them occur at the same place, and say why that is the mirror image of the timelike case.
Solution
(a) , so the pair is spacelike. Nothing at can influence , because a signal would need speed .
(b) Simultaneous frame. Set in :
Then
and ✓, the proper distance (2.3.30). Check the interval: ✓. An observer moving at says these two events happen at the same moment, light-years apart.
(c) Reversed frame. Any works. Take , :
In , came first by years. In this frame comes first by years. Both are correct, and nothing is at stake.
(d) No common-place frame. Setting would need , which exceeds and is not a velocity. More robustly, from invariance in every frame, so can never vanish and in fact can never fall below .
That is the exact mirror of the timelike result (2.3.27), where and no frame can make the events simultaneous. Swap the roles of space and time and the two statements map onto each other, which is what you should expect of a geometry whose only asymmetry between them is a sign. In slogan form: a timelike pair has an invariant time and a negotiable separation. A spacelike pair has an invariant separation and a negotiable time.
Problem 3 — the accelerated worldline is a hyperbola, and it has a horizon
A rocket has constant proper acceleration . Chapter 2.2's Problem 4 gave its rapidity as and its worldline, in the frame where it is momentarily at rest at , as
(a) Show that this is an invariant hyperbola of §3 and identify which one. (b) Find its asymptotes. (c) A beacon sits at the spatial origin and flashes at times . Show that flashes with never reach the rocket, and that flashes with always do. (d) Identify the resulting Rindler horizon and say precisely what it is a horizon for.
Solution
(a) It is an invariant hyperbola. Use :
a constant. So the worldline lies on , which is one of §3's spacelike hyperbolae, the right-hand branch, with , so that the proper distance from the origin is . Write for brevity. It is a length, and it is the rocket's distance from the origin at .
This is a striking fact worth pausing on. A boost slides points along these curves ((2.3.19)), so the uniformly accelerated worldline is an orbit of the boost group. Boosting the rocket by gives you the same worldline, reparametrised. That is the exact analogue of "a circle is carried to itself by rotation", and it is why this trajectory is the spacetime counterpart of uniform circular motion: the curve of constant "curvature", generated by a one-parameter symmetry. It is also why Chapter 2.2's crew feel a constant forever. Their situation is literally unchanged by the passage of proper time, up to a boost.
(b) Asymptotes. As , , so the asymptotes are : the light cone of the origin. The rocket approaches the speed of light without reaching it, forever, hugging the cone from outside.
(c) Which flashes arrive. A flash emitted at travels along . It reaches the rocket when
Square both sides, which is permissible provided , a condition we check afterwards:
Note that cancels, so the equation is linear in and there is exactly one candidate meeting time.
Case . Require :
which is impossible. Case . The formula has zero in the denominator, so there is no finite solution. Case . Dividing by now reverses the inequality, and the condition becomes , which is always true. So the flash does arrive, after a delay that grows without bound as .
Numbers, with and Chapter 2.2's : a flash at is caught at ; one at is caught at ; one at at . The last light to make it takes forever to arrive, arbitrarily redshifted.
(d) The horizon. The dividing surface is the null line , which is the future light cone of the origin and simultaneously the asymptote from (b). Generalise the calculation to a flash from any event . The ray eventually overtakes the asymptote if and only if . So
Events on the far side of are permanently invisible to it. That surface is the Rindler horizon.
What it is a horizon for, stated carefully. It is not a property of spacetime. This is flat Minkowski spacetime, with no matter, no curvature, and no singularity anywhere. It is a property of the observer. It is the boundary of the region from which that particular eternally accelerating worldline can ever receive news. An inertial observer sails across it without noticing. Different accelerations give different horizons, and if you stop accelerating then yours vanishes and all the delayed signals arrive.
⚑ Three forward pointers, all quoted. (i) The region that the rocket can explore is called the Rindler wedge, and the coordinates adapted to the family of such observers are Rindler coordinates. (ii) Chapter 3.8 finds a horizon at the Schwarzschild radius of a black hole with exactly the same local character, a one-way surface with no local marker on it, and the argument that it is not a singularity is essentially the one above. (iii) Chapter 3.1 uses the equivalence principle to argue that a uniformly accelerated observer is locally indistinguishable from one at rest in a gravitational field, at which point the fact that acceleration alone can manufacture a horizon in empty flat space stops being a curiosity and becomes a clue.
Problem 4 — counting the Lorentz group
(a) Verify that a spatial rotation with satisfies (2.3.13), so rotations are Lorentz transformations. (b) By counting independent equations in , show that the solutions form a six-parameter family. (c) Identify the six as three boosts and three rotations by writing for infinitesimal and showing that is antisymmetric. (d) Connect the count to Chapter 2.4's arithmetic for antisymmetric tensors.
Solution
(a) Rotations qualify. In block form with ,
That makes sense in hindsight. A rotation leaves alone and preserves , so it preserves . The Lorentz group therefore contains the rotation group as a subgroup, namely the transformations relating two observers who are at rest with respect to each other but have turned their axes.
(b) Six parameters. is a matrix, so it has unknown entries. The condition is an equation between two matrices, so naively that is equations. But both sides are symmetric. The left side is symmetric because , using , and the right side is . A symmetric matrix equation carries only
independent equations. Hence
(The counting assumes the ten constraints are independent, which they are; (c) confirms it by exhibiting a six-dimensional solution space explicitly at the linear level.)
(c) Three and three. Put with small and keep first order:
Define the matrix , whose entries are , so the index has been lowered, in the sense Chapter 2.4 §4.1 makes precise. Then
So : antisymmetric. Its independent entries are the strictly-upper-triangular ones, and they split naturally:
- for , mixing time with a space direction. Three boosts, one per axis.
- for , mixing two space directions. Three rotations, one per plane , , .
Total , matching (b) ✓. Now check the shape against Chapter 2.2's boost. Expanding to first order in gives , and lowering the first index with turns the second of these into while , which is antisymmetric ✓. Note that itself is not antisymmetric. The property appears only after lowering, which is a first hint that index position carries meaning.
(d) The connection to Chapter 2.4. Chapter 2.4 §7.2 counts the independent components of an antisymmetric rank-2 tensor in dimensions as , giving in spacetime. It remarks there that six is not four, so such an object cannot masquerade as a four-vector the way masquerades in three dimensions. Here is the same six, arrived at completely independently: not by counting slots in a tensor, but by counting how many ways there are to move while preserving .
They are the same six for a reason. The generators of the Lorentz group are an antisymmetric rank-2 object, and Chapter 2.6 will find that the electromagnetic field tensor is another one with the same six slots: three of and three of . ⚑ That the match is structural rather than coincidental is Chapter 6.1's business. The field strength of a gauge theory lives in the Lie algebra of its symmetry group, and for Lorentz symmetry that algebra is exactly the antisymmetric matrices counted above.
You turned Chapter 2.2's algebra into a geometry. The interval is invariant, proved by direct substitution. The Lorentz transformations are defined as the linear maps preserving it, exactly as rotations are the maps preserving . And the whole subject is Euclidean geometry with one sign changed.
You met , , the summation convention, and , all of it informally, with the debt to Chapter 2.4 recorded. You saw that boosts are hyperbolic rotations, that the rapidity is the hyperbolic angle and adds because areas add, that the orbits are the invariant hyperbolae, and that those hyperbolae calibrate every observer's axes and explain why the page misleads.
You classified separations as timelike, null or spacelike, proved the classification and the causal order are invariant, proved that spacelike order is always reversible, and built a closed causal loop from two superluminal signals. That is why is a limit on causation and not a fact about light.
You defined proper time as the length of a worldline, proved that the straight worldline maximises it, both directly and variationally, and identified the twin paradox as the reverse triangle inequality. And you met the four-velocity, an object of permanently fixed length .
Where this gets spent. Chapter 2.4 takes every piece of notation introduced here on trust, meaning , , upper and lower indices, and "transforms like ", and turns all of it into definitions, with the transformation law as the definition of a tensor. Its opening sentence is about the debt this chapter incurred. Chapter 2.5 feeds the four-velocity of §7 into , gets from (2.3.47) in one line, and reads off it. Chapter 2.6 finds the six components counted in Problem 4 sitting inside the electromagnetic field tensor. Chapter 3.1 keeps the light cone and throws away everything else, which is what "spacetime is locally Minkowski" means. Chapter 3.2 builds the manifold on which that statement can be made. Chapter 3.3 takes (2.3.36) unchanged, replaces by , and calls the extremal paths gravity. That is the single most important thing this chapter sets up. Chapter 3.8 meets Problem 3's horizon again, this time around a black hole. And Chapter 7.3 returns to the observation that the invariance group of a metric is where the physics is, in two dimensions, where that group turns out to be infinite-dimensional and the string becomes solvable because of it.