Part II · Special Relativity — Chapter 2.5
Relativistic Dynamics
Conservation laws have to be statements about four-vectors, or they are not laws. Everything else in this chapter is the bill for that sentence.
Chapters 2.2 to 2.4 rebuilt kinematics. We know how coordinates transform, we know what is invariant, and we know what a tensor is. We also know why writing a law as an equation between tensors makes it frame-independent by inspection.
What we do not yet have is any mechanics. Momentum is still , energy is still , and force is still . All three came over unchanged from Chapter 1.1, and all three rest on a Galilean picture that Chapter 2.2 destroyed.
This chapter rebuilds them. It has exactly one organising idea, and it is worth stating before any algebra so you can watch it do all the work:
Momentum conservation, as you were taught it, is a statement about three numbers. Three numbers are not a frame-independent object. So the law as stated cannot be fundamental.
Repairing that is not optional, and the repair is not free. There is essentially one four-vector you can build out of a particle's mass and motion. Once you insist that it is the conserved thing, a fourth conserved quantity arrives along with the three you wanted. You did not ask for it, no experiment forced it on you, and you have no way to refuse it.
That fourth quantity has a low-speed limit of plus a constant, and the constant is . Everything famous in this chapter follows from that one structural demand. By §3 you will have watched arrive as an unwanted guest rather than as a revelation.
Tools you'll need — Chapter 0.3 §2: the binomial series for arbitrary real , and its worked example 2, which already expanded without saying what it was. Chapter 0.6 §5: the multivariable chain rule, used constantly below to convert into . Chapter 1.2 §3.5 and §8.1: the Euler–Lagrange equation, and the table of actions this book will write down. Section 5 here supplies the second entry in that table. Chapter 1.4 §3: Noether's theorem, in particular that invariance under time translation gives a conserved energy. We use it in §5 as an independent check. Chapter 2.3 §5 (proper time as the length of a worldline), §6 (straight worldlines maximise it), §4 (the causal classification of four-vectors), and §7 (four-velocity, sketched there and built properly here). Chapter 2.4 §4 (raising and lowering), §5 (the operations preserving tensor character, and the quotient theorem), and above all §6, the invariance theorem, which is the load-bearing wall of this entire chapter.
1 · Velocity is not a four-vector, and the repair
Start with the object we would most like to keep. A particle traces a worldline through spacetime, and the obvious thing to call its velocity is the rate at which its coordinates change with time:
That is four numbers, built from a four-vector, containing exactly the information you want. It is also not a four-vector. The reason is worth stating precisely rather than gesturing at, so let's do that first.
1.1 · Exactly what goes wrong
Being a four-vector means one thing and one thing only, by Chapter 2.4 §2: the four components must transform as . So let's test (2.5.1) against that requirement.
The numerator is fine. is the prototype four-vector, so . The denominator is the problem. is not a scalar. It is the zeroth component of a four-vector, and under a boost along it becomes
where the second form divides out and uses . Now put the transformed numerator over that transformed denominator and see what we are left with:
Let's look at what that line is actually saying. transforms almost correctly. It picks up the right matrix , and then it is spoiled by an extra scalar factor out front.
Worse, that factor depends on the particle's own velocity . Two different particles observed in the same boost are therefore multiplied by two different numbers, and no rescaling of the definition can repair that.
This single line is also the whole reason Chapter 2.2's velocity-addition law came out as an awkward quotient rather than a matrix acting on a column. The awkward denominator in is the factor in (2.5.3).
Now state the diagnosis in a form that tells you the cure. We divided a four-vector by a frame-dependent quantity. So divide by a frame-independent one instead. Chapter 2.3 §5 supplied exactly one natural candidate along a worldline: the proper time , which every observer computes and every observer agrees on.
1.2 · The bridge:
Before we can use we need to know how it relates to coordinate time. This is a two-line recomputation of a result from Chapter 2.3 §5. It is worth having in front of you, because we will use it perhaps thirty times.
Proper time along a worldline is defined by . So our first job is to evaluate between two neighbouring events on the worldline, and then to pull a factor of out to the front:
The bracket is , so . Take the positive root, since both and increase toward the future, and we have the relation we came for:
Two remarks here will save confusion later.
First, in this formula is the particle's own , evaluated with its instantaneous speed in whatever frame you are using. It is a function of , and it has nothing to do with the of any boost between frames.
Second, (2.5.5) is the reason is computable at all. Whenever a -derivative turns up, you convert it into the -derivative you already know how to take:
which is nothing but the chain rule of Chapter 0.6 §5 applied to a one-parameter reparametrisation.
1.3 · The four-velocity
Now define what we were forced to define:
This is a four-vector, and the proof is one sentence. The numerator transforms with and the denominator is a scalar, so with no leftover factor. That is the entire content of the definition, and it is why relativity keeps replacing by everywhere.
Next we want the components of in terms of quantities a laboratory can measure, so apply the conversion rule (2.5.6) to the definition:
Note what happened, because it is the pattern of the whole chapter. The repaired object is the old object times .
At low speed and , so nothing familiar is lost. At high speed the components run off to infinity. That is exactly the behaviour that lets a fixed-length object point in ever more extreme directions without ever exceeding the speed limit.
Which brings us to the property that makes special.
1.4 · , by explicit contraction
We want the Minkowski length of , so contract it with itself using the metric. Chapter 2.4 §4.3 tells us how the bookkeeping goes. The time component keeps its sign and the spatial components flip:
The answer is , identically. Not for special particles, and not in special frames. It holds for every particle, at every instant, in every frame, whatever that particle happens to be doing. The four-velocity is a four-vector of fixed Minkowski length , and all it can do is point in different future-timelike directions.
There is a second proof, shorter, which explains why rather than merely confirming. It is the one to remember. Write the two -derivatives as a single ratio of differentials and the answer falls out of the definition of :
So the constraint is not a fact about particles at all. It is the definition of , rearranged.
We chose to parametrise the worldline by its own arc length, and a curve parametrised by arc length has a unit tangent vector. Here "unit" means length , because we measured arc length in seconds and tangent vectors in metres per second. Everything in this chapter that looks like a miracle is this sentence in disguise.
1.5 · Four-acceleration, and an orthogonality you get for free
Differentiate again with respect to the invariant:
Same argument as before, run twice over: is a four-vector.
Now do something that costs one line and pays for the rest of the chapter. We want to know what the fixed length of forces on , so differentiate the identity (2.5.9) with respect to . Its right-hand side is the constant , so whatever the left-hand side works out to has to vanish:
where the two terms merged because is a constant array and is symmetric, so the second sum is the first with the dummy names swapped. Divide by the 2 and what remains is
and it holds always, for every worldline. In Minkowski language, four-acceleration is orthogonal to four-velocity. The word "orthogonal" is doing honest work here. It is the same bilinear form, and the same definition of perpendicularity, that Chapter 0.5 gave for Euclidean space, with one sign changed.
In Euclidean geometry, if a point moves on a sphere of fixed radius, its velocity is tangent to the sphere, which is to say perpendicular to the radius vector. That is the same statement as (2.5.13), one dimension up and one sign over. The four-velocity lives on the hyperboloid , and acceleration can only move it along that surface, never off it.
So there is no such thing as accelerating "faster through spacetime". A force can turn your four-velocity, tilting it further from your original time axis and thereby increasing without limit. What it cannot do is lengthen it.
The speed limit is therefore not a barrier the particle runs into. It is the statement that the only motion on offer is rotation of a fixed-length vector, and a rotation never gets you off the surface you started on. Chapter 2.3 §3.2 drew this surface as the invariant hyperbolae, and here it is again, doing dynamics.
It also does immediate accounting work. says the four-acceleration has three independent components rather than four, since one of them is always fixed by the other three. In §6 that single constraint turns into the work–energy theorem, which you will therefore not have to assume.
Grind box — components of , and checked the hard way
We will need the components in §6, and computing them is a good drill in (2.5.6). First we need the derivative of . With and , the chain rule gives
using , with the ordinary three-acceleration. Write for this. Now apply (2.5.6) to (2.5.8):
Now contract, carefully, keeping every term:
Substituting from the first line makes the two terms cancel identically. ✓ The abstract one-liner (2.5.12) and the brute-force component calculation agree, as they must. The point of doing both once is to see how much labour the index notation is saving.
One consequence worth extracting now. Go to the particle's instantaneous rest frame, where and . Then forces , so in that frame , where is the acceleration the particle itself feels. Call it the proper acceleration. Its invariant square is
so four-acceleration is always a spacelike vector, and its invariant magnitude is the proper acceleration, meaning what an accelerometer bolted to the particle reads. That number is frame-independent. This is why "the rocket burns at " is a meaningful statement, while "the rocket's acceleration is " is a statement about somebody's coordinates. Problem 2 uses this.
The object sketched at the close of the last chapter now gets built, and the diagnosis is sharper than the sketch allowed. Divide a displacement between two events by the time one observer's clock assigns to it: the numerator behaves impeccably, being the prototype everything else is measured against, while the denominator belongs to whoever holds the clock. What results is spoiled by a stray factor depending on the particle's own speed, so two particles watched through one change of frame are spoiled by different numbers and no rescaling rescues either.
Divide instead by the reading of the clock the particle carries, which everybody computes and everybody agrees about, and the defect has nowhere to come from. What comes out has the same length for every particle at every instant, and that constraint is not a discovery about matter but the definition of the carried clock rearranged: a curve measured by its own length has a tangent of fixed size.
Differentiate once more and something arrives for nothing. Since the length never changes, its rate of change stands perpendicular to it in the interval's geometry, so a force can turn a fixed-length arrow and can never stretch it. The speed limit stops being a barrier a particle runs into and becomes the observation that turning is the only motion on offer, and turning keeps you on the surface you began on.
2 · Four-momentum, and the argument this chapter turns on
We have two ingredients on the table. One is a four-vector that describes a particle's motion. The other is a scalar that describes the particle itself, namely its rest mass , which by the conventions of this Part always means the rest mass and never anything else.
Multiplying a four-vector by a scalar gives a four-vector (2.4 §5.1, operation 2), so we are entitled to define the four-momentum
Its spatial part is , which reduces to the Newtonian when . So we have at least a candidate for "momentum", and we will call it .
The time component is, at this stage, an unnamed number that came along for the ride. We are not going to name it until §3, and we are not going to name it by guessing. The name will be forced on us.
Everything now rests on a single question: why should this be the conserved quantity? The answer is the intellectual centre of the chapter, so we take it slowly.
2.1 · What a conservation law has to be
A conservation law is a claim about a physical process. Some quantity computed before the process equals the same quantity computed after it. Consider a collision, with particles coming in and particles going out, and the assertion
Now ask the question that Chapter 2.4 taught us to ask about every equation. Who is asserting this, and does the assertion survive translation into another observer's language?
The question is not pedantry. If the assertion does not survive, then (2.5.15) is not a law of nature. It is a report from one laboratory, of no more fundamental standing than "the train is moving".
Chapter 2.4 §6 gave the tool for deciding. If the two sides of an equation are tensors of the same type, then the equation holds in every frame as soon as it holds in one. If they are not, all bets are off.
Chapter 2.4's closing warning was blunt about which side of that line statements like fall on. They are statements about components, and therefore about a frame, and they must never be carried across a boost unexamined.
So the demand is clear. Whatever is conserved must be a tensor, and the conservation statement must be an equation between tensors of the same type. Three numbers are not a tensor under the Lorentz group. Four numbers that transform with are.
2.2 · Newtonian momentum conservation does not survive a boost
That is an argument from principle. Now let's see the failure in arithmetic, in a collision you can hold in your head.
Two identical lumps of putty, each of rest mass , approach each other along the -axis in frame with speeds and . They collide and stick. By symmetry the composite is at rest in , and we will call its rest mass , whatever that turns out to be.
Start by checking Newtonian momentum conservation in , where the arithmetic is easiest:
Perfect. But note how little that check actually tested. By symmetry it would have passed for any definition of momentum of the form , with any function whatever of the speed, because the two terms cancel no matter what is.
A symmetric collision in a single frame is almost completely uninformative. The information is in the boost, so let's boost.
Move to a frame travelling at along relative to , so that in everything is drifting in the direction. Chapter 2.2's velocity-addition law gives the new velocities of the two lumps:
and the composite, at rest in , moves at in . Our next goal is the incoming Newtonian momentum in , so add those two velocities. Write for compactness and put the two fractions over the common denominator :
where the last step used , so that .
Meanwhile the Newtonian momentum after the collision is , and Newtonian mechanics also insists that mass is additive, . So here are the two sides of the conservation law as reads them:
These are equal only if or , which is to say only if there was no collision or no boost. Put in numbers. For the fraction is , so an observer in finds the outgoing momentum 36% larger than the incoming momentum. The books balance in and fail badly in .
It is tempting to say "Newtonian momentum isn't conserved at high speed". That is not what the calculation showed. In frame it was conserved exactly, to all orders, as (2.5.16) shows. What failed is agreement between observers about whether it was conserved. One frame says the law holds, and another says it is violated by 36%.
That is a much worse disease than being approximately wrong, and no small correction term can cure it. A quantity whose conservation is frame-dependent is not describing anything about the collision. It is describing the laboratory.
This is precisely the defect Chapter 2.4 §6.2 diagnosed in . That chapter said its content changes when you change frames, which disqualifies it as a statement about nature rather than about a laboratory. Here the same defect has been shown rather than asserted.
2.3 · The four-vector version cannot fail
Now run the same collision with and watch the disease vanish for structural reasons.
Suppose that in frame the total four-momentum is conserved, meaning for all four components. To test that claim in another frame we want a single object to transform, so define the shortfall
Each is a four-vector, and sums and differences of tensors of the same type are tensors of that type (2.4 §5.1, operation 1). So is itself a four-vector.
Our hypothesis is now the tensor equation , and 2.4 §6 says a tensor equation true in one frame is true in all of them. Written out explicitly, that reads
That is the whole proof, and it is deliberately anticlimactic. The transformation law is linear and homogeneous, so it maps zero to zero, and there is nowhere for a violation to come from. If four-momentum is conserved for one observer, it is conserved for every observer, automatically, with no further conditions.
Grind box — the same collision, done relativistically, term by term
The abstract argument is airtight, but you should see the numbers come out once. We need one identity, and it is worth having permanently.
The velocity-addition identity for . If , then
Expand the numerator, being careful with every term:
The terms cancel and what remains factorises. Taking reciprocals and square roots,
This is just the statement that rapidities add, , written in velocities. Chapter 2.2 §5 proved the rapidity version.
Momentum before, in . Using the identity with and then :
The messy denominators cancel exactly, which is the identity earning its keep. Adding the two and multiplying by :
Momentum after, in . The composite has rest mass and moves at , so its momentum is . Equality demands
Read that twice. Relativistic momentum conservation in the boosted frame does not merely permit the composite's rest mass to exceed . It requires it, and fixes the excess exactly. The lump of putty is heavier than its parts, by a factor , and nothing in the calculation gave us a choice. Section 9 is about that number.
The time components, for free. Do the same with :
while after the collision , using the just derived. Equal ✓, and note that we did not impose this. It came out. That is (2.5.21) being true in coordinates.
And in itself. Before the collision , and after it ✓, while the spatial parts vanish on both sides by symmetry. So the four-vector statement holds in , holds in , and by (2.5.21) holds in every frame there is.
2.4 · The part you cannot refuse: momentum forces energy
So far we have shown that four-momentum conservation is self-consistent across frames. Now comes the sharper statement, and it is the one that makes this chapter inevitable.
Suppose you are a minimalist. You accept that the conserved object must be built from the four-vector , since you have seen §2.2 and you are not going back. But you want to conserve only the three spatial components, because those are the ones you have experimental evidence for. You want the law " in every frame", and you decline to say anything at all about .
You cannot have it. Here is why, in three lines.
Let be a four-vector. If its three spatial components vanish in every inertial frame, then its time component vanishes too.
Proof. By hypothesis in the original frame. Boost along with speed . By Chapter 2.4's transformation law (2.5.21) the new first spatial component is
The hypothesis says this must vanish for every . Since for , we need .
Look at what that argument did. It used no dynamics, no experiment, and no assumption about forces or collisions. It used only the transformation law.
And what it says is this. A universe in which momentum is conserved for every observer is a universe in which is conserved too. You are not permitted to conserve three components of a four-vector. The boost mixes the fourth into the first three and drags it into the law whether you invited it or not.
That is the whole chapter in one paragraph. The demand that conservation laws be frame-independent selects out of all the candidate momenta. Having selected it, you are stuck with its time component, a fourth conserved quantity you did not ask for, did not derive from any experiment, and cannot discard. Section 3 asks what on earth it is.
This is not a proof that anything is conserved. Nothing so far says nature conserves . That is a physical claim, and Chapter 1.4 gave its real origin: Noether's theorem, applied to the invariance of the laws under translations in space, which gives momentum, and in time, which gives energy.
What we have proved here is conditional, and it comes in two parts. If some four-vector quantity is conserved in one frame, it is conserved in all of them. And if a frame-independent momentum conservation law exists at all, then the conserved object must be a four-vector and its time component comes along with it.
Relativity narrows the field of candidates to essentially one. Experiment then confirms that this one is right, and it does so every day in every particle detector on Earth.
The choice of as the scalar was not forced by tensor algebra alone. is a four-vector, but so is for any function , and so is for any dimensionless combination you can invent. What fixes the choice is the correspondence limit. As the spatial part must become the Newtonian that three centuries of experiment supports, and that pins the scalar to itself. Covariance narrows the field and correspondence chooses within it. That pair of moves is the method of the entire second half of this book.
Every conservation law is somebody's claim about a process, and what matters is whether the claim survives translation into another observer's language, since one that does not is a report about a laboratory rather than a fact about the world. Three numbers cannot survive, and the failure is not gentle. Two identical lumps of putty approaching at equal speeds and sticking conserve the old momentum exactly in the symmetric frame, while an observer drifting past finds the outgoing total larger than the incoming one by more than a third.
Notice precisely what that is. It is not that the old momentum is slightly wrong at high speed, since in that frame it was right to every order. What fails is agreement about whether the law holds at all, and a disease of that kind admits no correction term. Conserve the four-part object instead and the disease has nowhere to live, since the transformation between observers is linear and carries zero to zero.
Then comes the part that cannot be declined. Suppose you accept the object but wish to conserve only its three familiar entries. A change of frame mixes the fourth into the first three, so demanding that the three vanish for every observer forces the fourth to vanish as well. Frame-independence has selected the object out of all the candidates, and having taken three of its components you are stuck with the fourth.
3 · What the time component is
We have an unnamed conserved quantity . Its dimensions are those of energy, since it is a mass times a velocity squared. Dimensions are suggestive and prove nothing, so let's find out what it actually is by the only honest route available. We expand it at low speed and see what it turns into.
3.1 · The expansion
Apply the binomial series of Chapter 0.3 §2, which for arbitrary real exponent reads with . Take and , where , and the first few coefficients come out as
Substituting flips the sign of every odd power, so all the signs become positive:
That is the expansion of alone. The quantity we actually want is , so multiply through by and restore , keeping an eye on what each term becomes:
3.2 · The identification, made carefully
Now the part that deserves care, because it is usually skated over. What (2.5.24) shows is a mathematical fact about a function, not a physical identification. It says that the conserved quantity equals a constant, plus the Newtonian kinetic energy, plus terms that vanish as .
Getting from there to " is the energy" takes one further argument, and here it is.
Newtonian mechanics has its own conserved quantity for elastic collisions, , verified to exhaustion at low speeds. Our new law says instead that is conserved. We want to see how the two claims are related, so expand the new one particle by particle at low speed:
Now suppose the collision does not change what the particles are, so the same species go in as come out. That is the only kind of collision Newtonian mechanics ever contemplated.
Then the sum is the same before and after, and it cancels out of the conservation statement entirely. What is left is exactly Newtonian kinetic-energy conservation.
So the new law contains the old one, and the quantity it conserves must be what the old theory called energy. That, and only that, is what licenses the name:
The four-momentum is the energy and the momentum, welded into one object by the boost that mixes them. From here on is called the energy–momentum four-vector, and the two conservation laws you learned separately are one law with four components.
In Newtonian mechanics the zero of energy is arbitrary. Adding a constant to changes no equation of motion, no conservation law, and no prediction. That freedom is why nobody ever asked what the "true" energy of a stationary object was. The question had no content.
Relativity removes the freedom, and the mechanism is visible in (2.5.25). The would-be constant is , and it cancels only if the rest masses do not change.
Processes exist that change them. A nucleus binds. A particle decays. Two lumps of putty stick together and end up with rest mass rather than , as §2.3's grind box proved. In any such process the sum is different before and after, and the difference has to be paid for out of kinetic energy.
So is not an additive constant one may drop. It is a reservoir, and processes exist that draw on it. Setting in (2.5.26) names it:
That is the rest energy. Note carefully what it is not. It is not a formula for the energy of a moving body, which is . It is not a claim that mass "is" energy in some substantial sense. It is the value of the conserved quantity for a particle that happens to be at rest, and its only physical content is the one just derived: changes in rest mass show up as changes in kinetic energy, at an exchange rate of .
3.3 · The first correction, and where you can see it
The term in (2.5.24) is not decoration. To judge its size we need something to judge it against, so compare it with the Newtonian term:
The correction is quadratically small in . That is why Newtonian mechanics survived so long, and it is also why the correction becomes visible the moment you have a system with a characteristic speed and a spectroscope.
The electron in a hydrogen atom has . Chapter 0.3 §5 got that from dimensional analysis alone, since the natural speed divided by is exactly . Putting that number into the ratio above gives
a relative shift of four parts in in the atom's energy levels. On a scale that is of order , which is comfortably resolvable.
It is one of the three contributions to the fine structure of hydrogen that Chapter 0.3's worked example 2 flagged and could not yet explain. Chapter 4.16 computes the full splitting. The point here is only that the leading piece of it is a term in (2.5.24) and nothing more exotic.
Grind box — the same correction in momentum, and why the coefficient changes sign
Open a quantum mechanics text and you will find the relativistic correction to the hydrogen Hamiltonian written as , a negative term with coefficient . We just derived a positive term with coefficient . Both are correct. Chapter 0.3's worked example 2 noticed the clash and named the culprit. Now that we have actually derived and the mass shell, we can close it arithmetically.
The kinetic energy in terms of . Anticipating §4's mass shell, . Factor out and expand with the binomial series again, now with and , for which and :
There it is, sign and coefficient. Quantum mechanics uses this form because is the natural operator, not .
Why the two agree. The variable is different, because rather than . So expand to the order needed,
so the two terms become
Add them: ✓, which is (2.5.24). The apparent discrepancy was bookkeeping. A of the quartic term comes from not being , and a comes from the square root, and .
The moral is one you will need again in Part V: a perturbative coefficient is meaningless until you say which variable is being held fixed. Terms move between orders when you change variables. Only the physical prediction is invariant, which here means the size of the fine-structure shift, and that is either way.
A fourth conserved quantity stands about unnamed, and the only honest route to its identity is to expand it for a slow particle and watch what it becomes. Out comes a constant belonging to the particle, then the Newtonian kinetic energy, then corrections falling with the square of the speed ratio. Nothing there is a concession: nearly everything is an approximation, and the controlled kind names its next term. What licenses the name is that when a collision leaves the participants the species it found them, the constant cancels and kinetic energy is conserved as before.
The constant is where the interest lies. Newtonian mechanics let you shift the zero of energy by any amount, which is why nobody asked a stationary object's energy; the question had no content. That freedom is withdrawn, since the constant cancels only while the rest masses hold, and processes change them: a nucleus binds, a particle decays, two lumps of putty stick and come out heavier than their parts.
Rest energy is a reservoir rather than an offset, at a fixed exchange rate against motion. Notice how little was requested and how much arrived. Nobody went looking for it, no experiment produced it, no argument about matter suggested it. It was conscripted: the demand that a law read alike for everybody left a fourth component lying about, and what it turned out to be is the equation everyone recites.
4 · The mass shell
Every four-vector has an invariant square, and the invariant square of turns out to contain everything.
4.1 · , and what it says
Contract with itself. The mass is a scalar, so it comes straight out of the contraction, and (2.5.9) finishes the job:
That is the answer in terms of the mass. We also want it in terms of the things a detector actually reports, so compute the same quantity from the components in (2.5.26), remembering that lowering flips the sign of the spatial parts:
Two expressions for one quantity. Set them equal and multiply through by :
This is the mass-shell relation, and it is the most useful equation in particle physics. Three observations before we use it.
It is an invariant statement. The left side involves and the right involves . Both change under a boost, and the particular combination does not. Different observers assign different energies and momenta to the same electron, and all of them compute the same .
So mass is a Lorentz invariant. It is not a quantity that grows with speed. It is the fixed label attached to the four-vector, exactly as the interval is the fixed label attached to a pair of events.
It is Chapter 2.3's geometry, in different variables. Plot against and (2.5.31) is a hyperbola with asymptote . Compare 2.3 §3.2, where the orbits of a boost in the plane were the hyperbolae with asymptotes at .
Same curves, and for the same reason. A boost is a hyperbolic rotation, and it slides any four-vector along the level surface of its own invariant square. Momentum space is spacetime with the axes relabelled.
It has no reference to . That matters more than it looks, and §4.3 collects the debt.
4.2 · Velocity from the four-momentum
Before we take the massless case we need one more result, and the order matters, because this result is what makes the massless case work at all. We want the particle's velocity written in terms of and . Both components of carry the same factor , so dividing one by the other kills it:
Read it as a recipe. Given a particle's energy and momentum, its velocity is the ratio. No , no , no square roots.
Two checks. At low speed, and , so the ratio returns ✓. And in general by (2.5.31) whenever , so the formula automatically gives . The speed limit is not imposed on (2.5.32). It falls out of it.
4.3 · Massless particles
Set in (2.5.31). Taking the positive root, since energies of real particles are positive:
That fixes the relation between energy and momentum. To get the speed, feed it into the recipe (2.5.32) and watch the momentum cancel:
Exactly , for every energy, with no limiting process and no approximation. A massless particle does not travel near the speed of light. It travels at it, always, and it has no rest frame in which to be at rest.
It is natural to read as the end of a sequence: particles get lighter and lighter, and in the limit you have a photon. That reading is wrong, and it is wrong in a way that will cost you.
The construction does not survive the limit. As at fixed energy, , so becomes . That is an indeterminate form, not a limit.
Worse, needs , and a null worldline has at every point, so identically and there is no proper time to divide by. Photons do not have worldlines parametrised by proper time. A photon has no clock, and "the time experienced by a photon" is not a small number. It is a meaningless phrase.
The correct definition takes as primary. A massless particle is one whose four-momentum is a null vector, . That is a perfectly good statement with no , no , and no multiplying anything.
Everything else follows from it. (2.5.31) gives , and (2.5.32) gives , and Chapter 2.3 §4's causal classification already told you where such a vector points, namely along the light cone. Nothing new had to be assumed. The massless case is the null case of a classification we already had.
This inversion is worth naming, because it is the version that survives into quantum field theory. The four-momentum is fundamental, and mass is the label carries. In field theory particles are excitations of fields and appears as a parameter in a Lagrangian long before anything is moving.
Grind box — why the Newtonian curve is so good for so long
The figure claims the Newtonian energy is within 1% of the truth up to , which is surprising if you expect errors of order . Here is the reason, and it is a nice piece of Chapter 0.3 bookkeeping.
Work in units of and . Write , so the exact energy is and the Newtonian one is . The fractional error is
Expand with the binomial series, :
The terms cancel exactly. That cancellation is not luck. The Newtonian formula was constructed to reproduce the first two terms of (2.5.24), so the leading error is necessarily the third. Hence , which is quartic rather than quadratic, and a quartic is very small for a while and then is not.
Numbers. Setting exactly, numerically rather than from the expansion, gives , whence and . Some more, exact:
| Newtonian | error | |||
|---|---|---|---|---|
This is the quantitative answer to "why didn't anyone notice?". Everything in nineteenth-century mechanics ran at , where . You do not find relativity by doing careful mechanics. You find it by doing electromagnetism, which is Chapter 2.1's story, or by building something that goes fast.
Look back over §§1–4 and notice what is not in them. There is no axiom saying "speeds cannot exceed ". The prohibition is a consequence, and it arrives twice, by different routes.
Causally (Chapter 2.3 §4.5). A signal travelling faster than can be boosted into a frame where it arrives before it left, and combining two such signals builds a closed causal loop. That argument came from the geometry alone and applies to influences, not just to objects.
Energetically (this chapter). From (2.5.26), pushing a massive particle to requires and hence . The barrier is not a wall. It is a bill. And (2.5.32) shows the same thing from the other side: no finite and satisfying (2.5.31) with can give .
The two arguments are independent and they agree, which is the sort of over-determination that makes a structure believable.
Note also what neither of them forbids. A hypothetical particle with , meaning a spacelike four-momentum and an "imaginary mass", would satisfy happily. Nothing in the algebra excludes it. What excludes it is causality, together with the fact that in quantum field theory such a field signals an unstable vacuum rather than a fast particle. Chapter 6.6 makes that precise, because the Higgs field before symmetry breaking is exactly such a case, and the resolution is not that anything travels faster than light.
Contract the welded energy and momentum with itself and the mass falls out as a number nobody disputes. Observers assign the same electron different energies and different momenta and all compute the same mass, which makes mass a label the object carries rather than something growing with speed, exactly as the interval is the label a pair of events carries. Plot energy against momentum and the hyperbolae drawn two chapters ago for time and position reappear, since a change of frame slides any such object along its own invariant square.
One consequence deserves extracting first, because it makes the massless case work: the velocity is the ratio of momentum to energy, with no dilation factor and no mass in it anywhere. That recipe does not care whether a mass exists. Set the mass to zero and it returns the speed limit exactly, at every energy, with no approximation.
Reading the massless case as the end of a sequence of ever lighter particles is the expensive mistake to avoid. The construction that built momentum from a mass and a carried clock does not survive that limit, since a light ray has no carried clock to divide by, and the phrase naming the time a light ray experiences names nothing. The repair is to invert the order: take the four-part momentum as primary and let mass be the label it carries.
5 · The action is proper time
We have the repaired momentum , and we expect the equation of motion to read . That is Newton's second law in the form he actually wrote it, with the repaired momentum substituted for the old one. (Section 6 shows this is the spatial part of a proper tensor equation. For now, take it as the target.)
Chapter 1.2 established that respectable dynamics comes from an action principle. So the question in front of us is this: what Lagrangian produces this?
We are not going to guess. We impose the answer and solve for . That is the honest way round, and it is exactly how Chapter 1.2 §4 checked that reproduces Newton, run backwards.
5.1 · Solving for
Take a particle in a potential, , with the free part depending only on the velocity. We need the Euler–Lagrange equation for the coordinate , which gives one equation per degree of freedom, , from Chapter 1.2 §3.5. Here it reads
We want the left-hand side to be with . Matching the two demands
Now solve this for . The right-hand side is times a function of the speed alone, so can depend on only through . The chain rule then gives . Comparing with (2.5.36), the direction cancels and we are left with a single ordinary differential equation:
All that remains is to integrate it. The substitution that clears the square root is , for which and therefore :
The constant is genuinely free. Chapter 1.2's Problem 4 showed that adding a constant to adds a term linear in to , which is a total derivative and changes no equation of motion. Set it to zero, for a reason that will be apparent in ten lines:
Check the limit before going on. Expanding the square root with the binomial series, , so
which is Chapter 1.2's plus the irrelevant constant plus small corrections ✓.
Note in passing that is emphatically not any more. The quantity is not the kinetic energy of anything. Chapter 1.2 §5 warned that "" is a special case rather than a definition, and here is the first place that warning cashes out.
5.2 · The rewrite that makes it obvious
Now form the action and use (2.5.5) in the direction we have not yet used it. Until now we have converted -derivatives into -derivatives. This time we go the other way and turn into . For a free particle (),
Stop here. This is the most important equation in the chapter after (2.5.31), and it deserves to be read as a sentence rather than a formula.
Up to a constant , the action of a free particle is the reading of the clock it carries. Not a quantity resembling it, and not a quantity proportional to it in some limit. It is that number.
So the two principles this book has stated separately are one principle. Chapter 1.2 said that nature makes stationary. Chapter 2.3 §6 proved that among all worldlines joining two events, the straight one has the greatest proper time.
Put those side by side and the minus sign does its job. Maximising is minimising , because . The minus sign in front of is not a convention chosen for tidiness. It is what converts 2.3's maximum into 1.2's stationary point.
Chapter 1.2's grind box on the second variation already checked the Legendre condition for exactly this Lagrangian and found , confirming that the stationary point is a genuine minimum of .
Two loops close at once. First, the free particle's trajectory is the longest worldline, which is why the travelling twin comes back younger. She took a shorter path in the only sense of "length" spacetime recognises. Second, this is entry two in Chapter 1.2 §8.1's table of actions, quoted forward there and derived here. Whatever it is worth, the derivation cost eleven lines.
And the forward view. Chapter 3.3 keeps (2.5.41) verbatim and changes only how is computed, replacing by a position-dependent metric . Extremising the same functional then gives the geodesic equation, and general relativity's statement that free particles fall along geodesics is this statement, with a different ruler. Nothing about the principle changes at all.
Grind box — three independent checks on , including Noether's
Check 1: the canonical momentum. Chapter 1.2 §6 defines the momentum conjugate to as . From (2.5.39), with ,
Which it had better, since (2.5.36) is what we solved. But it is reassuring that the canonical momentum of Lagrangian mechanics and the spatial part of the four-momentum are literally the same object. They did not have to be.
Check 2: Noether's energy. Chapter 1.4 §3 showed that if has no explicit time dependence, the conserved Noether charge is . Compute it for the free case:
where the third line used . So the Noether charge of time-translation invariance is exactly , the quantity §3 identified by an entirely different argument, namely matching a Taylor expansion to Newtonian kinetic energy. Two independent derivations, one answer. ✓
Check 3: the equation of motion, in full. Confirm that (2.5.39) really does give the promised law and not something that merely resembles it. With , Euler–Lagrange says , so is constant, so is constant, since is a monotonic function of and fixing therefore fixes . The path is a straight line traversed uniformly. Free particles move in straight lines at constant speed, which is Newton's first law, recovered from an action whose Lagrangian is nothing like . ✓
An aside on manifest covariance. (2.5.41) is manifestly a scalar, since is a scalar and is a scalar and is a scalar. So the action is frame-independent by inspection, in the sense of Chapter 2.4 §6. The Lagrangian (2.5.39) is not, and neither is . Only their product is. This is the standard situation in relativistic mechanics, and it is why the action formulation is preferred. The object with the good transformation properties is the integral, not the integrand-in-.
Guessing which number to attach to each history would be a poor way to proceed, so the calculation runs backwards: demand that the machinery of the first part return the repaired equation of motion, and solve for it. What comes back for a free particle, once a harmless constant is set aside, is startling in its plainness. The number attached to a history is the reading of the clock carried along it, multiplied by the mass and the square of the speed limit, with a minus sign in front.
Two statements this book has made separately are therefore one statement. The first part said nature selects the history at which the number stops changing; the chapter before last proved that among all histories joining two events the unaccelerated one carries the most time. The minus sign reconciles them, since making a quantity largest is making its negative smallest.
Something else is settled on the way past. The number being attached is no longer kinetic energy less potential energy, resembling that combination in nothing but its low-speed limit, so the warning issued when the action principle first arrived, that the familiar difference is a special case and not a definition, collects here. Notice too which object is well behaved: neither the integrand nor the element of time multiplying it is the same for every observer, and only their product is.
6 · Force, and the death of relativistic mass
Newton's second law is the one piece of the old mechanics we have not yet replaced. Chapter 2.4 §6.2 already announced the verdict, which is that is not a tensor equation while is. Announcing is not deriving. Here is the object, and here is what it costs.
6.1 · The covariant force
Define the four-force by the only construction available: differentiate the four-momentum by the invariant.
where the last step holds for particles of constant rest mass. Both sides are tensors, so by Chapter 2.4 §6 this is a legitimate law. Written once, it holds in every frame.
It also inherits a constraint at no cost. We know from (2.5.13), so contract the definition with and let the mass ride along:
So a four-force has only three independent components. Whatever three you specify, the fourth is determined. That is not a defect. It is a theorem with a name you already know, as we are about to see.
6.2 · The three-force, and the work–energy theorem for free
Connect to the laboratory. Define the ordinary three-force as the rate of change of the ordinary momentum with respect to ordinary time, . That is Newton's second law in its correct form, the form he actually wrote. Using (2.5.6) on (2.5.42), component by component:
Now we cash in the constraint. We have in laboratory quantities, and (2.5.43) says its contraction with vanishes. So write out that contraction, using , and see what it forces:
and since , the bracket itself must vanish, which leaves
That is the work–energy theorem. The rate of change of energy is the power delivered by the force. We did not assume it, define it, or import it. It is the time component of a four-vector equation, forced by the geometric identity .
In Newtonian mechanics the work–energy theorem is a separate derivation. Here it is the statement that the four-force is orthogonal to the four-velocity, which is the statement that the four-velocity has fixed length, which is the definition of proper time. Everything is the same thing.
6.3 · and are not parallel
Now the damage. Expand with the product rule, using from §1.5's grind box:
The bracket contains a piece along and a piece along . Unless those two directions coincide, or unless the second term vanishes, the force and the acceleration point in different directions.
Push a fast particle sideways and it does not accelerate sideways. It accelerates partly forward, or partly backward, depending on the sign of . Nothing in Newtonian mechanics prepares you for this, and it is not a small effect.
The two special cases where they are parallel are worth having explicitly.
Longitudinal, meaning a push along the motion. Then , so and the bracket is . Use the identity , that is , and the whole factor collapses:
Transverse, meaning a push across the motion. Then , the second term dies, and what is left is
Two different laws, differing by , for the same particle at the same instant. At a longitudinal push is a hundred times less effective at producing acceleration than a transverse one of the same magnitude.
This is not an exotic regime. It is routine at any electron accelerator, and it is why the transverse focusing of a beam and its longitudinal acceleration are engineered as separate problems with separate hardware.
Grind box — inverting the relation: in terms of
(2.5.47) gives from . For solving actual problems you want the other direction, and inverting it is a small exercise in projection that is worth doing once, because the same trick appears in Chapter 3.3.
Step 1, get the component along . Dot (2.5.47) with :
using again. So
Step 2, substitute back. Put that into (2.5.47) and solve for :
hence
Check it. For the bracket is , so ✓ matching (2.5.48). For the second term vanishes, so ✓ matching (2.5.49).
What the formula says. The bracket removes from the piece , a projection along the motion weighted by . So the acceleration is a distorted image of the force, squashed along the direction of travel.
At the longitudinal response is suppressed entirely. You can push a nearly-light-speed particle as hard as you like along its motion and it barely speeds up, even though its energy rises steadily at the rate (2.5.46) says it must. The energy has to go somewhere, and it goes into rather than into .
One consequence worth naming. Set constant, along , starting from rest. Then integrates immediately to , so
approaching but never reaching , whereas the Newtonian sails past it at . Same constant force, same duration, entirely different destination. And the relativistic answer required no new physical assumption beyond .
6.4 · Why this book will not say "relativistic mass"
There is a tempting way to make (2.5.48) and (2.5.49) look like , which is to absorb the 's into the mass. Define and you can write and , and the formulas look Newtonian again. Textbooks did this for sixty years. It is a mistake, and the calculation above is the cleanest way to see why.
It would have to be two different numbers at once. To keep you need for a longitudinal push and for a transverse one. A "mass" that depends on which way you shove is not a property of the particle. It is a badly chosen name for the components of a tensor relation.
And for a general push, by (2.5.47), no scalar whatsoever works, because and are not even parallel. You would need a matrix, and calling a matrix "the mass" abandons the only thing the word was supposed to mean.
It makes empty. With , the celebrated equation says . All the physical content lives in the rest mass and is invisible in the relativistic one: that a particle at rest still has energy , that binding changes rest mass, and that mass is not additive.
It hides the invariant. The real structure is (2.5.31). Here is what you get by contracting with itself, a Lorentz scalar, the same for all observers, in exactly the way the interval is the same for all observers.
Redefining "mass" to mean replaces an invariant with a component and throws away the geometry. It is the momentum-space version of saying that a metre stick "really" gets shorter when it moves.
So means rest mass, always, everywhere in this book, exactly as the conventions for Part II state. Energy is and momentum is , and the 's stay where they are, visible, attached to the motion rather than smuggled into the particle. The phrase "relativistic mass" does not appear again.
The second law costs more to replace than it appears to. Differentiate the four-part momentum by the carried clock and you have a force with the right credentials, carrying a constraint nobody imposed: since the four-velocity has fixed length, the force stands perpendicular to it, so only three of its entries are free. Write out the entry that is not free and it says energy changes at the rate the ordinary force does work, which in Newtonian mechanics was a separate derivation and here is the fixed length restated.
Then the damage. Force and acceleration stop pointing the same way. Push a fast particle along its motion and it responds far more sluggishly than to an equal push across it: at a dilation factor of ten, a hundred times less so. That is why transverse focusing of a beam and forward acceleration are engineered as separate problems with separate hardware.
A habit sixty years of textbooks kept dies with it. Absorbing the dilation factors into the mass makes the formulas look Newtonian again, at the price of the mass being two numbers at once, one for a push along the motion and one across. For any other direction no single number works, the two vectors not being parallel, so a matrix is needed, and calling a matrix the mass abandons what the word was for. The phrase will not appear again.
7 · Light: the wave four-vector, and Doppler
Everything so far has been about particles. Waves need one more four-vector, and it comes from an argument so simple it is easy to undervalue.
7.1 · Why has to be a four-vector
A plane wave has a phase
with the wave vector (, pointing along propagation) and the angular frequency. Here is the physical claim, and it is the whole argument:
A crest is not a thing that moves. It is a set of events at which the field is momentarily maximal. So "is this event a crest?" is a yes-or-no question about a single point of spacetime, and both observers are being asked about the same point. They must give the same answer.
Now count. Put a detector at a fixed place and let it click once per crest between two events and on its own worldline. The number of clicks is an integer.
An observer flying past at watches the same detector and counts the same integer. Each click is an event on one worldline, and 2.3 §4.3 proved that the ordering of events on a single timelike worldline is invariant. Nobody can gain or lose a click by changing frames.
That count is . An integer cannot transform, and and were arbitrary, so is a Lorentz scalar.
Now use the scalar to manufacture the four-vector. We want (2.5.50) written as a contraction, since a contraction is the shape the quotient theorem knows how to act on. With and the metric lowering the spatial indices,
So . We know is a four-vector, and we have just argued that is a scalar for every event , not merely for one.
That is precisely the hypothesis of the quotient theorem (Chapter 2.4 §5.1's grind box). If is a scalar for every four-vector , then is a covector, and raising with makes a four-vector. Hence
For light in vacuum the dispersion relation is , which Chapter 2.1 §2 derived from Maxwell. That says precisely , so the wave four-vector of light is null. Set that beside from §4.3 and the shape of Chapter 4.1's punchline is already visible. We come back to it below.
7.2 · Longitudinal Doppler
Everything about the Doppler effect is now one matrix multiplication. Let a light wave travel in the direction in frame , so
the spatial part having magnitude as the null condition requires. Now boost to moving at along , which is an observer running away from the source, in the same direction the light is going. Chapter 2.4 §2.1's boost matrix acts on any four-vector, so it acts on this one:
Both components pick up the same factor, as they must, since has to be null too. The frequency we want is , so read that off and simplify with :
For a receding observer, and , which is a redshift. For an approaching one, replace and get a blueshift. At this reduces to , the classical result, so nothing familiar is lost. What is new is the exactness, and the second-order term, which is where the physics is.
7.3 · Transverse Doppler: time dilation, seen directly
Now a case with no classical counterpart at all. Let a source move at along past a detector, and consider the light that reaches the detector travelling in the direction in the detector's frame. That is the light emitted, as the lab sees it, at the moment of closest approach, so that the source's velocity is entirely perpendicular to the line of sight. Classically there is no Doppler shift whatsoever in that configuration, because the distance is momentarily not changing.
In the lab frame , the received wave has
What we want to compare it with is the frequency the source emits, so transform to the source's rest frame , which moves at along . Only the time component is needed:
But is the frequency in the source's own frame, which is the frequency the source actually emits. Call it , a property of the atom. Then , which rearranges to
A redshift, of exactly the factor , in a geometry where the classical prediction is precisely nothing. And its content is unmistakable. The factor is the time-dilation factor of Chapter 2.2. The source is a clock, the clock is moving, so it runs slow, and you see its emission slowed.
The transverse Doppler shift is time dilation with no admixture of anything else, which is exactly why it was worth measuring.
Grind box — the general Doppler formula, aberration, and what Ives and Stilwell actually measured
Arbitrary angle. Let the light travel at angle to the -axis in , so . Boosting along by :
The first line is the general Doppler formula. Check the special cases. Setting gives , which is (2.5.55) ✓. Setting gives , which is (2.5.58) ✓, with .
Aberration. Divide the second line by the first to get the direction in :
which is the relativistic aberration of light. It is the same formula that turns the rain on your windscreen into a forward-slanting streak, but exact, and with the crucial difference that the aberration of light does not depend on any medium. This is Problem 3, where it becomes the headlight effect.
⚑ Ives and Stilwell, 1938: quoted experiment, derived analysis. The transverse configuration is hard to arrange, because "transverse" must be specified in a definite frame and a tiny angular error contaminates the result with a first-order longitudinal shift, which is times larger in the relevant sense. Ives and Stilwell got around this by measuring longitudinally in both directions at once. Let a source recede with and approach with , so that in wavelengths . Average them:
The first-order shifts cancel exactly and what survives is pure , which is the transverse effect extracted from a longitudinal measurement. Ives and Stilwell ran hydrogen canal rays at , where , and found the predicted second-order displacement of the mean. It remains one of the cleanest direct tests of time dilation, and the derivation above is the entire theory of the experiment. We quote the experiment, and the algebra is ours.
Two null four-vectors have now appeared in this chapter for entirely different reasons: for a massless particle (§4.3), and for a light wave (§7.1). Nothing so far connects them. But notice that they have the same transformation law and the same null condition. Compare (2.5.33) with and they also have the same relation between time and space components.
Chapter 4.1 supplies the missing constant. The relation is
and the reason it must be exactly this, with a single constant and no angle-dependent fudge, is that both sides are four-vectors. One scalar relates them, or none does.
So the relativity of this chapter forces and to stand or fall together. You cannot have the Planck relation for energy without the de Broglie relation for momentum. That is a strong structural statement, and it was available years before anyone believed either half of it.
Waves need one more four-part object, and the argument producing it is counting rather than calculation. A crest is not a thing that travels but a set of events at which the field is momentarily largest, so whether a given event is a crest is a question about one point of spacetime that both observers are asked about. Let a detector click once per crest between two events on its own history: the count is a whole number, nobody gains or loses one by moving, and whole numbers cannot transform. So the phase is agreed, and whatever pairs with position to produce it must be four-part too.
Frequency and wavelength are thereby welded together as energy and momentum are, and the whole Doppler effect becomes one multiplication. Light from a source passing at closest approach, where the separation is momentarily unchanging and the old theory predicts no shift, arrives reddened by exactly the dilation factor.
Two null objects have now appeared for unrelated reasons, one for a massless particle and one for a light wave, sharing a transformation law, a null condition and the same relation between time and space parts. One constant relating them relates both halves at once, since a single scalar between two such objects covers everything or nothing. Planck's relation for energy and de Broglie's for momentum stand or fall together, which was available years before anybody believed either.
8 · Collisions: invariant mass, and why colliders exist
The whole practical value of §2 is that it turns collision problems into linear algebra. You write down the four-momenta, you add them, you square whatever combination kills the unknowns, and you are done. This section builds the three tools and then uses them to say something quantitative and slightly outrageous about accelerator design.
8.1 · The invariant mass of a system
Take any collection of particles, whatever they are doing, and add up their four-momenta:
A sum of four-vectors is a four-vector, so transforms properly and has an invariant square. That square is the object we want, so give it a name. Define the system's invariant mass by
Everyone agrees on . And the system's total four-momentum is conserved, by §2, so is conserved as well. That is a strong constraint on what a system may turn into.
8.2 · The centre-of-momentum frame
Suppose is timelike, . Then by exactly the construction of Chapter 2.3 §6.1 there is a frame in which its spatial part vanishes. Boost with , which has magnitude less than precisely because is timelike. In that frame, called the centre-of-momentum frame or CM frame,
So the invariant mass of a system is its total energy in the frame where it is collectively at rest, divided by . Notice that this is the direct generalisation of from one particle to many.
Notice also that is exactly (2.5.32) applied to the system's four-momentum, which is a small piece of evidence that we defined things sensibly.
Take two photons of energy flying in opposite directions along . Each is massless. Add their four-momenta:
so and the pair has invariant mass . That is nonzero, and it is built entirely out of massless constituents. There is even a CM frame, and it is the one you are already in, since .
Now aim the same two photons in the same direction. Then , , and . Same two objects, same energies, different invariant mass.
So mass is not stuff that the parts have and the whole inherits. It is a property of the total four-momentum, and specifically of how much the constituent four-momenta point in different directions in spacetime. Randomly directed momentum inside a box shows up as mass of the box.
This is not an analogy and not a special case. It is where most of your own mass comes from, as §9 explains, and it is why "matter is made of massive things" is the wrong picture at the bottom.
8.3 · Mandelstam
For a two-body collision the invariant mass of the initial state gets its own name. Define
(The factor of is bookkeeping so that is an energy in SI units. In the natural units of Part V, where , the definition is simply and this parenthesis is unnecessary. Chapter 5.9 introduces the companions and . Here we need only .)
Our goal is to see what depends on, so expand the product once, using on the two square terms:
Everything now depends on the single cross term, and is an invariant we can evaluate in whichever frame is convenient. Two configurations matter.
Fixed target. Particle 1 has energy and hits particle 2 at rest, . Then , so
Collider. Two beams of energy meet head-on with equal and opposite momenta. Then , so you are already in the CM frame and
Compare the two scalings. Doubling the beam energy of a collider doubles . Doubling the beam energy of a fixed-target machine multiplies by .
Useful energy is , not . In the fixed-target case the rest of it is spent dragging the centre of mass forward, and is unavailable for making anything. The consequence compounds viciously with energy, and Worked example 2 puts a number on it that is hard to believe until you check it.
8.4 · Threshold energies, and the antiproton
Here is the technique that makes worth defining. A reaction can proceed only if the initial state carries enough invariant mass to build the final state.
The minimum case is unambiguous. At threshold the products have no kinetic energy left over in the CM frame, so they are all at rest there and moving together as one. The final invariant mass is then just the sum of the final rest masses. Since is conserved and invariant,
Take the reaction that mattered historically, which is producing an antiproton by smashing a proton beam into a hydrogen target,
Charge and baryon number both force the extra proton to accompany the antiproton, since you cannot make alone, so four particles of mass must emerge. (Antiparticles have the same rest mass as their partners. ⚑ We quote that here, and Chapter 5.5 derives it, where it falls out of the Dirac equation.) Now set in (2.5.64) with and solve for the beam energy:
That is the projectile's total energy, and what an accelerator is rated by is the kinetic energy it delivers. So subtract the rest energy the proton had before anyone switched the machine on:
Six proton rest energies to make two particles' worth of new rest mass. The factor of three overhead is exactly the energy locked up in the forward motion of the centre of mass, which (2.5.64) says you cannot spend. ⚑ Historically: the Bevatron at Berkeley was designed to reach , comfortably above this threshold and for this reason, and the antiproton was found there in 1955. The number in (2.5.69) is a piece of accelerator engineering that came out of four lines of four-vector algebra.
Grind box — the general threshold formula, and a second worked case
Do the algebra once in general so you never have to repeat it. A projectile of mass and total energy strikes a stationary target of mass and produces final particles of total rest mass . Equate (2.5.64) with (2.5.66):
Check against (2.5.68): , gives ✓.
A second case, pion photoproduction, . Now for the photon, , and with (⚑ quoted mass). Then
Numerically . The photon must supply about 7% more than the pion's rest energy, the surplus being the recoil the proton is obliged to take. Note how the formula degrades gracefully. For a very heavy target, and the overhead vanishes, which is why a nucleus makes a better anvil than a single proton.
The same physics, from the CM frame. There is a shortcut worth knowing. The overhead in fixed-target work is the kinetic energy of the centre of mass, which by (2.5.32) moves at . Everything above can be re-derived by boosting to that frame and demanding the products be at rest, but the invariant route is shorter precisely because it never needs at all. That is the habit to acquire: find the invariant, evaluate it in the easy frame, use it in the hard one.
Add the four-part momenta of any collection of particles, whatever they are doing, and the total is another object of the same kind with an invariant square of its own. That square defines the collection's mass; everybody agrees on it, and since the total is conserved it is conserved too. In the frame where the collection is collectively at rest it is the total energy over the square of the speed limit, which is the single-particle statement about rest energy carried over to many.
What this does to the word mass repays dwelling on. Two pulses of light of equal energy flying apart possess a mass, built entirely from constituents having none, while the same two aimed the same way possess none. Nothing changed but where they pointed. Mass is a property of the total, measuring how far the constituent momenta point in different directions, so randomly directed motion inside a box shows up as mass of the box.
The practical consequence is a piece of accelerator engineering. Only that invariant mass is available for making anything new, and energy tied up in the forward motion of the whole cannot be spent. Doubling the beam energy of a machine whose beams meet head-on doubles what is available; doubling it in one firing into a stationary target multiplies it by the square root of two. That difference in scaling is the whole argument for building colliders.
9 · Mass is not additive
Section 2.3's grind box already produced the result that makes this section necessary. Two lumps of putty of rest mass each, colliding and sticking, form a body of rest mass , not . The kinetic energy did not disappear. It became rest mass. Run that backwards and you have the most consequential fact in applied physics.
9.1 · Binding energy and the mass defect
Consider a bound system: a nucleus, an atom, a planet. Assemble it from constituents that start far apart and at rest, and let it settle by radiating away the excess. The total four-momentum is conserved throughout, so in the frame where the final object sits at rest,
where is the energy carried away, called the binding energy. The bound system is lighter than its parts, by exactly the energy you would have to put back to take it apart, over .
That difference is the mass defect. It is not a small correction to a picture in which mass is additive. It is a demonstration that mass was never additive.
Real numbers, since the whole point is that this is measurable. ⚑ The masses below are quoted from standard tables, and everything done with them is ours.
| System | Constituents () | Bound mass | ||
|---|---|---|---|---|
| Deuteron H | ||||
| Helium-4 | ||||
| Hydrogen atom |
Read the last column. Chemistry runs at or of the rest energy. That is why Lavoisier could weigh reactants and products and conclude that mass is conserved. His balance was many orders of magnitude too coarse to see the defect, and he was right to the precision available.
Nuclear binding runs at , which is five to seven orders of magnitude larger. That gap is the entire reason nuclear energy is a different kind of thing rather than a better kind of chemistry.
Grind box — the arithmetic, and a fusion reaction end to end
Deuteron. A proton and a neutron bound together. The constituent rest energies are , and the measured deuteron rest energy is . Difference:
So a deuteron weighs about less than the parts you made it from. That is a fractional mass difference roughly times larger than any chemical bond, and comfortably measurable with a mass spectrometer.
Helium-4. The constituents come to and the measured mass to , so . That is per nucleon and of the total. Helium-4 is unusually tightly bound for its size, which is why particles exist as a distinct thing and why stellar nucleosynthesis piles up at helium.
A complete reaction: D–T fusion. . Using nuclear rest energies in MeV:
As a fraction of the input rest energy, , which is . Per kilogram of fuel that is , against roughly for burning hydrocarbons. The ratio is about seven million.
The lesson to extract. Every one of these is the same calculation: add the rest energies before, add them after, and the difference is kinetic energy. Nothing was "converted into energy" in a way that requires new physics. The bookkeeping quantity was conserved throughout, and all that changed was how much of it was sitting in the rest-mass column.
The standard telling is that in a nuclear reaction "some mass is converted into energy, according to ". Nothing in this chapter supports that reading, and it causes real confusion.
What is actually true. The conserved quantity is , and it never changes. In a fusion reaction the rest masses of the products are smaller than those of the reactants, so the rest-energy share of the total falls, and the difference reappears in the kinetic-energy share. Nothing was created or destroyed and nothing turned into anything. Energy was reallocated between two columns of the same ledger.
And the system's mass? Do the reaction inside a sealed, perfectly reflecting box and weigh the box. Its invariant mass is unchanged, because the kinetic energy and radiation are still inside, and by §8.2 the box's invariant mass is its total energy in its rest frame over . The mass only drops once you let the heat out.
So the honest sentence is this. Rest energy is released as kinetic energy, and if the kinetic energy escapes, the system's rest mass drops by exactly the energy that left, over .
A related trap. "" is about rest energy. The energy of a moving body is , and the two differ by every factor that matters at an accelerator. When a physicist writes they mean the case of (2.5.26), and the famous equation is famous partly because that qualification is usually dropped.
9.2 · Where your mass actually comes from
Push the argument one step further than nuclear physics does and it stops being a correction and becomes the main effect.
A proton has rest energy . It is made of three light quarks whose rest energies sum to roughly . ⚑ That figure is quoted from Chapter 6.5's tables, and note that even defining a quark mass takes care, since quarks are never found alone. It accounts for about of the proton.
The other is exactly what §8.2's warning described. It is the energy of quarks moving relativistically inside a small volume, plus the energy of the gluon field binding them, all of it appearing as invariant mass because the constituent four-momenta point in many different directions and their spatial parts cancel while their energies add.
So the mass of ordinary matter is, to better than , not a property of its ingredients. It is confined kinetic and field energy, measured in the frame where the total momentum vanishes. That is (2.5.60), applied to a bound state of a strongly coupled field theory.
Chapter 6.5 does that calculation, or rather explains why it can only be done numerically. The structural statement is available now, and it belongs to this chapter rather than to quantum chromodynamics.
Everybody has had to explain that variances add and standard deviations do not. The reason is a norm identity, , whose cross term is the covariance. When the parts are uncorrelated the cross term vanishes and the squares add. Section 8 is that identity in a space with a different inner product.
Write (2.5.60) out for two particles, using (2.5.59):
The squares add and a cross term measures alignment, exactly as before. Two photons flying apart have and therefore a mass built entirely out of massless parts. The same two aimed the same way have a cross term of zero and no mass at all.
Nothing changed but the correlation between the directions, which is why §9.2's proton weighs what it does. Its constituents' momenta point every way, the cross terms survive in full, and the mass of the whole is mostly the misalignment of the parts.
Now the place the analogy stops, which is worth more than the place it holds. A covariance may be negative, so can fall below . The Minkowski cross term cannot do the corresponding thing. For two future-pointing momenta,
with equality only when the two move together, so for any collection of free particles. There is no anti-correlated arrangement that lightens a system. A bound system does weigh less than its pieces, and §9.1 is careful about why: energy physically left, carried away as radiation, which is not this identity at all.
The algebra of the norm of a sum is shared. The sign structure that makes this one a one-sided bound is not, and it comes from the same minus sign that produced the causal structure two chapters ago.
A bound system, assembled from constituents that began far apart and settled by radiating the excess away, weighs less than its parts by exactly what left. Chemistry does this at one part in a hundred million, far under what Lavoisier's balance could resolve, which is why he pronounced mass conserved. Nuclear binding does it at one part in a thousand, and that gap of five orders is why nuclear energy differs from chemistry in kind.
The usual telling, in which mass converts into energy, is not what happens. The conserved total never changes; only which column it sits in, rest mass or motion. Run the reaction in a sealed reflecting box and weigh it: the mass is unchanged, everything released still inside. It falls only when the heat is let out.
One step further and the correction becomes the main effect. The three quarks in a proton account for about one hundredth of its mass; the rest is their confined motion and the binding field, appearing as mass because their momenta point every way and cancel while the energies add. That number is the one the chapter on expansion said no series would find, lying so flat near the origin that every term reports it as zero. Mass entered as the amount of stuff in a body and leaves as a label on a four-part object, better than ninety-nine per cent of yours borrowed motion.
10 · Worked examples
A photon of wavelength strikes a free electron at rest. The photon scatters through an angle and emerges with wavelength , while the electron recoils in some unknown direction with some unknown speed. Find .
Why this is the model calculation. There are four unknowns after the collision, namely the photon's new wavelength and the electron's three momentum components, and there are only four conservation equations. But we are not asked about the electron at all.
So here is the technique. Isolate the four-momentum you do not care about on one side, and square it. Squaring converts an unknown four-vector into the one number you do know about it, namely . That single move removes the recoil direction from the problem without ever computing it.
Setup. ⚑ One quoted input, and only one: a photon of wavelength has energy . That is Planck and Einstein, and Chapter 4.1 is where it is argued for. Everything else below is this chapter's. Given it, §4.3's gives the photon's momentum , so with a unit vector along its travel,
and the electron starts at rest, , with unknown final .
Step 1. Conservation, rearranged to isolate the unknown.
Step 2. Square both sides. The left side is known, since by (2.5.29), whatever the electron ended up doing. Expanding the right side needs six terms, so take them one at a time:
The first two vanish because photons are massless. Setting the whole thing equal to , the terms cancel and we are left with a relation among three dot products:
The electron's final state has vanished entirely. That is the whole trick.
Step 3. Evaluate the three dot products. Each is elementary. Keep the metric signs straight.
Step 4. Assemble.
Now multiply through by . The product was the only place the two wavelengths appeared together, and it cancels completely:
What the answer says. The shift does not depend on . It is bounded, between at and at . And it is a pure length times a pure number.
The Compton wavelength, and a payoff from Chapter 0.3. The combination
is the only length you can build from , and . Check that against Chapter 0.3 §5's method. With , and , seeking of dimension gives , , , whence , , , uniquely.
So dimensional analysis alone guaranteed that if the answer is a wavelength shift built from these three constants, it must be times a dimensionless function of . The four-vector calculation supplied the function, and it is .
Numbers, and why anyone believed it. Compton used molybdenum X-rays, . At the shift is , a change, easily resolved. At it is , or .
Crucially, the shift is independent of the incident wavelength and of the target material, which is impossible for any classical scattering mechanism. A classical wave shakes an electron at the driving frequency and the electron re-radiates at that same frequency, giving no shift at all.
Chapter 4.1 uses exactly this calculation as evidence that light carries momentum in discrete parcels obeying (2.5.33). The collision is between two particles, and the bookkeeping that works is this chapter's.
The LHC collides protons head-on at per beam, giving . What beam energy would a fixed-target machine need to reach the same against a stationary proton? And conversely, what does the LHC's own beam achieve if you point it at a stationary target instead? (⚑ The machine parameters are quoted, and the physics is ours.)
Part (a), the fixed-target equivalent. Rearrange (2.5.64) with :
With and , the is utterly negligible against , so
Compare that with the each LHC beam actually carries. The fixed-target machine would need
times more energy per particle, which is about fourteen and a half thousand. The LHC's ring is around and its bending magnets are near the limit of what superconductors will do. Scaling the same technology by means a ring on the order of , which is about the distance to the Moon. That is the scaling, made visceral.
Part (b), the LHC beam on a fixed target. Same formula, other direction, with :
So a proton hitting a stationary proton makes only of invariant mass available, which is 1.7% of the beam energy. The other 98.3% is spent dragging the centre of mass down the beam pipe and is gone.
The overhead is not a constant. Measure the waste by the ratio of beam energy to available energy, . Rearranging (2.5.64) for equal masses gives
which grows linearly with the energy you are trying to reach. At the antiproton threshold of §8.4, and the ratio is . You buy of invariant mass with a beam, which is tolerable. At the LHC's beam energy the same ratio is .
So there is no fixed inefficiency to design around. The waste compounds. That is why every energy-frontier machine built since the 1970s has been a collider, and why fixed-target experiments now go looking for rare processes at modest rather than for new mass.
11 · Your turn
Problem 1 — a photon cannot decay into an electron–positron pair
In empty space, consider the proposed process . Show, using invariant mass alone, that it is impossible, no matter how energetic the photon is. Then explain in one sentence why the same process does occur routinely near a heavy nucleus.
Solution
Four-momentum conservation would require . Squaring both sides is legitimate because both are four-vectors, and the square is an invariant, so if the two sides are equal their squares are equal in every frame.
Left side. The photon is massless, so .
Right side. Let . Each electron four-momentum is timelike and future-pointing, since energies are positive, so their sum is too, and by §8.2 there is a frame in which . Evaluate the invariant there, which is the whole art:
In that frame the two particles have equal and opposite momenta, so equal energies, and each satisfies by (2.5.31). Hence
So the process demands , a contradiction. The photon's energy never entered the argument, because is an invariant and we were free to compute it in the pair's own rest frame.
Restated geometrically, which is the version to remember: a null vector cannot be the sum of two timelike future-pointing vectors. The sum of future-timelike vectors is future-timelike, strictly. The light cone is the boundary, and you cannot get back onto it by adding things from the inside.
Near a nucleus. The nucleus absorbs four-momentum, so the reaction is really , and the initial invariant mass is no longer zero. By (2.5.63) it is , which exceeds once is large enough. Applying the grind box's threshold formula with , , :
which for a heavy nucleus is barely above . The nucleus takes almost no energy but supplies the momentum balance. It is there to break the kinematics, not to pay for it. This is why is the number quoted for pair production, and why positron emission tomography works at per photon in the reverse process.
Problem 2 — the relativistic rocket, and its energy budget
A ship accelerates with constant proper acceleration , meaning the accelerometer bolted to its floor always reads , starting from rest at .
(a) Using , and , show that with , and hence that and .
(b) With , find and after one year and after ten years of ship-time, and the energy per kilogram of payload required in each case.
Solution
(a) The constraint says the four-velocity is confined to a hyperbola in the plane, and the general parametrisation of is , for some function . That is precisely Chapter 2.3 §3.2's hyperbolic parametrisation, which is why rapidity was worth defining. So with no loss of generality. Differentiate:
Check the orthogonality first, as a sanity test: ✓, automatically. The constraint (2.5.13) is built into the parametrisation. Now the magnitude:
Setting this equal to gives , so with for a start from rest. Finally, from we read off and , hence
Constant proper acceleration is uniform growth of rapidity. That is the clean statement, and it is why appears. Velocities do not add, rapidities do, and a constant push adds rapidity at a constant rate. It also settles the speed limit question for good, since approaches and never reaches it, however long you burn.
(b) The natural timescale is . So one year of ship-time is .
| ship-time | kinetic energy per kg | |||
|---|---|---|---|---|
The kinetic energy is , with . After one year it is per kilogram, which is more than half the payload's rest energy, and that is only the payload, ignoring the fuel needed to carry the fuel. After ten years it is times the rest energy. To deliver one kilogram you must supply the total annihilation energy of about fifteen tonnes, and no engine is perfect.
The point. Relativity does not forbid interstellar travel. It prices it, and it prices it exponentially, because grows exponentially in ship-time, and so does the distance covered, . Set light-years, the distance to the galactic centre, and you get , hence and years of ship-time. That is the romance, and it is real. The bill is per kilogram delivered. And since the ship never quite reaches , some years will have passed on Earth by the time it arrives.
Problem 3 — the headlight effect
A source at rest in frame emits photons uniformly in all directions. moves at along relative to the lab. Using the aberration formula from §7's grind box, show that the photons emitted into the forward hemisphere in (those with ) are all compressed, in the lab, into a cone about the -axis of half-angle , and that for this is approximately . Evaluate for and , and comment on what this does to the apparent brightness of a relativistic jet pointed at you.
Solution
The grind box in §7 derived , mapping lab angle to source angle. We want the inverse, which is the same formula with , since that swaps the roles of the frames:
The boundary ray. Put , so that . Then . So the photon emitted exactly sideways in the source frame arrives in the lab at angle to the direction of motion. Since the map is monotonic in , every ray with lands inside that cone.
The small-angle form. For small, , so radians. Neat, and worth remembering as a rule of thumb.
| solid angle fraction | |||
|---|---|---|---|
(The solid-angle fraction of a cone of half-angle is .)
Brightness. Half of all the emitted photons, meaning the forward hemisphere, which is steradians in the source frame, end up inside a lab cone of solid angle . The photon flux per unit solid angle is therefore boosted by roughly .
Each of those photons is also blueshifted by the longitudinal Doppler factor from (2.5.55), and they arrive at a compressed rate for the same reason. Multiplying the effects gives an apparent brightness enhanced by a large power of . The standard estimate for a continuously emitting jet is in the observed flux, though the exact exponent depends on the source's spectral shape.
Consequence. A jet pointed within of your line of sight looks enormously brighter than the identical jet pointed away. That is why blazars, which are active galaxies whose jets happen to aim at Earth, dominate the catalogues of the brightest gamma-ray sources despite being a tiny fraction of active galaxies. It is also why radio astronomers see one-sided jets from manifestly two-sided sources, the receding jet being de-beamed by the same large factor. Neither observation requires any asymmetry in the source. Both are (2.5.52), boosted.
Problem 4 — mass defect and energy release in D–D fusion
Consider . Using nuclear rest energies , , (all in MeV), compute (a) the energy released ; (b) as a fraction of the input rest energy; (c) the energy released per kilogram of deuterium fuel, and compare with burning an equal mass of methane (). (d) Finally, the neutron and the He share as kinetic energy; use momentum conservation in the CM frame to find how it splits, and check that the non-relativistic treatment is justified.
Solution
(a) Total four-momentum conservation, evaluated in the CM frame where the initial nuclei are brought together with negligible kinetic energy, gives :
(b) , which is . That is under a tenth of a percent of the rest energy, and yet:
(c) Energy per kilogram is that fraction times :
Against methane's , that is a factor of , a million and a half. One kilogram of deuterium releases what about 1400 tonnes of methane would.
The entire difference between chemistry and nuclear physics is the difference between the last column of §9.1's table for a hydrogen atom () and for a nucleus (). The mechanism is identical and only the scale of the binding differs.
(d) In the CM frame the products have equal and opposite momenta, . Non-relativistically the kinetic energies are , so they are inversely proportional to the masses:
With , this gives
The light one takes most of the energy. That is a general feature of two-body decay, and it is the reason fusion reactors must cope with fast neutrons.
Justifying the approximation. The neutron's kinetic energy is of its rest energy, so and . By the grind box in §4, the fractional error in using at this speed is about with , which comes to . Three parts per million is far below the precision of the input masses, so the non-relativistic split is fine. Note that we needed relativity to compute at all, and then did not need it to divide up. That combination is typical of nuclear physics.
You have mechanics that survives a boost. The route was one demand, that a conservation law must be an equation between tensors, and everything else was consequence.
Differentiate by the invariant to get with . Multiply by to get . Observe that conserving three components of a four-vector in every frame forces the fourth, and that the fourth expands to . From the invariant square came , from that the massless case as the null case, and from the fact that null four-momentum means exactly .
You also have the action. was solved for, not guessed, and it closed Chapter 2.3's maximisation theorem. Extremising the action and maximising proper time are the same instruction, differing by the minus sign that converts one into the other. That is entry two in Chapter 1.2 §8.1's table, now derived.
And you have the tools that make relativistic problems tractable: contract to kill unknowns, boost to the frame where the answer is easy, and remember that invariant mass is a property of a four-momentum rather than a substance carried by parts.
Where this gets spent. Chapter 2.6 takes and asks what four-vector the electromagnetic field puts on the right. The answer is , and its time component is the power (2.5.46). That chapter also finds the missing momentum that Chapter 1.1 left unaccounted for, and it is in the field.
Chapter 3.6 builds , the object that sources gravity, out of exactly the energy and momentum densities defined here. Energy gravitates, not mass, which is why light bends.
Chapter 4.1 reuses Worked example 1 as evidence that photons carry momentum, and promotes from a structural suspicion to a law. Chapter 5.1 shows that plus quantum mechanics makes particle number unfixable, since you can always find somewhere, which is why fields replace particles. Chapter 5.9 computes a real cross-section in the Mandelstam variables §8.3 introduced. And Chapter 6.5 explains the of the proton's mass that §9.2 could locate but not account for.