Part II · Special Relativity — Chapter 2.6
Electromagnetism Is Relativity
Maxwell's equations never needed fixing. They were relativistic before anyone knew what that meant. That is exactly why they produced a frame-independent and started the crisis in the first place.
Chapter 2.1 left you at a fork. Three things were believed in 1900: the principle of relativity, Maxwell's equations, and the Galilean transformation. Any two of them sat together comfortably. All three at once were impossible.
Three branches led away from that fork.
- Branch (A) said relativity fails for optics, and there is an ether frame after all.
- Branch (B) said Maxwell is wrong.
- Branch (C) said the Galilean transformation is wrong.
Chapters 2.2 to 2.5 took branch (C) and rebuilt kinematics and dynamics from the wreckage.
This chapter collects the payment. Rewritten in the language of Chapter 2.4, Maxwell's four equations become two, and one of those two is not physics at all but bookkeeping.
The paradox of 2.1 §3 evaporates as well. That paradox was the question with respect to what? It goes away not because we answer it, but because we finally see that it was malformed. Maxwell's equations single out a speed. They never singled out a frame. Those are two different things. Only after 2.2 do you have a geometry in which a speed can come out the same for everyone without there being anything for it to be measured relative to.
Two promises also come due here.
The first is from Chapter 2.4 §7.2. That section counted six independent components in an antisymmetric tensor and told you they would turn out to be and . Sections 2 and 3 below deliver them, by construction, with every component checked.
The second is from Chapter 1.1 §3.3. That section showed Newton's third law failing for two moving charges, left an explicit IOU marked "the field carries momentum", and told you Chapter 2.6 would pay it. Section 10 pays it, down to the last factor of .
Tools you'll need — Chapter 0.7: divergence and curl in Cartesian components (§3.3, §4.2), the continuity equation (§6), and the two second-derivative identities and (§7), together with §7.3 on why the vector potential exists and is not unique. Chapter 1.1 §3.3: the two moving charges that break the third law. Section 10 here is the resolution, so it is worth rereading the setup. Chapter 1.2 §8: the field Euler–Lagrange equation, and §8.1's table of actions. Section 9 here supplies the third entry. Chapter 1.4 §6: Noether's theorem for fields, and its grind box on improvement terms, which §10 uses without apology. Chapter 2.1 §2 (Maxwell produces ), §4 (the wave equation is not Galilean invariant) and the fork of §4.3. Chapter 2.4 §2.1 (the boost matrix ), §4 (raising and lowering), §5 (the tensor transformation law and the quotient theorem), §6 (the invariance theorem), §7.2 (the count of six), §8.1 (the Levi-Civita symbol). Chapter 2.5 §1.3 (four-velocity), §2 (four-momentum), §6.1 (the covariant force with ).
This chapter does not derive electromagnetism from nothing. Two things are taken as given.
(i) Maxwell's four equations, quoted in Chapter 2.1 §2 as the empirical input of that chapter, in SI units and with sources present:
(ii) Charge invariance. The electric charge of a body is the same number in every inertial frame. This is an experimental fact rather than a theorem, and it has been tested very sharply. Atoms and molecules come out neutral to better than one part in , even though the electrons inside them move at wildly different speeds from the nuclei. Chapter 6.3 will make charge invariance a consequence of gauge symmetry. Here it is an input, and §6 is careful about exactly how much work it does.
Everything else below is derived. What the chapter produces is three observations. The first is that (i) is already a Lorentz-tensor equation in disguise. The second is that the disguise is the only thing that ever made it look incompatible with relativity. The third is that Chapter 6.3 can then run the whole construction backwards and obtain (i) from a symmetry principle.
1 · The four-current
We start with the simplest object in the theory. We are going to build it by requirement rather than by decree, which means we will state what we need of it and let that fix it completely.
Chapter 0.7 §6 derived, from nothing but "charge is neither created nor destroyed", the continuity equation
Here is charge per unit volume and is current density. Now look at (2.6.1) with the eyes of Chapter 2.4. It is a sum of four terms: one time derivative of one function, and three space derivatives of three functions. That is exactly the shape of a four-divergence . The catch is the word if: it is that shape only if the four functions can be arranged into a four-vector.
So let us not guess. Let us demand. Suppose there exists a four-vector whose four-divergence reproduces (2.6.1), and ask what its components would have to be. Before we can compare anything we need the derivative operator in the right form. Recall from Chapter 2.4 §8 that carries a lower index, and that , so
Our goal now is to write out the four-divergence in ordinary three-dimensional language, so that we can set it side by side with the continuity equation. Write the unknown four-vector as , split the sum over into its time part and its three space parts, and use the two derivatives we just wrote down:
Now compare term by term with (2.6.1). The spatial part matches immediately if . The temporal part matches if , that is, if . There is no freedom left anywhere. The two expressions agree as functions, for arbitrary and , only for that one choice. So the four-vector we demanded exists, and it is this one:
The factor of is not cosmetic. It is dimensional bookkeeping forced on us by , exactly as the in was.
Now let's look at what has happened to charge conservation. It arrived as a four-term partial differential equation whose behaviour under a change of frame was completely obscure. It now reads . That is a contraction of a index against a index, so it is a scalar by Chapter 2.4 §5.1, so it is the same statement in every frame by the invariance theorem of §6. Charge conservation is manifestly relativistic, and it takes one line to say so.
You should find this unsurprising. If you do not, here is the physical version. A line of static charges has and . Run past it and the charges are moving, so now .
Charge density and current density are the same thing seen from different frames. That is exactly what happened to and in Chapter 2.2, and to and in Chapter 2.5. Section 6 of this chapter turns that one observation into the entire explanation of magnetism.
Chapter 1.4 §6 proved that a continuous symmetry of a field theory produces a conserved current satisfying , and §6.2 of that chapter ran the machinery on a global phase rotation of a complex field. That current, once the phase is allowed to vary from point to point, is (2.6.4). We are not in a position to show that yet, since it needs the gauge principle of Chapter 6.3. So for now is built out of measured densities.
The logical order is still worth knowing in advance. Conservation of charge is not an extra postulate bolted onto electromagnetism. It is a theorem about a symmetry.
Grind box — why alone is not a scalar, and how the 's work out
A common first guess is that charge density should be an invariant, since charge is. It is not, and the reason is instructive.
Take charges sitting at rest in a box of volume , so in their rest frame. View them from a frame in which the box moves at speed along . The charge is unchanged, since that is the invariance assumption (ii). The volume is not unchanged. The box is contracted along by , so , and therefore
Now assemble the four-vector and compare with the four-velocity of Chapter 2.5 §1.3:
So the four-current of a moving cloud of charge is the invariant rest-frame density times the four-velocity. That is manifestly a four-vector, being a scalar multiplying a tensor.
Notice how the single factor of from length contraction is exactly the factor of that already carries. It is the same , arriving twice for the same reason. Had charge density been invariant instead, would not have been a four-vector at all, and (2.6.4) would be false.
One caution. The relation holds for a single species of charge all moving together. A copper wire has two species, a stationary lattice and a drifting electron gas, and you add their four-currents. That addition is the whole of §6, and the fact that the two species carry different factors of is the whole of magnetism.
Charge conservation arrived long ago as a four-term equation whose behaviour under a change of frame was completely opaque: one derivative in time of one function and three in space of three others. Set beside the machinery of the last two chapters, that is the shape of a four-dimensional divergence, provided the four functions assemble into one object. Rather than guess at the arrangement, demand it, and no freedom is left: the charge density, carrying one factor of the speed limit so the entries are commensurable, sits above the three of current.
Charge density is not itself an agreed number, although the charge is. The same charge occupies a contracted volume when you run past it, so the density picks up the dilation factor the four-velocity already carried. Density of charge and density of current are one thing seen from different states of motion, exactly as elapsed time and displacement were, and as energy and momentum were a chapter ago.
The machine assembled three chapters back for manufacturing agreement between observers was promised a velocity, then a momentum, then a current, and it now has all three and wants nothing further. What it returns on this last feeding is charge conservation as a single contraction, the same statement for everybody at a glance, where before it was four terms whose fate under a change of frame nobody could see.
2 · The four-potential and the field tensor
Now for the six components that Chapter 2.4 promised. The plan of this section is short. We recall the two potentials that Maxwell's equations force into existence, we stack them into one four-component array, and we then find that and are what you get by differentiating that array.
2.1 · Recalling the potentials
Chapter 0.7 §7.3 established the first half. Since identically, writing
makes automatic. Feed that into Faraday's law:
The middle step there is the commutation of with the spatial derivatives inside the curl, which is Clairaut's theorem again, as in Chapter 2.1 §2.2.
Now use what that last line tells us. Chapter 0.7 §2.3 showed that a curl-free field on a simply connected region is a gradient. So for some scalar . The minus sign is the standard convention, chosen so that reduces to the electrostatic potential in the static case. Rearranging that for , and carrying the curl formula for along beside it, we have both fields written in terms of potentials:
Two of Maxwell's four equations have now been used up entirely. They are the reason the potentials exist, and once you work with and they hold identically, with nothing left to check. Hold on to that thought. Section 3.4 makes it the punchline of the whole section.
2.2 · Assembling
We now have four functions in hand: one scalar and the three components of . The pattern of §1 suggests trying for some constant that dimensions will fix. It also suggests checking afterwards that the result transforms as a four-vector, rather than assuming that it does.
Dimensions first. In SI, has units of and has units of . So has the units of , and is the only choice that makes the four components commensurable. It is the same , appearing for the same reason, as in and . With that settled, define
One warning about the letter. Chapters 2.2 to 2.5 used for rapidity, and from here to the end of the book it means the electric scalar potential. That is an unfortunate clash of symbols, and the whole subject lives with it. Rapidity does not appear again in this chapter, so nothing is ambiguous on this page. When Part V has to write both in one line, it writes the rapidity as .
Is a four-vector? Not by fiat. The check comes in §4. There we will find that in the Lorenz gauge satisfies , where is a scalar operator and is a four-vector by §1. That forces to be a four-vector too, since otherwise the equation could not hold in every frame. Until we have that argument in hand, treat (2.6.8) as the definition of an array of four numbers and nothing more.
2.3 · The field tensor
Look at (2.6.7) again. Every entry is one derivative of one potential component, minus another derivative of another potential component. In four-dimensional language there is exactly one object with that shape:
It is antisymmetric, , and you can see that by inspection. Swapping and swaps the two terms. That is not a design choice anyone made. It is forced by the construction, and the callout at the end of §2 explains why the construction itself is forced.
The upper-index derivative is , and since this flips the sign of the spatial components:
That minus sign is where readers get hurt. Write it out once and keep it.
2.4 · The six components, one at a time
The three . Put , into (2.6.9):
the last step being (2.6.7) read backwards. So the first row of the matrix is , and the first column is by antisymmetry.
The three . Both indices spatial, so both derivatives pick up the minus sign from (2.6.10):
Our goal now is to recognise that bracket as the curl, so that appears. Chapter 0.7 §4.2 wrote the curl in components. In Levi-Civita notation, which is Chapter 2.4 §8.1 restricted to three indices, it reads . To get at the bracket we contract that with another epsilon and use the standard identity :
The right-hand side there is precisely the bracket sitting inside (2.6.12), so we can substitute it and read off the purely spatial part of the field tensor:
Three entries, since the diagonal vanishes and the array is antisymmetric. Six in total. Assemble them:
Check one entry against (2.6.14) to be sure the epsilon signs landed where they should: ✓, and ✓ since .
That chapter counted the components. An antisymmetric rank-2 tensor in four dimensions has of them. Six is too many for a four-vector, so such a tensor has to be its own kind of object. Chapter 2.4 then told you the six would turn out to be the three components of and the three of , and said nothing more.
Here they are. This is not an analogy, and it is not a coincidence of counting. The object was built out of the potentials that Maxwell's own equations forced into existence, and its six independent entries came out as and with no room to choose otherwise.
and are not two fields. They are one tensor, sliced up by an observer. Which slices you call "electric" and which "magnetic" depends on your state of motion, exactly as which part of you call "time" does. Section 5 computes the slicing rule and §6 shows you a wire where the whole of magnetism is nothing but a re-slicing.
2.5 · Both indices down, and the signs that flip
We will need constantly, and this is exactly the step at which sign errors are manufactured. Lower with two metrics, per Chapter 2.4 §4.1:
Since is diagonal, each entry simply picks up the factor , with no sum implied. For a entry that factor is . For an entry it is . Applying those two factors to the entries we found above gives the rule we will be using for the rest of the chapter:
Lowering both indices of reverses the sign of the entries and leaves the entries alone. One index is spatial in the first case and both are spatial in the second, so one minus sign survives in the first and two cancel in the second. If a calculation below ever comes out with and carrying the wrong relative sign, this is where to look first.
It is tempting to read (2.6.9) as a clever guess. The tempting story goes like this. Somebody noticed that an antisymmetric tensor has six slots, and that and have six components between them, so they packed them in.
That reading gets the logic backwards. The correct order matters, because Chapter 6.3 runs on it.
Here is the chain, in order.
- (1) and Faraday's law force the existence of potentials (§2.1).
- (2) Those potentials are not unique, since you may add to without changing a field. So the physical fields have to be built from in a way that kills that freedom.
- (3) The only combination of one derivative and one potential that is annihilated by is the antisymmetric one. The reason is that is symmetric, so only the antisymmetric part of can avoid it.
- (4) Therefore is antisymmetric, therefore it has six components, and therefore those components are and .
Gauge invariance is upstream of everything. The six-component count is a consequence of it, which makes Chapter 2.4's promise a prediction rather than a coincidence noticed after the fact.
Notice also what happens to the other half of . The symmetric part , with its ten components, is not gauge invariant and carries no physics here. It is nonetheless exactly the object that does carry physics in Chapter 3.6, where the potential is a symmetric and the field is the linearised Riemann tensor. Same construction, different symmetry, and that difference is the difference between a spin-1 force and a spin-2 one.
Grind box — every entry of from , by hand
Nothing here is subtle. The point is to have watched all sixteen numbers appear once, so that the matrix above is a fact rather than a memory. Write and .
Diagonal. for each fixed (no sum). Four zeros.
.
using . Identically and .
. Both spatial, so both derivatives carry the minus:
since from Chapter 0.7 §4.2. Cycling gives and , and the last of these is .
Lower half. By antisymmetry, , , and so on. Sixteen numbers, six of them independent, exactly as (2.6.15) says.
Sanity check on units. Every entry of must have the same dimensions, or the object is not a tensor. is in tesla. And is ✓. So the factor of in is not decoration. Without it the six components could not be six components of one thing.
Two of the four equations are spent before the real work starts, and what they buy is the existence of the potentials: one number and one three-part object, out of which both fields are then built. Look at how they are built and every entry is one derivative of one potential component minus a different derivative of a different one. In four dimensions exactly one object has that shape, and it reverses its sign when its two labels are exchanged.
The antisymmetry is forced rather than noticed. Potentials are not unique, since a whole function's worth of freedom may be added without altering any field, and the only way to build something blind to that freedom from one derivative and one potential is to keep the part reversing sign, the leftover being symmetric and impossible to dodge otherwise. Six independent entries follow, and that count was performed a chapter ago with the answer promised and withheld.
Here is the answer. The six are the three electric components and the three magnetic ones, arriving by construction with no room to choose otherwise. They are not two fields that happen to fit inside one container. They are one object sliced by an observer, and which slices somebody calls electric depends on that observer's motion in exactly the way that which part of a separation between events somebody calls time depends on it.
3 · Four equations become two
This is the centrepiece of the chapter, and nothing in it will be waved through. We are going to write down two tensor equations and then expand every single component of both, so that you can see Maxwell's familiar four emerge one at a time.
3.1 · The inhomogeneous half
What can we build from , one derivative, and the current? The index structure almost writes the answer for us. Contracting the derivative into the first slot gives , which has one free upper index. So does . Two objects with matching index structure can be set equal, and there is only one way to do it, with a single dimensionful constant left undetermined:
We now show that (2.6.19) is Gauss's law together with the Ampère–Maxwell law, and that the constant is . There are four components in it, and we will do all four.
Component . The sum over runs over , but , so only the spatial terms survive:
reading off the first column of (2.6.15). The right-hand side is . Equate and multiply by :
The last step there uses . That is not an extra input. It is Chapter 2.1's result rearranged.
That is Gauss's law. It also fixes the undetermined constant in (2.6.19) to be , because any other choice would give the wrong Coulomb force.
Component (spatial). Now both the time term and the space terms contribute:
The first piece is already in ordinary language. The second piece is not, so let's handle it. We want to see a curl there, so swap the first two indices of the epsilon, which costs one sign, and compare with the component form of the curl:
Putting that back into (2.6.22), and setting the whole thing equal to , the component equation reads . Collecting the three values of back into vectors and moving the time derivative to the right,
That is the Ampère–Maxwell law, displacement current and all. One tensor equation, four components, two of Maxwell's four laws.
Look at where the displacement term came from. Maxwell had to add it by hand, and Chapter 0.7's Problem 4 showed that charge conservation forces it. Here it is not added at all. It is the term of a sum that has to run over all four values of , because the index is contracted and a contracted index sums over everything.
3.2 · The homogeneous half
Two of Maxwell's equations are left: and Faraday's law. Neither contains a source, so the covariant version must have zero on the right. The object that works is the totally antisymmetrised derivative of , written out as a cyclic sum:
Before expanding, notice two structural facts, because they save all the work.
It is totally antisymmetric in . Swap any two of the three indices and the three terms permute into one another with an overall sign. Check for yourself: the first term becomes , which is minus the second, and so on around. So the left-hand side is an antisymmetric rank-3 object.
Therefore there are only four independent components. By the count of Chapter 2.4 §7.2, generalised, a totally antisymmetric rank-3 array in four dimensions has independent entries. There is one entry for each way of choosing three distinct indices from , namely , , and . Any triple with a repeated index vanishes identically.
Four components, and four is exactly the number of scalar equations we need. is one of them, and Faraday's law is the other three. The bookkeeping fits before we compute anything at all.
Triple . All three indices spatial. Using from (2.6.17), so that , , :
No magnetic monopoles.
Triple . One time index. Now , , , and :
Multiply by and you have the -component of Faraday's law, . The triples and give the - and -components by the identical computation with the labels cycled. The grind box does them explicitly, so that "by symmetry" is not doing any hidden work here.
Grind box — the other two Faraday components, written out
Triple . The entries needed are , , :
Triple . Entries , , :
Setting each to zero gives the - and -components of Faraday's law. Together with (2.6.27) and (2.6.26) that is all four independent components of (2.6.25), and all four of Maxwell's homogeneous scalar equations. Nothing is left over and nothing is missing.
Why the cyclic pattern and not ? Because and are the same component up to a sign, and picking the cyclic representative keeps every equation's sign uniform. Choosing the other one is not an error. It gives the same equation multiplied by .
3.3 · The scoreboard
| Covariant statement | Component | Traditional name |
|---|---|---|
| Gauss's law | ||
| Ampère–Maxwell | ||
| Faraday's law |
Four equations, eight scalar equations in the traditional bookkeeping, compressed into two lines.
There is a second payoff, and it is the larger one. Both lines are equations between tensors of the same type, so the invariance theorem of Chapter 2.4 §6 applies to each of them. Each holds in every inertial frame the moment it holds in one. No chain-rule computation, no cross terms, none of the carnage of Chapter 2.1 §4.2. That is the whole purpose of the tensor apparatus, and this is the first place the book collects on it.
3.4 · Two of Maxwell's equations are not physics
Now for the structural point of the section. It is easy to prove and it deserves emphasis out of all proportion to that ease.
Our goal is to see what the homogeneous half says once the potentials are put back in. So substitute the definition (2.6.9) into (2.6.25), with all indices down so that :
Six terms. Let's pair them off. The first cancels the fourth, the second cancels the fifth, and the third cancels the sixth. Each pair differs only by the order of two partial derivatives, and those commute by Clairaut's theorem. So the sum is identically zero.
Given that and come from potentials, and Faraday's law cannot fail. They carry no information about how electromagnetism works. They are the statement that mixed partial derivatives commute, wearing a hat.
You have met both of them before, in Chapter 0.7 §7. There, was six terms cancelling in pairs by Clairaut, and that identity is exactly (2.6.26). Alongside it, was the same cancellation with fewer terms, and that identity is exactly the content of Faraday's law once is written in terms of and .
So the six-term cancellation in (2.6.28) is both of Chapter 0.7's identities at once. In four dimensions they are one identity.
The physical content of electromagnetism therefore sits entirely in (2.6.19), which says that charges and currents source the field. The other two equations are the price of using potentials, and they are free.
Chapter 0.7 §7.4 lined up gradient, curl and divergence and observed that composing two consecutive arrows gives zero, both times, and promised that in the language of differential forms both statements are the single equation . Chapter 3.5 supplies that language. In it, is a one-form , the definition is (2.6.9), and (2.6.25) collapses to the single symbol . The inhomogeneous half becomes .
The asymmetry you have just noticed between the two halves, that one is an identity and the other is a field equation, then becomes the visible asymmetry between and . We flag this as a promise rather than a derivation. You now know what it is a promise about.
Written in the new language four equations become two, and the compression is not typographical. One of them carries a source and admits a single sensible arrangement of labels; expanding its four entries returns the law relating field to charge and the one relating circulating magnetic field to current. The term Maxwell inserted by hand, whose absence made the set inconsistent with charge conservation, is not inserted at all: it is one entry of a sum obliged to run over all four values because the label is contracted.
The other equation is not physics. Write the field tensor in terms of the potentials and six terms cancel in pairs, each pair differing only in the order of two derivatives. Given potentials, the absence of magnetic sources and the law of induction cannot fail, and carry no information about how electromagnetism works. They are the toolkit's two identities, the vanishing swirl of a gradient and the vanishing outflow of a curl, which in four dimensions are one.
Both lines relate objects of the same type, so each holds in every frame the moment it holds in one, with no chain rule, no cross terms and none of the carnage that transforming the wave equation by hand produced earlier. The content of the subject is thereby one sentence: charges and currents make fields. The rest is the price of using potentials, and the price is nothing.
4 · Gauge invariance
4.1 · The freedom
The potentials are not unique, and Chapter 0.7 §7.3 already said why. Adding a gradient to leaves alone, because the curl of a gradient vanishes. In four-dimensional language the statement is cleaner, and it covers as well. Let be any smooth scalar function of spacetime, and transform the four-potential by
The question is what that does to the fields, so our next step is to build the field tensor out of the new potential and see what survives. Feed straight into the definition (2.6.9):
That is Clairaut again, the third appearance in as many sections, and the last. The fields are untouched. Since the fields are what exert forces, no experiment can distinguish from .
It is worth seeing what this looks like in the older notation, so let's unpack (2.6.29) into three-vector language. The one thing to watch is the minus sign in :
You may have seen the familiar textbook form, which has and . That is the same transformation with . Nothing whatever depends on which sign convention you adopt.
Chapter 1.2's Problem 4 established that adding a total time derivative to a Lagrangian, , changes the action only by a boundary term and therefore leaves the equations of motion untouched. Chapter 1.4 §1.2 then built the definition of a symmetry around exactly that loophole. Gauge freedom is that freedom.
Section 9 makes the correspondence exact. The interaction term in the electromagnetic Lagrangian is , and under (2.6.29) it changes by
That is a pure four-divergence, which is the field-theory version of . It integrates to a boundary term by the divergence theorem, so it changes nothing.
Note what had to be true for that to work: charge conservation. Gauge invariance of the action and conservation of charge are the same statement seen from two sides. That is the first hint of the structure Chapter 6.3 turns into a machine.
4.2 · Choosing a gauge, and recovering 2.1's wave equation
Freedom is a nuisance if you want to solve for something, so we spend it. Impose the Lorenz condition
The equivalence there follows from .
Notice that (2.6.32) is itself a scalar equation, being a contraction, so imposing it in one frame imposes it in all of them. That is not true of the other common choice, the Coulomb gauge . That one singles out a frame, which makes it useless for our purposes here. It is convenient elsewhere, and §10 uses it once, deliberately.
It is always attainable. Suppose you are handed potentials with . Gauge-transform by some and compute the new divergence:
So the Lorenz condition holds for the new potentials precisely when . That is an inhomogeneous wave equation for with a known source.
For any reasonable source the equation has solutions. This is the standard existence theorem for the wave operator (the solution is built from the retarded Green's function, ), and we quote it rather than prove it. The proof is a chapter of analysis rather than of physics, and Chapter 5.4 constructs the Green's functions properly when it needs the propagator. What matters here is only that a solution exists, so the Lorenz gauge is always available.
Note also that the solution is not unique. Any with may be added to it, and that leftover freedom is called the residual gauge. It is precisely why the photon has two polarisation states rather than four, and Chapter 5.8 spends real effort on it.
Now the payoff. Substitute the definition of into the field equation (2.6.19) and impose (2.6.32):
The middle step moved past , which is legitimate because partial derivatives commute. That let the Lorenz condition kill the second term outright. So all of electromagnetism, in the Lorenz gauge, has come down to one line:
Four uncoupled wave equations, one per component, each with its own source. In vacuum () this is , that is,
That is the wave equation of Chapter 2.1 §2. It was derived there by taking the curl of Faraday's law and grinding through . Same equation, same , obtained here in three lines because the bookkeeping was done first.
Chapter 2.1 §3 asked the question that broke nineteenth-century physics. The wave equation contains a speed . Every other wave equation's speed is measured relative to a medium. So what is measured relative to?
Look at (2.6.35) and you can see why the question has no answer. The operator is a scalar, with two indices contracted, so by Chapter 2.4 §5.1 it takes the same form in every inertial frame. And enters it only through , which is to say only through the metric of spacetime itself.
So the constant in Maxwell's equations is a property of the geometry in which the fields live, rather than of any substance they live in. Asking what it is measured relative to is like asking what the number is measured relative to.
Branch (A) of 2.1's fork is now worse off than merely unsupported by Michelson and Morley. It is structurally unavailable. Maxwell's equations do pick out a preferred speed. They never picked out a preferred frame, and only the assumption that made those two look like the same claim.
Electromagnetism was never the theory that needed fixing. It was relativistic from birth. What had to give way was mechanics, with its Galilean addition of velocities, and Chapter 2.5 has already paid that bill.
In this chapter gauge invariance is a convenience. It lets you choose (2.6.32) and turn a coupled mess into four wave equations. That is a wild understatement of its importance.
Chapter 6.3 reverses the logic. Instead of noticing that electromagnetism happens to have this redundancy, it demands that the phase of a charged quantum field be adjustable independently at every point of spacetime. That is a symmetry with a whole function's worth of parameters, and it is a heavy demand. Chapter 6.3 then finds that the demand cannot be met unless a vector field exists, transforming exactly as (2.6.29) and coupling exactly as . The entire content of this chapter comes back out as a consequence.
Repeat the trick with a non-commuting symmetry group and out come the weak and strong interactions (Chapter 6.4). Gauge freedom stops being a convenience and becomes the generating principle of every known force except gravity. And gravity turns out to be the same trick applied to the Lorentz group.
The natural reading of §4.1 is that is a convenient fiction and is the real thing. The argument runs like this. The fields are gauge invariant, the potentials are not, so only the fields can be physical. That reading is almost right, and the exception to it is one of the most instructive facts in physics.
Chapter 0.7 §2.4 built a vector field on the punctured plane that is curl-free everywhere and yet has circulation around every loop enclosing the puncture. Its potential exists locally and is the polar angle, which is multivalued. Now read that as electromagnetism. Take a long solenoid with inside and everywhere outside. Outside, is curl-free. It still cannot be set to zero, because around a loop enclosing the solenoid equals the enclosed flux, by Stokes' theorem, and that is not zero. The region outside the solenoid is not simply connected, and Chapter 0.7's counterexample is exactly this situation.
Classically nothing follows, since no charge outside ever feels a force. Quantum mechanically something does follow. Chapter 5.6 shows that a charged particle's amplitude picks up a phase along its path. So two paths passing on opposite sides of the solenoid differ in phase by , and the interference pattern shifts, even though the particle never enters a region where the field is nonzero.
This is the Aharonov–Bohm effect. It was measured in 1960, and it settles the question. The potential carries information that the field does not, and that information is topological. What is physical is neither , which is gauge dependent, nor , which is too little. It is the gauge-invariant loop integral , called the holonomy. Chapter 6.3 builds the whole of gauge theory on that object, and Chapter 6.5 finds that in the strong interaction it is essentially all there is.
Fit a model with a categorical predictor of several levels and you have already met (2.6.29). Write the fitted value for group as . The parameters are not identified, because replacing
for any constant leaves every fitted value, every residual, every contrast and the entire likelihood exactly as they were. Software does not announce this. It silently imposes a constraint of its own, either dropping a reference level so that or using sum-to-zero coding so that , and then prints whichever coefficients that constraint produces. A coefficient reported without its coding scheme is not a quantity.
Every clause of that paragraph is this section. The potential is the parameter, (2.6.29) is the reparametrisation, and (2.6.30) is the statement that nothing measurable moves. The Lorenz condition (2.6.32) is a coding scheme, adopted in §4.2 because it makes the equations uncouple, exactly as sum-to-zero coding is adopted because it makes a table symmetric.
The test for what is real is the same test in both subjects. A quantity is physical if and only if the reparametrisation leaves it alone, which is estimability under another name. In both cases what survives are the contrasts.
Three differences, and the third is why Part VI exists.
- First, the redundancy here is a whole function's worth. Here is arbitrary at every point of spacetime, rather than being one number for the whole model, and §4.2's counting is what that costs.
- Second, the potentials are not pure bookkeeping. Chapter 0.7's solenoid already showed that carries something does not, and a design matrix has no counterpart to that.
- Third, and this is the one that matters, in statistics the redundancy is a nuisance and nothing whatever is lost by constraining it away. Here it turns out to be generative.
Take that third difference seriously for a moment. Demand that the freedom hold separately at each point rather than once for the whole universe, and a field is forced into existence to enforce it. Every force in the Standard Model is that demand made of a different group, and Chapter 6.3 is the one place in this book where a choice of reference category writes down a law of nature.
A whole function's worth of freedom sits inside the potentials and no experiment can see it. Add the four-dimensional gradient of anything smooth and both fields come out unaltered, for the reason that has now done this work three times, that two derivatives are indifferent to their order. Since the fields are what push charges about, potentials differing that way describe the same world. Freedom is an obstacle when you want to solve for something, so it gets spent, and one condition spends enough.
The condition chosen is itself a contraction, so imposing it in one frame imposes it in all, which the other common choice cannot claim. With it in force the four components stop talking to one another and each satisfies a wave equation with its own source. Remove the sources and what remains is the equation that opened this part, obtained in three lines because the bookkeeping was done first instead of last.
That dissolves the question which broke nineteenth-century physics. The operator in that equation is a contraction, so it reads the same for everybody, and the speed enters it only through the array of signs, which is to say through the geometry rather than through any substance filling space. Maxwell's equations do single out a speed. They never singled out a frame, and only the assumption of one universal clock made those two look like the same claim.
5 · and mix
If and really are the components of one tensor, then a boost must shuffle them into each other, and Chapter 2.4 tells us exactly how. For a tensor the law is
In matrix notation, with the array of (2.6.15), that reads . There is one acting on the rows and one acting on the columns, which is what two upper indices buys you. To make it concrete we take the standard boost of Chapter 2.4 §2.1, with moving at along :
Carrying out the two matrix multiplications, which the grind box does entry by entry, and then reading the primed fields off the result, gives the transformation rules in full.
Splitting each field into the component along the boost () and the two perpendicular to it ():
Or, compactly, with the boost velocity:
Three things are worth noticing immediately.
The components along the boost do not change. That is the opposite of what happens to a four-vector, whose parallel component is the one that mixes with time. The reason is visible in (2.6.15). Here sits in the slot, so both of its indices lie in the block the boost acts on, and it acquires two factors of that undo each other. Meanwhile sits in the slot, which the boost does not touch at all.
The combination is already familiar. It is the Lorentz force per unit charge. Section 8 shows that is not a coincidence.
If in one frame, in another. The one exception is when happens to be parallel to . So a pure electric field is not a frame-independent notion at all. Magnetism is what an electric field looks like from a moving frame, and §6 makes that quantitative in the one case where you can check it against a laboratory measurement.
Grind box — the boost of , entry by entry
The transformation is . Since is symmetric here, , and it acts nontrivially only on the – block. Write as the block on indices and the identity on .
Entries with both indices in . Only is nonzero there. The block of is , which is times the antisymmetric symbol. For any matrix and the antisymmetric we have , a two-dimensional special case of the determinant identity in Chapter 2.4 §8.1. Here . Hence
Entries with both indices in . is the identity there, so , i.e. .
Mixed entries. These carry one factor of the block and one identity. For :
Substitute and :
For , with and :
For , now the first index sits in the block as a row:
and since , . Identically, gives .
Six components, six rules, no others. Every one of them was checked symbolically against before this box was written.
A consistency check worth doing. Apply the rules twice, once with and once with , and you must get back what you started with. For : ✓. The that makes it work is the same identity that made above.
It does not mean you can turn an electric field into a magnetic field by running. The invariants of §7 forbid that in general. If in one frame it is nonzero in every frame, so a field configuration with both fields present and non-orthogonal has no frame at all in which either one vanishes.
What it does mean is more precise and weaker. and are components of a single object, and the split between them is observer-dependent in exactly the way the split of a four-vector into time and space parts is. Two observers disagree about how much of is electric in the same sense that they disagree about how much of is energy. Neither of them is confused. The word "electric" names a slice, and they are slicing differently. What they do agree about is itself, and the two scalars §7 builds from it.
If the two fields are entries of one object, changing frames must shuffle them into each other, and the rule needs nothing beyond the transformation law in hand. The entries along the direction of motion come through untouched, the reverse of what happens to a four-part vector, whose parallel component is the one mixing with time; a boost acting on both labels of one entry undoes itself. Across the motion the fields trade, and the combination appearing in the trade is the familiar force per unit charge.
The immediate consequence is that a purely electric field is not a notion anybody can defend as absolute. Let one observer find no magnetism anywhere and another moving past will find some. Magnetism is what an electric field looks like from a moving frame, a slogan about to be turned into arithmetic checkable against a laboratory bench.
What the slogan does not license is the belief that either field can always be transformed away. Some configurations refuse, and two agreed numbers built by contracting every label decide which. The exact statement is weaker and better: the split between electric and magnetic is observer-dependent in precisely the way the split of a displacement into time and space is. Two observers disagree about how much of the object is electric in the same sense that they disagree about how much of a momentum is energy.
6 · The wire — magnetism as electrostatics in disguise
This is the argument that makes the chapter's title literal, and it can be checked against a current balance on a laboratory bench.
6.1 · The setup, in the lab
Take a long straight wire along the -axis. We model it as Chapter 1.1's grind boxes model everything, with the crudest structure that still has the right physics. That means a rigid lattice of positive ions at rest, and a gas of conduction electrons drifting through it.
- Lattice: at rest in the lab, linear charge density .
- Electrons: drifting with velocity , linear charge density as measured in the lab.
The two densities are equal and opposite in the lab, so the wire is electrically neutral there. That is an experimental fact about wires rather than an assumption we are making. A current-carrying wire does not attract a stationary pith ball. With the densities set, the conventional current is
flowing in the direction, because negative charge moving one way is positive current the other.
Now put a test charge at perpendicular distance from the wire, at position , moving parallel to the wire with velocity .
In the lab the wire is neutral, so and there is no electric force whatever. That leaves the magnetic force to account for everything. The magnetic field of a long straight wire at distance is . In our geometry and , so the field and the force on the test charge come out as
The force is attractive. It points from the charge toward the wire, since . The force is purely magnetic and there is no ambiguity about that.
6.2 · The same situation, from the test charge's frame
Now board the test charge. Transform to the frame moving at . In the test charge is at rest, so it cannot feel a magnetic force at all. The Lorentz force on a stationary charge has no magnetic term, whatever may be.
And yet the charge is certainly still accelerating toward the wire. Whether it hits the wire is not a matter of opinion. So some other force must be doing the work, and there is only one candidate.
Our next step is to compute the two charge densities in . This is the crux of the whole argument, so we go slowly.
The lattice. It is at rest in the lab with density , so its proper density, meaning the density in its own rest frame, is . In it moves at speed , so the lattice is length-contracted by and the same charge occupies less length:
The electrons. These were already moving in the lab, so we must first back out their proper density. In the lab they move at and have density . Running the logic of (2.6.41) backwards, their proper density is with . To contract that proper density into we need their speed in , which the velocity-addition law of Chapter 2.2 §5 supplies:
What the contraction actually needs is not that velocity but the Lorentz factor built from it, so substitute the last line into and simplify. The grind box does the algebra. The result is remarkably tidy:
Now we have everything the electrons need. Their density in is their proper density multiplied by the contraction factor we just computed:
The two densities have been contracted by different factors. Both of them picked up . The electrons picked up an extra on top of that, because they were already moving before we boosted. Adding the two densities gives the net charge on the wire as sees it:
The wire is not neutral in this frame. It carries a net negative charge per unit length. The test charge is positive, so it is attracted, and that is the direction we already know the force to have. To leading order in we may set , which gives .
6.3 · The force, and the comparison
The field of an infinite line charge at perpendicular distance is radial, with magnitude . Distances perpendicular to the boost are unchanged, so . The test charge is stationary in , so the force on it is purely electric:
The last step there used . Now set that beside the lab result (2.6.40), which is what we came here to compare it with:
That factor is exactly the factor by which a transverse force is expected to change between the two frames. Here is why, in three lines.
Momentum transverse to the boost is unchanged, , because is a four-vector and the boost matrix of Chapter 2.4 §2.1 acts as the identity on the and rows. Time is not unchanged. For a particle instantaneously at rest in we have , so the time transformation gives . Dividing the first fact by the second,
which is (2.6.47) rearranged. The two calculations agree identically, not approximately.
One observer says this. The wire is neutral, the charge is moving, there is a magnetic field, and the force is .
The other says this. There is no magnetic force here, because nothing is moving. The wire carries a net negative charge, and what you are watching is Coulomb attraction.
They are both right, they compute the same trajectory, and neither description is more fundamental than the other. What one observer calls magnetism, another calls electrostatics acting on a wire that is not neutral.
So magnetism is a relativistic correction to Coulomb's law, the piece of it. The only reason it is not a fantastically small effect is that the zeroth-order term has been cancelled to fantastic precision by the neutrality of matter.
6.4 · How small is , and why you can still pick up a nail
Let's put numbers in, because they are startling. Copper has one conduction electron per atom, density and molar mass , so the number density of mobile electrons is
In a wire of cross-section carrying , the drift speed follows from :
That is , slower than a growing fingernail. Expressed as a fraction of the speed of light, which is the form the physics above cares about, it is
So magnetism is an effect of relative order , and yet electromagnets lift cars. How can both of those be true?
The resolution is in (2.6.45). The surviving effect is not on its own. It is the difference between two enormous numbers, and the enormous number is itself. In that same wire,
thirteen thousand coulombs of mobile charge per metre, exactly cancelled by thirteen thousand coulombs of lattice charge. A tiny fractional imbalance in a colossal cancellation is still a substantial charge. Worked example 1 does the arithmetic all the way to a force.
Grind box — the composition identity
Let a particle move at velocity along in frame , and let move at . The velocity in is . Compute :
Expand the numerator, writing , and factoring out :
The cross terms cancel exactly, and that is the only step in this box with any content. Putting that factorised numerator back and taking the reciprocal square root,
For the electrons, , and the minus sign in the bracket becomes a plus: , which is (2.6.43). ✓
Bonus. Read the same identity in reverse and it says that is the -component of the boosted four-velocity. That is to say, it is written out. Nothing new was needed here. Velocity addition is the transformation of , as Chapter 2.5 remarked.
And the same identity again. Notice that the spacings drawn in the figure obey , which is this identity divided by itself. The two rows of charges compress by different amounts because they enter the boost with different 's, and for no other reason.
"Magnetism is just relativity applied to electrostatics" is a true and beautiful slogan, and it is routinely overstated. Here is precisely what the argument above used.
- Coulomb's law, for the field of a static line charge. That is an input.
- Charge invariance. Without it, is false and the whole calculation collapses. This is an experimental fact, flagged in the ⚑ callout at the head of the chapter. It is not a consequence of relativity.
- The field concept, meaning the idea that the force is mediated by something local, so that "the field at the test charge" is a meaningful quantity to transform. A pure action-at-a-distance theory would need a different argument.
- The Lorentz force law in the lab, to have something to compare against.
What relativity supplies is the relation between the two descriptions, and it supplies it with no freedom at all. Given Coulomb, charge invariance and the field concept, the magnetic force is not optional and its magnitude is fixed.
That is a very strong statement. It means the constant is not independent of , which is Chapter 2.1's arriving from the other direction. It is still not the same as saying "magnetism can be derived from electrostatics alone", which is false.
The argument making the claim literal is checkable against a current balance. A current-carrying wire is electrically neutral in the laboratory, which is experiment and not assumption, since a live wire does not attract a pith ball, so a charge moving alongside feels a purely magnetic force. Now board that charge. In its own frame nothing moves, so no magnetic force is available, yet it still accelerates towards the wire, whether it strikes not being a matter of opinion.
Something else is doing the work, and there is one candidate. The wire holds two populations of charge, a lattice at rest and a gas of electrons drifting through it, contracting by different factors when frames change, the electrons having been moving already and the lattice not. The cancellation that made the wire neutral is spoiled, the wire carries a net charge, and the attraction is ordinary electrostatics. Both compute the same trajectory and neither description is deeper.
The numbers make it startling. Electrons drift along a copper wire at under a tenth of a millimetre per second, two and a half parts in ten million million of light speed, so magnetism is an effect of that size. Electromagnets lift cars because the unbalanced quantity is colossal: thirteen thousand coulombs of mobile charge per metre, cancelled to that precision by the lattice. A minute fractional imbalance in an enormous cancellation is still a substantial charge.
7 · The invariants
If and separately are observer-dependent, what do all observers agree about? Chapter 2.4 §5.1 answers that in general terms. What everybody agrees about is whatever you can build by contracting every index away. From a single antisymmetric rank-2 tensor there are exactly two such quantities, and this section constructs both of them.
7.1 · The first invariant
The simplest thing you can do with two copies of is contract every index of one against every index of the other, which leaves nothing free:
To see what that is in terms of and , split the double sum by index type. The diagonal contributes nothing, since the diagonal of is zero. The and entries contribute twice over, once each way round, and by (2.6.17) we have . So those terms give
That accounts for the electric entries. Now for the purely spatial ones. Lowering both indices leaves them alone, so , and their contribution is a sum over two epsilons:
using the contraction identity (sum on ), which is the previous identity with one more index contracted. Adding:
7.2 · The second invariant
The other contraction available uses the Levi-Civita symbol of Chapter 2.4 §8.1:
Total antisymmetry means the only surviving terms are those in which and are complementary pairs. There are three such pair-splittings, namely , and . Each of the three accounts for eight of the nonzero permutations: two orderings inside the first pair, two inside the second, and two choices of which pair comes first.
Flipping any one of those three things costs a sign in and a sign in one factor of , so the two signs cancel and all eight terms are equal. That reduces the whole sum to three representative terms with a factor of eight in front:
using , (one transposition) and (two transpositions), together with the entries of (2.6.15). All three terms come out positive:
That chapter proved and noted that for Lorentz transformations. So (2.6.59) is genuinely invariant under boosts and rotations () but changes sign under a spatial reflection. It is a pseudoscalar. That is not a defect. It is a fact about , which is the dot product of a vector with a pseudovector.
Chapter 2.4 also flagged, and we re-flag here, that for general coordinate changes is a tensor density and needs a factor of . Within Part II, where , that never bites.
One thing worth knowing in advance. A term proportional to (2.6.59) may legally be added to the Lagrangian of §9, and it is the notorious -term. Chapter 6.5 explains why the strong interaction's version of it is measured to be astonishingly close to zero.
7.3 · What the invariants forbid
We now have two numbers that everyone agrees on. They are strong constraints, and it is worth being exact about what each of them rules out.
A plane light wave stays a plane light wave. Chapter 2.1 §2 found that in an electromagnetic wave and . Then both invariants vanish:
Both are zero in every frame, so and for every observer, no matter how fast. You may Doppler-shift a light wave to any frequency you like and change its amplitude by any factor, and Chapter 2.5 §7.2 computed exactly how. What you can never do is make it into something other than a light wave.
In particular, you cannot bring light to rest. A stationary field configuration would have to have and independent of each other, and there is no frame in which the ratio is anything but .
When can you kill the magnetic field? Suppose there exists a frame with . Then in that frame the invariants read and . Since invariants are invariants, that requires
Those two conditions must hold in every frame, and in particular in yours. So the condition is checkable without leaving home. The fields must be perpendicular, and the electric one must dominate. These conditions are also sufficient, and Worked example 2 constructs the frame explicitly in the mirror-image case.
When can you kill the electric field? Run the identical argument with the roles swapped and the conditions come out as
Worked example 2 does this case in full and finds the frame velocity, which turns out to be the drift velocity every mass spectrometer relies on.
And if , neither field can be removed in any frame, ever. The two fields are then irreducibly both present, and the best you can do is find a frame where they are parallel.
The pair classifies field configurations into frame-independent types, exactly as the sign of classified pairs of events in Chapter 2.3 into timelike, spacelike and null.
| Invariants | Name | Simplest frame |
|---|---|---|
| , | electric | a frame with : pure electrostatics |
| , | magnetic | a frame with : pure magnetostatics |
| , | null | none simpler, a radiation field |
| generic | a frame with |
The null case is the boundary between the other two, and it is exactly the case that cannot be simplified. That happens in the same way that a null interval sits on the boundary of the light cone and cannot be transformed into either a pure time separation or a pure space separation. Light is the null case of electromagnetism in precisely the sense that a light ray is the null case of a worldline.
That table has a shape, and it is worth seeing it. Take and . That is a configuration on which a boost along closes in two dimensions, so the picture stays flat, and the four rows of the table become four regions of one plane.
Since each field separately depends on who is looking, the quantities worth having are those nobody can dispute, which means whatever survives contracting every label away. From one antisymmetric object there are exactly two such numbers: the difference between the squared magnitudes of the fields, and the amount by which they overlap. Everybody computes the same pair, whatever their motion.
The pair sorts field configurations into types not open to argument, as the sign of the interval sorted pairs of events into three kinds. Where the fields stand perpendicular and the electric one dominates, some observer finds no magnetism at all; where the magnetic dominates, some observer finds no electricity; and where the two are not perpendicular, neither can be removed by anybody ever, the best available simplification being a frame in which they lie parallel. The test is run without leaving home, because the quantities tested are agreed.
Light is the case where both numbers vanish, and since they vanish for everybody, no change of frame turns a light wave into anything else. Its frequency may be shifted to whatever you like and its amplitude scaled by any factor, and it can never be brought to rest, because no frame exists in which the two fields become independent. Light sits at the boundary between the electric and magnetic types exactly as a light ray sits on the boundary of the light cone.
8 · The Lorentz force, covariantly
Chapter 2.5 §6.1 established the shape any relativistic force law must have:
What can be for a charge in an electromagnetic field? Three requirements narrow it almost to nothing. It must be a four-vector, it must be linear in the charge, and it must be built from the only field object we have. With the particle's four-velocity, there is essentially one candidate:
Chapter 2.4's grind box on the quotient theorem already flagged this construction. Requiring to be a four-vector for every four-velocity is what certifies as a tensor in the first place. Our job now is to verify that (2.6.64) is the force law you already know.
8.1 · The spatial components
Lower the index on the four-velocity of Chapter 2.5 §1.3, which flips the sign of the spatial part:
Take and split the sum over into its time and space parts:
The left-hand side is , since (Chapter 2.5 §1.2). Cancel the common :
That is the Lorentz force law, the one Chapter 1.1 §3.3 had to quote without justification. Here it arrives as the spatial part of a four-vector equation. Notice also that here rather than , so the relativistic correction rides along for free.
8.2 · The time component is the power
Take . The term dies because :
And the left-hand side is , where is now the particle's energy of Chapter 2.5 §3 (an unfortunate clash of symbols that the whole subject lives with). Cancelling :
The magnetic field does no work. That familiar fact is not a separate observation here. It is a consequence of where sits in the matrix. The entries are purely electric, so only can appear in the energy equation at all.
And (2.6.69) is Chapter 2.5's work–energy theorem with the Lorentz force, since automatically.
8.3 · The constraint is automatic, and why
Chapter 2.5 §6.1 proved that any four-force must satisfy , because is constant. That is a strong condition, and it is not obvious that an arbitrary guess for would satisfy it. Check (2.6.64):
Why zero? Because is antisymmetric under while is symmetric, and a symmetric object fully contracted with an antisymmetric one vanishes.
Here is that argument once, in case it has not been made explicit before. Relabel the dummy indices in the sum, which changes nothing, and then use the two symmetry properties to put the labels back the way they were:
Let's look at what just happened. The constraint is a statement about the geometry of spacetime. It says a four-force can only push a four-velocity sideways, because four-velocities all have the same length. It came from Chapter 2.5 with no reference to electromagnetism whatsoever.
And it is satisfied by (2.6.64) because is antisymmetric. That fact came from , with no reference to spacetime geometry whatsoever. Two independent constraints, arriving from two unrelated directions, match exactly.
The physical content of the match is (2.6.69). The reason the magnetic force does no work is the reason a four-force is orthogonal to a four-velocity. Had carried any symmetric part, charges would gain energy from nothing.
Three requirements and a shortage of materials pin down the force on a charge. It must be a four-part object, proportional to the charge, and assembled out of the particle's four-velocity and the one field object available; essentially one candidate meets all three. Expanding it returns the force law the first part of this book had to quote without justification, with the repaired momentum riding along at no extra cost, while its time entry says the particle's energy changes at the rate the electric field does work.
That the magnetic field does no work is therefore not a separate fact to be remembered. It is a remark about where the magnetic entries sit, since they occupy the purely spatial part of the array and so cannot appear in the equation governing energy. The familiar statement and the arrangement of the object are one thing.
Then a check that had to work, and does. Any four-force must stand perpendicular to the four-velocity, a constraint out of the geometry of the interval with no reference to electricity. The candidate satisfies it because the field tensor reverses sign when its labels are swapped, a property out of the freedom in the potentials with no reference to geometry. Two constraints from unrelated directions, matching exactly. Had the field tensor carried any part that did not reverse sign, charges would gain energy from nothing.
9 · The action
Chapter 1.2 §8.1 listed eight actions and promised each would be constructed in its own chapter. Line three said electromagnetic field, Chapter 2.6. Here it is, with the source term and the SI constants restored:
Two terms, one line, and it contains every result of §3. The claim we have to verify is that feeding (2.6.72) to the field Euler–Lagrange equation of Chapter 1.2 §8 returns .
9.1 · Varying it
The dynamical variable here is , which is four fields, one for each value of . So the Euler–Lagrange equation carries a free index:
The second term is immediate. Only contains undifferentiated, so . The first term is where the work is, and the grind box does it slowly. The answer is
We now have both pieces the Euler–Lagrange equation asks for, so put them into (2.6.73) and see what field equation comes out:
That is (2.6.19). Gauss's law and Ampère–Maxwell, both of them, out of one scalar. The homogeneous pair needs no varying at all. It holds identically the moment you write in terms of , which was the point of §3.4.
Grind box — differentiating with respect to
This is pure index gymnastics and it is worth doing once with every step visible, because the same manipulation reappears in Chapters 5.2, 6.3 and 6.4 with more indices.
Step 1. The chain rule on a product. Write and note , so it is quadratic in the with constant coefficients. Hence
Step 2. The inner derivative. Since and the sixteen quantities are treated as independent variables,
The two deltas in each product are what turn "differentiate with respect to a particular component" into "select that component". The two terms are the two places appears in .
Step 3. Contract. Now feed Step 2 back into Step 1. The deltas rename indices:
where the last step used antisymmetry. There is the factor of that the in (2.6.72) was put there to cancel. Carrying the prefactor along,
which is (2.6.74). The source term contains no derivatives of and contributes nothing here. ✓
Why the sign of the whole Lagrangian is what it is. Expand (2.6.72) into three-vector language using (2.6.56):
That is kinetic minus potential. The electric term carries the time derivatives of and plays the role of , while the magnetic term carries the spatial derivatives and plays the role of . The overall sign of (2.6.72) is fixed by demanding that come out positive, which is Chapter 1.2 §5's discussion transplanted verbatim to a field. Get it wrong and the field has negative kinetic energy, which Chapter 5.2 shows is a catastrophe rather than a sign convention.
9.2 · Why this Lagrangian and not another
The variation above is a verification rather than a derivation. We wrote down (2.6.72) and checked it. So where did it come from? The honest answer is a short list of demands.
- Lorentz invariance. must be a scalar, or the action is not the same number for every observer and the theory has a preferred frame. That immediately restricts us to contractions: , , , , and products and derivatives of these.
- Gauge invariance. must be unchanged, up to a four-divergence, under (2.6.29). This kills outright, since that term is not gauge invariant. It would have been a photon mass term. The photon is massless because a mass term is not gauge invariant. The same demand permits , and only because charge is conserved, as §4.1 showed.
- Second-order equations of motion. should contain at most first derivatives of , so that Euler–Lagrange returns something second order and the initial-value problem is the usual one.
- Parity, if you want it. passes tests 1–3 but is a pseudoscalar, and moreover is a total derivative, so it does not affect the classical field equations at all.
What survives at lowest order is and , with two constants in front. One of those constants is fixed by matching Coulomb's law, and the other is absorbed into the normalisation of . The form is very nearly forced.
Chapter 6.4, Yang–Mills. Take (2.6.72), let carry an internal index and become matrix-valued, and define . Then write down , which is the same expression, entry seven in Chapter 1.2's table. The commutator is the only new ingredient, and it is what makes gluons carry colour charge and interact with each other. The strong and weak interactions are this chapter with a non-commuting symmetry group.
Chapter 5.11, effective field theory. The argument of §9.2 is the beginning of a general method. List the fields, list the symmetries, write down every term allowed, and order them by how many derivatives they have. The leading term is the theory, and the rest are corrections suppressed by powers of energy. Electromagnetism looks fundamental partly because the leading allowed term is so simple. Higher terms do exist. One of them, , is the Euler–Heisenberg term, which makes light scatter off light, and it is suppressed by four powers of the electron mass. That is why nobody noticed it for a century.
Chapter 5.2, quantisation. A Lagrangian is the input to a path integral. Once you have (2.6.72), quantising the electromagnetic field is a well-posed problem, and its answer is the photon. The vertex of quantum electrodynamics is the term (Chapter 5.8), which is the entire interaction between light and matter and which you have just written down.
The factors of , and cluttering this chapter are SI, and they carry real information here. They are what let you compare a derived force against a laboratory measurement in newtons, which §6 and the worked examples do.
From Part V onward the book switches to natural units, , in the Heaviside–Lorentz convention where the is moved out of Coulomb's law and . In those units (2.6.72) reads , which is how you will see it written everywhere else, and every equation in this chapter loses its decoration.
Nothing physical changes when you do that, because the constants can always be restored by dimensional analysis. We keep them here because Part II is about measurable consequences, and dropping while still learning what means would be perverse.
Compress everything so far and one line is left: one number formed by contracting the field tensor with itself, and one term coupling the current to the potential. Hand that line to the machinery of the first part and out comes the equation saying charges and currents make fields, constants and all. The other half needs no varying, holding the instant the field is written in terms of potentials.
How nearly the form is forced is what to carry away. Demand a number every observer computes alike and only contractions are admitted. Demand that the redundancy in the potentials stay invisible and one candidate is killed outright, the one that would have given the force's carrier a mass, which is the whole reason light has none. Demand no more than first derivatives and almost nothing is left. Two terms survive, their constants fixed by matching the static force and by a choice of normalisation.
The method is worth more than the result. List the ingredients, list the symmetries, write down every term they permit and order these by how many derivatives they carry, whereupon the leading one is the theory and the rest are corrections. Let the potential carry an internal label and become a matrix, add the one ingredient non-commuting labels supply, and the same line written again is the strong interaction, whose carriers act on one another because that ingredient says they must.
10 · Field momentum, and Newton's third law repaired
Chapter 1.1 §3.3 put two moving charges alone in the universe and showed that the total force on them does not vanish. Momentum appeared from nowhere at a computable rate.
That chapter offered two escapes. Either momentum conservation is false, or momentum is stored somewhere that is not a particle. It took the second escape, named the electromagnetic field as the only candidate, and left an IOU. This section pays it, and the payment is exact.
10.1 · The energy–momentum tensor, from Noether
Chapter 1.4 proved that every continuous symmetry yields a conserved current. The symmetry we have not yet used is the most basic one there is. The laws are the same here as there, and now as then, which is invariance under spacetime translations . Since has four components, we get four conserved currents, one per direction, and those four assemble into a rank-2 object.
Run Chapter 1.4 §6's machinery. Under a translation the field changes by and the Lagrangian density by That last expression is a four-divergence, which is exactly the loosened invariance condition of Chapter 1.4 §1.2. So the symmetry qualifies, and the Noether current for the -th translation is
using (2.6.74) and setting (we are asking what the free field carries). It satisfies by Noether's theorem.
And it is wrong in two ways. It is not symmetric in . It is also not gauge invariant, because it contains bare, so different gauges would assign different energy densities to the same physical field.
Both defects are fixed by the freedom that Chapter 1.4's grind box identified. Adding , with antisymmetric in its first two indices, changes no conserved charge. So we are free to choose a that cleans both problems up at once. Take
Its divergence is, using in the source-free region,
Add it to (2.6.76) and the two terms combine into an :
One last cosmetic step puts the indices in their conventional places. Since , the minus sign in front can be absorbed, and we are left with the object we were after:
Now it is manifestly gauge invariant, since only appears in it. It is also symmetric, and that is worth checking rather than believing. The first term is . Swap and relabel the dummies , and you get it back, with the two minus signs from antisymmetry cancelling each other.
That symmetry is not cosmetic. Chapter 3.6 puts on the right-hand side of Einstein's equation, and the left-hand side of that equation is symmetric by construction.
10.2 · What its components are
. Only spatial contributes to , and lowering the index costs a minus:
using and . With and (2.6.56) for the second term,
The last step used . That is the energy density of the electromagnetic field. You have probably met that expression as a stated fact. Here it is derived, as the conserved Noether charge density of time translation, which is what the word "energy" has meant since Chapter 1.4 §3.1.
. Now , so only the first term survives:
Only contributes there, since . The metric supplies one minus sign and the antisymmetry supplies the other, so the two cancel. Putting that into the definition of ,
is the Poynting vector. Now recall what is by construction. It is the density of , with the -index labelling the conserved density and the labelling which component of momentum we mean. Reading the spatial off, the momentum density of the field is
Which is precisely the expression Chapter 1.1 promised, arriving here as a component of a Noether current rather than as an assertion.
10.3 · The conservation law, with sources
Now restore the charges, because it is the exchange between field and matter that we care about. Take the divergence of (2.6.80):
Term (1) is the field equation itself, , so it is . Terms (2) and (3) cancel identically, by the Bianchi identity, and the grind box does that cancellation in full. Dividing through by ,
The right-hand side is not zero, and it should not be. The field is exchanging energy and momentum with the charges. Let's see what that exchange term actually is. For , using and ,
That is the Lorentz force per unit volume, which is (2.6.67) written for a continuous distribution. Meanwhile the left-hand side, with and , is . Putting the two sides together,
Here is the spatial block , called the Maxwell stress tensor.
We want a statement about totals rather than densities, so integrate that over all space and take the three terms one at a time. The middle term becomes a surface integral at infinity by the divergence theorem. For a system of charges confined to a finite region the fields fall off at least as fast as , so while the area grows only as , and the surface term vanishes. The last term integrates to the total force on the matter, which is . What is left is
The component of (2.6.87) gives, by the identical route, . That is Poynting's theorem. It says the field energy in a region falls by the flux of out of it, plus the work done on the charges inside. The two statements are one four-vector equation.
Grind box — why terms (2) and (3) cancel
Write term (2) with cleaner dummy names, and :
Because is antisymmetric, only the antisymmetric part of whatever it multiplies survives. So we may replace the second factor by half of it minus its swap:
Now bring in the Bianchi identity (2.6.25), with the third index raised. Raising is legitimate here because we raise it on every term with the same :
Use on the middle term and rearrange:
Substitute:
Note what did the work there. It was the homogeneous half of Maxwell, the half §3.4 called bookkeeping. It is not decorative after all. Without it would not be conserved, and energy and momentum would not balance.
10.4 · Back to Chapter 1.1
Here is the configuration again, unchanged. Charge sits at the origin moving with . Charge sits at , directly ahead of it, moving with . Both are in uniform motion at the instant considered.
What 1.1 found. Charge 1 produces no magnetic field directly ahead of itself, so charge 2 feels no magnetic force. Charge 2 does produce one at charge 1's location, so charge 1 does feel a magnetic force. The magnetic residue was
What 1.1 was entitled to drop, and we are not. That chapter said the electric forces "are equal and opposite along the line joining them, so they cancel". That is true to zeroth order in and false at order , which is the very order the magnetic residue lives at.
Chapter 1.1 was making a qualitative point, so the omission cost it nothing. We are about to make a quantitative one, so we need the correction. It is three lines with §5's transformation rules.
Charge 1's field at charge 2. Boost to charge 1's rest frame . The field point has , so in the charge sits at a Coulomb distance and the field there is , pointing along . That direction is parallel to the boost, and by §5 the parallel component is unchanged on transforming back. So
Charge 2's field at charge 1. Boost to charge 2's rest frame, which moves along . The displacement from charge 2 to the origin, , is perpendicular to that boost, so it is unchanged and the Coulomb distance is . The field there is . That direction is also perpendicular to the boost, so by §5 the field gets multiplied by on transforming to the lab:
The two electric forces are and , and they no longer cancel. One is weakened by and the other is strengthened by , because charge 1 is moving along the line joining them while charge 2 is moving across it. Adding them, and expanding to order with and , gives the electric residue:
The last step used . So the third law fails in the direction too, and by a larger amount than it fails in . Adding the two residues, the total rate at which mechanical momentum is being created is
10.5 · And where it went
Now we compute the field momentum, to see whether it is disappearing at the same rate. Two simplifications are available at the order we need. Each charge's magnetic field is , which §5 gives exactly, by boosting a rest frame where . And each may be taken as Coulomb, since is already first order in while corrections to are second. With the fields written as sums of the two charges' contributions, (2.6.90) expands to
The two self terms depend only on , which is constant. So however large they are, and they are in fact infinite, as the caution below explains, they contribute nothing to . That leaves the cross terms. Expanding each of those with ,
Two integrals are needed to finish this, and both are elementary. The grind box evaluates them. Writing , and , they are
Now substitute those two results back. Write . The two tensor terms combine into , using the symmetry of , and the prefactors collapse with :
That is the general answer. For the configuration at hand, at the instant , we have , and , so
There is momentum in the empty space between the charges.
What we actually want is the rate at which that momentum changes, so differentiate. The separation vector evolves as , since charge 1 moves in and charge 2 moves in . At that gives and . Grinding through the product rule, as the grind box does,
Now set that beside (2.6.95), which is the rate at which the particles were gaining momentum from nowhere. The two match term by term, including the awkward that Chapter 1.1 never saw. So the books balance exactly:
Chapter 1.1 §3.3 wrote down a momentum non-conservation of and offered you a choice. Either abandon momentum conservation, or accept that the field is a physical object with momentum of its own.
The second was right, and now it is computed. The missing momentum is (2.6.100), it lives in the field, and its rate of change is exactly minus the rate at which the particles are gaining momentum from nowhere. That holds in both components, including the that only shows up if you are honest about the electric forces at order .
Notice how the repair works. Newton's third law presumed two things: that all the momentum in the universe is carried by particles, and that particles exchange it instantaneously. Relativity forbids the second of those, and once it goes the first cannot survive either. If charge 1 moves now and charge 2 will not know for seconds, the momentum has to be somewhere during the interval.
So the third law is not a fundamental principle that electromagnetism violates. It is a low-velocity approximation to a four-vector conservation law, and it fails exactly where the approximation does.
The conceptual bill is large and worth paying explicitly. A field can no longer be read as a bookkeeping device, a table of where the force would be if you put a test charge there. Something that stores energy, stores momentum, exerts stresses, and can carry momentum across a room during the interval when neither particle has it, is not a table. It is a physical system with its own degrees of freedom and its own dynamics.
That is why Part V quantises it, and why the resulting quanta, photons, are as real as electrons rather than being a way of talking about electrons.
The self terms are infinite. diverges at the location of a point charge, as does the self-energy . We dodged it above by noting that these terms are constant when is, so they drop out of the rate.
That dodge is legitimate here and evasive in general. The divergence is the classical electron self-energy problem. It is genuinely unsolved in classical physics, and it is the ancestor of the renormalisation programme of Chapter 5.11. Point charges are an idealisation, and they bite.
What we computed and what we did not. The general theorem (2.6.90) is exact, derived from with no approximation. The two-charge check is narrower. It keeps the velocity-dependent terms to order and drops the acceleration fields, the radiation ones. Their contribution is balanced separately against the terms in (2.6.99), which is a second and independent piece of bookkeeping, controlled by a different small parameter: the ratio of the charges' Coulomb energy to their kinetic energy.
Nothing was assumed in order to make (2.6.102) come out. Every coefficient was computed first and compared afterwards.
Grind box — the two field integrals of (2.6.98)
Both are done with the same two moves. Integrate by parts, then use Gauss's law with . Surface terms vanish because the fields fall off as and the integrands as or faster.
The scalar one. Write and integrate by parts:
Multiplied by , that is the familiar interaction energy , which is a useful check that the normalisation is right.
The tensor one. Call it . Three facts pin it down.
(i) It is symmetric. Invert through the midpoint, . Under that map , so and , and the Jacobian is . Hence .
(ii) Its form. The only vectors in the problem are and the coordinate axes, so by rotational symmetry about and (i),
with and constants. The is there because has the dimensions of the scalar integral, which scales as .
(iii) Two equations. The trace is the scalar integral just computed:
For the second, differentiate with respect to and contract. Since depends on and , we have , so
Evaluate the same derivative on the parametrised form, using and :
Comparing, . Solve the two equations: and . Hence
which is (2.6.98). Note the structure. Here is the projector onto directions perpendicular to the line joining the charges, so , and the longitudinal part of the integral vanishes identically. This was also checked by direct numerical integration before being written down.
Grind box — differentiating (2.6.99)
Write and , a constant. Then
The kinematics at . , so and
Note , as it must be for a unit vector.
First term.
Second term. Three pieces by the product rule:
Total.
which is (2.6.101) after pulling out a factor of : . ✓ Every term of this box, and of (2.6.94), was verified symbolically. The sum (2.6.102) is zero identically rather than numerically.
You have built a symmetric, conserved, traceless rank-2 tensor whose component is energy density and whose components are momentum density. Chapter 3.6 will take Einstein's field equation
and put this object on the right-hand side. That is what "energy gravitates" means technically. It is not mass but that sources spacetime curvature, so the electromagnetic field itself gravitates, with a strength you can now compute. That includes light, a magnetic field, and the energy stored in a capacitor.
It is also why the equation had to be symmetric in , which is why §10.1's improvement term was not optional.
And the tracelessness , which you can verify in one line from (2.6.80), is the statement that the photon is massless. It reappears in Chapter 5.11 as the conformal symmetry of classical electromagnetism, a symmetry that quantum corrections break.
The oldest debt in the book is settled by computation rather than assertion. Two charges in uniform motion, alone in the universe, were seen in the first part to push on each other unequally, momentum appearing from nowhere. Work out what the empty space between them holds and differentiate: the field loses momentum at exactly the rate the particles gain it, awkward factor of three halves and all.
The diagnosis matters more than the arithmetic. The third law assumed all momentum belongs to particles, handed over the instant either moves. Relativity forbids the second clause: once one charge moves and the other cannot learn of it until light crosses the gap, the momentum must be somewhere meanwhile. What holds it there is no table of where a force would be but a physical system with its own degrees of freedom, which is why a later part must quantise it.
Part II has done its work. You hold a geometry where one speed is the same for everybody with nothing to measure it against, a language making frame-independence visible in an equation's shape, a mechanics of four-entry objects, which is where energy came from, and one field with electricity and magnetism as its slices. One thing has not moved: gravity is still a force reaching across empty space, arriving the moment it is sent, and nothing in this part permits an influence with no delay.
11 · Worked examples
A copper wire of cross-section carries . A proton () travels parallel to it at , at a perpendicular distance . Compute the force on it twice, first magnetically in the lab and then electrostatically in its own rest frame. Then say how big the charge imbalance is.
The lab. The field of the wire at is
about a third of the Earth's field. The magnetic force is
directed toward the wire. There is no electric force, because the wire is neutral in this frame.
The wire's insides. From (2.6.49), , so the mobile charge per metre is
and the drift speed is , i.e. .
The proton's frame. Here , so . That is utterly negligible, and we keep it anyway, because the whole effect we are chasing is of this size. From (2.6.45),
The electric field of that line charge at is
pointing toward the wire, and the force on the proton is
Compare. ✓, exactly as (2.6.47) requires. Notice also the clean intermediate result to this accuracy. That is with , which is the field transformation rule of §5 arriving by a completely independent route.
How big is the imbalance? Count it in electrons per metre:
That is a fractional imbalance of , or one part in twelve million billion. And that is the entire magnetic force on this proton, a discrepancy in the seventeenth significant figure of a cancellation. If the positive and negative charge in a wire did not cancel to far better than that, the residual electrostatic force would swamp every magnetic effect ever measured, and nobody would have discovered magnetism in a laboratory at all.
In the lab, and are uniform, crossed, and satisfy . Find a frame in which the electric field vanishes, describe the motion of a charge in that frame, and identify the frame velocity in a form that does not refer to the axes.
Check first that it is possible. The invariants of §7 are and , which is precisely condition (2.6.62). So a frame with is permitted. Now construct it.
The construction. Boost along at speed . From §5, and automatically. The only surviving component is
and since is never zero, requires
That speed is subluminal precisely when . There is the invariant condition again, arriving this time as a kinematic constraint rather than as a separate assumption. Note in passing what happens if . The required then exceeds , so there is no such frame, and the invariant told you so in advance without any construction at all.
Axis-free form. With and ,
So the frame in question moves with the drift velocity
That formula makes no reference to any choice of axes, and it is valid whenever and .
What the motion looks like. In the primed frame there is only a magnetic field, of magnitude
using and . The last expression is the square root of the first invariant divided by two, as it has to be. With the invariant (2.6.56) reads , and that must agree with .
A charge in a pure magnetic field feels a force always perpendicular to its velocity and does no work, by (2.6.69). So it moves in a circle at constant speed, which is the cyclotron gyration.
Back in the lab, that circle is being carried along at , so the trajectory is a cycloid-like drift. Here is the point to take away. The drift velocity does not depend on the charge, the sign of the charge, or the mass. Every particle drifts at .
Why every mass spectrometer contains one. Turn the argument around. A particle that enters crossed fields moving at exactly is at rest in the primed frame, and a charge at rest in a pure magnetic field feels no force at all. So it passes through undeflected, while anything faster or slower is bent aside.
That is a velocity selector: a slit, two plates, a magnet, and one equation. It selects regardless of mass and charge, which is exactly what you want in front of a mass analyser that will then separate by . The same drift governs charged particles in the magnetosphere, the confinement of tokamak plasmas, and the Hall effect.
12 · Your turn
Problem 1 · the field of a charge in uniform motion
A charge moves with constant velocity , passing the origin at . By boosting the Coulomb field from its rest frame, show that at the lab field at position is
where is the angle between and . Show the field is still radial from the present position, that it is weakened by directly ahead and strengthened by broadside, and that for the magnetic field reduces to the Biot–Savart expression Chapter 1.1 had to quote. Comment on what the field looks like as .
Solution
Set-up. Let be the charge's rest frame, moving at . There and .
Coordinates. The lab event is . By the Lorentz transformation, , , , so
Fields. Transform back from to the lab. Parallel components are unchanged and perpendicular ones acquire a (with the terms drop):
All three carry the same factor , so
The field points radially away from the charge's present position. That is a small miracle, given that information travels at and the charge has moved since the field "left". It works only for uniform motion. Accelerate the charge and the radial structure breaks, which is what radiation is.
The angular factor. Write and :
using and . Then and
The two special directions. Ahead or behind (): the factor is , so the field is weaker than Coulomb by . Broadside (): the factor is , so the field is stronger by . These are exactly (2.6.92) and (2.6.93), which §10 needed and derived by the same argument in two special cases.
The magnetic field. Transforming from a frame where it vanishes gives and , which in view of the expressions above is exactly . That is exact rather than a leading order statement. For we may replace by its Coulomb value:
using . That is the Biot–Savart law for a point charge, which was the second of the two results Chapter 1.1 §3.3 had to quote. It is now derived, so both of that chapter's ⚑ items are discharged.
The pancake. As the forward field dies like while the broadside field grows like . The angular width over which the field is appreciable is set by , that is, by . So the field collapses into a disc perpendicular to the motion, of angular thickness , travelling with the charge. At the LHC, for protons, so each beam's field is squashed into a pancake about radians thick. That is why an ultrarelativistic charged particle acts on a target like a pulse rather than a slowly growing force, and it is the starting point for the equivalent-photon approximation used to describe ultraperipheral collisions in Part V.
Problem 2 · a plane wave has nothing to give up
For a plane electromagnetic wave in vacuum, and . Show that both invariants of §7 vanish. Then use that to prove three things. No boost can eliminate either field. No boost can make the fields non-perpendicular. And no boost can bring the wave to rest. What can a boost do to a light wave?
Solution
The invariants. , and since the fields are perpendicular. Both are zero. By the invariance theorem they are zero in every frame.
(i) Neither field can be eliminated. Suppose some frame had . Then in that frame , which must equal zero, forcing as well. So the only way to lose the magnetic field is to lose the whole wave. And you cannot do that, since a tensor that vanishes in one frame vanishes in all of them (Chapter 2.4 §6), so a wave that exists at all exists for everyone. The same argument runs identically for .
(ii) The fields stay perpendicular and stay in ratio. in every frame because the pseudoscalar is invariant, and in every frame because the scalar is. A light wave looks like a light wave to everybody.
(iii) It cannot be brought to rest. "At rest" would mean a static field configuration, for which and are independent and generically . Here is the sharper version. The wave's four-wavevector satisfies (Chapter 2.5 §7.1), and a null four-vector cannot be boosted to a purely timelike one. That is the same statement as " is preserved" in Chapter 2.3. The vanishing of both field invariants is the field-theoretic face of the same fact, which is that null is a Lorentz-invariant category.
What a boost can do. Change the amplitude and the frequency, together and in the same ratio. From §5, boosting along the propagation direction multiplies both and by the Doppler factor , exactly the factor Chapter 2.5 §7.2 derived for . So a light wave can be made as weak and as red as you like, approaching zero without ever reaching it. It can also be aimed differently, since a boost transverse to changes the propagation direction, which is aberration. What survives every boost is the wave's identity as a wave.
The physical moral. The vanishing of both invariants is what makes light structurally different from every static field. It is also why the photon is massless. The null condition on is (Chapter 2.5 §4.3), and that means .
Problem 3 · charge conservation as an output, not an input
Show that identically, for any antisymmetric . Deduce that follows from the field equation rather than being an extra assumption. Then explain what this says about a hypothetical universe in which charge is not conserved, and connect it to Chapter 0.7's Problem 4 on the displacement current.
Solution
The identity. The object is symmetric under , because mixed partials commute, and is antisymmetric. Contracting a symmetric object with an antisymmetric one over both indices gives zero, which is the argument of (2.6.71) repeated. Explicitly, relabel the two dummy indices:
The consequence. Take of the field equation (2.6.19):
Charge conservation is a theorem rather than a postulate. Section 1 assumed it in order to build , and the theory then hands it straight back.
What it forbids. This is stronger than it looks. It says you cannot write down a consistent electrodynamics with a non-conserved source. If someone hands you a with , the equation has no solutions at all. Not "solutions with strange behaviour". None. A universe in which charge could appear from nowhere could not have Maxwell's equations, in any frame, even approximately. The rigidity comes entirely from the antisymmetry of , which came from gauge invariance (the callout at the end of §2), which is why Chapter 6.3 can say that gauge symmetry and charge conservation are two descriptions of one fact.
The displacement current. Chapter 0.7's Problem 4 asked you to show that is inconsistent with charge conservation, because taking the divergence gives , which contradicts (2.6.1) whenever . It also showed that the unique minimal repair is Maxwell's .
This problem is the four-dimensional version of that argument, and it is now visibly the same argument. The three-dimensional identity that forced the repair is the spatial part of . In the covariant formulation the displacement current is not a repair. It was never absent.
Problem 4 · two parallel wires, both ways
Two long parallel wires a distance apart carry currents and in the same direction. Model each of them as in §6, with lattice at rest and electrons drifting at , so that . (a) Compute the force per unit length magnetically, in the lab. (b) Now boost to the rest frame of wire 2's conduction electrons and recompute it electrostatically, being careful about how "per unit length" transforms. (c) Check that wire 2's lattice feels no net force in that frame either, as it must not. (d) Put in and .
Solution
(a) The lab. Wire 1 produces at wire 2 (taking wire 2 to lie in the direction from wire 1). Only wire 2's electrons move, with charge per unit length and velocity :
Attractive, magnitude . The lattice of wire 2 is at rest and wire 1 is neutral, so nothing else contributes.
(b) The co-moving frame. Boost at (so wire 2's electrons are at rest). By (2.6.45) with , wire 1 acquires
That is a positive charge density, and wire 2's electrons are negative, so the force is attractive, agreeing with (a) in direction.
Wire 2's electrons are now at rest, so their density is their proper density, (they were contracted by in the lab). The electric field from wire 1 at distance is , so the force per unit length measured in this frame is
The "per unit length" bookkeeping. That looks like the same number as (a), which would be wrong, because transverse forces are not equal in the two frames. They differ by , as (2.6.48) says. The resolution is that lengths differ too. Take a lab segment of length containing electrons. Those same electrons are at rest in the primed frame, so they occupy there. Hence the total force on them is
Two factors of appeared there, one from length contraction and one from force transformation, and they cancel. That is exactly why the force per unit length happens to come out numerically the same in both frames. The coincidence is worth noticing precisely so that you do not mistake it for a proof. The frame-covariant statement is the total force on a given set of charges.
(c) The lattice of wire 2. In the primed frame it moves at , has density , and sits in both an electric and a magnetic field. From §5, since in the lab. So
They cancel exactly. The lattice feels nothing, as it must, since in the lab it feels nothing and "zero force" is a frame-independent statement.
Notice how the cancellation works. In this frame the lattice is a current, and the repulsion from the now-charged wire 1 is precisely balanced by the magnetic attraction between the two currents. Neither description is more true than the other.
(d) Numbers.
Twenty micronewtons per metre. That is small, and it is measurable with a torsion balance. Until 2019 this configuration defined the ampere, as the current which, in two infinite parallel wires one metre apart, produces a force of . An entire base unit of the SI was defined by a relativistic correction to Coulomb's law of relative order .
Six objects, and Part II is finished.
The four-current , built by requiring to reproduce Chapter 0.7's continuity equation. After that, charge conservation is one contracted index, and therefore true in every frame.
The field tensor , whose six independent components are in the first row and in the spatial block. Those are the six that Chapter 2.4 §7.2 counted and promised.
Maxwell's four equations as two. Here expands to Gauss and Ampère–Maxwell, and expands to and Faraday. The second of the two is not physics at all. It is six terms cancelling by Clairaut, which is Chapter 0.7's and fused into one identity.
Gauge freedom , which is Chapter 1.2's total-derivative freedom in field form, spent on the Lorenz gauge to give .
Two invariants, and , which classify field configurations the way the sign of classifies intervals.
The Lagrangian , entry three in Chapter 1.2's table, which on variation returns the field equations and which Chapter 6.4 will copy verbatim with a non-abelian .
And the energy–momentum tensor, whose is and whose is , giving field momentum density .
And two long-standing debts, both settled by name.
Chapter 2.1 posed a contradiction between Maxwell and Galileo and offered a fork. The resolution is that Maxwell's equations single out a speed, which lives in the metric, and never singled out a frame. So branch (A) was not merely unsupported by Michelson and Morley. It was structurally impossible.
Chapter 1.1 showed Newton's third law failing for two moving charges by and wrote an IOU. Section 10 found the missing momentum sitting in the field at , and showed that its rate of change cancels the mechanical one exactly, in both components, including a that 1.1 never saw. That forced you to accept that a field is not a bookkeeping device but a physical system.
And the thesis, which is the reason the chapter exists. Electromagnetism did not need to be made compatible with relativity. It was relativistic before the word existed, which is exactly why it produced a frame-independent and broke the physics of 1900. What needed fixing was mechanics, and Chapters 2.2 to 2.5 fixed it.
The magnetic field is not a second force alongside the electric one. It is what an electric field looks like from a moving frame, an effect of relative order , which for two ordinary currents in copper is about . It is visible at all only because the positive and negative charge in matter cancel to at least that precision. A discrepancy in the twenty-sixth significant figure, and it defines a base unit of the SI.
Two fields, one tensor. Four equations, two lines. And a that belongs to spacetime rather than to any medium. That is the whole of Part II in one chapter, and it is the last thing you need before geometry starts to move.
Where this gets spent.
- → Chapter 3.6, where it is what sits on the right-hand side of the Einstein field equations. Energy gravitates, and now you know what "energy" means as a tensor and why it had to be symmetric.
- and → Chapter 3.5, where they become and literally, and stops being an analogy.
- → Chapter 5.2, quantised, where the field's normal modes become photons. Then Chapter 5.8, where becomes the single vertex of QED, and every Feynman diagram in the theory is built from copies of it.
- Gauge freedom → Chapter 6.3, promoted from a convenience to the generating principle. Demand it locally and this entire chapter is the output rather than the input. Then Chapter 6.4 repeats the construction with a non-commuting group and gets Yang–Mills, using §9's Lagrangian unchanged except for a commutator.
- Charge invariance and the – mixing → Chapter 5.5, where the Dirac field's coupling to is fixed by exactly these transformation properties.
- The whole structure → Chapter 7.5, where the gauge fields of the Standard Model appear as excitations of an open string, and Chapter 7.7, where turns out to be the leading term of an expansion on a D-brane whose higher terms are the Born–Infeld action.