Part II · Special Relativity — Chapter 2.4

Tensors, Honestly

An index is not decoration. It is a promise about what happens when someone else picks different coordinates.

Where we are

Chapter 2.3 wrote xμx^{\mu} and ημν\eta_{\mu\nu} and ds2=ημνdxμdxν\dd s^{2}=\eta_{\mu\nu}\dd x^{\mu}\dd x^{\nu} and got away with it, because the formulas came out right. This chapter pays for that. Everything in 2.3 was legitimate. None of it was defined. We go back now and supply the definitions, and along the way the index notation stops being a bookkeeping habit and becomes an idea.

Here is the destination, stated in advance so you know what we are aiming at:

A tensor is defined by how its components change when you change coordinates.

Not by a picture, not by "a thing with indices", not by "a matrix with more slots". By a transformation law, and by nothing else. That sounds like a technicality, and it is in fact the whole point, because it is what makes the following true: any equation whose two sides transform the same way is automatically true in every frame if it is true in one. That single sentence is why the rest of this book is written in indices. Section 6 proves it. Everything before Section 6 is the machinery needed to state it, and everything after it is what to do with it.

One more thing to hold onto. The definition we are about to give never mentions Lorentz transformations at all. It works for any smooth change of coordinates: rotations, polar coordinates, an accelerating frame, a coordinate patch on a curved manifold. That generality is not showing off. It is the reason Part III can be about geometry rather than about notation. Chapter 3.2 reuses this chapter's definition word for word with only one change, and that one change is the whole of general relativity.

Tools you'll need.  Chapter 0.4 §4: change of basis, the fact that components transform with P1P^{-1} while basis vectors transform with PP, and A=P1APA'=P^{-1}AP. That section is the entire skeleton of this one. Chapter 0.6 §4: the gradient is really a one-form, gijg_{ij}, and (f)i=gijjf(\nabla f)^{i}=g^{ij}\partial_{j}f. This chapter is the promised payoff of that section, and we pick the thread up by name in §3. Also 0.6 §5, the multivariable chain rule, which does almost all of the work below. Chapter 0.7 §4.3, the unique split of a matrix into symmetric plus antisymmetric parts, and §6, the continuity equation. Chapter 2.2 for the boost and rapidity. Chapter 2.3 for the interval, ημν\eta_{\mu\nu}, and the statement that a four-vector is something transforming like dxμ\dd x^{\mu}, which is the statement we are about to turn into a definition.

1 · The problem, stated sharply

Start with the demand that the whole chapter answers to. Physics is not allowed to depend on the coordinates a human being chose. Nothing about the electron knows whether you set up your axes pointing north or north-north-east, or whether you are drifting past the laboratory at 0.6c0.6c. Whatever is real must be the same for everyone.

But components are not the same for everyone, and nothing could make them so. That is what components are: numbers you get by comparing a physical thing to a set of reference directions you invented. Chapter 0.4 §4 made this precise for an ordinary vector. Change the basis by PP and the components change by P1P^{-1}. Chapter 2.2 made it precise for spacetime, where changing to a moving frame reshuffles the four numbers (ct,x,y,z)(ct,x,y,z) by a boost.

So we have a tension. The things we can write down are components, and those are not invariant. The things we want to assert are laws, and those must be. The resolution is not to abandon components. It is to insist on a particular class of objects.

The whole strategy

Work only with objects whose components change in a controlled, known way, meaning a way fixed entirely by the coordinate change and not at all by the object. Then build statements out of them in which the changes cancel. The pieces are frame-dependent. The statement is not.

That is a design constraint, and like all good design constraints it rules things out. Let's watch it rule something out, because here the failure teaches more than the success would.

1.1 · Four numbers that are not a four-vector

Here is an object that has an index and is not a tensor. Define

Nμ    (1,0,0,0)in every inertial frame. N^{\mu} \;\equiv\; (1,\,0,\,0,\,0) \qquad\text{in every inertial frame.} (2.4.1)

Nothing in that definition is ill-formed. It is a perfectly good rule: hand me a frame, I hand you four numbers. You could imagine it standing for "the direction of time". You could also imagine it standing for the rest frame of the ether, which is what people historically did, and which is exactly the mistake Chapter 2.1 dismantled. It has an upper index, it has four components, and it looks like everything else we write. Watch what happens.

Take a particle with four-momentum pμ=(E/c,px,py,pz)p^{\mu}=(E/c,\,p^{x},\,p^{y},\,p^{z}), a genuine four-vector. Build the number

E    ημνNμpν  =  N0p0Np  =  Ec. \mathcal{E} \;\equiv\; \eta_{\mu\nu}\,N^{\mu}p^{\nu} \;=\; N^{0}p^{0} - \vv N\cdot\vv p \;=\; \frac{E}{c}. (2.4.2)

It looks like a scalar, with the indices all matched and summed and nothing left dangling. So let's evaluate it in two frames using specific numbers. Let pμ=(5,3,0,0)p^{\mu}=(5,\,3,\,0,\,0) in units of MeV/c\mathrm{MeV}/c, and let SS' move at β=3/5\beta=3/5, so γ=5/4\gamma=5/4. Chapter 2.2's boost gives pμ=(4,0,0,0)p'^{\mu}=(4,\,0,\,0,\,0). We will grind that arithmetic out in §9, and in the meantime you can check it in your head: 54(5353)=54165=4\tfrac54(5-\tfrac35\cdot3)=\tfrac54\cdot\tfrac{16}{5}=4. Now feed both sets of components into the recipe.

ES  =  15  =  5,ES  =  14  =  4. \mathcal{E}\big|_{S} \;=\; 1\cdot 5 \;=\; 5, \qquad\qquad \mathcal{E}\big|_{S'} \;=\; 1\cdot 4 \;=\; 4. (2.4.3)

Two observers, one recipe, two answers. Look carefully at what did not go wrong. No algebra was botched and no sign was dropped. The recipe ημνNμpν\eta_{\mu\nu}N^{\mu}p^{\nu} is unambiguous, and both computations are correct. What failed is the claim that the recipe defines a number. It defines a number per frame, which is to say it defines nothing at all. You cannot write "E=5\mathcal{E}=5" and mean anything by it.

The damage gets worse if you write an equation rather than a single number. Consider what happens to this one.

pμ  =  mcNμ. p^{\mu} \;=\; mc\,N^{\mu}. (2.4.4)

In frame SS', where our particle is at rest, that equation is true: pμ=(4,0,0,0)p'^{\mu}=(4,0,0,0) and mc=4mc=4. Now suppose NμN^{\mu} were a four-vector. The theorem of §6 would then let us conclude that (2.4.4) holds in every frame. And since NμN^{\mu} has been declared to be (1,0,0,0)(1,0,0,0) everywhere, that would say pμ=(mc,0,0,0)p^{\mu}=(mc,0,0,0) everywhere. This particle would be at rest for every observer, including the one watching it go past at 0.6c0.6c. The conclusion is false, so one of the premises must be false. It is not the theorem. It is the assumption that NμN^{\mu} is a four-vector, and the theorem's one hypothesis is exactly the one NμN^{\mu} fails.

⚠ Why this isn't obvious

The array (1,0,0,0)(1,0,0,0) is a perfectly respectable set of components. There is a genuine four-vector with those components, and in fact there are many. The four-velocity of any observer, divided by cc, has exactly those components in that observer's own frame. The error in (2.4.1) is not the numbers. The error is the phrase "in every frame."

Specify a four-vector by giving its components in one frame and you have said something. The transformation law then tells everyone else what they will measure, and nobody is free to disagree. Specify it instead by decreeing the components in all frames at once and you have specified the object twice over. Fixing four components in one frame uses up every freedom the object has, and the transformation law then determines what they must be everywhere else. Declaring them to be (1,0,0,0)(1,0,0,0) again in a frame moving at β\beta along xx contradicts that prediction, (γ,γβ,0,0)(\gamma,-\gamma\beta,0,0), in two of its four components. In every frame but the one you started in, it contradicts the prediction in something.

This is the single most common way that beginners' index expressions go wrong, and it survives into professional work in subtler disguises. Any time you find yourself saying "in this frame the components are simply ..." and then using that form in another frame, you have made this mistake.

So here is the slogan for the chapter. Read it as a definition rather than as a warning.

What a tensor is not

A tensor is not "a thing with indices". It is a thing with indices that transform correctly. The indices are not a filing system for numbers. They are a claim, and the claim is checkable.

In plain terms 2.4.1

Numbers on their own carry no context. A set of components means something only relative to the framework used to measure them, which is to say relative to a choice of axes, so shifting that choice must shift the numbers describing one unchanged reality. That is not a defect in the description. It is what describing anything physical consists of.

The fatal error is the opposite instinct, which is to freeze one set of numbers in place and require everybody to use it. Doing so does not make the object more real. It produces observers computing conflicting values for what was supposed to be a single fact about the world, and this section stages exactly that collapse with four entirely innocent-looking numbers.

So the goal of everything that follows is not to hunt for quantities that refuse to change, since almost nothing does. It is to find quantities that change in a fully predictable way when the choice of axes changes, so that the physical statements assembled out of them come out the same for everybody, whatever axes each of them chose.

2 · Vectors, defined by how they transform

Now let's build the definition. What we need first is a prototype: one object whose transformation law we can compute rather than choose. There is exactly one natural candidate, and Chapter 2.3 already leaned on it.

Let xμx^{\mu} be one set of coordinates and xμx'^{\mu} another, related by four functions xμ=xμ(x0,x1,x2,x3)x'^{\mu}=x'^{\mu}(x^{0},x^{1},x^{2},x^{3}) that are smooth and invertible. Take two nearby events separated by an infinitesimal displacement dxμ\dd x^{\mu}, and ask what that separation is in primed coordinates. The answer is not a matter of convention. It is the multivariable chain rule of Chapter 0.6 §5, applied to each of the four functions xμx'^{\mu} in turn.

dxμ  =  xμxνdxν, \dd x'^{\mu} \;=\; \pdv{x'^{\mu}}{x^{\nu}}\,\dd x^{\nu}, (2.4.5)

summed over ν=0,1,2,3\nu=0,1,2,3. Nothing was assumed there. Both sides are the same physical displacement between the same two events, described twice over. That array of partial derivatives is going to appear on every page from here on, so let's give it a name.

Λμν    xμxν,so thatdxμ=Λμνdxν. \Lambda^{\mu}{}_{\nu} \;\equiv\; \pdv{x'^{\mu}}{x^{\nu}}, \qquad\text{so that}\qquad \dd x'^{\mu} = \Lambda^{\mu}{}_{\nu}\,\dd x^{\nu}. (2.4.6)

This is the Jacobian matrix of Chapter 0.6 §8, wearing index notation. The upper index labels the row and the new coordinate. The lower index labels the column and the old one. Keep the horizontal offset in Λμν\Lambda^{\mu}{}_{\nu}, because it records which slot comes first, and by §4 that will start to matter.

Now make the leap that organises everything after it. We have one object, dxμ\dd x^{\mu}, that we know transforms by (2.4.6). We define the class of objects that behave the same way.

Definition — contravariant vector

A contravariant vector (or just vector, or in four dimensions a four-vector) is an object with four components VμV^{\mu} per coordinate system, such that under any change of coordinates

Vμ  =  ΛμνVν,Λμν=xμxν. V'^{\mu} \;=\; \Lambda^{\mu}{}_{\nu}\,V^{\nu}, \qquad \Lambda^{\mu}{}_{\nu} = \pdv{x'^{\mu}}{x^{\nu}}.

That is the entire definition. There is no additional requirement, and no picture is part of it.

Two things should feel unsatisfying about that definition, and both are worth confronting before we go on.

"It's circular. You defined vectors by copying dx\dd x." Yes, deliberately. The definition is ostensive: it points at a known object and says "these ones". What makes it non-vacuous is that many other objects turn out to belong to the class, and their membership is a theorem rather than a stipulation. Take three of them in turn. The four-velocity uμ=dxμ/dτu^{\mu}=\dd x^{\mu}/\dd\tau belongs, because dτ\dd\tau is a scalar (Chapter 2.3), and dividing a vector by a scalar leaves the law (2.4.6) intact. The four-momentum pμ=muμp^{\mu}=mu^{\mu} belongs, since mm is a scalar by construction. In this book mm always means rest mass, the number every observer computes from pμpμ=m2c2p_{\mu}p^{\mu}=m^{2}c^{2}. The current Jμ=(ρc,J)J^{\mu}=(\rho c,\vv J) belongs as well, and §9 shows what that buys. Each of those is a claim about physics, checkable and falsifiable, and each of them fails for NμN^{\mu} of (2.4.1).

"Where did the arrow go?" Nowhere. It is still available, and Chapter 0.4 tells you exactly how to recover it. There, a vector was an element of a vector space and its components were xjx_{j} in a basis {ej}\{e_{j}\}. Changing basis by ej=Pijeie'_{j}=P_{ij}e_{i} sent the components to x=P1x\vv x'=P^{-1}\vv x. Compare that with the definition above and read off the dictionary between the two accounts.

Λ  =  P1. \Lambda \;=\; P^{-1}. (2.4.7)

The components transform with Λ\Lambda and the basis vectors transform with Λ1=P\Lambda^{-1}=P, so their product, which is the actual arrow, does not move at all. Chapter 0.4 §4 put it in one sentence: make your ruler twice as long and every measurement in rulers is halved. "Contra-variant" is literally that. The components vary contrary to the basis. So the arrow is exactly as real as it ever was. What has changed is which end of the relationship we take as the definition. We take the components' end, because that is the end that still makes sense when there is no global vector space to draw arrows in. From Chapter 3.2 onward, that is always the situation.

2.1 · What Λ\Lambda is for a boost

Before going further, let's see what all of this looks like for the one transformation we already know. Specialise to the boost of Chapter 2.2, with SS' moving along +x+x at speed v=βcv=\beta c and xμ=(ct,x,y,z)x^{\mu}=(ct,x,y,z). Chapter 2.2 gave us the coordinate change itself.

ct=γ(ctβx),x=γ(xβct),y=y,z=z. ct' = \gamma(ct-\beta x), \quad x' = \gamma(x-\beta ct), \quad y'=y, \quad z'=z. (2.4.8)

What we want from it is Λ\Lambda, and getting it means differentiating each new coordinate with respect to each old one. The coefficients here are constants, so the partial derivatives are just those constants.

Λμν  =  (γγβ00γβγ0000100001),γ=11β2. \Lambda^{\mu}{}_{\nu} \;=\; \begin{pmatrix} \gamma & -\gamma\beta & 0 & 0\\ -\gamma\beta & \gamma & 0 & 0\\ 0 & 0 & 1 & 0\\ 0 & 0 & 0 & 1 \end{pmatrix}, \qquad \gamma=\frac{1}{\sqrt{1-\beta^{2}}}. (2.4.9)

That is the matrix Chapter 2.3 wrote down. It is no longer a postulate. It is a computed Jacobian, and every entry of it is independent of position. That constancy is a special feature of Lorentz transformations rather than a general fact about coordinate changes. Hold onto the observation. It is the hinge on which §9's second worked example turns, and it is the single technical difference between this chapter and Chapter 3.3.

Grind box — the inverse boost, and why rapidity is the sane variable

Invert (2.4.8) by solving for ctct and xx. Multiply the first by γ\gamma and the second by γβ\gamma\beta and add:

γct+γβx=γ2(1β2)ct=ct, \gamma\,ct' + \gamma\beta\,x' = \gamma^{2}(1-\beta^{2})\,ct = ct,

using γ2(1β2)=1\gamma^{2}(1-\beta^{2})=1, which is the definition of γ\gamma rearranged. The same manoeuvre gives x=γ(x+βct)x=\gamma(x'+\beta ct'). So

(Λ1)μν  =  xμxν  =  (γ+γβ00+γβγ0000100001)  =  Λ(β), (\Lambda^{-1})^{\mu}{}_{\nu} \;=\; \pdv{x^{\mu}}{x'^{\nu}} \;=\; \begin{pmatrix} \gamma & +\gamma\beta & 0&0\\ +\gamma\beta & \gamma & 0&0\\ 0&0&1&0\\ 0&0&0&1\end{pmatrix} \;=\; \Lambda(-\beta),

which is the physically unsurprising statement that if SS' moves at +v+v relative to SS then SS moves at v-v relative to SS'. Here that statement is derived rather than assumed.

Chapter 2.2 introduced the rapidity ϕ\phi with tanhϕ=β\tanh\phi=\beta, whence γ=coshϕ\gamma=\cosh\phi and γβ=sinhϕ\gamma\beta=\sinh\phi (divide cosh2sinh2=1\cosh^{2}-\sinh^{2}=1 by cosh2\cosh^{2} to see the first). In that variable

Λ(ϕ)=(coshϕsinhϕ00sinhϕcoshϕ0000100001), \Lambda(\phi) = \begin{pmatrix}\cosh\phi & -\sinh\phi & 0&0\\ -\sinh\phi & \cosh\phi & 0&0\\ 0&0&1&0\\0&0&0&1\end{pmatrix},

and the hyperbolic addition formulas give, in two lines of matrix multiplication,

Λ(ϕ2)Λ(ϕ1)  =  Λ(ϕ1+ϕ2). \Lambda(\phi_{2})\,\Lambda(\phi_{1}) \;=\; \Lambda(\phi_{1}+\phi_{2}).

Rapidities add. That is why ϕ\phi is the right variable and β\beta is not. It is also a consistency check on the whole framework, because the chain rule guarantees that composing coordinate changes composes their Jacobians,

xμxν  =  xμxρ  xρxν, \pdv{x''^{\mu}}{x^{\nu}} \;=\; \pdv{x''^{\mu}}{x'^{\rho}}\;\pdv{x'^{\rho}}{x^{\nu}},

so the transformation law of §2 is automatically consistent under composition. Nothing extra had to be imposed. Chapter 6.1 will call this "the transformations form a group and the components carry a representation of it". Here it is nothing more than the chain rule.

In plain terms 2.4.2

Objects here are defined by their behaviour rather than by their appearance, which feels backwards the first time it is met. Nothing dictates what a vector is made of or what it should look like, and the definition says only how its description must adapt when the frame of reference changes.

The prototype for that behaviour is the physical gap between two nearby events. Since the two events happen whether or not anybody is watching, the rule converting one observer's measurement of that gap into another's is fixed by the change of frame alone and by nothing else. Any object transforming by that same rule is a vector, and there is no further qualification to meet.

This behaviour-based definition is worth its initial awkwardness because it abandons any need for flat backgrounds, straight arrows, or a fixed grid on which to draw them. Nothing in it depends on the geometry being simple. That is precisely why it will survive perfectly intact into Part III, when the coordinates begin to curve and the comfortable picture of an arrow stops making sense.

3 · Covectors — the other kind of index

Chapter 0.6 §4 ended with a promise: differentiation naturally produces lower indices, vectors naturally carry upper ones, the metric converts between them, and when the metric is the identity you cannot see the conversion happening. It said that Chapter 2.4 would make this a formal convention. Here we are. This section builds the second species of index, and it builds it without using a metric at all. The metric does not arrive until §4.

Take a scalar field ff, meaning a function on spacetime whose value at an event is the same number for everyone, so that f(x)=f(x)f'(x')=f(x). Temperature at a point, say, or the proper time elapsed along some fixed worldline. Its four partial derivatives

μf    fxμ \partial_{\mu}f \;\equiv\; \pdv{f}{x^{\mu}} (2.4.10)

are perfectly good components of something. The question is what. We want their transformation law, so apply the chain rule to ff regarded as a function of the primed coordinates.

μf  =  fxμ  =  xνxμfxν  =  xνxμνf. \partial'_{\mu}f \;=\; \pdv{f}{x'^{\mu}} \;=\; \pdv{x^{\nu}}{x'^{\mu}}\,\pdv{f}{x^{\nu}} \;=\; \pdv{x^{\nu}}{x'^{\mu}}\,\partial_{\nu}f. (2.4.11)

Let's look at what appeared. Not xμxν\pdv{x'^{\mu}}{x^{\nu}}, but xνxμ\pdv{x^{\nu}}{x'^{\mu}}, the derivatives of the old coordinates with respect to the new. That is the inverse Jacobian. Calling it the inverse is a claim about matrices, so let's establish it. We can do that without computing anything, again by the chain rule.

xμxν  xνxρ  =  xμxρ  =  δμρ, \pdv{x'^{\mu}}{x^{\nu}}\;\pdv{x^{\nu}}{x'^{\rho}} \;=\; \pdv{x'^{\mu}}{x'^{\rho}} \;=\; \delta^{\mu}{}_{\rho}, (2.4.12)

since the primed coordinates are independent of one another: x0/x0=1\partial x'^{0}/\partial x'^{0}=1 and x0/x1=0\partial x'^{0}/\partial x'^{1}=0, and so on. So write (Λ1)νμxνxμ(\Lambda^{-1})^{\nu}{}_{\mu}\equiv\pdv{x^{\nu}}{x'^{\mu}}. The notation is honest. That array really is the matrix inverse of (2.4.6), and (2.4.12) reads ΛΛ1=I\Lambda\Lambda^{-1}=I.

So μf\partial_{\mu}f transforms with the inverse of Λ\Lambda, while dxμ\dd x^{\mu} transforms with Λ\Lambda itself. These are different laws. They give different numbers. Objects that obey the inverse law deserve their own name and their own index position.

Definition — covariant vector (one-form, covector)

A covector is an object with four components ωμ\omega_{\mu} per coordinate system such that under any change of coordinates

ωμ  =  (Λ1)νμων,(Λ1)νμ=xνxμ. \omega'_{\mu} \;=\; (\Lambda^{-1})^{\nu}{}_{\mu}\,\omega_{\nu}, \qquad (\Lambda^{-1})^{\nu}{}_{\mu} = \pdv{x^{\nu}}{x'^{\mu}}.

We write its index down. The gradient μf\partial_{\mu}f of any scalar is the prototype, exactly as dxμ\dd x^{\mu} was the prototype for vectors.

3.1 · Why the two species exist: the pairing

Now for the structural fact that justifies having invented a second kind of object at all. Take a vector VμV^{\mu} and a covector ωμ\omega_{\mu} and contract them, which means summing over the repeated index with one copy up and one down. Watch what the two transformation laws do to each other.

Vμωμ  =  (ΛμνVν)((Λ1)ρμωρ)=  Λμν(Λ1)ρμ=  δρν  Vνωρ  =  δρνVνωρ  =  Vνων. \begin{aligned} V'^{\mu}\omega'_{\mu} \;&=\; \Big(\Lambda^{\mu}{}_{\nu}V^{\nu}\Big)\Big((\Lambda^{-1})^{\rho}{}_{\mu}\,\omega_{\rho}\Big)\\[3pt] &=\; \underbrace{\Lambda^{\mu}{}_{\nu}\,(\Lambda^{-1})^{\rho}{}_{\mu}}_{=\;\delta^{\rho}{}_{\nu}}\;V^{\nu}\omega_{\rho} \;=\; \delta^{\rho}{}_{\nu}V^{\nu}\omega_{\rho} \;=\; V^{\nu}\omega_{\nu}. \end{aligned} (2.4.13)

The two Jacobians met and collapsed to a Kronecker delta, and the delta then did nothing except rename an index. So VμωμV^{\mu}\omega_{\mu} takes the same value in every coordinate system. It is a scalar, meaning a real number rather than a number-per-frame. It is the first genuinely invariant thing we have built, and building it took no metric.

Let's read the two definitions again in the light of that calculation. The vector law and the covector law are not two arbitrary conventions that happen to be inverse to one another. Covectors are defined to be the objects that pair with vectors to give invariants. That is what they are for, and the inverse Jacobian is forced on them by that job. In Chapter 0.4's language, the covectors at a point form the dual space VV^{*}, which is the space of linear maps from vectors to numbers.

3.2 · The picture: arrows and stacks

There is a genuine geometric image behind all of this, and it is not the one you would guess.

A vector is an arrow. A covector is not an arrow. A covector ω\omega is a machine that eats an arrow and returns a number, linearly, and the honest picture of such a machine is a family of evenly spaced parallel surfaces. They are the level sets ω=0,1,2,3,\omega=0,1,2,3,\ldots of the linear function it computes. In two dimensions they are parallel lines. In four they are parallel hyperplanes. A "big" covector has closely spaced surfaces and a "small" one has widely spaced surfaces, so the covector's magnitude lives in the density of the stack. That is why doubling ω\omega halves the spacing.

With that picture in hand, the contraction VμωμV^{\mu}\omega_{\mu} acquires a completely concrete meaning.

What contraction counts

VμωμV^{\mu}\omega_{\mu} is the number of ω\omega-surfaces the arrow VV pierces.

An arrow either crosses a given surface or it does not. No coordinate system has an opinion about that. This is why the pairing is invariant, in one sentence and with no algebra.

For the gradient, this is a picture you have already seen without knowing what you were looking at. The level sets of df\dd f are the contours of ff, and VμμfV^{\mu}\partial_{\mu}f is the number of contour lines you cross walking along VV, which is the change in ff. Chapter 0.6's contour figure was a picture of a one-form all along.

90.0°
1.00
1.00
V¹ = 2.60000 V² = 1.50000
ω₁ = 1.00000 ω₂ = 0.60000
Vⁱωᵢ = 3.50000000 (drift 0.0e+0)
lines pierced = 3 (tip sits at ω = 3.50)
Two species, one number. The green arrow is a vector VV; the orange stack is a covector ω\omega, drawn as its level lines ω=0,1,2,\omega=0,1,2,\ldots. Both are fixed geometric objects and neither moves when you touch a slider. There is deliberately no background grid to anchor them to — the only coordinate system on the page is the blue one, {e1,e2}\{\vv e_{1},\vv e_{2}\} and its lattice, and that one is your arbitrary choice. The readouts are the components of VV and of ω\omega in that basis, and their contraction. Start with e1|\vv e_{1}|. Double it: V1V^{1} halves, from 2.62.6 to 1.31.3, while ω1\omega_{1} doubles, from 11 to 22. That is Λ\Lambda and Λ1\Lambda^{-1}, visible in two numbers — measure in longer units and a vector's components shrink while a covector's grow, because a longer step crosses more of ω\omega's lines. Their product is untouched. Now skew the basis and rescale e2\vv e_{2} as well: all four components wander, some of them through zero and into negative values, and ViωiV^{i}\omega_{i} sits at 3.500000003.50000000 throughout, with a residual of at most a few parts in 101610^{16} that is floating-point roundoff and nothing else. That is (2.4.13), measured rather than asserted. Meanwhile the ringed dots mark where the arrow crosses a level line: there are always exactly three, because the tip sits at ω=3.5\omega=3.5, and no re-coordinatising can change how many lines a fixed arrow crosses. Now press "show the wrong pairing". That readout treats ω\omega as though it were a vector — takes its contravariant components ω~i\tilde\omega^{i} in the same basis and forms the naive sum iViω~i\sum_{i}V^{i}\tilde\omega^{i}, in which both factors transform the same way. At the Cartesian start it agrees exactly, 3.5000003.500000, which is precisely why nobody notices the distinction in a first course. Then move a single slider and it goes: 1.5500001.550000 when you double e1|\vv e_{1}|, 2.3333082.333308 at 6060^{\circ}, and — with both basis vectors at their shortest and the angle at its widest — 77.14499177.144991. The up/down distinction is not decoration. It is the difference between a number and a number-per-frame.

3.3 · Why nobody told you this before

Because in Euclidean space with Cartesian coordinates the two species have numerically identical components, which makes the distinction invisible. Chapter 0.6 §4 showed why. Converting between them requires the metric, and in Cartesian coordinates on flat Euclidean space gij=δijg_{ij}=\delta_{ij}, so the conversion multiplies everything by 11. Every dot product you have ever computed as iaibi\sum_{i}a_{i}b_{i} was silently a contraction of a vector with a covector, and you were spared the bookkeeping because the metric was doing nothing.

The invisibility ends the moment either of two things happens, and in relativity both of them do. Use non-Cartesian coordinates and gijg_{ij} stops being δij\delta_{ij}, which is where Chapter 0.6 §4 got the polar gradient's notorious 1/r1/r. Or use a metric that is not positive-definite, such as ημν=diag(1,1,1,1)\eta_{\mu\nu}=\mathrm{diag}(1,-1,-1,-1), and the conversion starts flipping signs. That second case is what §4 is about.

⚠ Two different species of object

VμV^{\mu} and ωμ\omega_{\mu} are not two notations for the same thing. They live in different vector spaces, obey different transformation laws, and have different geometric pictures. An expression like Vμ+ωμV^{\mu}+\omega_{\mu} is not merely bad style. It is meaningless, in the way that "3 metres + 3 kilograms" is meaningless. The sum would be one thing in one frame and something else in another.

The reason this is worth being strict about now is that in Part III the metric acquires position dependence, and then no coordinate system anywhere makes the two coincide. Any habit of thought that relies on "they're the same really" fails there, permanently and silently.

In plain terms 2.4.3

Spacetime turns out to be inhabited by two distinct but complementary types of object. The first is the ordinary vector, which you can visualise as an arrow representing a physical displacement, pointing from here to there. The second is the covector, which is better understood as a measuring device than as a thing being measured. Rather than an arrow, picture a covector as a series of evenly spaced parallel sheets, much like the contour lines on a map.

Pair the two together and the covector counts how many of its sheets the arrow punches through, and that count is the whole of the pairing. The single image explains why their numbers move in opposite directions whenever the units change. If you stretch your measuring ruler, the arrow's numerical components shrink, and so the covector's sheets must crowd closer together in order that the total count of pierced sheets comes out exactly as it did before. That count is an indisputable physical fact, and no choice of ruler is entitled to alter it.

The only reason these two species are so frequently conflated is that their numbers happen to mirror each other perfectly in a simple, flat, Cartesian grid. That agreement is a mathematical coincidence of the flat ruler rather than a law of nature, and it shatters the moment either the geometry or the coordinate system becomes interesting.

4 · The metric as the dictionary

So far there has been no metric. Vectors, covectors and the pairing between them all exist without one. Now we add the object that Chapter 2.3 built out of the invariance of the interval, and look at exactly what it buys.

Chapter 2.3's central claim was that the quantity

ds2  =  ημνdxμdxν,ημν=diag(1,1,1,1), \dd s^{2} \;=\; \eta_{\mu\nu}\,\dd x^{\mu}\dd x^{\nu}, \qquad \eta_{\mu\nu}=\mathrm{diag}(1,-1,-1,-1), (2.4.14)

is the same for all inertial observers. Let's write that out as an equation between two computations of the same thing, using (2.4.6) on the primed side.

ημνdxμdxν  =  ημνΛμρΛνσdxρdxσ  =!  ηρσdxρdxσ. \eta_{\mu\nu}\,\dd x'^{\mu}\dd x'^{\nu} \;=\; \eta_{\mu\nu}\,\Lambda^{\mu}{}_{\rho}\Lambda^{\nu}{}_{\sigma}\,\dd x^{\rho}\dd x^{\sigma} \;\overset{!}{=}\; \eta_{\rho\sigma}\,\dd x^{\rho}\dd x^{\sigma}. (2.4.15)

Since dxρ\dd x^{\rho} is arbitrary, the coefficients at the two ends must agree. (They must agree after symmetrising, and η\eta is already symmetric.) That gives us the defining property of a Lorentz transformation.

  ημνΛμρΛνσ  =  ηρσ.   \boxed{\;\eta_{\mu\nu}\,\Lambda^{\mu}{}_{\rho}\,\Lambda^{\nu}{}_{\sigma} \;=\; \eta_{\rho\sigma}.\;} (2.4.16)

In matrix form that reads ΛTηΛ=η\Lambda^{\mathsf T}\eta\,\Lambda=\eta. Let's pause on what it is saying. The Lorentz transformations are exactly the linear maps that preserve η\eta, in the same way that rotations are exactly the linear maps preserving δij\delta_{ij}, which is the condition RTR=IR^{\mathsf T}R=I. Chapter 2.3 derived the boost from the two postulates. Equation (2.4.16) is that same result restated as an algebraic condition, and it is the form Chapter 6.1 generalises when it names the group SO(1,3)\mathrm{SO}(1,3).

4.1 · Lowering

Define, for any vector VμV^{\mu},

Vμ    ημνVν. V_{\mu} \;\equiv\; \eta_{\mu\nu}V^{\nu}. (2.4.17)

The name for that operation is "lowering the index". Writing the result with a down index carries an implicit claim, which is that VμV_{\mu} is a genuine covector obeying the law of §3. That is a claim, so let's prove it. Start from the definition in the primed frame and push.

Vμ  =  ημνVν  =  ημνΛνρVρ. V'_{\mu} \;=\; \eta_{\mu\nu}V'^{\nu} \;=\; \eta_{\mu\nu}\Lambda^{\nu}{}_{\rho}\,V^{\rho}. (2.4.18)

We want that to equal (Λ1)σμVσ=(Λ1)σμησρVρ(\Lambda^{-1})^{\sigma}{}_{\mu}V_{\sigma}=(\Lambda^{-1})^{\sigma}{}_{\mu} \eta_{\sigma\rho}V^{\rho}. So the whole question is whether the following identity holds.

ημνΛνρ  =?  (Λ1)σμησρ. \eta_{\mu\nu}\,\Lambda^{\nu}{}_{\rho} \;\overset{?}{=}\; (\Lambda^{-1})^{\sigma}{}_{\mu}\,\eta_{\sigma\rho}. (2.4.19)

Take (2.4.16) and contract it with (Λ1)σα(\Lambda^{-1})^{\sigma}{}_{\alpha} on the index σ\sigma:

ηρσ(Λ1)σα  =  ημνΛμρΛνσ(Λ1)σα=  ημνΛμρδνα  =  ημαΛμρ, \begin{aligned} \eta_{\rho\sigma}\,(\Lambda^{-1})^{\sigma}{}_{\alpha} \;&=\; \eta_{\mu\nu}\,\Lambda^{\mu}{}_{\rho}\,\Lambda^{\nu}{}_{\sigma}\,(\Lambda^{-1})^{\sigma}{}_{\alpha}\\[3pt] &=\; \eta_{\mu\nu}\,\Lambda^{\mu}{}_{\rho}\,\delta^{\nu}{}_{\alpha} \;=\; \eta_{\mu\alpha}\,\Lambda^{\mu}{}_{\rho}, \end{aligned} (2.4.20)

which is (2.4.19) with the free indices renamed and the symmetry ημν=ηνμ\eta_{\mu\nu}=\eta_{\nu\mu} used on both sides. So yes: VμV_{\mu} is a covector, and (2.4.17) is a legitimate map from one species to the other. Note precisely what made it work. It was (2.4.16), the invariance of η\eta. Lowering with an arbitrary array of numbers would not have produced a covector.

4.2 · Raising

Going the other way needs the inverse of η\eta. Since η\eta is a non-degenerate matrix it has one. Call its entries ημν\eta^{\mu\nu}, defined by

ημρηρν  =  δμν. \eta^{\mu\rho}\eta_{\rho\nu} \;=\; \delta^{\mu}{}_{\nu}. (2.4.21)

For our signature the inverse is easy to find, because diag(1,1,1,1)\mathrm{diag}(1,-1,-1,-1) squares to the identity. So ημν=diag(1,1,1,1)\eta^{\mu\nu}=\mathrm{diag}(1,-1,-1,-1) as an array of numbers, the same four numbers down the diagonal. That coincidence is a property of this particular metric and not a general fact. In Chapter 3.3 the inverse metric gμνg^{\mu\nu} is a genuinely different array from gμνg_{\mu\nu}, and confusing the two is a classic and expensive error. With the inverse in hand, we can raise an index.

Vμ  =  ημνVν. V^{\mu} \;=\; \eta^{\mu\nu}V_{\nu}. (2.4.22)

That is consistent, since ημνVν=ημνηνρVρ=δμρVρ=Vμ\eta^{\mu\nu}V_{\nu}=\eta^{\mu\nu}\eta_{\nu\rho}V^{\rho}=\delta^{\mu}{}_{\rho}V^{\rho}=V^{\mu}. Raising after lowering returns you to where you started, which is the least one can ask of a dictionary.

4.3 · What it actually does to the numbers

Let's see what lowering does to actual numbers. Write (2.4.17) out component by component. Since η\eta is diagonal, only one term survives in each sum.

V0=η00V0=+V0,V1=η11V1=V1,V2=V2,V3=V3. \begin{aligned} V_{0} &= \eta_{00}V^{0} = +V^{0}, & V_{1} &= \eta_{11}V^{1} = -V^{1},\\ V_{2} &= -V^{2}, & V_{3} &= -V^{3}. \end{aligned} (2.4.23)

So with signature (+,,,)(+,-,-,-), lowering flips the sign of the three spatial components and leaves the time component alone. Take the four-momentum of §1 as a concrete case,

pμ=(5,3,0,0)pμ=(5,3,0,0), p^{\mu} = (5,\,3,\,0,\,0) \qquad\Longrightarrow\qquad p_{\mu} = (5,\,-3,\,0,\,0), (2.4.24)

and the invariant length is then the contraction of the two arrays against each other,

pμpμ  =  (5)(5)+(3)(3)+0+0  =  259  =  16  =  (mc)2, p_{\mu}p^{\mu} \;=\; (5)(5) + (-3)(3) + 0 + 0 \;=\; 25-9 \;=\; 16 \;=\; (mc)^{2}, (2.4.25)

so m=4 MeV/c2m=4\ \mathrm{MeV}/c^{2}. Note that pμpμp_{\mu}p^{\mu} is a contraction of a covector with a vector, so by (2.4.13) it is invariant with no further argument needed. Section 9 checks that by brute force in three different frames.

⚠ Raising an index is not moving a letter

The notation is so light that it invites you to think VμV^{\mu} and VμV_{\mu} are the same object written two ways, and that shifting the index up or down is a typographical act. It is not. It is the application of a specific linear map, namely the matrix η\eta, and that map is not the identity. With our signature it multiplies three of the four components by 1-1.

This is the single most productive source of sign errors in relativity, and they are the worst kind of error: dimensionally consistent, index-consistent, and off by a sign. Three habits protect you.

  • Never write an expression with two indices in the same position summed. VμWμV^{\mu}W^{\mu} is not a thing. It is not "VWV\cdot W in a hurry". It is a frame-dependent number, as §9 measures. If you want VWV\cdot W, write ημνVμWν\eta_{\mu\nu}V^{\mu}W^{\nu} or VμWμV_{\mu}W^{\mu}, which are the same thing.
  • Track where the minus signs are. VμWμ=V0W0VWV_{\mu}W^{\mu}=V^{0}W^{0}-\vv V\cdot\vv W. The three-vector dot product enters with a minus. Every time.
  • μ\partial^{\mu} is not μ\partial_{\mu} with the index moved. It is ημνν\eta^{\mu\nu}\partial_{\nu}, so 0=0\partial^{0}=\partial_{0} but i=i\partial^{i}=-\partial_{i}. Chapter 2.6's Fμν=μAννAμF^{\mu\nu}=\partial^{\mu}A^{\nu}-\partial^{\nu}A^{\mu} has signs in it that come from precisely here.
Grind box — index gymnastics, and a bonus identity

Three working rules, each a one-line consequence of (2.4.21).

(i) You may slide a contracted pair up and down together.

VμWμ=ημνVνWμ=Vν(ηνμWμ)=VνWν. V_{\mu}W^{\mu} = \eta_{\mu\nu}V^{\nu}W^{\mu} = V^{\nu}\big(\eta_{\nu\mu}W^{\mu}\big) = V^{\nu}W_{\nu}.

So VμWμ=VμWμV_{\mu}W^{\mu}=V^{\mu}W_{\mu}: it does not matter which factor carries the lowered index, only that exactly one does.

(ii) η\eta with one index up and one down is the Kronecker delta. By (2.4.21), ημν=ημρηρν=δμν\eta^{\mu}{}_{\nu}=\eta^{\mu\rho}\eta_{\rho\nu}=\delta^{\mu}{}_{\nu}. So you never need to write ημν\eta^{\mu}{}_{\nu} at all. It is δ\delta.

(iii) The metric knows the inverse transformation. Define Λμν\Lambda_{\mu}{}^{\nu} the only way the rules allow, by lowering the first index and raising the second: ΛμνημαηνβΛαβ\Lambda_{\mu}{}^{\nu}\equiv\eta_{\mu\alpha}\,\eta^{\nu\beta}\,\Lambda^{\alpha}{}_{\beta}. Now contract the identity (2.4.20), namely ημαΛμρ=ηρσ(Λ1)σα\eta_{\mu\alpha}\Lambda^{\mu}{}_{\rho}=\eta_{\rho\sigma}(\Lambda^{-1})^{\sigma}{}_{\alpha}, with ηρτ\eta^{\rho\tau}.

left=ηρτημαΛμρ=Λατ,right=ηρτηρσ(Λ1)σα=δτσ(Λ1)σα=(Λ1)τα. \begin{aligned} \text{left} &= \eta^{\rho\tau}\eta_{\mu\alpha}\Lambda^{\mu}{}_{\rho} = \Lambda_{\alpha}{}^{\tau},\\[3pt] \text{right} &= \eta^{\rho\tau}\eta_{\rho\sigma}(\Lambda^{-1})^{\sigma}{}_{\alpha} = \delta^{\tau}{}_{\sigma}(\Lambda^{-1})^{\sigma}{}_{\alpha} = (\Lambda^{-1})^{\tau}{}_{\alpha}. \end{aligned}

Hence Λατ=(Λ1)τα\Lambda_{\alpha}{}^{\tau}=(\Lambda^{-1})^{\tau}{}_{\alpha}. Lowering and raising the indices of Λ\Lambda produces its inverse. This is why most books never write Λ1\Lambda^{-1} at all. They write Λμν\Lambda_{\mu}{}^{\nu} and let the index positions carry the information. We have kept Λ1\Lambda^{-1} explicit up to this point, because the covector law of §3 is true with no metric in sight and it would have been dishonest to define it using one.

(iv) detΛ=±1\det\Lambda=\pm1. Take determinants of ΛTηΛ=η\Lambda^{\mathsf T}\eta\Lambda=\eta. By Chapter 0.4 §5, det\det is multiplicative and detΛT=detΛ\det\Lambda^{\mathsf T}=\det\Lambda, so (detΛ)2detη=detη(\det\Lambda)^{2}\det\eta=\det\eta, and detη=10\det\eta=-1\neq0 gives (detΛ)2=1(\det\Lambda)^{2}=1. Boosts and rotations have det=+1\det=+1. A spatial reflection has det=1\det=-1. That sign will matter exactly once in this chapter, in §8.

Two last remarks on η\eta, both pointing forward.

It is what makes "length" mean anything. Without a metric there is no way to ask how long a four-vector is. You can pair a vector with a covector, but not a vector with a vector. η\eta supplies exactly the missing ingredient, and the price is that "length squared" can come out positive, negative or zero. That is the timelike, spacelike and null trichotomy of Chapter 2.3. The metric does not merely measure. It is what defines the causal structure.

It is about to become a field. Everything in this section used only that η\eta is a symmetric, invertible, position-independent array satisfying (2.4.16). Now drop the words "position-independent", and replace ημν\eta_{\mu\nu} by gμν(x)g_{\mu\nu}(x), ten functions of position. Every formula above survives unchanged. What changes is that Λ\Lambda becomes position-dependent, and then exactly one thing breaks. What breaks is the derivative, as §9's second worked example shows. Repairing it is Chapter 3.3, and the repair is the gravitational field. That is the entire difference between special and general relativity, and it is a difference in one adjective.

Familiar ground — why a regression coefficient is not a number

You have fitted y^=jβjxj\hat y=\sum_{j}\beta_{j}x_{j} a thousand times, and (2.4.17) has been in it the whole while under another name.

Report creatinine in mg/dL\mathrm{mg/dL} rather than μmol/L\mathrm{\mu mol/L} and a column is rescaled, xjaxjx_{j}\to a\,x_{j}. The fitted coefficient moves the other way, βjβj/a\beta_{j}\to\beta_{j}/a, and the product βjxj\beta_{j}x_{j}, which is the only part anybody ever uses, is untouched. That is §3's arrow and stack of sheets exactly. The data column carries an upper index, the coefficient carries a lower one, and the contraction (2.4.13) is the invariant. A predicted value is a scalar. A coefficient is a component.

The consequence is a mistake that appears in print every week. Answering "which predictor matters most" by ranking raw coefficients is the precise error this chapter's notation exists to catch, since the ranking changes when somebody changes a unit, and a statement that does that was never a statement about the patients.

The standard repair is standardisation, βjstd=βjsj/sy\beta_{j}^{\mathrm{std}}=\beta_{j}s_{j}/s_{y} with the ss sample standard deviations. Look at what that operation is. It converts a lower index into an upper one by contracting with a symmetric array assembled out of the data's own second moments, which is (2.4.17) run backwards with the covariance structure playing the part of the metric. This section's warning then applies unchanged: the translation depends on which converter you hand it. Standardise against the sample's spread, against a reference population's spread, or against an interquartile range, and the most important predictor can change. It changes not because the biology moved but because a different metric was used.

What breaks, and it is the interesting half. Nothing obliges a statistician to be frame-independent. Choosing whichever standardisation makes a table readable is legitimate, and the choice gets reported rather than derived. In physics the converter is not chosen. The metric is handed over by the world, it is the same one for every observer, and §6's theorem then makes the frame-independence of an entire equation checkable by inspection. No such theorem exists for a regression table, because the demand that produced it was never made.

In plain terms 2.4.4

Because arrows and sheets are genuinely different species of object, something has to translate between them, and the metric is that translator. Feed the metric a displacement arrow and it hands back the particular stack of sheets that measures lengths in the same way. This is the physical reality behind the notation of moving an index up or down: it is a literal act of translation, performed by a specific piece of machinery, rather than a rearrangement of symbols on the page.

The distinction is not a typographical flourish. Depending on the rules of the space you are working in, the translation can flip the mathematical signs of your components, which means that an index moved carelessly is a sign error already in flight, and typically one that will not be discovered for several lines yet.

Most importantly of all, the translation depends entirely on which metric you hand it. Here that metric is a rigid, unchanging table of values, identical at every point in spacetime. Later it will evolve into a flexible field that varies from place to place, taking different values here than it does over there — and that single conceptual shift, from a fixed table to a field, is the entire foundation of gravity.

5 · Tensors in general

Vectors carry one upper index and one factor of Λ\Lambda. Covectors carry one lower index and one factor of Λ1\Lambda^{-1}. The generalisation writes itself, and it is a definition rather than a theorem.

Definition — tensor of type (p,q)(p,q)

A tensor of type (p,q)(p,q) has pp upper and qq lower indices, hence 4p+q4^{p+q} components per coordinate system, and transforms with one factor of Λ\Lambda per upper index and one factor of Λ1\Lambda^{-1} per lower index:

Tμ1μpν1νq  =  Λμ1α1Λμpαp×  (Λ1)β1ν1(Λ1)βqνq  Tα1αpβ1βq. \begin{aligned} T'^{\mu_{1}\ldots\mu_{p}}{}_{\nu_{1}\ldots\nu_{q}} \;=\; &\Lambda^{\mu_{1}}{}_{\alpha_{1}}\cdots\Lambda^{\mu_{p}}{}_{\alpha_{p}}\\[2pt] &\times\;(\Lambda^{-1})^{\beta_{1}}{}_{\nu_{1}}\cdots(\Lambda^{-1})^{\beta_{q}}{}_{\nu_{q}}\; T^{\alpha_{1}\ldots\alpha_{p}}{}_{\beta_{1}\ldots\beta_{q}}. \end{aligned}

The law is linear and homogeneous in the components of TT. Remember that. Section 6 is nothing but that one observation.

Here is everything we have met so far, sorted by type.

TypeObjectExample
(0,0)(0,0)scalarmm, dτ\dd\tau, ds2\dd s^{2}, pμpμp_{\mu}p^{\mu}
(1,0)(1,0)vectordxμ\dd x^{\mu}, uμu^{\mu}, pμp^{\mu}, JμJ^{\mu}
(0,1)(0,1)covectorμf\partial_{\mu}f, pμp_{\mu}, AμA_{\mu}
(0,2)(0,2)rank-2, both downημν\eta_{\mu\nu}, gμνg_{\mu\nu} (Ch 3.3)
(1,1)(1,1)mixedδμν\delta^{\mu}{}_{\nu}, νVμ\partial_{\nu}V^{\mu}
(2,0)(2,0)rank-2, both upFμνF^{\mu\nu} (Ch 2.6), TμνT^{\mu\nu} (Ch 3.6)

A scalar has no indices and therefore no factors of Λ\Lambda, so ϕ(x)=ϕ(x)\phi'(x')=\phi(x), the same number for everyone. That is what the general law says at p=q=0p=q=0, and it is why "scalar" and "invariant" are the same word here.

5.1 · The three operations that preserve tensor character

Tensors are useful because you can compute with them without leaving the class. Exactly three basic moves are permitted, and the shortness of that list is itself worth knowing.

(1) Linear combination, same type. If SS and TT are both type (p,q)(p,q) and a,ba,b are scalars, then aS+bTaS+bT is type (p,q)(p,q). That follows at once from the linearity of the transformation law. Apply the law to each term and factor out the common string of Λ\Lambdas. Note the restriction on the type. You cannot add a (1,0)(1,0) to a (0,1)(0,1), for the reason given in §3.3.

(2) Outer product. If AμA^{\mu} is type (1,0)(1,0) and BνB_{\nu} is type (0,1)(0,1), then the sixteen numbers CμνAμBνC^{\mu}{}_{\nu}\equiv A^{\mu}B_{\nu} form a type (1,1)(1,1) tensor. Proof: multiply the two transformation laws together and read off. In general, multiplying a (p,q)(p,q) by an (r,s)(r,s) component-by-component with all indices distinct gives a (p+r,q+s)(p+r,q+s).

(3) Contraction. Set one upper index equal to one lower index and sum. This is the interesting one, so we prove it.

Theorem — contraction

If TμνT^{\mu}{}_{\nu} is a tensor of type (1,1)(1,1), then TμμT^{\mu}{}_{\mu} (summed) is a scalar. More generally, contracting any upper index of a type-(p,q)(p,q) tensor against any lower index yields a tensor of type (p1,q1)(p-1,q-1).

Proof of the basic case. Transform, then set the indices equal:

Tμμ  =  Λμα(Λ1)βμ  Tαβ=  δβαTαβ  =  Tαα. \begin{aligned} T'^{\mu}{}_{\mu} \;&=\; \Lambda^{\mu}{}_{\alpha}\,(\Lambda^{-1})^{\beta}{}_{\mu}\;T^{\alpha}{}_{\beta}\\[3pt] &=\; \delta^{\beta}{}_{\alpha}\,T^{\alpha}{}_{\beta} \;=\; T^{\alpha}{}_{\alpha}. \qquad\blacksquare \end{aligned} (2.4.26)

It is the same collapse as (2.4.13), and it is the same collapse every time: a Λ\Lambda and a Λ1\Lambda^{-1} sharing a summed index annihilate into a δ\delta. Once you have seen that, you have seen all of tensor algebra. The rest is bookkeeping about which indices are spectators.

Grind box — contraction in the general case, with spectators

Take a (2,2)(2,2) tensor TμνρσT^{\mu\nu}{}_{\rho\sigma} and contract ν\nu against σ\sigma. Define UμρTμλρλU^{\mu}{}_{\rho}\equiv T^{\mu\lambda}{}_{\rho\lambda} (sum on λ\lambda). Claim: UU is a (1,1)(1,1) tensor. Write the transformation law and set ν=σ\nu=\sigma, calling the common value λ\lambda:

Uμρ=Tμλρλ=ΛμαΛλβ(Λ1)γρ(Λ1)δλ  Tαβγδ. \begin{aligned} U'^{\mu}{}_{\rho} &= T'^{\mu\lambda}{}_{\rho\lambda}\\[2pt] &= \Lambda^{\mu}{}_{\alpha}\,\Lambda^{\lambda}{}_{\beta}\,(\Lambda^{-1})^{\gamma}{}_{\rho}\,(\Lambda^{-1})^{\delta}{}_{\lambda}\;T^{\alpha\beta}{}_{\gamma\delta}. \end{aligned}

Now the only factors carrying λ\lambda are Λλβ\Lambda^{\lambda}{}_{\beta} and (Λ1)δλ(\Lambda^{-1})^{\delta}{}_{\lambda}, and λ\lambda is summed. Group them:

Λλβ(Λ1)δλ  =  δδβ, \Lambda^{\lambda}{}_{\beta}\,(\Lambda^{-1})^{\delta}{}_{\lambda} \;=\; \delta^{\delta}{}_{\beta}, Uμρ  =  Λμα(Λ1)γρ  δδβTαβγδ  =  Λμα(Λ1)γρ  Uαγ, U'^{\mu}{}_{\rho} \;=\; \Lambda^{\mu}{}_{\alpha}\,(\Lambda^{-1})^{\gamma}{}_{\rho}\;\delta^{\delta}{}_{\beta}\,T^{\alpha\beta}{}_{\gamma\delta} \;=\; \Lambda^{\mu}{}_{\alpha}\,(\Lambda^{-1})^{\gamma}{}_{\rho}\;U^{\alpha}{}_{\gamma},

which is exactly the (1,1)(1,1) law. The spectator indices μ\mu and ρ\rho never participated. They carried their own Λ\Lambda and Λ1\Lambda^{-1} through untouched. The argument is identical for any (p,q)(p,q) and any pair of indices, which is why nobody writes it out twice.

Two consequences worth stating. Contracting a vector with a covector is the case (1,0)(0,1)(0,0)(1,0)\otimes(0,1)\to(0,0), which is (2.4.13) again. And pμpμ=ημνpμpνp_{\mu}p^{\mu}=\eta_{\mu\nu}p^{\mu}p^{\nu} is (0,2)(1,0)(1,0)(0,2)\otimes(1,0)\otimes(1,0) contracted twice down to (0,0)(0,0). That is why it is a scalar. Not because someone declared mass invariant, but because of the index structure.

⚠ Why the summation convention only ever pairs up with down

Because that is the only pairing whose result is a tensor. Λ\Lambda against Λ1\Lambda^{-1} collapses to δ\delta. Λ\Lambda against Λ\Lambda does not collapse to anything, and what you get depends on the frame. So an expression like μVμWμ\sum_{\mu}V^{\mu}W^{\mu}, or a repeated index appearing twice upstairs, is not a slightly informal notation for something correct. It is a symptom. Either an index has been raised or lowered without saying so, or the expression is wrong.

This is a genuinely useful diagnostic. When you have made an error in a long index computation, the odds are excellent that you can find it without checking any algebra at all, just by scanning for an index that is repeated in the same position or that appears three times. Treat the convention as a type system, and let it fail loudly.

Grind box — the quotient theorem, or how to prove something is a tensor without transforming it

Often you meet an array and want to know whether it is a tensor, and transforming it directly is painful. There is a shortcut, and it is used constantly.

Claim. Suppose KμνK_{\mu\nu} is an array of sixteen numbers per frame, and suppose that for every vector VνV^{\nu} the sixteen-fold sum KμνVνK_{\mu\nu}V^{\nu} is a covector. Then KμνK_{\mu\nu} is a (0,2)(0,2) tensor.

Proof. By hypothesis WμKμνVνW_{\mu}\equiv K_{\mu\nu}V^{\nu} is a covector, so

KμνVν  =  Wμ  =  (Λ1)αμWα  =  (Λ1)αμKαβVβ. K'_{\mu\nu}V'^{\nu} \;=\; W'_{\mu} \;=\; (\Lambda^{-1})^{\alpha}{}_{\mu}W_{\alpha} \;=\; (\Lambda^{-1})^{\alpha}{}_{\mu}K_{\alpha\beta}V^{\beta}.

On the left substitute Vν=ΛνβVβV'^{\nu}=\Lambda^{\nu}{}_{\beta}V^{\beta}:

[KμνΛνβ(Λ1)αμKαβ]Vβ  =  0for every Vβ. \Big[K'_{\mu\nu}\Lambda^{\nu}{}_{\beta} - (\Lambda^{-1})^{\alpha}{}_{\mu}K_{\alpha\beta}\Big]V^{\beta} \;=\; 0 \qquad\text{for every }V^{\beta}.

A bracket that annihilates every vector is zero, as you can see by taking VV to be each of the four basis vectors in turn. Contract the resulting identity with (Λ1)βσ(\Lambda^{-1})^{\beta}{}_{\sigma} to free the index ν\nu, and you get

Kμσ  =  (Λ1)αμ(Λ1)βσKαβ, K'_{\mu\sigma} \;=\; (\Lambda^{-1})^{\alpha}{}_{\mu}(\Lambda^{-1})^{\beta}{}_{\sigma}K_{\alpha\beta},

the (0,2)(0,2) law. \blacksquare

The same argument works for any index structure. Why it matters: this is how you certify that the objects physics hands you are tensors. In Chapter 2.6 the field FμνF^{\mu\nu} is identified by the requirement that fμ=qFμνuν/cf^{\mu}=qF^{\mu\nu}u_{\nu}/c be a four-vector for every four-velocity uνu_{\nu}. The quotient theorem then forces FμνF^{\mu\nu} to be a tensor, which is most of the work of that chapter done in advance.

One caution, and it is the whole content of the phrase "for every": the hypothesis must hold for all vectors, not merely for the one you happen to care about. In §1, ημνNμpν\eta_{\mu\nu}N^{\mu}p^{\nu} was a number for each particular pp, and that never made NμN^{\mu} a vector.

In plain terms 2.4.5

A tensor with several slots is nothing more than those two objects layered, and the layering introduces no new idea. Whether a given index sits upstairs, behaving like an arrow, or downstairs, behaving like a stack of sheets, its transformation rule is the basic rule applied to each index in turn, one at a time and in any order. Nothing conceptually new arrives with the extra slots; the same idea is being stacked.

What genuinely matters is understanding which mathematical operations are safe to perform. You may safely add tensors of identical type, and you may safely multiply them together. You may also safely contract them, which means pairing one upper index against one lower index so that both are consumed and the resulting object is simpler than what you started with.

That is why the rules of this mathematics always demand pairing an up with a down. Summing two upper indices together is not a breach of house style, and calling it bad form understates the damage. It produces a statement meaning different things to different observers, which is another way of saying it means nothing at all, and the notation is built so that this particular failure is visible at a glance.

6 · The whole point, in one theorem

Everything so far has been apparatus. Here is what it was for.

Theorem — tensor equations are frame-independent

Let SS and TT be tensors of the same type (p,q)(p,q). If

Tμνρσ  =  Sμνρσ T^{\mu\nu\ldots}{}_{\rho\sigma\ldots} \;=\; S^{\mu\nu\ldots}{}_{\rho\sigma\ldots}

holds in one coordinate system, it holds in every coordinate system.

Proof. Define DTSD\equiv T-S. By operation (1) of §5.1, DD is a tensor of type (p,q)(p,q). The hypothesis says every component of DD vanishes in the unprimed frame. Now apply the transformation law. Each primed component of DD is a sum of terms, and every term contains a factor Dαβγδ=0D^{\alpha\beta\ldots}{}_{\gamma\delta\ldots}=0. A sum of terms each containing a factor of zero is zero. Hence D=0D'=0 in the primed frame too, which is to say T=ST'=S'. Since the primed frame was arbitrary, the equation holds everywhere. \blacksquare

That proof is four lines long and it is not deep. The transformation law is linear and homogeneous, meaning that every term has exactly one factor of DD and there is no constant term, so the law maps zero to zero. That is the entire mechanism. What is remarkable is not the proof but the consequence.

6.1 · What this buys

Consider what you would otherwise have to do. You have a candidate law of physics, and you want to know whether it is consistent with relativity, which means asking whether an observer sailing past at 0.9c0.9c would write the same law. The honest procedure is to express everything in the moving observer's coordinates, substitute, grind, and see whether the result has the same form. That is a long calculation, it must be redone for every law, and it is easy to get wrong.

The theorem replaces all of that with an inspection. If both sides of your equation are tensors of the same type, meaning the free indices match up with up and down with down on both sides, then you are done. There is no calculation to do. The frame-independence is visible in the shape of the equation. This property has a name, manifest covariance, and "manifest" is the operative word. The law is not merely true. It is true in a way you can see.

So the rest of this book has a writing convention that is also a physical principle:

The rule the rest of the book obeys

Every fundamental law is written as an equation between tensors of the same type. Anything that cannot be written that way is either not fundamental, or not right, or is hiding a choice of frame that will eventually cost you.

6.2 · The scoreboard

F=ma\vv F = m\vv a is not a tensor equation. Chapter 1.1 §4.2 already showed the damage from a different direction. Newton's second law changes form under an innocent change of spatial coordinates, sprouting terms that look like forces and are not. Relativity makes the failure sharper. The left side is a three-component object. The right side involves d2x/dt2\dd^{2}\vv x/\dd t^{2}, in which tt is a coordinate rather than a scalar, and Chapter 2.2 established that tt is not shared between frames. There is no way to read (2.4.9) as acting on this equation, because the equation is not built from objects the matrix knows how to act on. The trouble is not that F=ma\vv F=m\vv a is false. At 106c10^{-6}c it is superb. The trouble is that its content changes when you change frames, and that disqualifies it as a statement about nature rather than about a laboratory.

dpμdτ=fμ\dfrac{\dd p^{\mu}}{\dd\tau}=f^{\mu} is one. Both sides are (1,0)(1,0). Here pμp^{\mu} is a vector and dτ\dd\tau is a scalar, so the derivative is a vector, and fμf^{\mu} is defined to be whatever four-vector sits on the right. Chapter 2.5 builds it and shows that its spatial part reduces to F\vv F at low speed, so nothing is lost. The old law is recovered as a limit, exactly as 12mv2\tfrac12mv^{2} was recovered in Chapter 0.1.

μFμν=μ0Jν\partial_{\mu}F^{\mu\nu}=\mu_{0}J^{\nu} is one. Both sides are (1,0)(1,0), the left-hand one after its contraction. That single equation contains two of Maxwell's four equations. Chapter 2.6's punchline is that writing electromagnetism this way is not a repackaging but an explanation. The theory was relativistic before relativity was invented, which is why it broke Galilean invariance in Chapter 2.1.

Gμν=8πGc4TμνG_{\mu\nu}=\dfrac{8\pi G}{c^{4}}T_{\mu\nu} is one. Both sides are (0,2)(0,2) and symmetric. Chapter 3.6 derives it. The reason it can be guessed before it is derived is this chapter. Demand a (0,2)(0,2) symmetric tensor built from the metric and its first two derivatives, and there are almost no candidates. Tensor structure does not merely check laws. It drastically narrows the field of possible ones, and that is the method the whole second half of this book runs on.

⚠ Two things the theorem does not say

It does not say tensor equations are true. pμ=7uμp^{\mu}=7\,u^{\mu}, in whatever units you like, is a perfectly good tensor equation, and it is false for every particle whose rest mass is not 77. Covariance is a constraint on the form of a law, not evidence for its content. It narrows the search. Experiment still decides.

It does not license equations between components. "T00=5T^{00}=5" is not a tensor equation. The left side is one component of a tensor, the right side is a number, and they do not transform alike. Neither "V1=V2V^{1}=V^{2}" nor "p=0\vv p=0" is a tensor equation either. Such statements can be perfectly true, since a particle really can be at rest. But they are statements about a frame, they have to be labelled as such, and they must never be carried across a boost unexamined. That is precisely the sin of (2.4.1).

In plain terms 2.4.6

What all of this machinery is for is a single check, run quickly and run often: whether a proposed law of physics is true for everybody or is a quirk of one laboratory. The check has to be cheap, because it is going to be run on every equation in the rest of the book.

Without such a system you would have to rewrite every equation from the viewpoint of a moving observer, grind through the algebra and see whether the structure survived, and then repeat that exercise for every new law anyone proposed. Tensor notation replaces that calculation with an inspection of the page. If both sides of an equation are the same type of tensor, carrying matching indices at the same heights, the law holds in every frame and nothing further need be checked. The frame-independence has become visible in the shape of the equation itself, before a line of computation is performed.

This turns what looks like a writing convention into a design rule for physical law. It also explains the fate of Newton's second law, the one relating force to mass and acceleration. That law was never wrong, and at everyday speeds it remains superb. But its content changes with the observer's frame, and that alone disqualifies it as a fundamental statement about nature rather than a description of one laboratory.

7 · Symmetry, antisymmetry, and a promise

Chapter 0.7 §4.3 proved that any square matrix splits uniquely into a symmetric and an antisymmetric part, and then spent that result. Applied to the Jacobian of a velocity field, the symmetric part is strain and the antisymmetric part is rotation, at twice the local angular velocity. The same split applies to any rank-2 tensor, word for word.

Tμν  =  12(Tμν+Tνμ)T(μν)  +  12(TμνTνμ)T[μν], T^{\mu\nu} \;=\; \underbrace{\tfrac12\big(T^{\mu\nu}+T^{\nu\mu}\big)}_{\displaystyle T^{(\mu\nu)}} \;+\; \underbrace{\tfrac12\big(T^{\mu\nu}-T^{\nu\mu}\big)}_{\displaystyle T^{[\mu\nu]}}, (2.4.27)

with round brackets denoting the symmetric part and square brackets the antisymmetric part. That is standard notation, and it is worth learning now because Chapter 3.4 uses it heavily. Existence and uniqueness are Chapter 0.7's argument unchanged. Nothing in that argument used three dimensions or a metric.

7.1 · The split is frame-independent

Here is what is new, and it is the reason the split is physics rather than bookkeeping. Suppose Tμν=TνμT^{\mu\nu}=T^{\nu\mu} in one frame, and ask what every other frame sees. Transform both indices and compare.

Tνμ  =  ΛνβΛμαTβα  =  ΛμαΛνβTαβ  =  Tμν, T'^{\nu\mu} \;=\; \Lambda^{\nu}{}_{\beta}\Lambda^{\mu}{}_{\alpha}\,T^{\beta\alpha} \;=\; \Lambda^{\mu}{}_{\alpha}\Lambda^{\nu}{}_{\beta}\,T^{\alpha\beta} \;=\; T'^{\mu\nu}, (2.4.28)

The first equality is the (2,0)(2,0) law with the two dummy indices named β,α\beta,\alpha instead of α,β\alpha,\beta, which is a free choice by §8. The second equality uses only that numbers commute and that Tβα=TαβT^{\beta\alpha}=T^{\alpha\beta} by hypothesis. So symmetry survives the transformation. The identical computation with one minus sign inserted shows that antisymmetry survives too.

Why this matters

"TT is antisymmetric" is a frame-independent property of the object, not an accident of the coordinates you chose. Every observer agrees. So the decomposition (2.4.27) splits a tensor into two pieces that no coordinate change can mix, and pieces that cannot mix are pieces that can carry different physics.

⚠ Only same-position indices may be swapped

The statement "TμνT^{\mu\nu} is symmetric" is meaningful. The statement "TμνT^{\mu}{}_{\nu} is symmetric" is not. You would be comparing TμνT^{\mu}{}_{\nu} with TνμT^{\nu}{}_{\mu}, and those two arrays transform differently, so the comparison holds in one frame and fails in another. If you need such a statement, lower the upper index first with η\eta and then swap. Symmetry is a property of a pair of slots of the same kind.

7.2 · Counting, and what the counts are hiding

How many independent components does each piece have in nn dimensions? A symmetric array is fixed by its diagonal (nn entries) plus its strictly upper triangle (12n(n1)\tfrac12 n(n-1) entries):

#symmetric=n(n+1)2,#antisymmetric=n(n1)2, \#\text{symmetric} = \frac{n(n+1)}{2}, \qquad \#\text{antisymmetric} = \frac{n(n-1)}{2}, (2.4.29)

the antisymmetric count being the strictly upper triangle alone, since the diagonal is forced to zero by Aμμ=AμμA^{\mu\mu}=-A^{\mu\mu} for each fixed μ\mu (no sum). The two add to n2n^{2}, as they must.

Now put numbers in. A coincidence is about to appear, and it has caused a great deal of confusion.

In three dimensions, 1232=3\tfrac12\cdot3\cdot2=3. An antisymmetric 3×33\times3 array has exactly three independent components, which is the same number as a vector. So it can masquerade as a vector, and in three-dimensional physics it invariably does. Chapter 0.7 §4.4 showed you the mechanism. The curl is really the antisymmetric part of a Jacobian, and it is only because 3=12323=\tfrac12\cdot3\cdot2 that we can package it as an arrow at all. The magnetic field is the same trick. B\vv B is not a vector in the way v\vv v is. It is an antisymmetric rank-2 object wearing a disguise that fits only in three dimensions. That is why it behaves oddly under reflection, and why the right-hand rule needs a convention where the laws of physics should not.

In four dimensions, 1243=6\tfrac12\cdot4\cdot3=6. Six, not four, so the disguise fails. An antisymmetric rank-2 tensor in spacetime cannot possibly be a four-vector. It is its own kind of thing, with six slots.

Six independent numbers, sitting in the most natural rank-2 object spacetime admits. If you have been kept waiting since Chapter 0.7 §4.4, you should now be doing the arithmetic: three components of E\vv E, three components of B\vv B, six in total.

The promise

Chapter 2.6 will find exactly six components sitting in an antisymmetric FμνF^{\mu\nu}, and they will be the three components of E\vv E and the three of B\vv B. That comes by derivation rather than by analogy. E\vv E and B\vv B are not two fields that happen to be coupled. They are one tensor, split up differently by different observers, exactly as ρ\rho and J\vv J were one four-vector and tt and x\vv x were one set of coordinates.

Nothing more will be said about it here. But you should now find that outcome unsurprising, and finding it unsurprising in advance is the point of this section.

One more count before we leave the section: 1245=10\tfrac12\cdot4\cdot5=10. A symmetric rank-2 tensor in four dimensions has ten independent components. That is the metric gμνg_{\mu\nu} of Chapter 3.3, and it is why "the gravitational field" turns out to be ten functions rather than one. Problem 4 asks you to do these counts yourself, because doing them is how they stick.

In plain terms 2.4.7

Any object carrying two indices can be neatly divided into two distinct halves: a symmetric part, which remains identical if you swap the order of its indices, and an antisymmetric part, which flips its mathematical sign under that same swap.

The power of this division is that it survives any change of frame. What is symmetric to one observer remains symmetric to all of them, which proves that this is a genuine physical division of the object rather than an artefact of whichever notation someone happened to choose. It is a real seam in the thing itself, which is to say the object has fallen apart into independent pieces, and the two halves go on to have entirely separate careers.

Count how many independent numbers each part requires in four-dimensional spacetime and the symmetric part holds ten, the antisymmetric part six. Both counts turn out to matter enormously. The ten symmetric numbers will become the metric of curved spacetime, which is to say the ten functions that encode gravity itself. The six antisymmetric numbers will turn out to hold the electric and magnetic fields, three components each — the first substantial hint that electricity and magnetism were never two separate forces to begin with.

8 · Notation hygiene

This is a short section with a high yield. These rules are not style preferences. Each one is a consequence of something already proved, and each violation is a detectable error.

Free indices must match on both sides. A free index is one that appears exactly once in a term. It labels which component you are talking about, so both sides must be labelled the same way, with the same letters in the same positions and the same order of appearance. Aμ=BμA^{\mu}=B^{\mu} is fine. Aμ=BμA^{\mu}=B_{\mu} is not (different species, §3.3). Aμ=BμνA^{\mu}=B^{\mu\nu} is not (the right side is sixteen numbers per value of μ\mu). Every term in a sum must carry the same free indices too, so Aμ+BνA^{\mu}+B^{\nu} is meaningless.

Dummy indices are summed and freely renameable. A dummy index appears exactly twice, once up and once down, and is summed. Its name is invisible to the result, so Vμωμ=VαωαV^{\mu}\omega_{\mu}=V^{\alpha}\omega_{\alpha}. Use that freedom aggressively. In a long computation, renaming dummies to keep them distinct is the difference between an answer and a mess.

An index may not appear three times. If it does, either you meant two different indices and reused a letter, or you have a genuine error. There is no valid expression in which a letter appears three times in one term. This rule catches more mistakes than any other, precisely because it is purely mechanical: you can check it without understanding the expression at all.

The Kronecker delta is the mixed-index identity. δμν=1\delta^{\mu}{}_{\nu}=1 if μ=ν\mu=\nu and 00 otherwise. It has one index up and one down because it is a (1,1)(1,1) tensor. Problem 1 asks you to verify that, and the answer is more interesting than it sounds. Its two working properties are that it renames an index, δμνVν=Vμ\delta^{\mu}{}_{\nu}V^{\nu}=V^{\mu}, and that its trace counts dimensions.

δμμ  =  1+1+1+1  =  4. \delta^{\mu}{}_{\mu} \;=\; 1+1+1+1 \;=\; 4. (2.4.30)

Not 11. Getting that wrong is a rite of passage. The number 44 will appear all over Chapter 3.4 and, in the guise of dd, throughout Chapter 5.10's dimensional regularisation.

μ\partial_{\mu} carries a lower index. That is (2.4.11), and it is the most frequently forgotten fact in the subject, so here it is a second time. μ=xμ\partial_{\mu}=\pdv{}{x^{\mu}} carries a lower index even though xμx^{\mu} carries an upper one, because differentiating with respect to a thing inverts how it transforms. If you find yourself writing μ\partial^{\mu}, you must have raised it with the metric, and there is a minus sign on the spatial components waiting for you.

8.1 · The Levi-Civita symbol, and an honest caveat

One more array is worth naming. Define ϵμνρσ\epsilon_{\mu\nu\rho\sigma} to be totally antisymmetric, meaning that it changes sign under the exchange of any two indices, and fix ϵ0123=+1\epsilon_{0123}=+1. Total antisymmetry means it vanishes unless all four indices are different, so its 256256 entries are ±1\pm1 on the 2424 permutations of (0,1,2,3)(0,1,2,3) and zero elsewhere.

How does it transform? We need one fact about any 4×44\times4 matrix MM first, taken straight from Chapter 0.4 §5. The determinant is the alternating multilinear function of the columns normalised by detI=1\det I=1, so contracting a totally antisymmetric array with four copies of MM can only reproduce that array times the determinant.

MαμMβνMγρMδσ  ϵαβγδ  =  det(M)  ϵμνρσ. M^{\alpha}{}_{\mu}M^{\beta}{}_{\nu}M^{\gamma}{}_{\rho}M^{\delta}{}_{\sigma}\;\epsilon_{\alpha\beta\gamma\delta} \;=\; \det(M)\;\epsilon_{\mu\nu\rho\sigma}. (2.4.31)

Now ask what the (0,4)(0,4) tensor law would do to ϵ\epsilon. That law uses four factors of Λ1\Lambda^{-1}, so set M=Λ1M=\Lambda^{-1} in (2.4.31):

ϵμνρσ  =  det(Λ1)  ϵμνρσ  =  1detΛ  ϵμνρσ. \epsilon'_{\mu\nu\rho\sigma} \;=\; \det\big(\Lambda^{-1}\big)\;\epsilon_{\mu\nu\rho\sigma} \;=\; \frac{1}{\det\Lambda}\;\epsilon_{\mu\nu\rho\sigma}. (2.4.32)

By the grind box in §4, detΛ=±1\det\Lambda=\pm1 for Lorentz transformations, so 1/detΛ=detΛ1/\det\Lambda=\det\Lambda and the factor is ±1\pm1. Under boosts and rotations, where det=+1\det=+1, the symbol is genuinely invariant and may be used exactly like a (0,4)(0,4) tensor. Under a spatial reflection, where det=1\det=-1, it flips sign. That is where "pseudovector" and "pseudoscalar" come from, and Chapter 5.5 will find it decisive when the weak interaction turns out to care about the difference.

For a general coordinate change detΛ\det\Lambda is neither ±1\pm1 nor constant. Equation (2.4.32) then says that an ϵ\epsilon with fixed entries is not a tensor at all but a tensor density, an object that picks up a power of the Jacobian determinant. The repair is to multiply by g\sqrt{-g}, where g=detgμνg=\det g_{\mu\nu}, because that factor carries a compensating determinant. We quote the repair here and defer the derivation to Chapters 3.3 and 3.5, where the same factor is also what makes gd4x\sqrt{-g}\,\dd^{4}x the invariant volume element. Within Part II, where detΛ=±1\det\Lambda=\pm1, none of this bites.

In plain terms 2.4.8

The strict formatting rules governing these equations are not a matter of fussiness. They are an error detector built into the notation, and it catches almost everything.

For an equation to be valid, any free indices left over on one side must match, in both number and height, the free indices sitting on the other. Indices that have been summed away are internal bookkeeping and may be renamed freely, but a single index name may never appear three times within the same term.

Each of these rules follows from what the objects are rather than from a taste for tidiness, and each can be re-derived from the transformation law in a line. Taking them seriously catches fatal mistakes at once, with no physics computed at all. Almost every algebraic slip made over the coming chapters will announce itself as a mismatched index long before it can corrupt a final number. That makes the check of index balance a constant reflex, performed in the same way and for the same reasons as checking the units of an answer, and it costs about as much.

9 · Worked examples

Worked example 1 — boosting a four-momentum twice, and the number that will not move

Take pμ=(5,3,0,0)p^{\mu}=(5,3,0,0) in units of MeV/c\mathrm{MeV}/c. Boost it into two different frames and verify pμpμp_{\mu}p^{\mu} by hand each time.

Frame SS. Lower the index by (2.4.23): flip the spatial signs, leaving pμ=(5,3,0,0)p_{\mu}=(5,-3,0,0). Contract:

pμpμ=(5)(5)+(3)(3)+0+0=259=16. p_{\mu}p^{\mu} = (5)(5)+(-3)(3)+0+0 = 25-9 = 16.

So mc=4 MeV/cmc=4\ \mathrm{MeV}/c. For contrast, compute the forbidden pairing μpμpμ\sum_{\mu}p^{\mu}p^{\mu}, with both indices up: 25+9=3425+9=34. Remember that number.

Frame SS': β=3/5\beta=3/5, so γ=(19/25)1/2=5/4\gamma=(1-9/25)^{-1/2}=5/4. Apply (2.4.9) row by row:

p0=γ(p0βp1)=54(5353)=54165=4,p1=γ(p1βp0)=54(3355)=540=0, \begin{aligned} p'^{0} &= \gamma\big(p^{0}-\beta p^{1}\big) = \tfrac54\big(5-\tfrac35\cdot3\big) = \tfrac54\cdot\tfrac{16}{5} = 4,\\[3pt] p'^{1} &= \gamma\big(p^{1}-\beta p^{0}\big) = \tfrac54\big(3-\tfrac35\cdot5\big) = \tfrac54\cdot 0 = 0, \end{aligned}

with p2=p3=0p'^{2}=p'^{3}=0 untouched. So pμ=(4,0,0,0)p'^{\mu}=(4,0,0,0). This is the particle's rest frame, and that is no accident: the particle's velocity in SS is v/c=p1c/E=3/5v/c=p^{1}c/E=3/5, which is exactly the boost we chose. Contract as before.

pμpμ=(4)(4)000=16.   p'_{\mu}p'^{\mu} = (4)(4) - 0 - 0 - 0 = 16. \;\checkmark

The forbidden pairing here gives 16+0=1616+0=16.

Frame SS'': boost SS' again, by β2=4/5\beta_{2}=-4/5, so γ2=5/3\gamma_{2}=5/3.

p0=53(4(45)0)=203,p1=53(0(45)4)=53165=163. \begin{aligned} p''^{0} &= \tfrac53\big(4-(-\tfrac45)\cdot 0\big) = \tfrac{20}{3},\\[3pt] p''^{1} &= \tfrac53\big(0-(-\tfrac45)\cdot 4\big) = \tfrac53\cdot\tfrac{16}{5} = \tfrac{16}{3}. \end{aligned}

Contract, and note that the two large numbers must conspire:

pμpμ=(203)2(163)2=4002569=1449=16.   p''_{\mu}p''^{\mu} = \Big(\tfrac{20}{3}\Big)^{2} - \Big(\tfrac{16}{3}\Big)^{2} = \frac{400-256}{9} = \frac{144}{9} = 16. \;\checkmark

The forbidden pairing gives 400+2569=6569=72.888\tfrac{400+256}{9}=\tfrac{656}{9}=72.888\ldots

The scoreboard.

Framepμp^{\mu} pμpμp_{\mu}p^{\mu} μpμpμ\sum_{\mu}p^{\mu}p^{\mu} (wrong)
SS(5,3,0,0)(5,\,3,\,0,\,0)16163434
SS'(4,0,0,0)(4,\,0,\,0,\,0)16161616
SS''(203,163,0,0)(\tfrac{20}{3},\,\tfrac{16}{3},\,0,\,0)16166569\tfrac{656}{9}

The correctly contracted column is constant and the naively contracted one is not. That column is (2.4.13) and this chapter, in a table.

A check on the framework. Rapidities add, by the grind box in §2. Here ϕ1=artanh35=12ln1+3/513/5=12ln4=ln2\phi_{1}=\operatorname{artanh}\tfrac35=\tfrac12\ln\tfrac{1+3/5}{1-3/5}=\tfrac12\ln 4=\ln 2 and ϕ2=artanh(45)=12ln19=ln3\phi_{2}=\operatorname{artanh}(-\tfrac45)=\tfrac12\ln\tfrac19=-\ln 3, so the composite rapidity is ln2ln3=ln23\ln 2-\ln 3=\ln\tfrac23 and the composite velocity is

βtot=tanh(ln23)=(2/3)21(2/3)2+1=5/913/9=513,γtot=1312. \beta_{\rm tot} = \tanh\Big(\ln\tfrac23\Big) = \frac{(2/3)^{2}-1}{(2/3)^{2}+1} = \frac{-5/9}{13/9} = -\frac{5}{13}, \qquad \gamma_{\rm tot} = \frac{13}{12}.

Boost the original pμ=(5,3,0,0)p^{\mu}=(5,3,0,0) by that single transformation:

1312(5+5133)=13128013=203,1312(3+5135)=13126413=163, \begin{aligned} \tfrac{13}{12}\Big(5+\tfrac{5}{13}\cdot3\Big) &= \tfrac{13}{12}\cdot\tfrac{80}{13} = \tfrac{20}{3},\\[3pt] \tfrac{13}{12}\Big(3+\tfrac{5}{13}\cdot5\Big) &= \tfrac{13}{12}\cdot\tfrac{64}{13} = \tfrac{16}{3}, \end{aligned}

which is pμp''^{\mu} exactly. Two boosts composed equal one boost at the added rapidity, and the invariant is 1616 throughout, as it must be, since it never depended on Λ\Lambda at all.

Worked example 2 — μVμ\partial_{\mu}V^{\mu} is a scalar, and what that does to charge conservation

Show that νVμ\partial_{\nu}V^{\mu} is a (1,1)(1,1) tensor and hence that μVμ\partial_{\mu}V^{\mu} is a scalar. Then read Chapter 0.7's continuity equation in that light.

Step 1. Differentiate the vector law. Using the covector law (2.4.11) for ν\partial'_{\nu} and then the vector law for VμV'^{\mu}:

νVμ  =  (Λ1)βνβ(ΛμαVα). \partial'_{\nu}V'^{\mu} \;=\; (\Lambda^{-1})^{\beta}{}_{\nu}\,\partial_{\beta}\Big(\Lambda^{\mu}{}_{\alpha}V^{\alpha}\Big).

Now use the fact recorded under (2.4.9): for a Lorentz transformation every entry of Λ\Lambda is a constant, so β\partial_{\beta} passes straight through it:

νVμ  =  Λμα(Λ1)βν  βVα, \partial'_{\nu}V'^{\mu} \;=\; \Lambda^{\mu}{}_{\alpha}\,(\Lambda^{-1})^{\beta}{}_{\nu}\;\partial_{\beta}V^{\alpha},

which is exactly the (1,1)(1,1) transformation law. So νVμ\partial_{\nu}V^{\mu} is a tensor: sixteen numbers that transform correctly.

Step 2. Contract μ\mu with ν\nu. By §5.1's theorem, contracting a (1,1)(1,1) gives a (0,0)(0,0), which is a scalar. So

μVμis the same number in every inertial frame. \partial_{\mu}V^{\mu} \quad\text{is the same number in every inertial frame.}

The four-divergence is an invariant. Note what is not claimed: 0V0\partial_{0}V^{0} alone is not invariant, and neither is iVi\partial_{i}V^{i}. Only the sum, with the index correctly paired, is.

Step 3, cashing it in. Chapter 0.7 §6 derived the continuity equation ρt+J=0\pdv{\rho}{t}+\nabla\cdot\vv J=0 from nothing but "stuff is neither created nor destroyed", and noted at the end that ρ\rho and J\vv J assemble into Jμ=(ρc,J)J^{\mu}=(\rho c,\vv J). Take that seriously and expand μJμ\partial_{\mu}J^{\mu}, remembering x0=ctx^{0}=ct so 0=x0=1ct\partial_{0}=\pdv{}{x^{0}}=\tfrac1c\pdv{}{t}:

μJμ=0J0+iJi=1ct(ρc)+J=ρt+J. \partial_{\mu}J^{\mu} = \partial_{0}J^{0} + \partial_{i}J^{i} = \frac1c\pdv{}{t}\big(\rho c\big) + \nabla\cdot\vv J = \pdv{\rho}{t} + \nabla\cdot\vv J.

The factors of cc cancel exactly, which is the arithmetic reason JμJ^{\mu} is defined with the cc in the time slot. So

  μJμ=0   \boxed{\;\partial_{\mu}J^{\mu} = 0\;}

is the same equation as Chapter 0.7's continuity equation. Not an analogue: the same equation, rewritten.

But it now says more than it did. In Chapter 0.7 it was an equation about a particular observer's ρ\rho and J\vv J, in a particular frame, with a time derivative singled out. Here the left side is a scalar and the right side is zero, which makes the whole thing a scalar equation, so by §6 it holds in every frame with the same form. If charge is conserved for one observer it is conserved for all of them, and the two-term split into "rate of change here" and "flow away from here" is only how each observer happens to slice a single four-dimensional statement. Chapter 0.7 could not have told you that, because it did not have the machinery. The whole content of the upgrade is the index placement.

Grind box — where Step 1 breaks, and why Part III needs a new derivative

Step 1 above used one thing beyond the definitions: that Λμα\Lambda^{\mu}{}_{\alpha} is constant, so β\partial_{\beta} passes through it. Suppose it is not constant, as it is not for a change to polar coordinates, or an accelerating frame, or any coordinate patch on a curved manifold. Then the product rule gives an extra term.

νVμ  =  (Λ1)βν[ΛμαβVα  +  (βΛμα)Vαinhomogeneous]. \partial'_{\nu}V'^{\mu} \;=\; (\Lambda^{-1})^{\beta}{}_{\nu}\Big[\Lambda^{\mu}{}_{\alpha}\,\partial_{\beta}V^{\alpha} \;+\; \underbrace{\big(\partial_{\beta}\Lambda^{\mu}{}_{\alpha}\big)V^{\alpha}}_{\text{inhomogeneous}}\Big].

The first term is the tensor law. The second is not proportional to V\partial V at all. It is proportional to VV itself, so the whole expression is not of the form (matrix)×\times(components of V\partial V). νVμ\partial_{\nu}V^{\mu} is not a tensor under general coordinate changes.

Look closely at why this is fatal rather than annoying. The failure is an inhomogeneous term, and §6's theorem needed the transformation law to be homogeneous. With an inhomogeneous term, an expression that vanishes in one frame need not vanish in another. The proof of the central theorem collapses, and with it the guarantee that makes tensor equations worth writing.

The repair is to invent a new derivative μ\nabla_{\mu} which differs from μ\partial_{\mu} by a correction term chosen so that the two inhomogeneous pieces cancel. The correction is called the connection, it is built from derivatives of the metric, and constructing it is Chapter 3.3. Its failure to commute with itself is curvature, which is Chapter 3.4, which is gravity.

So the flat-space simplification that let this chapter be short is exactly the thing Part III gives up, and everything hard about general relativity is downstream of that one product rule.

10 · Your turn

Problem 1 — δ\delta versus η\eta

(a) Take δμν\delta^{\mu}{}_{\nu} to be defined as 11 when μ=ν\mu=\nu and 00 otherwise, in every coordinate system. Show that it really is a (1,1)(1,1) tensor, for any invertible coordinate change and not just Lorentz ones. Note that this is the same style of definition that failed disastrously for NμN^{\mu} in (2.4.1), and explain why it succeeds here.

(b) ημν\eta_{\mu\nu} also has the same components in every inertial frame. Show this is a special property of Lorentz transformations and not a general one, by exhibiting a coordinate change under which the components of η\eta change.

Solution

(a) Apply the (1,1)(1,1) law to the array δ\delta and see what comes out:

Λμα(Λ1)βν  δαβ  =  Λμα(Λ1)αν  =  δμν, \Lambda^{\mu}{}_{\alpha}\,(\Lambda^{-1})^{\beta}{}_{\nu}\;\delta^{\alpha}{}_{\beta} \;=\; \Lambda^{\mu}{}_{\alpha}\,(\Lambda^{-1})^{\alpha}{}_{\nu} \;=\; \delta^{\mu}{}_{\nu},

using (2.4.12) in the last step. The transformation law maps the array δ\delta to itself, so declaring it to be δ\delta in every frame is consistent with the tensor law. And no step used any property of Λ\Lambda beyond invertibility, so this holds for every smooth coordinate change whatsoever.

Why this succeeds where (2.4.1) failed. The two definitions have the same form, which is "these components, in every frame". Only one of them is compatible with the transformation law. Declaring components in all frames is not automatically wrong. It is an extra condition, and it is legitimate exactly when the transformation law happens to reproduce those components. For δ\delta it does. For (1,0,0,0)(1,0,0,0) it does not, since boosting that array gives (γ,γβ,0,0)(1,0,0,0)(\gamma,-\gamma\beta,0,0)\neq(1,0,0,0). The test is mechanical, and you should run it whenever you are tempted to fix components in all frames.

(b) By the same test: (2.4.16), ημνΛμρΛνσ=ηρσ\eta_{\mu\nu}\Lambda^{\mu}{}_{\rho}\Lambda^{\nu}{}_{\sigma}=\eta_{\rho\sigma}, says exactly that the (0,2)(0,2) law maps η\eta to itself. But that identity is the definition of a Lorentz transformation. It is not available for other coordinate changes.

Counterexample: rescale one spatial axis, x1=2x1x'^{1}=2x^{1}, everything else unchanged. Then Λ11=2\Lambda^{1}{}_{1}=2 and (Λ1)11=12(\Lambda^{-1})^{1}{}_{1}=\tfrac12, so

η11=(Λ1)α1(Λ1)β1ηαβ=(12)2η11=14    1. \eta'_{11} = (\Lambda^{-1})^{\alpha}{}_{1}(\Lambda^{-1})^{\beta}{}_{1}\,\eta_{\alpha\beta} = \Big(\tfrac12\Big)^{2}\eta_{11} = -\tfrac14 \;\neq\; -1.

The components changed, as they must. Measuring xx in half-metres has to show up somewhere, and it shows up in the metric. Here is a second and more consequential example. Switch the spatial part to plane polar coordinates, x=rcosθx=r\cos\theta, y=rsinθy=r\sin\theta. Chapter 0.6 §4's grind box computed dx2+dy2=dr2+r2dθ2\dd x^{2}+\dd y^{2}=\dd r^{2}+r^{2}\dd\theta^{2}, so

ημν=diag(1,1,r2,1) \eta'_{\mu\nu} = \mathrm{diag}\big(1,\,-1,\,-r^{2},\,-1\big)

in coordinates (ct,r,θ,z)(ct,r,\theta,z). That array is position-dependent, in flat spacetime, with no gravity anywhere. Here is the punchline worth carrying into Part III: a position-dependent metric does not by itself mean curvature. Telling the two apart requires an object that is zero for the polar metric and nonzero for a genuinely curved one, and constructing it is Chapter 3.4.

Problem 2 — transform, then check

In frame SS, two four-vectors have components Aμ=(3,1,2,0)A^{\mu}=(3,1,2,0) and Bμ=(2,1,0,1)B^{\mu}=(2,1,0,1). Compute ABημνAμBνA\cdot B\equiv\eta_{\mu\nu}A^{\mu}B^{\nu}. Then boost to SS' with β=3/5\beta=3/5, compute all eight primed components, and verify the invariance. Finally compute the naive sum μAμBμ\sum_{\mu}A^{\mu}B^{\mu} in both frames and comment.

Solution

In SS. Lower BB: Bμ=(2,1,0,1)B_{\mu}=(2,-1,0,-1). Then

AB=AμBμ=(3)(2)+(1)(1)+(2)(0)+(0)(1)=61=5. A\cdot B = A^{\mu}B_{\mu} = (3)(2)+(1)(-1)+(2)(0)+(0)(-1) = 6-1 = 5.

Boost, γ=5/4\gamma=5/4, β=3/5\beta=3/5. Only the 00 and 11 components mix:

A0=54(3351)=54125=3,A1=54(1353)=54(45)=1,B0=54(2351)=5475=74,B1=54(1352)=54(15)=14, \begin{aligned} A'^{0} &= \tfrac54\big(3-\tfrac35\cdot1\big) = \tfrac54\cdot\tfrac{12}{5} = 3, & A'^{1} &= \tfrac54\big(1-\tfrac35\cdot3\big) = \tfrac54\cdot\big(-\tfrac45\big) = -1,\\[3pt] B'^{0} &= \tfrac54\big(2-\tfrac35\cdot1\big) = \tfrac54\cdot\tfrac75 = \tfrac74, & B'^{1} &= \tfrac54\big(1-\tfrac35\cdot2\big) = \tfrac54\cdot\big(-\tfrac15\big) = -\tfrac14, \end{aligned}

with A2=2A'^{2}=2, A3=0A'^{3}=0, B2=0B'^{2}=0, B3=1B'^{3}=1 unchanged. So Aμ=(3,1,2,0)A'^{\mu}=(3,-1,2,0) and Bμ=(74,14,0,1)B'^{\mu}=\big(\tfrac74,-\tfrac14,0,1\big). Contract:

AB=(3)(74)(1)(14)(2)(0)(0)(1)=21414=5.   A'\cdot B' = (3)\big(\tfrac74\big) - (-1)\big(-\tfrac14\big) - (2)(0) - (0)(1) = \tfrac{21}{4}-\tfrac14 = 5. \;\checkmark

The naive sum. In SS: 6+1+0+0=76+1+0+0=7. In SS': 214+14+0+0=224=5.5\tfrac{21}{4}+\tfrac14+0+0=\tfrac{22}{4}=5.5. Different. The naive sum is not a scalar, it is not the "length" of anything, and the fact that it is what you would compute in Euclidean space is exactly why the metric has to be written explicitly.

Here is a useful habit. Notice that AμA'^{\mu} and BμB'^{\mu} are individually quite different from AμA^{\mu} and BμB^{\mu}, and that A1A^{1} even changed sign, while the contraction did not budge. That is the shape of every calculation in relativity worth doing.

Problem 3 — symmetric against antisymmetric

Let Sμν=SνμS^{\mu\nu}=S^{\nu\mu} and Aμν=AνμA_{\mu\nu}=-A_{\nu\mu}. Prove that SμνAμν=0S^{\mu\nu}A_{\mu\nu}=0 identically, with no conditions and no special frames. Then say where you expect this to be used.

Solution

Call the sum XSμνAμνX\equiv S^{\mu\nu}A_{\mu\nu}. Both indices are dummies, so rename them, swapping the letters: X=SνμAνμX = S^{\nu\mu}A_{\nu\mu}. That is not a manipulation of the object, only of the notation, and it leaves the same 1616 terms in a different order. Now use the two hypotheses on the renamed expression: Sνμ=SμνS^{\nu\mu}=S^{\mu\nu} and Aνμ=AμνA_{\nu\mu}=-A_{\mu\nu}, so

X  =  SνμAνμ  =  Sμν(Aμν)  =  X. X \;=\; S^{\nu\mu}A_{\nu\mu} \;=\; S^{\mu\nu}\big(-A_{\mu\nu}\big) \;=\; -X.

Hence 2X=02X=0 and X=0X=0. \blacksquare Note that the proof is pure index bookkeeping: no metric, no frame, no dimension. It holds in any number of dimensions and for any pair of contracted slots.

Where it gets used. Constantly, and usually to make an unwanted term disappear.

  • Chapter 2.6: FμνF^{\mu\nu} is antisymmetric, so contracting it with anything symmetric kills the term. In particular μνFμν=0\partial_{\mu}\partial_{\nu}F^{\mu\nu}=0, because μν\partial_{\mu}\partial_{\nu} is symmetric in μν\mu\nu by Clairaut's theorem (Chapter 0.6 §6). Applying ν\partial_{\nu} to μFμν=μ0Jν\partial_{\mu}F^{\mu\nu}=\mu_{0}J^{\nu} then forces νJν=0\partial_{\nu}J^{\nu}=0: Maxwell's equations require charge conservation, they do not merely permit it. That is this one-line lemma doing real work.
  • Chapter 3.3: the connection is symmetric in its lower indices, so contracting it with an antisymmetric object drops out. That is why the exterior derivative of Chapter 3.5 needs no connection at all.
  • Chapter 3.6: the Einstein tensor is symmetric, so only the symmetric part of any candidate source can appear. That is a constraint on what the right-hand side of the field equations is allowed to be.

Problem 4 — counting

In nn dimensions, count the independent components of a symmetric rank-2 tensor and of an antisymmetric one, and verify the two counts sum to n2n^{2}. Specialise to n=4n=4 and identify what each of the two numbers will turn out to be. Then do n=3n=3 and explain the historical accident that makes the magnetic field look like a vector.

Solution

Symmetric. Tμν=TνμT^{\mu\nu}=T^{\nu\mu}, so the array is determined by the entries with μν\mu\le\nu. There are nn diagonal entries and 12n(n1)\tfrac12 n(n-1) strictly-upper ones (half of the n2nn^{2}-n off-diagonal entries), giving

n+n(n1)2=2n+n2n2=n(n+1)2. n + \frac{n(n-1)}{2} = \frac{2n+n^{2}-n}{2} = \frac{n(n+1)}{2}.

Antisymmetric. Aμν=AνμA^{\mu\nu}=-A^{\nu\mu} forces Aμμ=AμμA^{\mu\mu}=-A^{\mu\mu} for each fixed μ\mu (no sum), hence Aμμ=0A^{\mu\mu}=0: the diagonal is empty. Only the strictly-upper triangle survives: 12n(n1)\tfrac12 n(n-1).

Sum. 12n(n+1)+12n(n1)=12n[(n+1)+(n1)]=n2\tfrac12 n(n+1)+\tfrac12 n(n-1)=\tfrac12 n\big[(n+1)+(n-1)\big]=n^{2}. ✓ That is as required, since (2.4.27) writes every rank-2 array as symmetric plus antisymmetric, uniquely, so the two pieces must account for all n2n^{2} numbers exactly once.

n=4n=4. Symmetric: 1245=10\tfrac12\cdot4\cdot5=\mathbf{10}. Antisymmetric: 1243=6\tfrac12\cdot4\cdot3=\mathbf{6}. And 10+6=16=4210+6=16=4^{2}. ✓

The ten is the metric gμνg_{\mu\nu} of Chapter 3.3. The gravitational field is ten functions, which is why general relativity is ten coupled equations and not one. (Four of the ten are pure coordinate freedom, leaving six genuine ones, and that count is Chapter 3.6's business.) The six is the electromagnetic field tensor FμνF^{\mu\nu} of Chapter 2.6: three components of E\vv E and three of B\vv B. You now know in advance that there is exactly enough room and not one slot to spare.

n=3n=3. Symmetric: 66. Antisymmetric: 33. The antisymmetric count coincides with the number of components of a vector. Solve 12n(n1)=n\tfrac12 n(n-1)=n and you get n=3n=3 (or the trivial n=0n=0), so this happens in three dimensions and nowhere else. That coincidence is why the curl of a vector field can be packaged as a vector (Chapter 0.7 §4.4) and why B\vv B is written as one.

The disguise is detectable, though. An antisymmetric object dressed as a vector reveals itself under reflection. Reflect space and a true vector flips its components. These objects do not, because they are built from a determinant-like construction and so carry an extra factor of detΛ=1\det\Lambda=-1, as in (2.4.32). Hence "axial vector", "pseudovector", and the right-hand rule, which is a convention we are forced to state because we have insisted on describing a six-slot object with three-slot notation. In four dimensions the problem does not arise at all, because there 646\neq4 and nobody is tempted.

The brick you just laid

You now have the definition the rest of the book runs on. A tensor of type (p,q)(p,q) is an object whose components transform with pp Jacobians and qq inverse Jacobians. Not a matrix, not an array, not a picture. Upper and lower indices are two different species: vectors transform like dxμ\dd x^{\mu}, covectors like μf\partial_{\mu}f, and they pair to give a number that every observer agrees on because the two Jacobians collapse to a δ\delta. You saw that measured rather than asserted, in the figure, with components moving in opposite directions while their contraction sat at 3.500000003.50000000 and the arrow went on piercing exactly three lines.

You have the metric as the dictionary between the species, Vμ=ημνVνV_{\mu}=\eta_{\mu\nu}V^{\nu}, which flips three signs and is a linear map rather than a typographical convenience. You know that η\eta having fixed components is a defining property of Lorentz transformations rather than a law of nature. You have the three operations that stay inside the class, with contraction proved. And you have the theorem the chapter exists for: a tensor equation true in one frame is true in all of them, because the transformation law is linear and homogeneous and therefore maps zero to zero. That is why every fundamental law from here on will be written with matched indices, and why F=ma\vv F=m\vv a cannot be.

Where this gets spent. Immediately, in Chapter 2.5, which writes dynamics as dpμ/dτ=fμ\dd p^{\mu}/\dd\tau=f^{\mu} and gets E=γmc2E=\gamma mc^{2} out of a contraction. Then Chapter 2.6, where §7's count of six becomes FμνF^{\mu\nu} and E\vv E and B\vv B stop being two fields. Then Part III, where the definition is reused verbatim with one change: Λ\Lambda acquires position dependence. That single change costs you the ordinary derivative (Chapter 3.3's connection), buys you curvature (Chapter 3.4), and produces Gμν=8πGTμν/c4G_{\mu\nu}=8\pi G\,T_{\mu\nu}/c^{4} (Chapter 3.6), which is a tensor equation and could not have been anything else. Chapter 5.2 writes field theory in this language because it must, and Chapter 6.4 writes Yang–Mills in it because the gauge field is a connection in the same sense as Chapter 3.3's.

Chapter 2.5 now takes the definition and does physics with it.