Part II · Special Relativity — Chapter 2.4
Tensors, Honestly
An index is not decoration. It is a promise about what happens when someone else picks different coordinates.
Chapter 2.3 wrote and and and got away with it, because the formulas came out right. This chapter pays for that. Everything in 2.3 was legitimate. None of it was defined. We go back now and supply the definitions, and along the way the index notation stops being a bookkeeping habit and becomes an idea.
Here is the destination, stated in advance so you know what we are aiming at:
A tensor is defined by how its components change when you change coordinates.
Not by a picture, not by "a thing with indices", not by "a matrix with more slots". By a transformation law, and by nothing else. That sounds like a technicality, and it is in fact the whole point, because it is what makes the following true: any equation whose two sides transform the same way is automatically true in every frame if it is true in one. That single sentence is why the rest of this book is written in indices. Section 6 proves it. Everything before Section 6 is the machinery needed to state it, and everything after it is what to do with it.
One more thing to hold onto. The definition we are about to give never mentions Lorentz transformations at all. It works for any smooth change of coordinates: rotations, polar coordinates, an accelerating frame, a coordinate patch on a curved manifold. That generality is not showing off. It is the reason Part III can be about geometry rather than about notation. Chapter 3.2 reuses this chapter's definition word for word with only one change, and that one change is the whole of general relativity.
Tools you'll need. Chapter 0.4 §4: change of basis, the fact that components transform with while basis vectors transform with , and . That section is the entire skeleton of this one. Chapter 0.6 §4: the gradient is really a one-form, , and . This chapter is the promised payoff of that section, and we pick the thread up by name in §3. Also 0.6 §5, the multivariable chain rule, which does almost all of the work below. Chapter 0.7 §4.3, the unique split of a matrix into symmetric plus antisymmetric parts, and §6, the continuity equation. Chapter 2.2 for the boost and rapidity. Chapter 2.3 for the interval, , and the statement that a four-vector is something transforming like , which is the statement we are about to turn into a definition.
1 · The problem, stated sharply
Start with the demand that the whole chapter answers to. Physics is not allowed to depend on the coordinates a human being chose. Nothing about the electron knows whether you set up your axes pointing north or north-north-east, or whether you are drifting past the laboratory at . Whatever is real must be the same for everyone.
But components are not the same for everyone, and nothing could make them so. That is what components are: numbers you get by comparing a physical thing to a set of reference directions you invented. Chapter 0.4 §4 made this precise for an ordinary vector. Change the basis by and the components change by . Chapter 2.2 made it precise for spacetime, where changing to a moving frame reshuffles the four numbers by a boost.
So we have a tension. The things we can write down are components, and those are not invariant. The things we want to assert are laws, and those must be. The resolution is not to abandon components. It is to insist on a particular class of objects.
Work only with objects whose components change in a controlled, known way, meaning a way fixed entirely by the coordinate change and not at all by the object. Then build statements out of them in which the changes cancel. The pieces are frame-dependent. The statement is not.
That is a design constraint, and like all good design constraints it rules things out. Let's watch it rule something out, because here the failure teaches more than the success would.
1.1 · Four numbers that are not a four-vector
Here is an object that has an index and is not a tensor. Define
Nothing in that definition is ill-formed. It is a perfectly good rule: hand me a frame, I hand you four numbers. You could imagine it standing for "the direction of time". You could also imagine it standing for the rest frame of the ether, which is what people historically did, and which is exactly the mistake Chapter 2.1 dismantled. It has an upper index, it has four components, and it looks like everything else we write. Watch what happens.
Take a particle with four-momentum , a genuine four-vector. Build the number
It looks like a scalar, with the indices all matched and summed and nothing left dangling. So let's evaluate it in two frames using specific numbers. Let in units of , and let move at , so . Chapter 2.2's boost gives . We will grind that arithmetic out in §9, and in the meantime you can check it in your head: . Now feed both sets of components into the recipe.
Two observers, one recipe, two answers. Look carefully at what did not go wrong. No algebra was botched and no sign was dropped. The recipe is unambiguous, and both computations are correct. What failed is the claim that the recipe defines a number. It defines a number per frame, which is to say it defines nothing at all. You cannot write "" and mean anything by it.
The damage gets worse if you write an equation rather than a single number. Consider what happens to this one.
In frame , where our particle is at rest, that equation is true: and . Now suppose were a four-vector. The theorem of §6 would then let us conclude that (2.4.4) holds in every frame. And since has been declared to be everywhere, that would say everywhere. This particle would be at rest for every observer, including the one watching it go past at . The conclusion is false, so one of the premises must be false. It is not the theorem. It is the assumption that is a four-vector, and the theorem's one hypothesis is exactly the one fails.
The array is a perfectly respectable set of components. There is a genuine four-vector with those components, and in fact there are many. The four-velocity of any observer, divided by , has exactly those components in that observer's own frame. The error in (2.4.1) is not the numbers. The error is the phrase "in every frame."
Specify a four-vector by giving its components in one frame and you have said something. The transformation law then tells everyone else what they will measure, and nobody is free to disagree. Specify it instead by decreeing the components in all frames at once and you have specified the object twice over. Fixing four components in one frame uses up every freedom the object has, and the transformation law then determines what they must be everywhere else. Declaring them to be again in a frame moving at along contradicts that prediction, , in two of its four components. In every frame but the one you started in, it contradicts the prediction in something.
This is the single most common way that beginners' index expressions go wrong, and it survives into professional work in subtler disguises. Any time you find yourself saying "in this frame the components are simply ..." and then using that form in another frame, you have made this mistake.
So here is the slogan for the chapter. Read it as a definition rather than as a warning.
A tensor is not "a thing with indices". It is a thing with indices that transform correctly. The indices are not a filing system for numbers. They are a claim, and the claim is checkable.
Numbers on their own carry no context. A set of components means something only relative to the framework used to measure them, which is to say relative to a choice of axes, so shifting that choice must shift the numbers describing one unchanged reality. That is not a defect in the description. It is what describing anything physical consists of.
The fatal error is the opposite instinct, which is to freeze one set of numbers in place and require everybody to use it. Doing so does not make the object more real. It produces observers computing conflicting values for what was supposed to be a single fact about the world, and this section stages exactly that collapse with four entirely innocent-looking numbers.
So the goal of everything that follows is not to hunt for quantities that refuse to change, since almost nothing does. It is to find quantities that change in a fully predictable way when the choice of axes changes, so that the physical statements assembled out of them come out the same for everybody, whatever axes each of them chose.
2 · Vectors, defined by how they transform
Now let's build the definition. What we need first is a prototype: one object whose transformation law we can compute rather than choose. There is exactly one natural candidate, and Chapter 2.3 already leaned on it.
Let be one set of coordinates and another, related by four functions that are smooth and invertible. Take two nearby events separated by an infinitesimal displacement , and ask what that separation is in primed coordinates. The answer is not a matter of convention. It is the multivariable chain rule of Chapter 0.6 §5, applied to each of the four functions in turn.
summed over . Nothing was assumed there. Both sides are the same physical displacement between the same two events, described twice over. That array of partial derivatives is going to appear on every page from here on, so let's give it a name.
This is the Jacobian matrix of Chapter 0.6 §8, wearing index notation. The upper index labels the row and the new coordinate. The lower index labels the column and the old one. Keep the horizontal offset in , because it records which slot comes first, and by §4 that will start to matter.
Now make the leap that organises everything after it. We have one object, , that we know transforms by (2.4.6). We define the class of objects that behave the same way.
A contravariant vector (or just vector, or in four dimensions a four-vector) is an object with four components per coordinate system, such that under any change of coordinates
That is the entire definition. There is no additional requirement, and no picture is part of it.
Two things should feel unsatisfying about that definition, and both are worth confronting before we go on.
"It's circular. You defined vectors by copying ." Yes, deliberately. The definition is ostensive: it points at a known object and says "these ones". What makes it non-vacuous is that many other objects turn out to belong to the class, and their membership is a theorem rather than a stipulation. Take three of them in turn. The four-velocity belongs, because is a scalar (Chapter 2.3), and dividing a vector by a scalar leaves the law (2.4.6) intact. The four-momentum belongs, since is a scalar by construction. In this book always means rest mass, the number every observer computes from . The current belongs as well, and §9 shows what that buys. Each of those is a claim about physics, checkable and falsifiable, and each of them fails for of (2.4.1).
"Where did the arrow go?" Nowhere. It is still available, and Chapter 0.4 tells you exactly how to recover it. There, a vector was an element of a vector space and its components were in a basis . Changing basis by sent the components to . Compare that with the definition above and read off the dictionary between the two accounts.
The components transform with and the basis vectors transform with , so their product, which is the actual arrow, does not move at all. Chapter 0.4 §4 put it in one sentence: make your ruler twice as long and every measurement in rulers is halved. "Contra-variant" is literally that. The components vary contrary to the basis. So the arrow is exactly as real as it ever was. What has changed is which end of the relationship we take as the definition. We take the components' end, because that is the end that still makes sense when there is no global vector space to draw arrows in. From Chapter 3.2 onward, that is always the situation.
2.1 · What is for a boost
Before going further, let's see what all of this looks like for the one transformation we already know. Specialise to the boost of Chapter 2.2, with moving along at speed and . Chapter 2.2 gave us the coordinate change itself.
What we want from it is , and getting it means differentiating each new coordinate with respect to each old one. The coefficients here are constants, so the partial derivatives are just those constants.
That is the matrix Chapter 2.3 wrote down. It is no longer a postulate. It is a computed Jacobian, and every entry of it is independent of position. That constancy is a special feature of Lorentz transformations rather than a general fact about coordinate changes. Hold onto the observation. It is the hinge on which §9's second worked example turns, and it is the single technical difference between this chapter and Chapter 3.3.
Grind box — the inverse boost, and why rapidity is the sane variable
Invert (2.4.8) by solving for and . Multiply the first by and the second by and add:
using , which is the definition of rearranged. The same manoeuvre gives . So
which is the physically unsurprising statement that if moves at relative to then moves at relative to . Here that statement is derived rather than assumed.
Chapter 2.2 introduced the rapidity with , whence and (divide by to see the first). In that variable
and the hyperbolic addition formulas give, in two lines of matrix multiplication,
Rapidities add. That is why is the right variable and is not. It is also a consistency check on the whole framework, because the chain rule guarantees that composing coordinate changes composes their Jacobians,
so the transformation law of §2 is automatically consistent under composition. Nothing extra had to be imposed. Chapter 6.1 will call this "the transformations form a group and the components carry a representation of it". Here it is nothing more than the chain rule.
Objects here are defined by their behaviour rather than by their appearance, which feels backwards the first time it is met. Nothing dictates what a vector is made of or what it should look like, and the definition says only how its description must adapt when the frame of reference changes.
The prototype for that behaviour is the physical gap between two nearby events. Since the two events happen whether or not anybody is watching, the rule converting one observer's measurement of that gap into another's is fixed by the change of frame alone and by nothing else. Any object transforming by that same rule is a vector, and there is no further qualification to meet.
This behaviour-based definition is worth its initial awkwardness because it abandons any need for flat backgrounds, straight arrows, or a fixed grid on which to draw them. Nothing in it depends on the geometry being simple. That is precisely why it will survive perfectly intact into Part III, when the coordinates begin to curve and the comfortable picture of an arrow stops making sense.
3 · Covectors — the other kind of index
Chapter 0.6 §4 ended with a promise: differentiation naturally produces lower indices, vectors naturally carry upper ones, the metric converts between them, and when the metric is the identity you cannot see the conversion happening. It said that Chapter 2.4 would make this a formal convention. Here we are. This section builds the second species of index, and it builds it without using a metric at all. The metric does not arrive until §4.
Take a scalar field , meaning a function on spacetime whose value at an event is the same number for everyone, so that . Temperature at a point, say, or the proper time elapsed along some fixed worldline. Its four partial derivatives
are perfectly good components of something. The question is what. We want their transformation law, so apply the chain rule to regarded as a function of the primed coordinates.
Let's look at what appeared. Not , but , the derivatives of the old coordinates with respect to the new. That is the inverse Jacobian. Calling it the inverse is a claim about matrices, so let's establish it. We can do that without computing anything, again by the chain rule.
since the primed coordinates are independent of one another: and , and so on. So write . The notation is honest. That array really is the matrix inverse of (2.4.6), and (2.4.12) reads .
So transforms with the inverse of , while transforms with itself. These are different laws. They give different numbers. Objects that obey the inverse law deserve their own name and their own index position.
A covector is an object with four components per coordinate system such that under any change of coordinates
We write its index down. The gradient of any scalar is the prototype, exactly as was the prototype for vectors.
3.1 · Why the two species exist: the pairing
Now for the structural fact that justifies having invented a second kind of object at all. Take a vector and a covector and contract them, which means summing over the repeated index with one copy up and one down. Watch what the two transformation laws do to each other.
The two Jacobians met and collapsed to a Kronecker delta, and the delta then did nothing except rename an index. So takes the same value in every coordinate system. It is a scalar, meaning a real number rather than a number-per-frame. It is the first genuinely invariant thing we have built, and building it took no metric.
Let's read the two definitions again in the light of that calculation. The vector law and the covector law are not two arbitrary conventions that happen to be inverse to one another. Covectors are defined to be the objects that pair with vectors to give invariants. That is what they are for, and the inverse Jacobian is forced on them by that job. In Chapter 0.4's language, the covectors at a point form the dual space , which is the space of linear maps from vectors to numbers.
3.2 · The picture: arrows and stacks
There is a genuine geometric image behind all of this, and it is not the one you would guess.
A vector is an arrow. A covector is not an arrow. A covector is a machine that eats an arrow and returns a number, linearly, and the honest picture of such a machine is a family of evenly spaced parallel surfaces. They are the level sets of the linear function it computes. In two dimensions they are parallel lines. In four they are parallel hyperplanes. A "big" covector has closely spaced surfaces and a "small" one has widely spaced surfaces, so the covector's magnitude lives in the density of the stack. That is why doubling halves the spacing.
With that picture in hand, the contraction acquires a completely concrete meaning.
is the number of -surfaces the arrow pierces.
An arrow either crosses a given surface or it does not. No coordinate system has an opinion about that. This is why the pairing is invariant, in one sentence and with no algebra.
For the gradient, this is a picture you have already seen without knowing what you were looking at. The level sets of are the contours of , and is the number of contour lines you cross walking along , which is the change in . Chapter 0.6's contour figure was a picture of a one-form all along.
3.3 · Why nobody told you this before
Because in Euclidean space with Cartesian coordinates the two species have numerically identical components, which makes the distinction invisible. Chapter 0.6 §4 showed why. Converting between them requires the metric, and in Cartesian coordinates on flat Euclidean space , so the conversion multiplies everything by . Every dot product you have ever computed as was silently a contraction of a vector with a covector, and you were spared the bookkeeping because the metric was doing nothing.
The invisibility ends the moment either of two things happens, and in relativity both of them do. Use non-Cartesian coordinates and stops being , which is where Chapter 0.6 §4 got the polar gradient's notorious . Or use a metric that is not positive-definite, such as , and the conversion starts flipping signs. That second case is what §4 is about.
and are not two notations for the same thing. They live in different vector spaces, obey different transformation laws, and have different geometric pictures. An expression like is not merely bad style. It is meaningless, in the way that "3 metres + 3 kilograms" is meaningless. The sum would be one thing in one frame and something else in another.
The reason this is worth being strict about now is that in Part III the metric acquires position dependence, and then no coordinate system anywhere makes the two coincide. Any habit of thought that relies on "they're the same really" fails there, permanently and silently.
Spacetime turns out to be inhabited by two distinct but complementary types of object. The first is the ordinary vector, which you can visualise as an arrow representing a physical displacement, pointing from here to there. The second is the covector, which is better understood as a measuring device than as a thing being measured. Rather than an arrow, picture a covector as a series of evenly spaced parallel sheets, much like the contour lines on a map.
Pair the two together and the covector counts how many of its sheets the arrow punches through, and that count is the whole of the pairing. The single image explains why their numbers move in opposite directions whenever the units change. If you stretch your measuring ruler, the arrow's numerical components shrink, and so the covector's sheets must crowd closer together in order that the total count of pierced sheets comes out exactly as it did before. That count is an indisputable physical fact, and no choice of ruler is entitled to alter it.
The only reason these two species are so frequently conflated is that their numbers happen to mirror each other perfectly in a simple, flat, Cartesian grid. That agreement is a mathematical coincidence of the flat ruler rather than a law of nature, and it shatters the moment either the geometry or the coordinate system becomes interesting.
4 · The metric as the dictionary
So far there has been no metric. Vectors, covectors and the pairing between them all exist without one. Now we add the object that Chapter 2.3 built out of the invariance of the interval, and look at exactly what it buys.
Chapter 2.3's central claim was that the quantity
is the same for all inertial observers. Let's write that out as an equation between two computations of the same thing, using (2.4.6) on the primed side.
Since is arbitrary, the coefficients at the two ends must agree. (They must agree after symmetrising, and is already symmetric.) That gives us the defining property of a Lorentz transformation.
In matrix form that reads . Let's pause on what it is saying. The Lorentz transformations are exactly the linear maps that preserve , in the same way that rotations are exactly the linear maps preserving , which is the condition . Chapter 2.3 derived the boost from the two postulates. Equation (2.4.16) is that same result restated as an algebraic condition, and it is the form Chapter 6.1 generalises when it names the group .
4.1 · Lowering
Define, for any vector ,
The name for that operation is "lowering the index". Writing the result with a down index carries an implicit claim, which is that is a genuine covector obeying the law of §3. That is a claim, so let's prove it. Start from the definition in the primed frame and push.
We want that to equal . So the whole question is whether the following identity holds.
Take (2.4.16) and contract it with on the index :
which is (2.4.19) with the free indices renamed and the symmetry used on both sides. So yes: is a covector, and (2.4.17) is a legitimate map from one species to the other. Note precisely what made it work. It was (2.4.16), the invariance of . Lowering with an arbitrary array of numbers would not have produced a covector.
4.2 · Raising
Going the other way needs the inverse of . Since is a non-degenerate matrix it has one. Call its entries , defined by
For our signature the inverse is easy to find, because squares to the identity. So as an array of numbers, the same four numbers down the diagonal. That coincidence is a property of this particular metric and not a general fact. In Chapter 3.3 the inverse metric is a genuinely different array from , and confusing the two is a classic and expensive error. With the inverse in hand, we can raise an index.
That is consistent, since . Raising after lowering returns you to where you started, which is the least one can ask of a dictionary.
4.3 · What it actually does to the numbers
Let's see what lowering does to actual numbers. Write (2.4.17) out component by component. Since is diagonal, only one term survives in each sum.
So with signature , lowering flips the sign of the three spatial components and leaves the time component alone. Take the four-momentum of §1 as a concrete case,
and the invariant length is then the contraction of the two arrays against each other,
so . Note that is a contraction of a covector with a vector, so by (2.4.13) it is invariant with no further argument needed. Section 9 checks that by brute force in three different frames.
The notation is so light that it invites you to think and are the same object written two ways, and that shifting the index up or down is a typographical act. It is not. It is the application of a specific linear map, namely the matrix , and that map is not the identity. With our signature it multiplies three of the four components by .
This is the single most productive source of sign errors in relativity, and they are the worst kind of error: dimensionally consistent, index-consistent, and off by a sign. Three habits protect you.
- Never write an expression with two indices in the same position summed. is not a thing. It is not " in a hurry". It is a frame-dependent number, as §9 measures. If you want , write or , which are the same thing.
- Track where the minus signs are. . The three-vector dot product enters with a minus. Every time.
- is not with the index moved. It is , so but . Chapter 2.6's has signs in it that come from precisely here.
Grind box — index gymnastics, and a bonus identity
Three working rules, each a one-line consequence of (2.4.21).
(i) You may slide a contracted pair up and down together.
So : it does not matter which factor carries the lowered index, only that exactly one does.
(ii) with one index up and one down is the Kronecker delta. By (2.4.21), . So you never need to write at all. It is .
(iii) The metric knows the inverse transformation. Define the only way the rules allow, by lowering the first index and raising the second: . Now contract the identity (2.4.20), namely , with .
Hence . Lowering and raising the indices of produces its inverse. This is why most books never write at all. They write and let the index positions carry the information. We have kept explicit up to this point, because the covector law of §3 is true with no metric in sight and it would have been dishonest to define it using one.
(iv) . Take determinants of . By Chapter 0.4 §5, is multiplicative and , so , and gives . Boosts and rotations have . A spatial reflection has . That sign will matter exactly once in this chapter, in §8.
Two last remarks on , both pointing forward.
It is what makes "length" mean anything. Without a metric there is no way to ask how long a four-vector is. You can pair a vector with a covector, but not a vector with a vector. supplies exactly the missing ingredient, and the price is that "length squared" can come out positive, negative or zero. That is the timelike, spacelike and null trichotomy of Chapter 2.3. The metric does not merely measure. It is what defines the causal structure.
It is about to become a field. Everything in this section used only that is a symmetric, invertible, position-independent array satisfying (2.4.16). Now drop the words "position-independent", and replace by , ten functions of position. Every formula above survives unchanged. What changes is that becomes position-dependent, and then exactly one thing breaks. What breaks is the derivative, as §9's second worked example shows. Repairing it is Chapter 3.3, and the repair is the gravitational field. That is the entire difference between special and general relativity, and it is a difference in one adjective.
You have fitted a thousand times, and (2.4.17) has been in it the whole while under another name.
Report creatinine in rather than and a column is rescaled, . The fitted coefficient moves the other way, , and the product , which is the only part anybody ever uses, is untouched. That is §3's arrow and stack of sheets exactly. The data column carries an upper index, the coefficient carries a lower one, and the contraction (2.4.13) is the invariant. A predicted value is a scalar. A coefficient is a component.
The consequence is a mistake that appears in print every week. Answering "which predictor matters most" by ranking raw coefficients is the precise error this chapter's notation exists to catch, since the ranking changes when somebody changes a unit, and a statement that does that was never a statement about the patients.
The standard repair is standardisation, with the sample standard deviations. Look at what that operation is. It converts a lower index into an upper one by contracting with a symmetric array assembled out of the data's own second moments, which is (2.4.17) run backwards with the covariance structure playing the part of the metric. This section's warning then applies unchanged: the translation depends on which converter you hand it. Standardise against the sample's spread, against a reference population's spread, or against an interquartile range, and the most important predictor can change. It changes not because the biology moved but because a different metric was used.
What breaks, and it is the interesting half. Nothing obliges a statistician to be frame-independent. Choosing whichever standardisation makes a table readable is legitimate, and the choice gets reported rather than derived. In physics the converter is not chosen. The metric is handed over by the world, it is the same one for every observer, and §6's theorem then makes the frame-independence of an entire equation checkable by inspection. No such theorem exists for a regression table, because the demand that produced it was never made.
Because arrows and sheets are genuinely different species of object, something has to translate between them, and the metric is that translator. Feed the metric a displacement arrow and it hands back the particular stack of sheets that measures lengths in the same way. This is the physical reality behind the notation of moving an index up or down: it is a literal act of translation, performed by a specific piece of machinery, rather than a rearrangement of symbols on the page.
The distinction is not a typographical flourish. Depending on the rules of the space you are working in, the translation can flip the mathematical signs of your components, which means that an index moved carelessly is a sign error already in flight, and typically one that will not be discovered for several lines yet.
Most importantly of all, the translation depends entirely on which metric you hand it. Here that metric is a rigid, unchanging table of values, identical at every point in spacetime. Later it will evolve into a flexible field that varies from place to place, taking different values here than it does over there — and that single conceptual shift, from a fixed table to a field, is the entire foundation of gravity.
5 · Tensors in general
Vectors carry one upper index and one factor of . Covectors carry one lower index and one factor of . The generalisation writes itself, and it is a definition rather than a theorem.
A tensor of type has upper and lower indices, hence components per coordinate system, and transforms with one factor of per upper index and one factor of per lower index:
The law is linear and homogeneous in the components of . Remember that. Section 6 is nothing but that one observation.
Here is everything we have met so far, sorted by type.
| Type | Object | Example |
|---|---|---|
| scalar | , , , | |
| vector | , , , | |
| covector | , , | |
| rank-2, both down | , (Ch 3.3) | |
| mixed | , | |
| rank-2, both up | (Ch 2.6), (Ch 3.6) |
A scalar has no indices and therefore no factors of , so , the same number for everyone. That is what the general law says at , and it is why "scalar" and "invariant" are the same word here.
5.1 · The three operations that preserve tensor character
Tensors are useful because you can compute with them without leaving the class. Exactly three basic moves are permitted, and the shortness of that list is itself worth knowing.
(1) Linear combination, same type. If and are both type and are scalars, then is type . That follows at once from the linearity of the transformation law. Apply the law to each term and factor out the common string of s. Note the restriction on the type. You cannot add a to a , for the reason given in §3.3.
(2) Outer product. If is type and is type , then the sixteen numbers form a type tensor. Proof: multiply the two transformation laws together and read off. In general, multiplying a by an component-by-component with all indices distinct gives a .
(3) Contraction. Set one upper index equal to one lower index and sum. This is the interesting one, so we prove it.
If is a tensor of type , then (summed) is a scalar. More generally, contracting any upper index of a type- tensor against any lower index yields a tensor of type .
Proof of the basic case. Transform, then set the indices equal:
It is the same collapse as (2.4.13), and it is the same collapse every time: a and a sharing a summed index annihilate into a . Once you have seen that, you have seen all of tensor algebra. The rest is bookkeeping about which indices are spectators.
Grind box — contraction in the general case, with spectators
Take a tensor and contract against . Define (sum on ). Claim: is a tensor. Write the transformation law and set , calling the common value :
Now the only factors carrying are and , and is summed. Group them:
which is exactly the law. The spectator indices and never participated. They carried their own and through untouched. The argument is identical for any and any pair of indices, which is why nobody writes it out twice.
Two consequences worth stating. Contracting a vector with a covector is the case , which is (2.4.13) again. And is contracted twice down to . That is why it is a scalar. Not because someone declared mass invariant, but because of the index structure.
Because that is the only pairing whose result is a tensor. against collapses to . against does not collapse to anything, and what you get depends on the frame. So an expression like , or a repeated index appearing twice upstairs, is not a slightly informal notation for something correct. It is a symptom. Either an index has been raised or lowered without saying so, or the expression is wrong.
This is a genuinely useful diagnostic. When you have made an error in a long index computation, the odds are excellent that you can find it without checking any algebra at all, just by scanning for an index that is repeated in the same position or that appears three times. Treat the convention as a type system, and let it fail loudly.
Grind box — the quotient theorem, or how to prove something is a tensor without transforming it
Often you meet an array and want to know whether it is a tensor, and transforming it directly is painful. There is a shortcut, and it is used constantly.
Claim. Suppose is an array of sixteen numbers per frame, and suppose that for every vector the sixteen-fold sum is a covector. Then is a tensor.
Proof. By hypothesis is a covector, so
On the left substitute :
A bracket that annihilates every vector is zero, as you can see by taking to be each of the four basis vectors in turn. Contract the resulting identity with to free the index , and you get
the law.
The same argument works for any index structure. Why it matters: this is how you certify that the objects physics hands you are tensors. In Chapter 2.6 the field is identified by the requirement that be a four-vector for every four-velocity . The quotient theorem then forces to be a tensor, which is most of the work of that chapter done in advance.
One caution, and it is the whole content of the phrase "for every": the hypothesis must hold for all vectors, not merely for the one you happen to care about. In §1, was a number for each particular , and that never made a vector.
A tensor with several slots is nothing more than those two objects layered, and the layering introduces no new idea. Whether a given index sits upstairs, behaving like an arrow, or downstairs, behaving like a stack of sheets, its transformation rule is the basic rule applied to each index in turn, one at a time and in any order. Nothing conceptually new arrives with the extra slots; the same idea is being stacked.
What genuinely matters is understanding which mathematical operations are safe to perform. You may safely add tensors of identical type, and you may safely multiply them together. You may also safely contract them, which means pairing one upper index against one lower index so that both are consumed and the resulting object is simpler than what you started with.
That is why the rules of this mathematics always demand pairing an up with a down. Summing two upper indices together is not a breach of house style, and calling it bad form understates the damage. It produces a statement meaning different things to different observers, which is another way of saying it means nothing at all, and the notation is built so that this particular failure is visible at a glance.
6 · The whole point, in one theorem
Everything so far has been apparatus. Here is what it was for.
Let and be tensors of the same type . If
holds in one coordinate system, it holds in every coordinate system.
Proof. Define . By operation (1) of §5.1, is a tensor of type . The hypothesis says every component of vanishes in the unprimed frame. Now apply the transformation law. Each primed component of is a sum of terms, and every term contains a factor . A sum of terms each containing a factor of zero is zero. Hence in the primed frame too, which is to say . Since the primed frame was arbitrary, the equation holds everywhere.
That proof is four lines long and it is not deep. The transformation law is linear and homogeneous, meaning that every term has exactly one factor of and there is no constant term, so the law maps zero to zero. That is the entire mechanism. What is remarkable is not the proof but the consequence.
6.1 · What this buys
Consider what you would otherwise have to do. You have a candidate law of physics, and you want to know whether it is consistent with relativity, which means asking whether an observer sailing past at would write the same law. The honest procedure is to express everything in the moving observer's coordinates, substitute, grind, and see whether the result has the same form. That is a long calculation, it must be redone for every law, and it is easy to get wrong.
The theorem replaces all of that with an inspection. If both sides of your equation are tensors of the same type, meaning the free indices match up with up and down with down on both sides, then you are done. There is no calculation to do. The frame-independence is visible in the shape of the equation. This property has a name, manifest covariance, and "manifest" is the operative word. The law is not merely true. It is true in a way you can see.
So the rest of this book has a writing convention that is also a physical principle:
Every fundamental law is written as an equation between tensors of the same type. Anything that cannot be written that way is either not fundamental, or not right, or is hiding a choice of frame that will eventually cost you.
6.2 · The scoreboard
is not a tensor equation. Chapter 1.1 §4.2 already showed the damage from a different direction. Newton's second law changes form under an innocent change of spatial coordinates, sprouting terms that look like forces and are not. Relativity makes the failure sharper. The left side is a three-component object. The right side involves , in which is a coordinate rather than a scalar, and Chapter 2.2 established that is not shared between frames. There is no way to read (2.4.9) as acting on this equation, because the equation is not built from objects the matrix knows how to act on. The trouble is not that is false. At it is superb. The trouble is that its content changes when you change frames, and that disqualifies it as a statement about nature rather than about a laboratory.
is one. Both sides are . Here is a vector and is a scalar, so the derivative is a vector, and is defined to be whatever four-vector sits on the right. Chapter 2.5 builds it and shows that its spatial part reduces to at low speed, so nothing is lost. The old law is recovered as a limit, exactly as was recovered in Chapter 0.1.
is one. Both sides are , the left-hand one after its contraction. That single equation contains two of Maxwell's four equations. Chapter 2.6's punchline is that writing electromagnetism this way is not a repackaging but an explanation. The theory was relativistic before relativity was invented, which is why it broke Galilean invariance in Chapter 2.1.
is one. Both sides are and symmetric. Chapter 3.6 derives it. The reason it can be guessed before it is derived is this chapter. Demand a symmetric tensor built from the metric and its first two derivatives, and there are almost no candidates. Tensor structure does not merely check laws. It drastically narrows the field of possible ones, and that is the method the whole second half of this book runs on.
It does not say tensor equations are true. , in whatever units you like, is a perfectly good tensor equation, and it is false for every particle whose rest mass is not . Covariance is a constraint on the form of a law, not evidence for its content. It narrows the search. Experiment still decides.
It does not license equations between components. "" is not a tensor equation. The left side is one component of a tensor, the right side is a number, and they do not transform alike. Neither "" nor "" is a tensor equation either. Such statements can be perfectly true, since a particle really can be at rest. But they are statements about a frame, they have to be labelled as such, and they must never be carried across a boost unexamined. That is precisely the sin of (2.4.1).
What all of this machinery is for is a single check, run quickly and run often: whether a proposed law of physics is true for everybody or is a quirk of one laboratory. The check has to be cheap, because it is going to be run on every equation in the rest of the book.
Without such a system you would have to rewrite every equation from the viewpoint of a moving observer, grind through the algebra and see whether the structure survived, and then repeat that exercise for every new law anyone proposed. Tensor notation replaces that calculation with an inspection of the page. If both sides of an equation are the same type of tensor, carrying matching indices at the same heights, the law holds in every frame and nothing further need be checked. The frame-independence has become visible in the shape of the equation itself, before a line of computation is performed.
This turns what looks like a writing convention into a design rule for physical law. It also explains the fate of Newton's second law, the one relating force to mass and acceleration. That law was never wrong, and at everyday speeds it remains superb. But its content changes with the observer's frame, and that alone disqualifies it as a fundamental statement about nature rather than a description of one laboratory.
7 · Symmetry, antisymmetry, and a promise
Chapter 0.7 §4.3 proved that any square matrix splits uniquely into a symmetric and an antisymmetric part, and then spent that result. Applied to the Jacobian of a velocity field, the symmetric part is strain and the antisymmetric part is rotation, at twice the local angular velocity. The same split applies to any rank-2 tensor, word for word.
with round brackets denoting the symmetric part and square brackets the antisymmetric part. That is standard notation, and it is worth learning now because Chapter 3.4 uses it heavily. Existence and uniqueness are Chapter 0.7's argument unchanged. Nothing in that argument used three dimensions or a metric.
7.1 · The split is frame-independent
Here is what is new, and it is the reason the split is physics rather than bookkeeping. Suppose in one frame, and ask what every other frame sees. Transform both indices and compare.
The first equality is the law with the two dummy indices named instead of , which is a free choice by §8. The second equality uses only that numbers commute and that by hypothesis. So symmetry survives the transformation. The identical computation with one minus sign inserted shows that antisymmetry survives too.
" is antisymmetric" is a frame-independent property of the object, not an accident of the coordinates you chose. Every observer agrees. So the decomposition (2.4.27) splits a tensor into two pieces that no coordinate change can mix, and pieces that cannot mix are pieces that can carry different physics.
The statement " is symmetric" is meaningful. The statement " is symmetric" is not. You would be comparing with , and those two arrays transform differently, so the comparison holds in one frame and fails in another. If you need such a statement, lower the upper index first with and then swap. Symmetry is a property of a pair of slots of the same kind.
7.2 · Counting, and what the counts are hiding
How many independent components does each piece have in dimensions? A symmetric array is fixed by its diagonal ( entries) plus its strictly upper triangle ( entries):
the antisymmetric count being the strictly upper triangle alone, since the diagonal is forced to zero by for each fixed (no sum). The two add to , as they must.
Now put numbers in. A coincidence is about to appear, and it has caused a great deal of confusion.
In three dimensions, . An antisymmetric array has exactly three independent components, which is the same number as a vector. So it can masquerade as a vector, and in three-dimensional physics it invariably does. Chapter 0.7 §4.4 showed you the mechanism. The curl is really the antisymmetric part of a Jacobian, and it is only because that we can package it as an arrow at all. The magnetic field is the same trick. is not a vector in the way is. It is an antisymmetric rank-2 object wearing a disguise that fits only in three dimensions. That is why it behaves oddly under reflection, and why the right-hand rule needs a convention where the laws of physics should not.
In four dimensions, . Six, not four, so the disguise fails. An antisymmetric rank-2 tensor in spacetime cannot possibly be a four-vector. It is its own kind of thing, with six slots.
Six independent numbers, sitting in the most natural rank-2 object spacetime admits. If you have been kept waiting since Chapter 0.7 §4.4, you should now be doing the arithmetic: three components of , three components of , six in total.
Chapter 2.6 will find exactly six components sitting in an antisymmetric , and they will be the three components of and the three of . That comes by derivation rather than by analogy. and are not two fields that happen to be coupled. They are one tensor, split up differently by different observers, exactly as and were one four-vector and and were one set of coordinates.
Nothing more will be said about it here. But you should now find that outcome unsurprising, and finding it unsurprising in advance is the point of this section.
One more count before we leave the section: . A symmetric rank-2 tensor in four dimensions has ten independent components. That is the metric of Chapter 3.3, and it is why "the gravitational field" turns out to be ten functions rather than one. Problem 4 asks you to do these counts yourself, because doing them is how they stick.
Any object carrying two indices can be neatly divided into two distinct halves: a symmetric part, which remains identical if you swap the order of its indices, and an antisymmetric part, which flips its mathematical sign under that same swap.
The power of this division is that it survives any change of frame. What is symmetric to one observer remains symmetric to all of them, which proves that this is a genuine physical division of the object rather than an artefact of whichever notation someone happened to choose. It is a real seam in the thing itself, which is to say the object has fallen apart into independent pieces, and the two halves go on to have entirely separate careers.
Count how many independent numbers each part requires in four-dimensional spacetime and the symmetric part holds ten, the antisymmetric part six. Both counts turn out to matter enormously. The ten symmetric numbers will become the metric of curved spacetime, which is to say the ten functions that encode gravity itself. The six antisymmetric numbers will turn out to hold the electric and magnetic fields, three components each — the first substantial hint that electricity and magnetism were never two separate forces to begin with.
8 · Notation hygiene
This is a short section with a high yield. These rules are not style preferences. Each one is a consequence of something already proved, and each violation is a detectable error.
Free indices must match on both sides. A free index is one that appears exactly once in a term. It labels which component you are talking about, so both sides must be labelled the same way, with the same letters in the same positions and the same order of appearance. is fine. is not (different species, §3.3). is not (the right side is sixteen numbers per value of ). Every term in a sum must carry the same free indices too, so is meaningless.
Dummy indices are summed and freely renameable. A dummy index appears exactly twice, once up and once down, and is summed. Its name is invisible to the result, so . Use that freedom aggressively. In a long computation, renaming dummies to keep them distinct is the difference between an answer and a mess.
An index may not appear three times. If it does, either you meant two different indices and reused a letter, or you have a genuine error. There is no valid expression in which a letter appears three times in one term. This rule catches more mistakes than any other, precisely because it is purely mechanical: you can check it without understanding the expression at all.
The Kronecker delta is the mixed-index identity. if and otherwise. It has one index up and one down because it is a tensor. Problem 1 asks you to verify that, and the answer is more interesting than it sounds. Its two working properties are that it renames an index, , and that its trace counts dimensions.
Not . Getting that wrong is a rite of passage. The number will appear all over Chapter 3.4 and, in the guise of , throughout Chapter 5.10's dimensional regularisation.
carries a lower index. That is (2.4.11), and it is the most frequently forgotten fact in the subject, so here it is a second time. carries a lower index even though carries an upper one, because differentiating with respect to a thing inverts how it transforms. If you find yourself writing , you must have raised it with the metric, and there is a minus sign on the spatial components waiting for you.
8.1 · The Levi-Civita symbol, and an honest caveat
One more array is worth naming. Define to be totally antisymmetric, meaning that it changes sign under the exchange of any two indices, and fix . Total antisymmetry means it vanishes unless all four indices are different, so its entries are on the permutations of and zero elsewhere.
How does it transform? We need one fact about any matrix first, taken straight from Chapter 0.4 §5. The determinant is the alternating multilinear function of the columns normalised by , so contracting a totally antisymmetric array with four copies of can only reproduce that array times the determinant.
Now ask what the tensor law would do to . That law uses four factors of , so set in (2.4.31):
By the grind box in §4, for Lorentz transformations, so and the factor is . Under boosts and rotations, where , the symbol is genuinely invariant and may be used exactly like a tensor. Under a spatial reflection, where , it flips sign. That is where "pseudovector" and "pseudoscalar" come from, and Chapter 5.5 will find it decisive when the weak interaction turns out to care about the difference.
⚑ For a general coordinate change is neither nor constant. Equation (2.4.32) then says that an with fixed entries is not a tensor at all but a tensor density, an object that picks up a power of the Jacobian determinant. The repair is to multiply by , where , because that factor carries a compensating determinant. We quote the repair here and defer the derivation to Chapters 3.3 and 3.5, where the same factor is also what makes the invariant volume element. Within Part II, where , none of this bites.
The strict formatting rules governing these equations are not a matter of fussiness. They are an error detector built into the notation, and it catches almost everything.
For an equation to be valid, any free indices left over on one side must match, in both number and height, the free indices sitting on the other. Indices that have been summed away are internal bookkeeping and may be renamed freely, but a single index name may never appear three times within the same term.
Each of these rules follows from what the objects are rather than from a taste for tidiness, and each can be re-derived from the transformation law in a line. Taking them seriously catches fatal mistakes at once, with no physics computed at all. Almost every algebraic slip made over the coming chapters will announce itself as a mismatched index long before it can corrupt a final number. That makes the check of index balance a constant reflex, performed in the same way and for the same reasons as checking the units of an answer, and it costs about as much.
9 · Worked examples
Take in units of . Boost it into two different frames and verify by hand each time.
Frame . Lower the index by (2.4.23): flip the spatial signs, leaving . Contract:
So . For contrast, compute the forbidden pairing , with both indices up: . Remember that number.
Frame : , so . Apply (2.4.9) row by row:
with untouched. So . This is the particle's rest frame, and that is no accident: the particle's velocity in is , which is exactly the boost we chose. Contract as before.
The forbidden pairing here gives .
Frame : boost again, by , so .
Contract, and note that the two large numbers must conspire:
The forbidden pairing gives
The scoreboard.
| Frame | (wrong) | ||
|---|---|---|---|
The correctly contracted column is constant and the naively contracted one is not. That column is (2.4.13) and this chapter, in a table.
A check on the framework. Rapidities add, by the grind box in §2. Here and , so the composite rapidity is and the composite velocity is
Boost the original by that single transformation:
which is exactly. Two boosts composed equal one boost at the added rapidity, and the invariant is throughout, as it must be, since it never depended on at all.
Show that is a tensor and hence that is a scalar. Then read Chapter 0.7's continuity equation in that light.
Step 1. Differentiate the vector law. Using the covector law (2.4.11) for and then the vector law for :
Now use the fact recorded under (2.4.9): for a Lorentz transformation every entry of is a constant, so passes straight through it:
which is exactly the transformation law. So is a tensor: sixteen numbers that transform correctly.
Step 2. Contract with . By §5.1's theorem, contracting a gives a , which is a scalar. So
The four-divergence is an invariant. Note what is not claimed: alone is not invariant, and neither is . Only the sum, with the index correctly paired, is.
Step 3, cashing it in. Chapter 0.7 §6 derived the continuity equation from nothing but "stuff is neither created nor destroyed", and noted at the end that and assemble into . Take that seriously and expand , remembering so :
The factors of cancel exactly, which is the arithmetic reason is defined with the in the time slot. So
is the same equation as Chapter 0.7's continuity equation. Not an analogue: the same equation, rewritten.
But it now says more than it did. In Chapter 0.7 it was an equation about a particular observer's and , in a particular frame, with a time derivative singled out. Here the left side is a scalar and the right side is zero, which makes the whole thing a scalar equation, so by §6 it holds in every frame with the same form. If charge is conserved for one observer it is conserved for all of them, and the two-term split into "rate of change here" and "flow away from here" is only how each observer happens to slice a single four-dimensional statement. Chapter 0.7 could not have told you that, because it did not have the machinery. The whole content of the upgrade is the index placement.
Grind box — where Step 1 breaks, and why Part III needs a new derivative
Step 1 above used one thing beyond the definitions: that is constant, so passes through it. Suppose it is not constant, as it is not for a change to polar coordinates, or an accelerating frame, or any coordinate patch on a curved manifold. Then the product rule gives an extra term.
The first term is the tensor law. The second is not proportional to at all. It is proportional to itself, so the whole expression is not of the form (matrix)(components of ). is not a tensor under general coordinate changes.
Look closely at why this is fatal rather than annoying. The failure is an inhomogeneous term, and §6's theorem needed the transformation law to be homogeneous. With an inhomogeneous term, an expression that vanishes in one frame need not vanish in another. The proof of the central theorem collapses, and with it the guarantee that makes tensor equations worth writing.
The repair is to invent a new derivative which differs from by a correction term chosen so that the two inhomogeneous pieces cancel. The correction is called the connection, it is built from derivatives of the metric, and constructing it is Chapter 3.3. Its failure to commute with itself is curvature, which is Chapter 3.4, which is gravity.
So the flat-space simplification that let this chapter be short is exactly the thing Part III gives up, and everything hard about general relativity is downstream of that one product rule.
10 · Your turn
Problem 1 — versus
(a) Take to be defined as when and otherwise, in every coordinate system. Show that it really is a tensor, for any invertible coordinate change and not just Lorentz ones. Note that this is the same style of definition that failed disastrously for in (2.4.1), and explain why it succeeds here.
(b) also has the same components in every inertial frame. Show this is a special property of Lorentz transformations and not a general one, by exhibiting a coordinate change under which the components of change.
Solution
(a) Apply the law to the array and see what comes out:
using (2.4.12) in the last step. The transformation law maps the array to itself, so declaring it to be in every frame is consistent with the tensor law. And no step used any property of beyond invertibility, so this holds for every smooth coordinate change whatsoever.
Why this succeeds where (2.4.1) failed. The two definitions have the same form, which is "these components, in every frame". Only one of them is compatible with the transformation law. Declaring components in all frames is not automatically wrong. It is an extra condition, and it is legitimate exactly when the transformation law happens to reproduce those components. For it does. For it does not, since boosting that array gives . The test is mechanical, and you should run it whenever you are tempted to fix components in all frames.
(b) By the same test: (2.4.16), , says exactly that the law maps to itself. But that identity is the definition of a Lorentz transformation. It is not available for other coordinate changes.
Counterexample: rescale one spatial axis, , everything else unchanged. Then and , so
The components changed, as they must. Measuring in half-metres has to show up somewhere, and it shows up in the metric. Here is a second and more consequential example. Switch the spatial part to plane polar coordinates, , . Chapter 0.6 §4's grind box computed , so
in coordinates . That array is position-dependent, in flat spacetime, with no gravity anywhere. Here is the punchline worth carrying into Part III: a position-dependent metric does not by itself mean curvature. Telling the two apart requires an object that is zero for the polar metric and nonzero for a genuinely curved one, and constructing it is Chapter 3.4.
Problem 2 — transform, then check
In frame , two four-vectors have components and . Compute . Then boost to with , compute all eight primed components, and verify the invariance. Finally compute the naive sum in both frames and comment.
Solution
In . Lower : . Then
Boost, , . Only the and components mix:
with , , , unchanged. So and . Contract:
The naive sum. In : . In : . Different. The naive sum is not a scalar, it is not the "length" of anything, and the fact that it is what you would compute in Euclidean space is exactly why the metric has to be written explicitly.
Here is a useful habit. Notice that and are individually quite different from and , and that even changed sign, while the contraction did not budge. That is the shape of every calculation in relativity worth doing.
Problem 3 — symmetric against antisymmetric
Let and . Prove that identically, with no conditions and no special frames. Then say where you expect this to be used.
Solution
Call the sum . Both indices are dummies, so rename them, swapping the letters: . That is not a manipulation of the object, only of the notation, and it leaves the same terms in a different order. Now use the two hypotheses on the renamed expression: and , so
Hence and . Note that the proof is pure index bookkeeping: no metric, no frame, no dimension. It holds in any number of dimensions and for any pair of contracted slots.
Where it gets used. Constantly, and usually to make an unwanted term disappear.
- Chapter 2.6: is antisymmetric, so contracting it with anything symmetric kills the term. In particular , because is symmetric in by Clairaut's theorem (Chapter 0.6 §6). Applying to then forces : Maxwell's equations require charge conservation, they do not merely permit it. That is this one-line lemma doing real work.
- Chapter 3.3: the connection is symmetric in its lower indices, so contracting it with an antisymmetric object drops out. That is why the exterior derivative of Chapter 3.5 needs no connection at all.
- Chapter 3.6: the Einstein tensor is symmetric, so only the symmetric part of any candidate source can appear. That is a constraint on what the right-hand side of the field equations is allowed to be.
Problem 4 — counting
In dimensions, count the independent components of a symmetric rank-2 tensor and of an antisymmetric one, and verify the two counts sum to . Specialise to and identify what each of the two numbers will turn out to be. Then do and explain the historical accident that makes the magnetic field look like a vector.
Solution
Symmetric. , so the array is determined by the entries with . There are diagonal entries and strictly-upper ones (half of the off-diagonal entries), giving
Antisymmetric. forces for each fixed (no sum), hence : the diagonal is empty. Only the strictly-upper triangle survives: .
Sum. . ✓ That is as required, since (2.4.27) writes every rank-2 array as symmetric plus antisymmetric, uniquely, so the two pieces must account for all numbers exactly once.
. Symmetric: . Antisymmetric: . And . ✓
The ten is the metric of Chapter 3.3. The gravitational field is ten functions, which is why general relativity is ten coupled equations and not one. (Four of the ten are pure coordinate freedom, leaving six genuine ones, and that count is Chapter 3.6's business.) The six is the electromagnetic field tensor of Chapter 2.6: three components of and three of . You now know in advance that there is exactly enough room and not one slot to spare.
. Symmetric: . Antisymmetric: . The antisymmetric count coincides with the number of components of a vector. Solve and you get (or the trivial ), so this happens in three dimensions and nowhere else. That coincidence is why the curl of a vector field can be packaged as a vector (Chapter 0.7 §4.4) and why is written as one.
The disguise is detectable, though. An antisymmetric object dressed as a vector reveals itself under reflection. Reflect space and a true vector flips its components. These objects do not, because they are built from a determinant-like construction and so carry an extra factor of , as in (2.4.32). Hence "axial vector", "pseudovector", and the right-hand rule, which is a convention we are forced to state because we have insisted on describing a six-slot object with three-slot notation. In four dimensions the problem does not arise at all, because there and nobody is tempted.
You now have the definition the rest of the book runs on. A tensor of type is an object whose components transform with Jacobians and inverse Jacobians. Not a matrix, not an array, not a picture. Upper and lower indices are two different species: vectors transform like , covectors like , and they pair to give a number that every observer agrees on because the two Jacobians collapse to a . You saw that measured rather than asserted, in the figure, with components moving in opposite directions while their contraction sat at and the arrow went on piercing exactly three lines.
You have the metric as the dictionary between the species, , which flips three signs and is a linear map rather than a typographical convenience. You know that having fixed components is a defining property of Lorentz transformations rather than a law of nature. You have the three operations that stay inside the class, with contraction proved. And you have the theorem the chapter exists for: a tensor equation true in one frame is true in all of them, because the transformation law is linear and homogeneous and therefore maps zero to zero. That is why every fundamental law from here on will be written with matched indices, and why cannot be.
Where this gets spent. Immediately, in Chapter 2.5, which writes dynamics as and gets out of a contraction. Then Chapter 2.6, where §7's count of six becomes and and stop being two fields. Then Part III, where the definition is reused verbatim with one change: acquires position dependence. That single change costs you the ordinary derivative (Chapter 3.3's connection), buys you curvature (Chapter 3.4), and produces (Chapter 3.6), which is a tensor equation and could not have been anything else. Chapter 5.2 writes field theory in this language because it must, and Chapter 6.4 writes Yang–Mills in it because the gauge field is a connection in the same sense as Chapter 3.3's.
Chapter 2.5 now takes the definition and does physics with it.