Part III · General Relativity — Chapter 3.2

Manifolds

The chapter that takes the arrow away, and puts something better in its place. It is the hardest reframe in the book, so we go slowly.

Where we are

Chapter 3.1 argued that gravity is eligible to be geometry: free fall deletes the field everywhere at once, what survives is the tidal deviation of nearby particles, and the paths of falling bodies belong to the arena rather than to the bodies. What it did not do is say what kind of arena can carry such a geometry. That is this chapter, and it is entirely mathematics.

The problem is easy to state and unpleasant to fix. Every arena so far was Rn\R^{n} with one global coordinate grid. That was true of the plane of Chapter 0.4, of the configuration space of Chapter 1.2, and of the Minkowski spacetime of Chapter 2.3. In each of them we drew vectors as arrows, because Rn\R^{n} supplies a background to draw them in. Neither of those luxuries survives. Spacetime need not be R4\R^{4} globally. And here is the harder loss: there is no ambient space around spacetime in which an arrow could be drawn. Every measurement is made from inside.

Here is the route, announced in advance. Sections 1 to 3 build the arena: a space that looks like Rn\R^{n} in patches, described by overlapping charts, with a rule for how the charts agree. Section 4 is the hard step and the reason this chapter exists. It removes the arrow and replaces it with a directional derivative, then proves that the replacement has exactly the right number of components. Section 5 states the single most consequential sentence in the chapter: the vectors at one point live in a different vector space from the vectors at any other point, so vectors at different points cannot be compared. Sections 6 and 7 rebuild covectors and tensors on the new footing, and find that Chapter 2.4's definitions survive verbatim with exactly one change. Section 8 builds vector fields and the one genuinely new operation available on them.

Section 5 opens a question, and this chapter leaves it open: how do you compare a vector here with a vector there? Chapter 3.3 exists to answer it, and the answer is the gravitational field.

Tools you'll need.  Chapter 0.4 §1: the axiomatic definition of a vector space, in which no arrow appears anywhere. That section was written for this chapter, and 2.4 §3 for covectors. Also 0.4 §4 (change of basis, components transforming with the inverse of the basis), and 2.4 §3 for covectors and the dual space. Chapter 0.6 §2 (the derivative as a linear map), §5 (the multivariable chain rule, which does most of the work below), §6.1 (Clairaut's theorem on mixed partials) and §8 (the Jacobian). Chapter 0.2 §2, the fundamental theorem of calculus, used once in a grind box. Chapter 0.7 §1.2, which warned that a vector field is not a vector and named the cheque this chapter has to cash. Chapter 2.4 in its entirety. Its §2, §3, §5 and §5.1 are re-used here with one alteration, and identifying that alteration is the point of §7.

1 · Why the old arena will not do

Two separate things go wrong with Rn\R^{n}, and it is worth keeping them apart because they have different remedies.

1.1 · The global objection

Physics has no licence to assume that spacetime, taken as a whole, is R4\R^{4}. Nothing we can measure is a global statement. Observations are made in a finite region over a finite time, and Chapter 3.1 §5 quantified just how finite. Whether the universe closes up on itself, or has more than one connected piece, or has a shape that no single set of four coordinates can cover, is a question to be answered rather than assumed. Chapter 3.9 will ask it seriously.

So the arena must be built to allow shapes other than R4\R^{4} while still supporting calculus. The recipe is the one you would guess: look like Rn\R^{n} in patches, and be honest about the seams.

1.2 · The local objection, which is worse

Even in a single patch there is trouble, and it is not about shape at all. Consider the sphere, and a vector tangent to it at a point pp. The picture in your head is an arrow lying in the plane tangent to the sphere at pp. That arrow is drawn in the surrounding three-dimensional space.

That picture uses R3\R^{3}. And R3\R^{3} is not part of the sphere. It is extra scaffolding, and three separate objections attach to leaning on it.

(i) The scaffolding is not unique. ⚑ Whitney's embedding theorem states that any smooth nn-dimensional manifold can be placed inside R2n\R^{2n}. We quote it without proof and will not use it. So an ambient space always exists. The trouble is that many of them exist, and a given surface can sit inside one in genuinely different ways. Any object defined by reference to the embedding would depend on a choice nobody made on physical grounds. Chapter 2.4 spent a whole chapter teaching us what to do with quantities that depend on an arbitrary choice, and the answer was never "keep them".

(ii) Physically there is no outside. Spacetime is not observed sitting inside anything, and no experiment reaches out of it. A definition requiring an exterior is a definition that cannot be checked from where we are standing. It also imports structure that no measurement constrains: the extra dimensions, their geometry, and the way the surface sits in them.

(iii) The arrow was never a subtraction of points anyway. In Rn\R^{n} the arrow from pp to qq is the difference qpq-p, and that difference is meaningful because Rn\R^{n} is a vector space as well as a place: points can be added and subtracted. A sphere has no addition. There is no point p+qp+q on a sphere, so there is no arrow from pp to qq in the sense that made the picture work.

⚠ The cheque Chapter 0.7 wrote

Chapter 0.7 §1.2 said this plainly and then asked to be trusted for four chapters. A vector field is not a vector. The arrow attached to pp lives in its own private copy of Rn\R^{n}, the space of displacements from pp. Nothing in the definition lets you compare it with the arrow at qq. In flat space with Cartesian coordinates the private copies are all identical, so you slide arrows around freely, form F(q)F(p)\vv F(q)-\vv F(p), divide by the separation, and call the result a derivative. That chapter's exact words were: that move is the one thing that stops working on a curved space.

This chapter cashes the cheque. Section 5 makes the private copies precise, and shows that the subtraction is not merely awkward but undefined. Chapter 3.3 repairs it, and Chapter 3.1 already told you what the repair will turn out to be.

In plain terms 3.2.1

Up to now every arena in this book has been a flat grid you could lay down once and use everywhere, and two separate things now make that untenable. The first is a matter of shape. Nothing anybody measures is a statement about the universe as a whole, since observations happen in bounded regions over bounded stretches of time, so whether space and time together form an endless flat sheet is to be settled rather than assumed. The remedy is to build the arena out of overlapping patches, each looking like the familiar grid, with careful books kept on how neighbours agree.

The second problem is worse and has nothing to do with shape. Picture a small arrow lying flat against the surface of a globe. That picture is drawn in the room the globe is sitting in, and the room is not part of the globe. Spacetime has no room around it and nothing outside from which it could be viewed, and every measurement anyone will ever make is made from within. Assuming an outside means assuming structure no experiment constrains.

There is a quieter version of the same complaint. An arrow from one place to another is a subtraction of positions, and subtracting positions requires an arena whose points can be added and taken away from each other. A curved surface is not that kind of thing.

2 · Charts, atlases, and why one chart is not enough

Start with the shape problem, which is the easy one. The fix is to stop demanding one description of the whole space and settle for a set of local ones that agree where they overlap.

Definition — chart, atlas, manifold

A chart on a set MM is a pair (U,φ)(U,\varphi) where UMU\subseteq M and φ:URn\varphi:U\to\R^{n} is a one-to-one map onto an open subset of Rn\R^{n}, continuous with a continuous inverse. The nn numbers φ(p)=(x1(p),,xn(p))\varphi(p)=\big(x^{1}(p),\ldots,x^{n}(p)\big) are the coordinates of the point pp in that chart.

An atlas is a collection of charts whose domains cover all of MM. Where two charts overlap, on UαUβU_{\alpha}\cap U_{\beta}, the composite

φβφα1: φα(UαUβ)    φβ(UαUβ) \varphi_{\beta}\circ\varphi_{\alpha}^{-1}:\ \varphi_{\alpha}(U_{\alpha}\cap U_{\beta}) \;\longrightarrow\; \varphi_{\beta}(U_{\alpha}\cap U_{\beta})

is a map from one open piece of Rn\R^{n} to another, called a transition map. It is the dictionary between the two coordinate systems, and it is the only object in the whole construction that ever gets differentiated.

An nn-dimensional manifold is a set with an atlas of nn-dimensional charts.

Notice what is and is not in that definition. There are coordinates, but no preferred ones. There is no distance, no angle, no straight line, no notion of one point being "between" two others. A manifold is the least structure on which calculus can be done. That austerity is deliberate. We will add the metric in Chapter 3.3 as a separate object, because in general relativity the metric is dynamical and must be able to vary. Structure baked into the arena cannot vary.

2.1 · One chart cannot cover a sphere

It is worth proving rather than asserting that the patchwork is necessary, because the proof is short and it explains exactly what goes wrong.

Claim

There is no chart whose domain is the whole of the sphere S2S^{2}.

Suppose there were: a map φ:S2R2\varphi:S^{2}\to\R^{2}, one-to-one, continuous, with continuous inverse, whose image V=φ(S2)V=\varphi(S^{2}) is an open subset of R2\R^{2}. We derive a contradiction in four steps.

Step 1. Write φ=(φ1,φ2)\varphi=(\varphi^{1},\varphi^{2}) and look at the first component alone. It is a continuous real-valued function on S2S^{2}.

Step 2. The sphere, sitting in R3\R^{3} as the set of points with x=1\abs{\vv x}=1, is closed and bounded. ⚑ The extreme value theorem therefore guarantees that φ1\varphi^{1} attains a largest value, at some point pS2p\in S^{2}. That theorem is the multivariable form of a familiar one-dimensional statement, namely that a continuous function on a closed interval attains its maximum.

Step 3. Now look at where pp goes. Its image φ(p)=(a,b)\varphi(p)=(a,b) lies in VV, and VV is open, so some small disc of radius ε\varepsilon around (a,b)(a,b) lies entirely inside VV. In particular the point (a+ε/2,  b)(a+\varepsilon/2,\;b) is in VV.

Step 4. Since VV is the image of φ\varphi, there is a point qS2q\in S^{2} with φ(q)=(a+ε/2,  b)\varphi(q)=(a+\varepsilon/2,\;b). But then φ1(q)=a+ε/2>a=φ1(p)\varphi^{1}(q)=a+\varepsilon/2 \gt a=\varphi^{1}(p), contradicting Step 2's statement that aa was the largest value. \blacksquare

Let's look at what did the work. The sphere is compact, meaning closed and bounded, so continuous functions on it attain their extremes. An open subset of R2\R^{2} has no extremes, because around every one of its points there is room to go further in every direction. Compactness and openness are incompatible under a continuous bijection, and the sphere is the smallest interesting example. The identical argument, run with S1S^{1} and R1\R^{1}, shows that a circle needs at least two charts. The same argument again shows that no compact manifold is covered by one chart.

2.2 · An explicit two-chart atlas

Two charts do suffice for the sphere, and the standard pair is worth computing in full because its transition map is used repeatedly below. Realise S2S^{2} as {(x,y,z):x2+y2+z2=1}\{(x,y,z):x^{2}+y^{2}+z^{2}=1\} and project stereographically.

The north chart. Delete the north pole N=(0,0,1)N=(0,0,1). For any other point PP, draw the straight line from NN through PP and record where it meets the equatorial plane z=0z=0. Parametrise that line as N+t(PN)N+t\,(P-N), whose third component is 1+t(z1)1+t(z-1). Setting that component to zero gives t=1/(1z)t=1/(1-z), which is finite exactly because z1z\neq1. What we want is where the line lands, so substitute that tt into the first two components:

φN(x,y,z)  =  (u,v)  =  (x1z,  y1z),UN=S2{N}. \varphi_{N}(x,y,z) \;=\; (u,v) \;=\; \left(\frac{x}{1-z},\;\frac{y}{1-z}\right), \qquad U_{N}=S^{2}\setminus\{N\}. (3.2.1)

The south chart. Now run the same construction from the other pole. Deleting S=(0,0,1)S=(0,0,-1) and projecting from there gives t=1/(1+z)t=1/(1+z) instead, and the second chart is:

φS(x,y,z)  =  (u~,v~)  =  (x1+z,  y1+z),US=S2{S}. \varphi_{S}(x,y,z) \;=\; (\tilde u,\tilde v) \;=\; \left(\frac{x}{1+z},\;\frac{y}{1+z}\right), \qquad U_{S}=S^{2}\setminus\{S\}. (3.2.2)

Their union is all of S2S^{2}, since no point is both poles, so this is an atlas. Each chart maps onto the whole of R2\R^{2}, which is open. The overlap UNUSU_{N}\cap U_{S} is the sphere minus both poles, and its image under φN\varphi_{N} is R2\R^{2} minus the origin.

The transition map. Take (u,v)(u,v) in the overlap and compute φS\varphi_{S} of the same point. One preliminary identity does the whole job. From (3.2.1),

u2+v2  =  x2+y2(1z)2  =  1z2(1z)2  =  (1z)(1+z)(1z)2  =  1+z1z, u^{2}+v^{2} \;=\; \frac{x^{2}+y^{2}}{(1-z)^{2}} \;=\; \frac{1-z^{2}}{(1-z)^{2}} \;=\; \frac{(1-z)(1+z)}{(1-z)^{2}} \;=\; \frac{1+z}{1-z}, (3.2.3)

where the second equality used x2+y2+z2=1x^{2}+y^{2}+z^{2}=1 and the third factorised the difference of squares. Now write u~\tilde u in terms of uu by inserting a factor of one:

u~  =  x1+z  =  x1z1z1+z  =  u1u2+v2, \tilde u \;=\; \frac{x}{1+z} \;=\; \frac{x}{1-z}\cdot\frac{1-z}{1+z} \;=\; u\cdot\frac{1}{u^{2}+v^{2}}, (3.2.4)

where the last step used (3.2.3). The same computation gives v~\tilde v. Putting the two components together, here is the dictionary between the charts:

  (φSφN1)(u,v)  =  (uu2+v2,  vu2+v2).   \boxed{\;\big(\varphi_{S}\circ\varphi_{N}^{-1}\big)(u,v) \;=\; \left(\frac{u}{u^{2}+v^{2}},\;\frac{v}{u^{2}+v^{2}}\right).\;} (3.2.5)

This is inversion in the unit circle: it keeps the direction from the origin and replaces the distance ϱ\varrho by 1/ϱ1/\varrho. It is defined and infinitely differentiable everywhere except at the origin. And the origin is exactly the point that is not in the overlap, since it is the image of the south pole. The seam is clean.

In plain terms 3.2.2

Covering everything at once with a single description is a luxury, and a globe is the cheapest demonstration that the luxury is often unavailable. Two overlapping charts will do instead, provided there is an explicit dictionary translating the coordinates of one into those of the other wherever both apply. That dictionary is the only thing in the whole construction that ever gets differentiated, which is why the apparatus is lighter than it first appears.

That one chart genuinely cannot suffice deserves a proof rather than an appeal to intuition, and the proof is four lines long. A globe is closed and bounded, so any continuous quantity spread over it has to reach a largest value somewhere. A flat open region has no largest anything, since around each of its points there is always room to move a little further in every direction. A faithful correspondence between the two would have to match a place where a quantity peaks against a place where nothing peaks, and no such correspondence exists.

The construction is also notable for what it lacks. Nothing so far measures a distance or an angle or picks out a straightest path, and this poverty is deliberate rather than an oversight. Distance is going to become the dynamical variable of the theory, the thing that responds to matter and changes with time, so it cannot be welded into the arena that carries it.

3 · Smoothness, and the only place it can live

Calculus needs derivatives, and derivatives need smoothness. Where does smoothness live on a manifold?

Not on MM itself. There is nothing there to differentiate, since MM is just a set of points with charts attached. Only one kind of object in the construction is a map between pieces of Rn\R^{n}, which means only one kind of object is something Chapter 0.6's machinery can be applied to. Those are the transition maps, and so that is where the requirement goes.

Definition — smooth structure

An atlas is smooth if every transition map φβφα1\varphi_{\beta}\circ \varphi_{\alpha}^{-1} is infinitely differentiable on its domain, in the ordinary several-variable sense of Chapter 0.6. A smooth manifold is a manifold with a smooth atlas.

A function f:MRf:M\to\R is then called smooth if its coordinate representative fφα1:RnRf\circ\varphi_{\alpha}^{-1}:\R^{n}\to\R is smooth for every chart.

The definition of a smooth function comes with an obligation, and discharging it is the point of requiring smooth transition maps. Suppose ff looks smooth in chart α\alpha. Does it look smooth in chart β\beta? Write the chart-β\beta representative in terms of the chart-α\alpha one by inserting the identity φα1φα\varphi_{\alpha}^{-1}\circ\varphi_{\alpha}:

fφβ1  =  (fφα1)smooth by hypothesis(φαφβ1)smooth by assumption on the atlas. f\circ\varphi_{\beta}^{-1} \;=\; \underbrace{\big(f\circ\varphi_{\alpha}^{-1}\big)}_{\text{smooth by hypothesis}}\circ\underbrace{\big(\varphi_{\alpha}\circ\varphi_{\beta}^{-1}\big)}_{\text{smooth by assumption on the atlas}}. (3.2.6)

A composition of smooth maps is smooth, which is the chain rule of Chapter 0.6 §5 applied repeatedly. So the answer is yes. Smoothness of a function on MM is a chart-independent property precisely because the transition maps are smooth. Let's see what that requirement is buying. Had we allowed a transition map with a kink, a perfectly innocent function could be smooth in one chart and not in another, and the word would mean nothing.

Check the sphere. Its transition map (3.2.5) is built from polynomials divided by u2+v2u^{2}+v^{2}, which is nonzero throughout the overlap, so it is infinitely differentiable there. The two-chart atlas is smooth, and S2S^{2} is a smooth two-dimensional manifold.

In plain terms 3.2.3

The rule for translating between neighbouring patches is what holds a patchwork together, and it is the only object in the construction with any calculus in it. The underlying space is a bare collection of points with nothing there to differentiate, the coordinates assigned to those points are labels, and the one genuinely mathematical object is the dictionary saying how one labelling turns into another. Requiring that dictionary to be infinitely differentiable is the whole of what the word smooth means here.

The requirement earns its keep immediately. Call a quantity spread over the space smooth if it looks smooth when written in the labels of some particular patch, and an obligation appears at once, since somebody using different labels must reach the same verdict or the word describes the bookkeeping rather than the quantity. Writing the second description as the first composed with the dictionary, and remembering that a composition of well-behaved maps is well behaved, settles the matter in one line.

This is the same move the book keeps making in different clothes. Ask what survives a change of description; whatever survives is a fact about the thing, and whatever does not was a fact about the description. Here it is applied to the mildest property imaginable, and having secured it we may talk about smooth quantities on a curved space without naming a preferred way of labelling it.

4 · The tangent vector, rebuilt from the inside

This is the hard step, and everything after it is comparatively easy. Read §1.2 again if the motivation has faded: the arrow has to go, because it was drawn in a space that is not there. Here is where we are heading.

Where we are going

We will replace "a vector at pp" by "an operator that differentiates functions at pp", for three reasons. It is definable using only things that exist on MM. It carries exactly the same information as the old arrow. And it has precisely nn independent components, which we will prove rather than assume. That takes three steps: what is available from inside (§4.2), the construction (§4.3), and the proof that it has the right size (§4.4).

4.1 · What a direction is for

Before building anything, ask what the arrow was ever used for. Not for decoration. In every physical application, a direction at a point is fed into something that measures a rate of change: the rate at which temperature rises as you walk that way, the rate at which a potential falls, the rate at which any measurable quantity varies. A measurable quantity spread over spacetime is a function on MM. So:

a direction at pp is a device that assigns, to each smooth function on MM, the rate at which it changes at pp in that direction.

That sentence mentions only MM, functions on MM, and real numbers. Nothing outside. If we can make it a definition, the ambient space is not needed, and the pieces we would have thrown away are exactly the pieces that were never measurable in the first place.

4.2 · What is available from inside

Exactly two kinds of object are available on a bare smooth manifold, and both are defined without reference to anything external.

Curves. A smooth map γ\gamma from an interval of R\R into MM, with γ(0)=p\gamma(0)=p. Concretely, a worldline. In a chart it has coordinate representatives xμ(λ)φμ(γ(λ))x^{\mu}(\lambda)\equiv\varphi^{\mu}\big(\gamma(\lambda)\big), which are ordinary functions of one real variable.

Functions. A smooth map f:MRf:M\to\R, in the sense of §3. Concretely, any scalar measurement made at each event.

That is the whole toolbox. Everything below is built from those two.

4.3 · The construction

Take a curve γ\gamma through pp and a function ff on MM. Compose them: fγf\circ\gamma takes a real number λ\lambda and returns a real number f(γ(λ))f(\gamma(\lambda)). That is a function of one real variable, and Chapter 0.1 tells us exactly what to do with one. Differentiate it at λ=0\lambda=0.

Definition — the tangent vector to a curve

The curve γ\gamma defines an operator XγX_{\gamma} acting on smooth functions:

Xγ(f)    ddλ(fγ)λ=0. X_{\gamma}(f) \;\equiv\; \left.\dv{}{\lambda}\Big(f\circ\gamma\Big)\right|_{\lambda=0}.

No chart, no ambient space and no metric appears anywhere in that expression.

Now compute it in a chart, to see what it is made of. Write f^fφ1\hat f\equiv f\circ\varphi^{-1} for the function expressed in coordinates, so that f(γ(λ))=f^(x1(λ),,xn(λ))f\big(\gamma(\lambda)\big)=\hat f\big(x^{1}(\lambda),\ldots,x^{n}(\lambda)\big). Differentiating a function of several variables each of which depends on one parameter is precisely the multivariable chain rule of Chapter 0.6 §5:

Xγ(f)  =  dxμdλ0  f^xμφ(p), X_{\gamma}(f) \;=\; \left.\dv{x^{\mu}}{\lambda}\right|_{0}\;\left.\pdv{\hat f}{x^{\mu}}\right|_{\varphi(p)}, (3.2.7)

summed over μ=1,,n\mu=1,\ldots,n. Look at how that separates. The first factor knows only about the curve. The second knows only about the function. Since ff was arbitrary, we may strip it off and read (3.2.7) as an identity between operators:

Xγ  =  dxμdλ0  xμp    Vμμp,Vμdxμdλ0. X_{\gamma} \;=\; \left.\dv{x^{\mu}}{\lambda}\right|_{0}\;\left.\pdv{}{x^{\mu}}\right|_{p} \;\equiv\; V^{\mu}\,\partial_{\mu}\big|_{p}, \qquad V^{\mu}\equiv\left.\dv{x^{\mu}}{\lambda}\right|_{0}. (3.2.8)

Let's stop and take stock, because the whole reframe is visible in that line. The numbers VμV^{\mu} are the coordinate velocities of the curve. They are exactly the components you would have written down for the old arrow. Nothing quantitative has been lost. What has changed is what those numbers are components of: no longer of an arrow in a surrounding space, but of an operator, expanded in the nn operators μp\partial_{\mu}|_{p}.

Two properties of XγX_{\gamma} are worth isolating, because they are about to become the definition.

Linearity. Xγ(af+bg)=aXγ(f)+bXγ(g)X_{\gamma}(af+bg)=aX_{\gamma}(f)+bX_{\gamma}(g) for constants a,ba,b, because differentiation is linear (Chapter 0.1 §4).

The Leibniz rule at pp. Apply the product rule of Chapter 0.1 §4 to (fg)γ=(fγ)(gγ)(fg)\circ\gamma=(f\circ\gamma)\,(g\circ\gamma) and evaluate at λ=0\lambda=0, where γ(0)=p\gamma(0)=p:

Xγ(fg)  =  f(p)Xγ(g)  +  g(p)Xγ(f). X_{\gamma}(fg) \;=\; f(p)\,X_{\gamma}(g) \;+\; g(p)\,X_{\gamma}(f). (3.2.9)

Now turn these into the definition. This is the same manoeuvre Chapter 2.4 §2 made when it defined a vector by the transformation law it obeys rather than by a picture: point at the properties, and admit everything that has them.

Definition — tangent vector (derivation)

A tangent vector at pp is an operator XX taking smooth functions on MM to real numbers, such that for all smooth f,gf,g and constants a,ba,b

X(af+bg)=aX(f)+bX(g),X(fg)=f(p)X(g)+g(p)X(f). X(af+bg)=aX(f)+bX(g), \qquad X(fg)=f(p)X(g)+g(p)X(f).

An operator with these two properties is called a derivation at pp. The set of all of them is written TpMT_{p}M, the tangent space at pp.

4.4 · The coordinate operators are a basis — proved, not assumed

The definition would be worthless if TpMT_{p}M turned out to be the wrong size. We now show it has dimension exactly nn, by exhibiting a basis. Here are the two halves of the job. The nn operators μp\partial_{\mu}|_{p} span TpMT_{p}M, and they are linearly independent.

Independence, which is three lines. Suppose some combination vanishes as an operator: aμμp=0a^{\mu}\partial_{\mu}|_{p}=0, meaning it gives zero on every function. Feed it the coordinate function xνx^{\nu}, which is a perfectly good smooth function, namely "the ν\nu-th coordinate of the point". The partial derivative of one coordinate with respect to another is xν/xμ=δνμ\partial x^{\nu}/\partial x^{\mu}=\delta^{\nu}{}_{\mu}, because the coordinates are independent of each other: differentiating x2x^{2} with respect to x2x^{2} gives one, and with respect to x1x^{1} gives zero. So

0  =  aμμ(xν)  =  aμδνμ  =  aν, 0 \;=\; a^{\mu}\,\partial_{\mu}\big(x^{\nu}\big) \;=\; a^{\mu}\,\delta^{\nu}{}_{\mu} \;=\; a^{\nu}, (3.2.10)

where the last step is the Kronecker delta doing its only job, which is to rename an index and collapse a sum. Every coefficient vanishes, so the operators are independent.

Spanning, which needs a genuine theorem. The claim is that every derivation at pp is of the form aμμpa^{\mu}\partial_{\mu}|_{p}, and that has to include the derivations which arise from no curve at all. The coefficients are given by what the derivation does to the coordinate functions:

X  =  X(xμ)  μpfor every XTpM. X \;=\; X(x^{\mu})\;\partial_{\mu}\big|_{p} \qquad\text{for every } X\in T_{p}M. (3.2.11)

The argument has two ingredients, and both are short. One is that a derivation annihilates constants. The other is a factorisation of smooth functions due to Hadamard. The logic runs like this: subtract the constant f(p)f(p), factor what is left as coordinates times smooth functions, apply the Leibniz rule, and watch all but one term die because the coordinates vanish at pp. The algebra is in the grind box, and the result is (3.2.11).

Grind box — proof that every derivation is X(xμ)μX(x^{\mu})\,\partial_{\mu}

Step 1. A derivation kills constants. Let 1\mathbf 1 denote the function that is equal to 11 everywhere. Since 11=1\mathbf 1\cdot\mathbf 1=\mathbf 1, the Leibniz rule gives

X(1)  =  X(11)  =  1(p)X(1)+1(p)X(1)  =  2X(1), X(\mathbf 1) \;=\; X(\mathbf 1\cdot\mathbf 1) \;=\; \mathbf 1(p)\,X(\mathbf 1) + \mathbf 1(p)\,X(\mathbf 1) \;=\; 2X(\mathbf 1),

so X(1)=0X(\mathbf 1)=0. By linearity X(c1)=cX(1)=0X(c\,\mathbf 1)=cX(\mathbf 1)=0 for any constant cc. Derivations do not see constants, exactly as ordinary derivatives do not.

Step 2. Hadamard's lemma. Choose a chart around pp and translate it so that pp sits at the origin, φ(p)=0\varphi(p)=0. That is legitimate, because composing a chart with a translation of Rn\R^{n} gives another chart. Let f^\hat f be the coordinate representative of ff on a ball around the origin. Then there exist smooth functions hμh_{\mu} on that ball with

f^(x)  =  f^(0)  +  xμhμ(x),hμ(0)  =  μf^(0). \hat f(x) \;=\; \hat f(0) \;+\; x^{\mu}\,h_{\mu}(x), \qquad h_{\mu}(0) \;=\; \partial_{\mu}\hat f(0).

Proof. Consider tf^(tx)t\mapsto \hat f(tx) for t[0,1]t\in[0,1], which walks along the straight segment from the origin to xx. By the fundamental theorem of calculus (Chapter 0.2 §2),

f^(x)f^(0)  =  01ddtf^(tx)  dt. \hat f(x)-\hat f(0) \;=\; \int_{0}^{1}\dv{}{t}\,\hat f(tx)\;\dd t.

Evaluate the integrand with the chain rule (Chapter 0.6 §5): the μ\mu-th coordinate of the point txtx is txμtx^{\mu}, whose tt-derivative is xμx^{\mu}, so

ddtf^(tx)  =  xμ(μf^)(tx). \dv{}{t}\hat f(tx) \;=\; x^{\mu}\,\big(\partial_{\mu}\hat f\big)(tx).

The factor xμx^{\mu} does not depend on tt and comes out of the integral:

f^(x)f^(0)  =  xμ01(μf^)(tx)  dt  hμ(x). \hat f(x)-\hat f(0) \;=\; x^{\mu}\underbrace{\int_{0}^{1}\big(\partial_{\mu}\hat f\big)(tx)\;\dd t}_{\displaystyle \equiv\;h_{\mu}(x)}.

Each hμh_{\mu} is smooth, because differentiating under the integral sign (the move Chapter 1.2 §3.1 used) turns derivatives of hμh_{\mu} into integrals of derivatives of f^\hat f, all of which exist. And setting x=0x=0 gives hμ(0)=01μf^(0)dt=μf^(0)h_{\mu}(0)=\int_{0}^{1}\partial_{\mu}\hat f(0)\,\dd t=\partial_{\mu}\hat f(0). \square

Step 3. Apply XX. Transporting Hadamard's identity back to MM, we have f=f(p)1+xμhμf = f(p)\,\mathbf 1 + x^{\mu}h_{\mu} as functions near pp, where now xμx^{\mu} means the coordinate function and hμh_{\mu} the corresponding function on MM. Apply XX, using linearity on the sum and the Leibniz rule on each product:

X(f)  =  f(p)X(1)=  0 by Step 1  +  xμ(p)=  0 since φ(p)=0X(hμ)  +  hμ(p)X(xμ). X(f) \;=\; \underbrace{f(p)\,X(\mathbf 1)}_{=\;0\ \text{by Step 1}} \;+\; \underbrace{x^{\mu}(p)}_{=\;0\ \text{since } \varphi(p)=0}\,X(h_{\mu}) \;+\; h_{\mu}(p)\,X(x^{\mu}).

Two of the three terms vanish outright. The first goes because derivations kill constants, and the second because we placed pp at the coordinate origin. What is left, using hμ(p)=μf^(0)h_{\mu}(p)=\partial_{\mu}\hat f(0) from Step 2, is

X(f)  =  X(xμ)  μf^p, X(f) \;=\; X(x^{\mu})\;\partial_{\mu}\hat f\big|_{p},

which is (3.2.11), since ff was arbitrary. \blacksquare

One honest caveat. Hadamard's factorisation holds on a ball around pp, not on all of MM, so strictly the argument shows that a derivation is determined by the behaviour of ff near pp. That is a feature rather than a gap: it says a tangent vector is a local object, which is exactly what one wants of a direction. Making it airtight requires a bump function that is 11 near pp and 00 far away, whose existence is a standard construction we shall not need again.

So TpMT_{p}M has dimension nn, and {μp}\{\partial_{\mu}|_{p}\} is a basis for it, called the coordinate basis of that chart. Every tangent vector is written

X  =  Vμμp,Vμ  =  X(xμ). X \;=\; V^{\mu}\,\partial_{\mu}\big|_{p}, \qquad V^{\mu} \;=\; X\big(x^{\mu}\big). (3.2.12)

The components are recovered by feeding the vector the coordinate functions. That is a pleasant inversion of the usual order: instead of the coordinates telling you the vector's components, the vector tells you its components by acting on the coordinates.

4.5 · Changing chart, and Chapter 2.4 collected

Two charts, coordinates xμx^{\mu} and xμx'^{\mu}, overlapping near pp. How do the two bases relate? Apply the chain rule (0.6 §5) to any function, differentiating with respect to a primed coordinate by going through the unprimed ones:

xμp  =  xνxμp  xνp,that isμ  =  (Λ1)νμ  ν, \left.\pdv{}{x'^{\mu}}\right|_{p} \;=\; \left.\pdv{x^{\nu}}{x'^{\mu}}\right|_{p}\;\left.\pdv{}{x^{\nu}}\right|_{p}, \qquad\text{that is}\qquad \partial'_{\mu} \;=\; \big(\Lambda^{-1}\big)^{\nu}{}_{\mu}\;\partial_{\nu}, (3.2.13)

with Λμνxμ/xν\Lambda^{\mu}{}_{\nu}\equiv\partial x'^{\mu}/\partial x^{\nu} exactly as in Chapter 2.4 §2. The basis vectors transform with the inverse Jacobian.

Now impose the one thing we insist on: the vector itself is a single operator and does not care which chart you expand it in. So Vμμ=VννV'^{\mu}\partial'_{\mu}=V^{\nu}\partial_{\nu}. Substitute (3.2.13) into the left side:

Vμ(Λ1)νμν  =  Vνν. V'^{\mu}\,\big(\Lambda^{-1}\big)^{\nu}{}_{\mu}\,\partial_{\nu} \;=\; V^{\nu}\,\partial_{\nu}. (3.2.14)

The ν\partial_{\nu} are linearly independent by §4.4, so their coefficients must agree index by index:

(Λ1)νμVμ  =  Vν. \big(\Lambda^{-1}\big)^{\nu}{}_{\mu}\,V'^{\mu} \;=\; V^{\nu}. (3.2.15)

What we want is VV' alone on the left, so multiply both sides by Λρν\Lambda^{\rho}{}_{\nu} and sum over ν\nu. The identity Λρν(Λ1)νμ=δρμ\Lambda^{\rho}{}_{\nu}(\Lambda^{-1})^{\nu}{}_{\mu}=\delta^{\rho}{}_{\mu} does the work, and it is the chain-rule identity of Chapter 2.4 §3 rather than a new assumption. The delta then renames μ\mu into ρ\rho on the left:

  Vμ  =  xμxν  Vν.   \boxed{\;V'^{\mu} \;=\; \pdv{x'^{\mu}}{x^{\nu}}\;V^{\nu}.\;} (3.2.16)
Chapter 2.4's promise, collected in full

Chapter 2.4's opening callout said: "Chapter 3.2 will reuse this chapter's definition word for word, with only one change, and that change is the whole of general relativity."

Here it is. Equation (3.2.16) is character-for-character Chapter 2.4 §2's definition of a contravariant vector. The one change: there, Λμν\Lambda^{\mu}{}_{\nu} was a constant matrix, because a Lorentz transformation is linear and its Jacobian is the same at every event. Here it is a function of position, because a change of chart is any smooth invertible map at all.

Every consequence of Part III flows from that single alteration. In particular, differentiating a vector field now differentiates Λ\Lambda as well, which is why Chapter 3.3 opens by showing that νVμ\partial_{\nu}V^{\mu} is not a tensor and spends the rest of itself repairing that.

In plain terms 3.2.4

Arrows need somewhere to live, and the somewhere has been quietly assumed all along. Take it away and the question becomes what a direction was ever for, and the answer is that a direction gets handed to something which reports a rate of change. Walk that way and the temperature rises; walk this way and it falls. Measurable quantities spread over the arena are functions on it, so a direction is a device taking any such quantity and returning the rate at which it changes here. Nothing in that sentence points outside.

The construction is then almost forced. A path through the point, composed with a quantity defined on the space, gives a function of one variable, which the first chapter taught us to differentiate. Doing that in coordinates splits the answer into a part belonging to the path and a part belonging to the quantity, and the first is exactly the list of numbers the discarded arrow would have carried. What was lost was the picture, and the picture was the one piece that needed an outside.

Two demands survive and become the definition: the device must be additive, and it must handle products the way differentiation does. What makes this worth the discomfort is a theorem rather than a hope. Such devices number exactly as many as the space has dimensions, so the replacement is the same size as the thing replaced.

5 · A different vector space at every point

First, TpMT_{p}M really is a vector space in the sense of Chapter 0.4 §1, the axiomatic sense, with no arrow in it. Given derivations X,YX,Y at pp and reals a,ba,b, define (aX+bY)(f)aX(f)+bY(f)(aX+bY)(f)\equiv aX(f)+bY(f). Linearity is immediate. What has to be checked is that the Leibniz rule survives the combination, and it does:

(aX+bY)(fg)  =  a[f(p)X(g)+g(p)X(f)]  +  b[f(p)Y(g)+g(p)Y(f)]=  f(p)(aX+bY)(g)  +  g(p)(aX+bY)(f), \begin{aligned} (aX+bY)(fg) \;&=\; a\big[f(p)X(g)+g(p)X(f)\big] \;+\; b\big[f(p)Y(g)+g(p)Y(f)\big]\\[4pt] &=\; f(p)\,\big(aX+bY\big)(g) \;+\; g(p)\,\big(aX+bY\big)(f), \end{aligned} (3.2.17)

where the second line is nothing but a regrouping of the four terms by whether they carry f(p)f(p) or g(p)g(p). So sums and multiples of derivations are derivations. The remaining axioms are inherited from the reals, and TpMT_{p}M is an nn-dimensional real vector space. Chapter 0.4's decision to define a vector space by axioms rather than by arrows is what makes this paragraph possible at all.

5.1 · The sentence

The most important statement in this chapter

For pqp\neq q, the tangent spaces TpMT_{p}M and TqMT_{q}M are different vector spaces, and there is no canonical way to identify one with the other.

Consequently a vector at pp and a vector at qq cannot be added, subtracted, or compared. The expression Xp+YqX_{p}+Y_{q} is not merely hard to evaluate. It is undefined, in the way that adding a temperature to a pressure is undefined.

Why it is true is immediate from the definition, and this is the payoff for having built tangent vectors the way we did. A derivation at pp is an operator that reports rates of change at pp. A derivation at qq reports rates of change at qq. They act on the same functions but return different kinds of information, and they satisfy different Leibniz rules. Look at (3.2.9) again: it evaluates ff at pp in one case and at qq in the other. There is no operation that turns one into the other, because there is no operation in the construction that relates the two points at all.

5.2 · Why this feels wrong, and what the flat case was hiding

The objection writes itself: in Rn\R^{n} we slide arrows around all the time and nobody complains. Let's diagnose that carefully, because it identifies precisely what a general manifold is missing.

Rn\R^{n} is two things at once. It is a manifold, which is to say a set with a chart. It is also a vector space, in which points can be added and subtracted. That second structure is extra, and it is what supplies the identification. Given a tangent vector at pp with components VμV^{\mu}, one declares it "the same as" the tangent vector at qq with the same components. The declaration is legitimate there because Rn\R^{n}'s vector-space structure singles out one way of doing it, and Cartesian coordinates make it invisible.

A sphere has no addition of points, so no such declaration is available. And the failure is not a formality. It has a concrete measurable signature, and the following figure exhibits it.

5.3 · "Are these two vectors parallel?" has no answer

Take two points pp and qq on a sphere, with a tangent vector at each. Ask whether they point the same way.

To answer, you must bring one of them to the other, and the only honest way to do that on a surface is to carry it along a path, keeping it as unchanged as the surface allows at every step. Along a great circle there is a rule for doing exactly that: keep the vector at a constant angle to the direction you are walking. That rule is entirely intrinsic, since it refers only to angles measured within the surface and never to any surrounding space. Build any path out of great-circle arcs and the rule extends to it.

That rule is well defined, natural, and the answer it gives depends on which path you took.

100°
60°
route A (equator): arrives on bearing 0.00°
route B (detour): arrives on bearing 60.27°
the two answers differ by 60.27° [tangency check at q: 0.0e+0]
The question this chapter cannot answer. A vector is planted at pp (left, black) and carried to qq (right) twice, by the intrinsic rule hold a constant angle to the direction of travel along each great-circle arc. The blue route goes straight along the equator; the purple route detours through a waypoint at the latitude you choose. Two arrows arrive at qq, and the readouts give their compass bearings there and the angle between them. Every rotation is computed numerically from the three-dimensional rotation carrying the start of each arc to its end — no formula from the text is used — and the arrival vectors are checked to be tangent at qq to fourteen digits. Set the detour to zero, or press the button, and the two routes become the same route: the disagreement falls to exactly 0.00°0.00°, which is the figure checking itself. Raise the detour and the two answers part company. There is no basis for preferring one arrival to the other, so "do these two vectors point the same way" is a question with no answer here. Chapter 3.3 supplies the structure that gives it one, and Chapter 3.4 turns the disagreement into a measurement of curvature.

5.4 · The consequence, which is why Chapter 3.3 exists

Everything so far has been about comparing two vectors. Now notice that differentiating a vector field is exactly that comparison, and therefore inherits the whole problem.

To differentiate a field VV along a curve you would write, following Chapter 0.1 §2 to the letter,

limε0  V(γ(ε))    V(γ(0))ε. \lim_{\varepsilon\to0}\;\frac{V\big(\gamma(\varepsilon)\big) \;-\; V\big(\gamma(0)\big)}{\varepsilon}. (3.2.18)

Look at what the numerator is asking for. It subtracts an element of Tγ(ε)MT_{\gamma(\varepsilon)}M from an element of Tγ(0)MT_{\gamma(0)}M, and by §5.1 that subtraction does not exist. Equation (3.2.18) is not a difficult limit. It is not a limit at all, because its numerator is not defined for any ε0\varepsilon\neq0.

One escape suggests itself, and it is worth seeing why it fails. Work in a chart, where the field is nn ordinary functions Vμ(x)V^{\mu}(x), and differentiate those. The array νVμ\partial_{\nu}V^{\mu} is a perfectly good collection of n2n^{2} numbers. The trouble is that those numbers describe the chart as much as they describe the field. Differentiating the transformation law (3.2.16) requires the product rule, so the derivative lands on the Jacobian as well as on the components. Unlike in Chapter 2.4, the Jacobian is no longer constant, so that extra term does not vanish. The result is that νVμ\partial_{\nu}V^{\mu} fails the tensor transformation law. Chapter 3.3 §4 does that computation in full and displays the offending term. Chapter 3.3 §5 then defines a correction whose only job is to cancel it, and the correction turns out to be the gravitational field.

5.5 · The tangent bundle

One piece of bookkeeping before we leave this section. All those separate tangent spaces are worth collecting into a single object, so that a vector field becomes a map into one place rather than a family of maps into many:

TM    pM  {p}×TpM, TM \;\equiv\; \bigcup_{p\in M}\;\{p\}\times T_{p}M, (3.2.19)

whose points are pairs: a place, and a vector there. It has 2n2n coordinates, nn of them to say where and nn to say which vector, and it is itself a smooth manifold. A vector field is a rule assigning to each pp an element of TpMT_{p}M, smoothly.

Two debts are settled by that paragraph. Chapter 0.7 §1.2 named TMTM as the object it was deferring. And Chapter 1.2 §4 observed that a Lagrangian L(x,x˙)L(x,\dot x) is properly a function on the tangent bundle of the configuration space, with position and velocity as independent coordinates. There the constraint "x˙\dot x is the derivative of xx" was imposed only when a path was substituted. That is exactly (3.2.19): the pair (place, vector there), with both free.

In plain terms 3.2.5

From here on, each point of the space carries its own private stock of directions, and no two stocks are the same stock. A direction at one place is a device reporting rates of change there, and a direction elsewhere reports rates of change elsewhere, so the two are not members of a common collection and cannot be added or subtracted. That is the most consequential thing said so far, and it deserves to be sat with rather than nodded through.

The reason it feels wrong is instructive. On a flat sheet we slide arrows about freely, and that habit works because a flat sheet is secretly two things at once, a place and also a system in which positions can be added and taken away from one another. The second provides a rule for saying an arrow here matches an arrow there, and square coordinates hide the fact that a rule was ever needed. A curved surface offers no such rule, and carrying an arrow across a globe by different routes delivers different verdicts.

This is why differentiating fails. Finding a rate of change means subtracting a value here from a value a short step away, and if the two values belong to separate systems the subtraction is not hard but meaningless. Repairing it is the whole business of the next chapter, and the leftover from the repair is going to be gravity.

6 · Covectors, and what a differential has been all along

Chapter 0.6 §4 named, and Chapter 2.4 §3 built, the dual VV^{*} of a vector space VV: the space of linear maps from VV to R\R. Neither chapter said how big that space is, and we will need to know, so this section settles it a few paragraphs below. Now apply that construction at each point of the manifold.

Definition — cotangent space

The cotangent space TpMT^{*}_{p}M is the dual of TpMT_{p}M: the set of linear maps ω:TpMR\omega:T_{p}M\to\R. Its elements are called covectors or one-forms at pp. It has dimension nn.

There is one natural way to manufacture covectors, and it needs nothing but a function.

Definition — the differential of a function

For a smooth ff on MM, define dfpTpM\dd f|_{p}\in T^{*}_{p}M by

df,  X    X(f)for every XTpM. \big\langle \dd f,\;X\big\rangle \;\equiv\; X(f) \qquad\text{for every } X\in T_{p}M.

That this is linear in XX is not a calculation. It is the first clause of the definition of a derivation, read from the other side.

Now feed the machine the coordinate functions themselves and see what falls out. Apply dxμ\dd x^{\mu} to the basis vector ν\partial_{\nu}:

dxμ,  ν  =  ν(xμ)  =  xμxν  =  δμν, \big\langle \dd x^{\mu},\;\partial_{\nu}\big\rangle \;=\; \partial_{\nu}\big(x^{\mu}\big) \;=\; \pdv{x^{\mu}}{x^{\nu}} \;=\; \delta^{\mu}{}_{\nu}, (3.2.20)

the same three-step computation used in §4.4's independence proof, run in the other direction. Equation (3.2.20) is exactly the condition that defines a dual basis: each member of the one set reads off one component of the other, and ignores the rest. So the nn covectors {dxμ}\{\dd x^{\mu}\} form the basis of TpMT^{*}_{p}M dual to {μ}\{\partial_{\mu}\}, and every covector expands as ω=ωμdxμ\omega=\omega_{\mu}\,\dd x^{\mu} with ωμ=ω,μ\omega_{\mu}=\langle\omega,\partial_{\mu}\rangle.

The covector we most want in that basis is df\dd f itself, so expand it. Its components are what it returns on the basis vectors, df,μ=μf\langle \dd f,\partial_{\mu}\rangle=\partial_{\mu}f, which gives

df  =  fxμ  dxμ. \dd f \;=\; \pdv{f}{x^{\mu}}\;\dd x^{\mu}. (3.2.21)
Familiar ground — the cheque from Chapter 0.1, finally cleared

Chapter 0.1 §7 warned that writing dy=f(x)dx\dd y=f'(x)\,\dd x and cancelling differentials was an abuse of notation that would be paid for later, and that Chapter 0.6 would begin the payment. Chapter 0.6 §4 duly defined df\dd f as a one-form on Rn\R^{n} and noted that on a manifold there would be no inner product at all, so the distinction would stop being optional.

Equation (3.2.21) is the final form. It is not an approximation, not an infinitesimal, and not a heuristic. It is an exact expansion of a covector in a basis, every symbol of which now has a definition. And it licenses the school-level manipulation rather than forbidding it. Pair both sides with a tangent vector XX and you get X(f)=μfX(xμ)X(f)=\partial_{\mu}f\cdot X(x^{\mu}), which is the chain rule.

One warning, and it matters for the next two chapters. df\dd f is not the gradient. The gradient is a vector. Chapter 0.6 §4 showed that producing it from df\dd f requires contracting with an inverse metric, (f)μ=gμννf(\nabla f)^{\mu}=g^{\mu\nu}\partial_{\nu}f. On the bare manifold of this chapter there is no metric, so df\dd f exists and f\nabla f does not. That is not a defect. It is the reason the metric will be introduced as a separate, physical, dynamical object in Chapter 3.3 rather than smuggled in as part of the arena.

6.1 · How covectors transform

The basis covectors follow from the chain rule, differentiating each new coordinate through the old ones:

dxμ  =  xμxνdxν  =  Λμνdxν. \dd x'^{\mu} \;=\; \pdv{x'^{\mu}}{x^{\nu}}\,\dd x^{\nu} \;=\; \Lambda^{\mu}{}_{\nu}\,\dd x^{\nu}. (3.2.22)

Compare with (3.2.13): the basis vectors carried Λ1\Lambda^{-1} and the basis covectors carry Λ\Lambda. The two bases transform oppositely, which is what "dual" means. Demanding that the covector ω=ωμdxμ\omega=\omega_{\mu}\dd x^{\mu} be one object independent of chart then forces its components the other way, by the same three-line argument as §4.5:

ωμ  =  xνxμ  ων  =  (Λ1)νμων, \omega'_{\mu} \;=\; \pdv{x^{\nu}}{x'^{\mu}}\;\omega_{\nu} \;=\; \big(\Lambda^{-1}\big)^{\nu}{}_{\mu}\,\omega_{\nu}, (3.2.23)

which is Chapter 2.4 §3's definition of a covariant vector, unaltered. As a check, the pairing is invariant:

ωμVμ  =  (Λ1)νμων  ΛμρVρ  =  δνρωνVρ  =  ωνVν, \omega'_{\mu}V'^{\mu} \;=\; \big(\Lambda^{-1}\big)^{\nu}{}_{\mu}\,\omega_{\nu}\;\Lambda^{\mu}{}_{\rho}\,V^{\rho} \;=\; \delta^{\nu}{}_{\rho}\,\omega_{\nu}V^{\rho} \;=\; \omega_{\nu}V^{\nu}, (3.2.24)

where the two Jacobians met and collapsed to a Kronecker delta, which then renamed an index. That is Chapter 2.4 §3.1's calculation, character for character, and it still works because the collapse ΛΛ1=δ\Lambda\Lambda^{-1}=\delta holds at each point separately and we never moved between points.

In plain terms 3.2.6

The measuring devices from the chapter on tensors reappear here, and this time they arrive with a definition rather than a picture. Alongside the directions at a point sit the machines that eat a direction and return a number, and those machines form a space of their own with as many dimensions. Any measurable quantity spread over the arena produces one automatically: hand it a direction and it reports how fast the quantity changes that way.

Feeding that construction the coordinate labels themselves produces exactly the objects physicists have been writing under integral signs since school and cancelling without justification. They turn out to be the machines paired with the coordinate directions, and the pairing of one with another gives one when the labels match and zero when they do not, which is the condition defining a dual basis. What was an abuse of notation four parts ago is now an exact statement about honest objects.

One distinction has to be kept sharp, because a later chapter will charge for confusing it. The machine built from a quantity is not the gradient of that quantity. Turning it into a genuine direction requires something that converts between the two species, and no such converter exists yet. That converter is the metric, it is a physical field rather than part of the furniture, and the whole of the next chapter is about it.

7 · Tensors, unchanged

With vectors and covectors in place at each point, Chapter 2.4 §5's definition can be transferred without modification.

Definition — tensor at a point, and tensor field

A tensor of type (k,l)(k,l) at pp is an object with nk+ln^{k+l} components per chart, transforming with one factor of Λ\Lambda per upper index and one of Λ1\Lambda^{-1} per lower index:

Tμ1μkν1νl  =  Λμ1α1Λμkαk  (Λ1)β1ν1(Λ1)βlνl  Tα1αkβ1βl, T'^{\mu_{1}\ldots\mu_{k}}{}_{\nu_{1}\ldots\nu_{l}} \;=\; \Lambda^{\mu_{1}}{}_{\alpha_{1}}\cdots\Lambda^{\mu_{k}}{}_{\alpha_{k}}\;\big(\Lambda^{-1}\big)^{\beta_{1}}{}_{\nu_{1}}\cdots\big(\Lambda^{-1}\big)^{\beta_{l}}{}_{\nu_{l}}\;T^{\alpha_{1}\ldots\alpha_{k}}{}_{\beta_{1}\ldots\beta_{l}},

with Λμν=xμ/xν\Lambda^{\mu}{}_{\nu}=\partial x'^{\mu}/\partial x^{\nu} evaluated at pp. Equivalently, a multilinear map taking kk covectors and ll vectors at pp to a real number. A tensor field assigns such an object to every point, smoothly.

Everything Chapter 2.4 proved about tensors used only two facts: that the law is linear and homogeneous in the components, and that ΛΛ1=δ\Lambda\Lambda^{-1}=\delta. Both hold here at each point. So the three operations of Chapter 2.4 §5.1 survive with their proofs intact. Those were linear combination of tensors of the same type, outer product, and contraction of an upper index against a lower one. So does the theorem of Chapter 2.4 §6, that an equation between tensors of the same type, true in one chart, is true in every chart. None of it needs redoing.

Three things are new, and they are the whole difference between Part II and Part III.

(a) Λ\Lambda depends on position. Stated already in §4.5, and it is the only alteration.

(b) Adding tensors at different points is forbidden. The definition is pointwise. Two tensors of the same type at the same point may be added. Two at different points may not, for exactly the reason of §5.1. In Chapter 2.4 this restriction was invisible, because the constant Λ\Lambda made components at different events directly comparable.

(c) Differentiation leaves the class. In Chapter 2.4's table, νVμ\partial_{\nu}V^{\mu} was listed as a type (1,1)(1,1) tensor. On a manifold it is not, by §5.4. That single entry going wrong is what Chapter 3.3 is for, and the repaired derivative is the connection.

⚠ Why this isn't obvious

It is tempting to think that once Λ\Lambda varies with position, everything must be rebuilt. It is worth being precise about why almost nothing is. The tensor transformation law is a statement made at one point, and at one point Λ\Lambda is just a matrix, indistinguishable from a Lorentz matrix. All of Chapter 2.4's algebra is pointwise algebra and therefore untouched.

Position dependence can only matter where a calculation compares one point with another. In the whole of Chapter 2.4 there is exactly one such operation, and that is differentiation. So the damage is narrow and locatable, which is why Part III can be four mathematics chapters rather than forty.

In plain terms 3.2.7

Nothing about the definition of a tensor had to be replaced, and that is the reward for having defined it by behaviour in the first place rather than by appearance. An object is a tensor if its description responds to a change of labels in one specific controlled way, and the definition never asked what the object looked like, whether the labels were straight, or whether the arena was flat. Every consequence drawn from it in the earlier chapter was drawn at a single point, so every consequence still holds.

Exactly one thing has changed. The matrix relating one labelling to another used to be the same everywhere, because the changes of description allowed in the previous part were rigid ones. Now it varies from place to place, since any smooth relabelling is permitted. At any single point that makes no difference whatever, because a matrix is a matrix. It makes a difference only where a calculation reaches across from one point to a neighbouring one.

There is precisely one such calculation in the whole earlier chapter, and it is taking a derivative. That is the entire damage, and locating it so narrowly is what keeps this part of the book to four mathematical chapters instead of forty. The next chapter repairs that one operation, and what has to be added in order to repair it turns out to be the gravitational field itself.

8 · Vector fields, and the one new operation

A vector field XX assigns a tangent vector to each point, so in a chart it is X=Xμ(x)μX=X^{\mu}(x)\,\partial_{\mu} with the components now functions rather than numbers. A vector field therefore acts on a function to give another function: (Xf)(p)Xp(f)(Xf)(p)\equiv X_{p}(f). This is new. A single tangent vector produced a number, but a field produces a function, and functions can be differentiated again.

Which raises the question: if XX and YY are vector fields, is the composite XYXY, meaning "apply YY, then apply XX", a vector field?

8.1 · It is not, and the failure is instructive

Test the Leibniz rule. Apply XYXY to a product, expanding one operator at a time and using Leibniz for YY first, then for XX:

XY(fg)  =  X(f(Yg)+g(Yf))=  (Xf)(Yg)+fXY(g)  +  (Xg)(Yf)+gXY(f). \begin{aligned} XY(fg) \;&=\; X\Big(f\,(Yg) + g\,(Yf)\Big)\\[3pt] &=\; (Xf)(Yg) + f\,XY(g) \;+\; (Xg)(Yf) + g\,XY(f). \end{aligned} (3.2.25)

The second and fourth terms are what the Leibniz rule requires. The first and third are surplus:

XY(fg)    fXY(g)    gXY(f)  =  (Xf)(Yg)+(Xg)(Yf)    0. XY(fg) \;-\; f\,XY(g) \;-\; g\,XY(f) \;=\; (Xf)(Yg) + (Xg)(Yf) \;\neq\; 0. (3.2.26)

So XYXY is not a derivation, and correspondingly not a vector field. In components the reason is visible: expanding XY(f)=Xμμ(Yννf)XY(f)=X^{\mu}\partial_{\mu}\big(Y^{\nu}\partial_{\nu}f\big) with the product rule gives

XY(f)  =  Xμ(μYν)νf  +  XμYνμνf, XY(f) \;=\; X^{\mu}\big(\partial_{\mu}Y^{\nu}\big)\,\partial_{\nu}f \;+\; X^{\mu}Y^{\nu}\,\partial_{\mu}\partial_{\nu}f, (3.2.27)

and the second term contains second derivatives of ff. A tangent vector is a first-order object, and second derivatives have no business in it.

8.2 · The repair: antisymmetrise

Look at (3.2.26) again. Its right-hand side is (Xf)(Yg)+(Xg)(Yf)(Xf)(Yg)+(Xg)(Yf), which is symmetric under exchanging XX and YY: swapping the two names turns the first term into the second and the second into the first. So computing YX(fg)fYX(g)gYX(f)YX(fg)-f\,YX(g)-g\,YX(f) gives the identical expression, and subtracting the two statements makes the surplus cancel exactly. Define

[X,Y]    XYYX. [X,Y] \;\equiv\; XY - YX. (3.2.28)

Then [X,Y](fg)f[X,Y](g)g[X,Y](f)=0[X,Y](fg)-f[X,Y](g)-g[X,Y](f)=0, so [X,Y][X,Y] obeys the Leibniz rule. It is linear because XX and YY are. So it is a derivation at each point. The commutator of two vector fields is a vector field. This object is called the Lie bracket.

8.3 · Its components, with every index step named

What we want are the components [X,Y]ν[X,Y]^{\nu}, and (3.2.27) already supplies half of them. Write that half out again as the first term of the commutator:

XY(f)  =  Xμ(μYν)νf  +  XμYνμνf, XY(f) \;=\; X^{\mu}\big(\partial_{\mu}Y^{\nu}\big)\,\partial_{\nu}f \;+\; X^{\mu}Y^{\nu}\,\partial_{\mu}\partial_{\nu}f, (3.2.29)

and that is the first of the two terms we need. The second is the same line with the letters XX and YY exchanged throughout, since that exchange is exactly what YXYX asks for:

YX(f)  =  Yμ(μXν)νf  +  YμXνμνf. YX(f) \;=\; Y^{\mu}\big(\partial_{\mu}X^{\nu}\big)\,\partial_{\nu}f \;+\; Y^{\mu}X^{\nu}\,\partial_{\mu}\partial_{\nu}f. (3.2.30)

Now deal with the two second-derivative terms, one manipulation at a time.

Move 1. Relabel the dummy indices in (3.2.30). In the second term of (3.2.30) both μ\mu and ν\nu are summed over, so the letters are private and may be exchanged with each other. Swapping μν\mu\leftrightarrow\nu throughout that one term turns YμXνμνfY^{\mu}X^{\nu}\partial_{\mu}\partial_{\nu}f into YνXμνμfY^{\nu}X^{\mu}\partial_{\nu}\partial_{\mu}f. Nothing has happened except a change of alphabet.

Move 2. Use the symmetry of mixed partials. Clairaut's theorem (Chapter 0.6 §6.1) says νμf=μνf\partial_{\nu}\partial_{\mu}f=\partial_{\mu}\partial_{\nu}f for smooth ff. The two numerical factors XμX^{\mu} and YνY^{\nu} are ordinary functions and multiply in either order, so they may be swapped as well. The term therefore becomes XμYνμνfX^{\mu}Y^{\nu}\partial_{\mu}\partial_{\nu}f.

Move 3. Subtract. The two second-derivative terms are now literally identical and cancel in (3.2.29) minus (3.2.30), leaving only the first-derivative terms:

[X,Y](f)  =  (XμμYν    YμμXν)νf. [X,Y](f) \;=\; \Big(X^{\mu}\,\partial_{\mu}Y^{\nu} \;-\; Y^{\mu}\,\partial_{\mu}X^{\nu}\Big)\,\partial_{\nu}f. (3.2.31)

Since ff was arbitrary, strip it off and read the components straight from the coefficient of ν\partial_{\nu}:

  [X,Y]ν  =  XμμYν    YμμXν.   \boxed{\;[X,Y]^{\nu} \;=\; X^{\mu}\,\partial_{\mu}Y^{\nu} \;-\; Y^{\mu}\,\partial_{\mu}X^{\nu}.\;} (3.2.32)
Recap — what went in, what came out

In: two vector fields, the product rule, the relabelling of dummy indices, and Clairaut's theorem on mixed partial derivatives.

Out: the composite XYXY is not a vector field because it contains second derivatives, but the antisymmetric combination XYYXXY-YX is, because the second-derivative terms are symmetric in XX and YY and therefore cancel.

Worth noticing: μYν\partial_{\mu}Y^{\nu} is not a tensor, by §5.4. Yet (3.2.32) is. Two non-tensorial objects have combined into a tensorial one. That is not a coincidence, and it happens once more, much more consequentially, in Chapter 3.4, where the non-tensorial pieces of the connection cancel in a commutator and leave the Riemann curvature tensor behind.

8.4 · What the bracket measures

Ask first what the bracket does on the basis itself. Set X=μX=\partial_{\mu} and Y=νY=\partial_{\nu} in (3.2.32). Their components are constants, namely Xρ=δρμX^{\rho}=\delta^{\rho}{}_{\mu} and Yρ=δρνY^{\rho}=\delta^{\rho}{}_{\nu}, so both derivative terms vanish and

[μ,ν]  =  0. \big[\partial_{\mu},\,\partial_{\nu}\big] \;=\; 0. (3.2.33)

A coordinate basis commutes. Geometrically, that says stepping ε\varepsilon along the x1x^{1} direction and then ε\varepsilon along x2x^{2} lands you at exactly the same point as doing it in the other order. That is precisely what it means for the two to be coordinates of one grid. So the bracket measures the failure of two flows to commute, and it vanishes for coordinate directions by construction.

Non-coordinate frames are a different story. Take plane polar coordinates and compare the coordinate basis {r,θ}\{\partial_{r},\partial_{\theta}\} with the normalised directions e^r=r\hat e_{r}=\partial_{r} and e^θ=1rθ\hat e_{\theta}=\tfrac1r\,\partial_{\theta}, which is the basis of unit vectors used in every mechanics course. Working out [e^r,e^θ][\hat e_{r},\hat e_{\theta}] requires the product rule once, since 1r\tfrac1r is a function:

[r,  1rθ]f  =  r ⁣(1rθf)1rθ(rf)  =  (1r2)θf  +  1rrθf1rθrf, \left[\partial_{r},\;\tfrac1r\,\partial_{\theta}\right]f \;=\; \partial_{r}\!\left(\tfrac1r\,\partial_{\theta}f\right) - \tfrac1r\,\partial_{\theta}\big(\partial_{r}f\big) \;=\; \left(-\tfrac{1}{r^{2}}\right)\partial_{\theta}f \;+\; \tfrac1r\,\partial_{r}\partial_{\theta}f - \tfrac1r\,\partial_{\theta}\partial_{r}f, (3.2.34)

What we want on the right is a single vector field, so apply Clairaut once more. The last two terms cancel against each other, leaving

[e^r,e^θ]  =  1r2θ  =  1re^θ    0. \big[\hat e_{r},\,\hat e_{\theta}\big] \;=\; -\frac{1}{r^{2}}\,\partial_{\theta} \;=\; -\frac{1}{r}\,\hat e_{\theta} \;\neq\; 0. (3.2.35)

Non-zero, so the unit vectors of elementary mechanics are not a coordinate basis: there are no coordinates whose partial derivatives they are. That single fact is what Chapter 1.1 §4 was struggling with when it had to differentiate rotating unit vectors by hand. That is what made the elementary calculation awkward. It is not yet where the centrifugal and Coriolis terms come from, and Worked example 2 below separates the two. ⚑ The general statement is Frobenius's theorem, which we quote and will not use: nn independent vector fields form a coordinate basis if and only if all their brackets vanish.

In plain terms 3.2.8

Assign a direction to every point rather than to one point and a question becomes askable that was not askable before. Follow one field briefly and then the other, then do it in the opposite order, and ask whether you end up in the same place. For directions belonging to a coordinate grid the answer is yes by construction, since going east then north and north then east is what having a grid means. For two arbitrary fields the answer is generally no, and the gap is itself a field of directions.

Getting there needs care. Applying one field and then another measures the bending of a quantity rather than its slope, which is too much machinery for a direction to carry. Doing it in both orders and subtracting cancels the excess, because the excess does not notice which field came first.

The arena is now complete, and it is worth being clear about what it still cannot do. There is no way to measure a length, no way to say two paths meet at a right angle, and above all no way to compare a direction here with a direction there. That last gap blocks differentiation outright, since a rate of change is a comparison between neighbouring places. Supplying the missing comparison is the whole content of the next chapter, and the thing that has to be supplied turns out to be gravity.

9 · Worked examples

Worked example 1 — one vector, two charts, on the sphere

Take the point PP at latitude 4545^{\circ}N on the meridian of longitude zero, so that in Cartesian terms P=(12,0,12)P=\big(\tfrac{1}{\sqrt2},\,0,\,\tfrac{1}{\sqrt2}\big). Find its coordinates in both stereographic charts. Take the curve that walks north along that meridian, parametrised by latitude λ\lambda, and compute its tangent vector's components in each chart. Then check (3.2.16) against the Jacobian of the transition map.

Coordinates. With x=z=1/2x=z=1/\sqrt2 and y=0y=0, equations (3.2.1) and (3.2.2) give

(u,v)=(1/211/2,  0)=(1+2,  0)=(2.41421,0), (u,v) = \left(\frac{1/\sqrt2}{1-1/\sqrt2},\;0\right) = \big(1+\sqrt2,\;0\big) = (2.41421,\,0), (u~,v~)=(1/21+1/2,  0)=(21,  0)=(0.41421,0). (\tilde u,\tilde v) = \left(\frac{1/\sqrt2}{1+1/\sqrt2},\;0\right) = \big(\sqrt2-1,\;0\big) = (0.41421,\,0).

A first check on (3.2.5): since v=0v=0, the transition map should send u1/uu\mapsto1/u, and indeed 1/(1+2)=211/(1+\sqrt2)=\sqrt2-1. ✓

The curve, in each chart. Walking north along the meridian means (x,y,z)=(cosλ,0,sinλ)(x,y,z)=(\cos\lambda,\,0,\,\sin\lambda) with λ\lambda the latitude. Feed that into each chart and differentiate. For the north chart, the quotient rule gives

ddλ(cosλ1sinλ)=sinλ(1sinλ)+cos2λ(1sinλ)2=1sinλ(1sinλ)2=11sinλ, \dv{}{\lambda}\left(\frac{\cos\lambda}{1-\sin\lambda}\right) = \frac{-\sin\lambda\,(1-\sin\lambda)+\cos^{2}\lambda}{(1-\sin\lambda)^{2}} = \frac{1-\sin\lambda}{(1-\sin\lambda)^{2}} = \frac{1}{1-\sin\lambda},

where the numerator collapsed because sin2λ+cos2λ=1\sin^{2}\lambda+\cos^{2}\lambda=1. At λ=45\lambda=45^{\circ} this is 2+22+\sqrt2. For the south chart the identical manipulation, with the sign of sinλ\sin\lambda flipped, gives 1/(1+sinλ)-1/(1+\sin\lambda), which at 4545^{\circ} is (22)-(2-\sqrt2). So

(Vu,Vv)=(2+2,  0)=(3.41421,0),(V~u~,V~v~)=(22,  0)=(0.58579,0). \big(V^{u},V^{v}\big) = \big(2+\sqrt2,\;0\big) = (3.41421,\,0), \qquad \big(\tilde V^{\tilde u},\tilde V^{\tilde v}\big) = \big(\sqrt2-2,\;0\big) = (-0.58579,\,0).

The two charts disagree about both the size and the sign, which is exactly what one expects of components.

Jacobian. Differentiate (3.2.5). Writing ϱ2=u2+v2\varrho^{2}=u^{2}+v^{2} and using the quotient rule on u/ϱ2u/\varrho^{2},

u~u=ϱ2u(2u)ϱ4=v2u2ϱ4,u~v=2uvϱ4, \pdv{\tilde u}{u} = \frac{\varrho^{2}-u(2u)}{\varrho^{4}} = \frac{v^{2}-u^{2}}{\varrho^{4}}, \qquad \pdv{\tilde u}{v} = \frac{-2uv}{\varrho^{4}},

and by the symmetry of the map under uvu\leftrightarrow v, v~/v=(u2v2)/ϱ4\partial\tilde v/\partial v=(u^{2}-v^{2})/\varrho^{4} and v~/u=2uv/ϱ4\partial\tilde v/\partial u=-2uv/\varrho^{4}. At our point v=0v=0 and ϱ2=(1+2)2=3+22\varrho^{2}=(1+\sqrt2)^{2}=3+2\sqrt2, so the off-diagonal entries vanish and

Λμν=(1/ϱ200+1/ϱ2)=(0.17157300+0.171573), \Lambda^{\mu}{}_{\nu} = \begin{pmatrix}-1/\varrho^{2} & 0\\[2pt] 0 & +1/\varrho^{2}\end{pmatrix} = \begin{pmatrix}-0.171573 & 0\\ 0 & +0.171573\end{pmatrix},

using 1/(3+22)=3221/(3+2\sqrt2)=3-2\sqrt2.

The check. Apply (3.2.16) to the north-chart components:

Λu~uVu=(322)(2+2)=(6+32424)=(22)=0.58579   \Lambda^{\tilde u}{}_{u}\,V^{u} = -\big(3-2\sqrt2\big)\big(2+\sqrt2\big) = -\big(6+3\sqrt2-4\sqrt2-4\big) = -\big(2-\sqrt2\big) = -0.58579\;\checkmark

which is V~u~\tilde V^{\tilde u} computed independently above. The second component is 0=00=0. The transformation law holds, and it held without anybody appealing to a picture.

Two things worth noticing. The determinant of the Jacobian is 1/ϱ4-1/\varrho^{4}, which is negative everywhere on the overlap. So the stereographic transition map is orientation-reversing, which is why some books flip a sign in one of the two charts so that the atlas is oriented. And on the equator, where ϱ=1\varrho=1, the Jacobian is a reflection: its determinant is 1-1, and at v=0v=0 it is diag(1,+1)\mathrm{diag}(-1,+1) exactly. The equator is the fixed circle of the inversion, so the two charts agree there on coordinates while still disagreeing on the sense of one axis.

Worked example 2 — where the centrifugal term went

Chapter 1.1 §4 wrote Newton's second law in plane polar coordinates and found two terms appearing from nowhere, which it attributed to "the basis directions turning as you move". Identify that statement in this chapter's language, and say precisely what was turning.

The two bases. The chart is (r,θ)(r,\theta), whose coordinate basis is {r,θ}\{\partial_{r},\partial_{\theta}\}. Mechanics courses instead use {e^r,e^θ}\{\hat e_{r},\hat e_{\theta}\}, chosen to have unit length. Since the metric on the plane gives θ\partial_{\theta} length rr, the two differ by

e^r=r,e^θ=1rθ. \hat e_{r}=\partial_{r}, \qquad \hat e_{\theta}=\tfrac1r\,\partial_{\theta}.

(This step is the only one needing a metric, which is why it belongs to Chapter 3.3 and is quoted here.)

The diagnosis. By (3.2.33) the coordinate basis commutes. By (3.2.35) the normalised basis does not, since [e^r,e^θ]=1re^θ[\hat e_{r},\hat e_{\theta}]=-\tfrac1r\hat e_{\theta}. So the unit vectors are not the partial derivatives of any coordinates whatever, and differentiating a vector expanded in them necessarily produces terms involving their variation.

What this does and does not explain. It explains why the elementary calculation was awkward. It does not yet explain the centrifugal term, because even in the coordinate basis the equation of motion for a free particle is not r¨=0\ddot r=0: Chapter 1.1 §4 found r¨rθ˙2=0\ddot r - r\dot\theta^{2}=0, and r,θ\partial_{r},\partial_{\theta} commute, so the extra term survives the change of basis.

The real source is the one this chapter has been circling. Differentiating the velocity means comparing the velocity vector at one point of the trajectory with the velocity vector at a neighbouring point, and by §5.1 those live in different tangent spaces. Whatever supplies the comparison will leave a residue in the equation of motion, and in polar coordinates that residue is the centrifugal and Coriolis terms. Chapter 1.1's own problem set said as much. It promised that the residue would be called Γijk\Gamma^{i}{}_{jk} from Chapter 3.3 onward, and that gravity would turn out to belong to the same category. Both halves of that promise are about to be kept. Chapter 3.1 §2.1 already showed why gravity is nonetheless different: no single change of coordinates removes it everywhere at once.

10 · Your turn

Problem 1 — the circle needs two charts

(a) Adapt §2.1's argument to prove that no single chart covers the circle S1={(x,y):x2+y2=1}S^{1}=\{(x,y):x^{2}+y^{2}=1\}. (b) Exhibit an explicit two-chart atlas using the angle, give the overlap, and write down the transition maps. (c) The overlap has two connected pieces. Show that the transition map is a different function on each, and say why that is allowed.

Solution

(a) Suppose φ:S1VR\varphi:S^{1}\to V\subseteq\R were a chart with VV open. The circle is closed and bounded in R2\R^{2}, so by the extreme value theorem the continuous function φ\varphi attains a maximum aa at some pp. Since VV is open, an interval (aε,a+ε)(a-\varepsilon,a+\varepsilon) lies in VV, so some qq has φ(q)=a+ε/2>a\varphi(q)=a+\varepsilon/2 \gt a. Contradiction. The argument used nothing about dimension.

(b) Let φ1\varphi_{1} be the angle taken in (0,2π)(0,2\pi), on U1=S1{(1,0)}U_{1}=S^{1}\setminus\{(1,0)\}. Let φ2\varphi_{2} be the angle taken in (π,π)(-\pi,\pi), on U2=S1{(1,0)}U_{2}=S^{1}\setminus\{(-1,0)\}. Together they cover the circle. The overlap U1U2U_{1}\cap U_{2} is the circle minus two points, so its φ1\varphi_{1}-image is (0,π)(π,2π)(0,\pi)\cup(\pi,2\pi).

(c) On (0,π)(0,\pi) the transition map is ϑ2=ϑ1\vartheta_{2}=\vartheta_{1}. On (π,2π)(\pi,2\pi) it is ϑ2=ϑ12π\vartheta_{2}=\vartheta_{1}-2\pi. Both are smooth on their own piece, and that is all the definition of a smooth atlas requires: the transition map must be smooth on each connected component of the overlap, not given by one formula. Trying to force one formula is exactly what fails, and the failure is the same one as in part (a).

Problem 2 — derivations on the line

Let M=RM=\R with the single chart xx. (a) Show directly from the two defining properties that any derivation XX at a point aa satisfies X(f)=X(x)f(a)X(f)=X(x)\cdot f'(a), so that TaRT_{a}\R is one-dimensional. You may use Hadamard's factorisation f(x)=f(a)+(xa)h(x)f(x)=f(a)+(x-a)h(x) with h(a)=f(a)h(a)=f'(a). (b) Deduce that the map sending a real number cc to the derivation cddxac\,\dv{}{x}\big|_{a} is an isomorphism RTaR\R\to T_{a}\R. (c) Explain in one sentence why this makes R\R the special case in which everyone's intuition about "sliding vectors around" is correct.

Solution

(a) First, XX kills constants: X(1)=X(11)=2X(1)X(\mathbf 1)=X(\mathbf 1\cdot\mathbf 1)=2X(\mathbf 1) forces X(1)=0X(\mathbf 1)=0, and linearity extends this to any constant. Now apply XX to Hadamard's factorisation, using linearity on the sum and Leibniz on the product:

X(f)=X(f(a))=0+(aa)=0X(h)+h(a)X(xa)=f(a)X(x), X(f) = \underbrace{X\big(f(a)\big)}_{=0} + \underbrace{(a-a)}_{=0}X(h) + h(a)\,X(x-a) = f'(a)\,X(x),

where X(xa)=X(x)X(x-a)=X(x) because XX kills the constant aa.

(b) The map is linear and, by (a), onto. It is also one-to-one, because cddxac\,\dv{}{x}|_{a} applied to the function xx returns cc, so distinct cc give distinct derivations. Hence an isomorphism, and dimTaR=1\dim T_{a}\R=1.

(c) Every TaRT_{a}\R is carried onto R\R by the same isomorphism, the one reading off the coefficient of ddx\dv{}{x}, so the tangent spaces at different points are canonically identified, which is exactly the extra structure §5.2 said Rn\R^{n} has and a general manifold lacks.

Problem 3 — the bracket is not a tensor in its arguments

On R2\R^{2} take X=xX=\partial_{x} and Y=xyY=x\,\partial_{y}. (a) Compute [X,Y][X,Y] from (3.2.32) and check it by acting on an arbitrary ff and watching the second derivatives cancel. (b) Now show in general that [hX,Y]=h[X,Y](Y(h))X[hX,Y]=h\,[X,Y]-\big(Y(h)\big)X for any smooth function hh. (c) A tensor's value at a point depends only on its arguments' values at that point. Use (b) to show the bracket is not a tensor in XX and YY, and say what it depends on instead.

Solution

(a) Components: Xμ=(1,0)X^{\mu}=(1,0) and Yμ=(0,x)Y^{\mu}=(0,x). Then

[X,Y]ν=XμμYνYμμXν=xYν0=(0,1), [X,Y]^{\nu} = X^{\mu}\partial_{\mu}Y^{\nu}-Y^{\mu}\partial_{\mu}X^{\nu} = \partial_{x}Y^{\nu} - 0 = (0,1),

since XX has constant components. So [X,Y]=y[X,Y]=\partial_{y}. Direct check:

XY(f)=x(xyf)=yf+xxyf,YX(f)=xyxf, XY(f)=\partial_{x}\big(x\,\partial_{y}f\big)=\partial_{y}f + x\,\partial_{x}\partial_{y}f, \qquad YX(f)=x\,\partial_{y}\partial_{x}f,

and the two second-derivative terms are equal by Clairaut, leaving [X,Y]f=yf[X,Y]f=\partial_{y}f. ✓

(b) Act on an arbitrary ff and use the product rule at the marked step:

[hX,Y]f=hX(Yf)Y(hXf)=prodhX(Yf)(Yh)(Xf)hY(Xf)=h[X,Y]f(Yh)Xf. [hX,Y]f = hX(Yf) - Y\big(hX f\big) \overset{\text{prod}}{=} hX(Yf) - \big(Yh\big)\big(Xf\big) - h\,Y(Xf) = h[X,Y]f - \big(Yh\big)Xf.

(c) A tensor is function-linear: multiplying an argument by hh multiplies the answer by hh, so that the value at a point depends only on the arguments there. The extra term (Yh)X-(Yh)X breaks that. Hence [X,Y][X,Y] at pp depends not only on XpX_{p} and YpY_{p} but on how the fields vary near pp. That is unsurprising, since the whole content of the bracket is a comparison of two nearby points. Chapter 3.5 gives this operation its proper name, the Lie derivative.

Problem 4 — the gradient is not free

On R2\R^{2} take f(x,y)=xf(x,y)=x. (a) Write down df\dd f and its components in the chart (x,y)(x,y). (b) Now suppose someone supplies an inner product with matrix gij=(1004)g_{ij}=\begin{pmatrix}1&0\\0&4\end{pmatrix}, and a second person supplies g~ij=(2111)\tilde g_{ij}=\begin{pmatrix}2&1\\1&1\end{pmatrix}. Compute the "gradient vector" gijjfg^{ij}\partial_{j}f in each case. (c) What does this show about the status of df\dd f versus f\nabla f on a bare manifold?

Solution

(a) df=xfdx+yfdy=1dx+0dy\dd f=\partial_{x}f\,\dd x+\partial_{y}f\,\dd y = 1\cdot\dd x + 0\cdot\dd y, so the components are ωi=(1,0)\omega_{i}=(1,0). No metric was used anywhere.

(b) The inverses are gij=diag(1,14)g^{ij}=\mathrm{diag}(1,\tfrac14) and, since detg~=2111=1\det\tilde g=2\cdot1-1\cdot1=1, g~ij=(1112)\tilde g^{ij}=\begin{pmatrix}1&-1\\-1&2\end{pmatrix}. Contracting:

(f)i=gijωj=(1,0),(~f)i=g~ijωj=(1,1). (\nabla f)^{i} = g^{ij}\omega_{j} = (1,\,0), \qquad (\tilde\nabla f)^{i} = \tilde g^{ij}\omega_{j} = (1,\,-1).

Two different vectors from one function.

(c) df\dd f is determined by ff alone and is an honest object on any smooth manifold. A gradient vector is determined by ff and a choice of metric, and different choices give genuinely different vectors. They differ in direction and not merely in length, as the second case shows. So on a bare manifold df\dd f exists and f\nabla f does not. Chapter 0.6 §4 said exactly this and promised that on a manifold the distinction would stop being optional.

The brick you just laid

You have an arena that supports calculus without assuming a shape and without assuming an outside. It is a set with overlapping charts whose transition maps are smooth, and smoothness of anything on it is defined through those transition maps and is chart-independent for that reason. A sphere needs two charts and you can prove it. The two-chart atlas and its transition map (3.2.5) are explicit.

The arrow is gone. In its place is a derivation, an operator that reports the rate of change of any function at a point. The coordinate operators μ\partial_{\mu} are a basis for the space of them, which was proved in §4.4 rather than assumed. The components transform by (3.2.16), which is Chapter 2.4's definition unaltered except that the Jacobian now varies with position. Covectors are the duals, dxμ\dd x^{\mu} is the dual basis, and df=μfdxμ\dd f=\partial_{\mu}f\,\dd x^{\mu} is now a theorem rather than an abuse. Tensors are Chapter 2.4's tensors, pointwise. Vector fields have a bracket (3.2.32), built by antisymmetrising away a second-derivative term.

What is deliberately missing. No length, no angle, no comparison of vectors at different points, and therefore no derivative of a vector field. Those are not oversights. They are the next chapter.

Where this gets spent. Chapter 3.3 adds two things. The first is the metric gμν(x)g_{\mu\nu}(x), an inner product on each tangent space that varies from point to point. The second is the connection, whose entire job is to repair (3.2.18). Chapter 3.4 finds that the repair fails to be path-independent and calls the failure curvature, at which point the figure in §5.3 becomes a measurement. Chapter 3.5 takes the bracket of §8 and generalises it into the Lie derivative, and turns df\dd f into a full exterior calculus. And in Part VI the same construction is rebuilt with an internal space replacing the tangent space, at which point the connection becomes the gauge field and the curvature becomes the field strength.