Part III · General Relativity — Chapter 3.3

Metric and Connection

Two objects are missing from the arena: one that measures, and one that compares. This chapter builds both, and finds that the second of them is the gravitational field.

Where we are

Chapter 3.2 finished by listing, deliberately, what the new arena could not do. There was no way to measure a length. There was no way to say that two directions meet at a right angle. And there was no way to compare a vector at one point with a vector at another, which is the gap that stops everything. A rate of change is exactly such a comparison, so the manifold of Chapter 3.2 does not support differentiating a vector field at all. Equation (3.2.18) there was not a hard limit. It was a limit with an undefined numerator.

Two separate objects repair the two separate gaps. They are worth keeping apart, because they are logically independent. The metric supplies measurement: an inner product on each tangent space, varying smoothly from point to point. The connection supplies comparison: a rule for carrying a vector from one point to a neighbouring one. On a bare manifold you may have either without the other. What makes general relativity possible is that once you have the metric, and once you make two modest demands, the connection is forced. There is exactly one, and we derive it.

Here is the route, announced in advance. Sections 1 to 3 build the metric and spend it: the line element, then a deliberately chosen warning example, then proper time and the action of a free particle. The warning example is flat space in polar coordinates, whose metric components vary wildly while nothing whatever is curved. Section 4 exhibits the obstruction in full. The coordinate derivative of a vector field is not a tensor, and we display the exact term that spoils it. Section 5 defines the connection as whatever cancels that term, and derives the strange transformation law it must therefore obey. Section 6 uses it to transport a vector along a curve. Section 7 is the set piece: two physical demands, and a formula for the connection in terms of the metric, derived with every step shown. Section 8 derives the geodesic equation twice, by two routes that share no assumptions, and finds one equation. That is where Chapter 1.2's promise is collected. Section 9 works the sphere and the plane in polar coordinates in full.

Conventions, stated once and used without further comment. The signature is (+,,,)(+,-,-,-), as in all of Part II, so that timelike intervals are positive. Greek indices run 0,1,2,30,1,2,3 and Latin indices 1,2,31,2,3. We write gμν(x)g_{\mu\nu}(x) for the general metric and reserve ημν=diag(1,1,1,1)\eta_{\mu\nu}=\mathrm{diag}(1,-1,-1,-1) for the flat one. The connection is Γλμν\Gamma^{\lambda}{}_{\mu\nu} and the covariant derivative μ\nabla_{\mu}. GG and cc are kept explicit throughout Part III. Nothing is harder on a reader than a constant that has quietly been set to one.

Tools you'll need.  Chapter 3.2 in its entirety: the tangent space, the coordinate basis μ\partial_{\mu}, the transformation law (3.2.16) for vector components, and above all §5, which is the problem this chapter solves. Chapter 2.4 §4: the metric as the dictionary between vectors and covectors, raising and lowering, and §6, the theorem that a tensor equation true in one chart is true in all of them. Chapter 0.5 §1 and §6: inner products defined by axioms, and the spectral theorem. Chapter 0.6 §5 (the multivariable chain rule, used constantly) and §6.1 (Clairaut's theorem on mixed partials, which does the essential work in §7). Chapter 1.2 §3 (the Euler–Lagrange equation) and its Worked example 2, which minimised path length in the plane and on a sphere and promised that this chapter would replace the flat line element with a curved one. Chapter 2.5 §5, the action S=mc2 ⁣dτS=-mc^{2}\!\int\dd\tau. Chapter 2.3 §5, proper time as the length of a worldline. Chapter 0.8 §1 for linear systems of ordinary differential equations, used once in §6.

1 · The metric: an inner product at every point

Chapter 0.5 defined an inner product on a vector space by axioms rather than by a formula: a symmetric bilinear map taking two vectors to a number. Chapter 2.4 §4 put one on Minkowski spacetime, with indefinite signature, and showed that it does two jobs at once. It measures intervals, and it converts vectors into covectors.

Both of those constructions apply to a single vector space. We now have one vector space per point, so we install one inner product per point. That is the whole idea, and the definition below is just careful bookkeeping about it.

Definition — the metric

A metric on a manifold MM assigns to each point pp a map

gp: TpM×TpM    R g_{p}:\ T_{p}M\times T_{p}M \;\longrightarrow\; \R

which is bilinear (linear in each slot separately), symmetric (gp(X,Y)=gp(Y,X)g_{p}(X,Y)=g_{p}(Y,X)), and non-degenerate (if gp(X,Y)=0g_{p}(X,Y)=0 for every YY, then X=0X=0), and which varies smoothly with pp. Its components in a chart are

gμν(x)    g(μ,ν), g_{\mu\nu}(x) \;\equiv\; g\big(\partial_{\mu},\,\partial_{\nu}\big),

an array of functions of position, symmetric in its two indices.

Three remarks on that definition, each of which will be used.

(i) It is a tensor field. The map eats two vectors and returns a number, linearly in each. That is Chapter 3.2 §7's definition of a type (0,2)(0,2) tensor, word for word. So its components obey the tensor law with two factors of the inverse Jacobian. We can also derive that law directly, and it is worth doing once, because the derivation is three lines and it fixes the index placement.

What we want is a relation between the components in two charts, so we start from the one thing that must hold whichever chart we use. Expand the same two vectors both ways, and demand that g(X,Y)g(X,Y) come out as one number:

gμνXμYν  =  gαβXαYβ. g'_{\mu\nu}\,X'^{\mu}Y'^{\nu} \;=\; g_{\alpha\beta}\,X^{\alpha}Y^{\beta}. (3.3.1)

Our next goal is to get the right-hand side written in primed components, so that both sides can be compared term by term. Substitute the component law Xα=(xα/xμ)XμX^{\alpha}=(\partial x^{\alpha}/\partial x'^{\mu})X'^{\mu}, which is (3.2.16) read backwards, and do the same for YY:

gμνXμYν  =  gαβxαxμxβxν  XμYν. g'_{\mu\nu}\,X'^{\mu}Y'^{\nu} \;=\; g_{\alpha\beta}\,\pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x^{\beta}}{x'^{\nu}}\;X'^{\mu}Y'^{\nu}. (3.3.2)

Both sides now carry the same factor XμYνX'^{\mu}Y'^{\nu}, and XX' and YY' were arbitrary. So the coefficients must agree index by index, and that gives the law we were after.

  gμν  =  xαxμ  xβxν  gαβ.   \boxed{\;g'_{\mu\nu} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\;\pdv{x^{\beta}}{x'^{\nu}}\;g_{\alpha\beta}.\;} (3.3.3)

(ii) The signature is fixed. At each point gμνg_{\mu\nu} is a real symmetric matrix. The spectral theorem (Chapter 0.5 §6) therefore gives it real eigenvalues and an orthogonal eigenbasis, and non-degeneracy says that none of those eigenvalues is zero. The count of positive and negative eigenvalues is the signature. On a connected manifold that count cannot change from point to point, because changing it would need an eigenvalue to pass through zero on the way, and non-degeneracy forbids exactly that. ⚑ Making the continuity argument airtight takes Sylvester's law of inertia together with a connectedness argument. We quote the conclusion and use it.

For spacetime we impose signature (+,,,)(+,-,-,-): three minus signs and one plus, everywhere. That is what makes each tangent space a copy of the Minkowski spacetime of Chapter 2.3. It is also the precise statement that special relativity holds locally.

(iii) The line element. The traditional notation for a metric is

ds2  =  gμν(x)  dxμdxν, \dd s^{2} \;=\; g_{\mu\nu}(x)\;\dd x^{\mu}\,\dd x^{\nu}, (3.3.4)

and it deserves one sentence of honesty, because Chapter 3.2 §6 taught us what dxμ\dd x^{\mu} really is. It is a basis covector, and the product of two covectors written side by side means their symmetric tensor product. So (3.3.4) is shorthand for the tensor g=gμνdxμdxνg=g_{\mu\nu}\,\dd x^{\mu}\otimes\dd x^{\nu}, and feeding it the same tangent vector twice returns the squared length of that vector.

There is an older reading of the same symbols: "the squared length of an infinitesimal displacement". It gives the right answers, because it is the statement above with the vector left implicit. We use the notation freely from here on, having said once what licenses it.

1.1 · Raising, lowering, and the gradient that finally exists

Non-degeneracy says the matrix gμνg_{\mu\nu} is invertible at each point. Write the inverse gμνg^{\mu\nu}, defined by

gμλgλν  =  δμν. g^{\mu\lambda}\,g_{\lambda\nu} \;=\; \delta^{\mu}{}_{\nu}. (3.3.5)

Everything Chapter 2.4 §4 built now transfers unchanged, because it was pointwise algebra: VμgμνVνV_{\mu}\equiv g_{\mu\nu}V^{\nu} lowers an index, ωμgμνων\omega^{\mu}\equiv g^{\mu\nu}\omega_{\nu} raises one, and the two are inverse operations by (3.3.5). The one alteration is the one Chapter 3.2 §7 identified: the dictionary is now different at every point.

Familiar ground — the gradient, four chapters late

Chapter 0.6 §4 distinguished df\dd f, which any smooth function produces for free, from f\nabla f, which requires an inner product. Chapter 3.2 §6 pointed out that on a bare manifold the first exists and the second does not, and made that the reason the metric would have to be introduced as a separate object. The debt is now settled:

(f)μ  =  gμννf. \big(\nabla f\big)^{\mu} \;=\; g^{\mu\nu}\,\partial_{\nu}f.

Chapter 3.2's Problem 4 showed two different metrics turning one function into two genuinely different gradient vectors, pointing in different directions. That was not a defect in the problem. It is the statement that the gradient is a joint property of the function and the geometry, and on a manifold the geometry is a physical field.

The length of a vector and the angle between two are read off in the usual way, with one warning inherited from Part II. In Lorentzian signature g(X,X)g(X,X) may be positive, negative or zero. So Chapter 2.3's three-way classification applies at each point separately: XX is timelike if g(X,X)>0g(X,X)>0, spacelike if g(X,X)<0g(X,X)<0, and null if g(X,X)=0g(X,X)=0 with X0X\neq0.

Since the signature is the same everywhere, so is the classification. That is why the light cone at each point is a genuine structure rather than an artefact of a chart.

In plain terms 3.3.1

The arena handed over at the end of the last chapter was deliberately unfurnished. Nothing in it could measure a distance, nothing could say whether two directions stood at right angles, and nothing could compare a direction here with one over there. The first two gaps close with one object installed at every point, taking two directions and returning a number. Out of that number come lengths, angles, and the sorting of directions into timelike, spacelike and lightlike that the previous part built its geometry on.

Two features of the installation matter. Because it happens point by point, the numbers describing it are functions of position rather than constants, and that dependence is where gravity will eventually live. Because the thing installed eats directions and returns numbers, it belongs to the species of the measuring machines of the tensor chapter, and responds to relabelling in the one controlled way that species does.

An old debt is settled almost in passing, and it is the charge the previous chapter warned would fall due. Ever since the multivariable chapter there has been a distinction between the object a function supplies for free, reporting how fast it changes in a direction, and the arrow of steepest ascent, undefinable until length exists. On the bare arena only the first existed. A length having been supplied, the second exists too, depending on the geometry as much as on the function.

2 · Flat space in polar coordinates — a warning delivered early

This section exists to destroy a belief before it forms. The belief is: if the metric components depend on position, the space is curved.

It is false. It is also the single most common misunderstanding at this stage, and it will wreck the next two chapters for anyone holding it. So we work the counterexample now, at length, before curvature has even been defined.

2.1 · The change of chart, one step at a time

Take the ordinary Euclidean plane, with Cartesian coordinates (x,y)(x,y) and the metric everyone agrees on,

ds2  =  dx2  +  dy2,sogij  =  (1001). \dd s^{2} \;=\; \dd x^{2} \;+\; \dd y^{2}, \qquad\text{so}\qquad g_{ij} \;=\; \begin{pmatrix}1&0\\0&1\end{pmatrix}. (3.3.6)

Now change chart to polar coordinates by x=rcosθx=r\cos\theta, y=rsinθy=r\sin\theta. This is a legitimate chart on the plane minus the origin: smooth, with a smooth inverse, exactly as Chapter 3.2 §2 requires.

We want the line element rewritten in the new coordinates, and the only way in is to express dx\dd x and dy\dd y in terms of dr\dd r and dθ\dd\theta. So differentiate each coordinate, using the product rule on rcosθr\cos\theta and treating rr and θ\theta as the two variables:

dx  =  cosθ  dr    rsinθ  dθ,dy  =  sinθ  dr  +  rcosθ  dθ. \dd x \;=\; \cos\theta\;\dd r \;-\; r\sin\theta\;\dd\theta, \qquad \dd y \;=\; \sin\theta\;\dd r \;+\; r\cos\theta\;\dd\theta. (3.3.7)

What the line element needs is dx2\dd x^{2} and dy2\dd y^{2}, so square each of those two expressions in turn. Here is the first one, expanded one term at a time:

dx2  =  cos2θ  dr2    2rsinθcosθ  drdθ  +  r2sin2θ  dθ2. \dd x^{2} \;=\; \cos^{2}\theta\;\dd r^{2} \;-\; 2r\sin\theta\cos\theta\;\dd r\,\dd\theta \;+\; r^{2}\sin^{2}\theta\;\dd\theta^{2}. (3.3.8)

The second one squares the same way. Watch the middle term, because its sign comes out opposite to the one above, and that is the whole reason the sum is going to be tidy:

dy2  =  sin2θ  dr2  +  2rsinθcosθ  drdθ  +  r2cos2θ  dθ2. \dd y^{2} \;=\; \sin^{2}\theta\;\dd r^{2} \;+\; 2r\sin\theta\cos\theta\;\dd r\,\dd\theta \;+\; r^{2}\cos^{2}\theta\;\dd\theta^{2}. (3.3.9)

Now add the two of them, which is what ds2\dd s^{2} asks for. Three things happen. The cross terms are equal and opposite, so they cancel outright. The two dr2\dd r^{2} coefficients combine by cos2θ+sin2θ=1\cos^{2}\theta+\sin^{2}\theta=1. The two dθ2\dd\theta^{2} coefficients combine by that same identity, once r2r^{2} has been factored out. The result is

  ds2  =  dr2  +  r2dθ2,gij  =  (100r2).   \boxed{\;\dd s^{2} \;=\; \dd r^{2} \;+\; r^{2}\,\dd\theta^{2}, \qquad g_{ij} \;=\; \begin{pmatrix}1&0\\0&r^{2}\end{pmatrix}.\;} (3.3.10)

2.2 · Nothing happened to the plane

Let's look at what changed and what did not. The component gθθ=r2g_{\theta\theta}=r^{2} is a function of position. It is 11 at r=1r=1, it is 100100 at r=10r=10, and it grows without bound. By the eye test the metric now looks dramatically more complicated than (3.3.6).

And yet the plane is the plane. The change of chart was a smooth bijection. No point was added, removed, stretched or bent. Every geometric statement one could make survives verbatim. Three of them are worth checking explicitly, because they are the ones people reach for when asked what "flat" means.

Circumference. Take the curve r=Rr=R constant, θ\theta running from 00 to 2π2\pi. Along it dr=0\dd r=0, so (3.3.10) gives ds=Rdθ\dd s = R\,\dd\theta, and the length is

L  =  02πR  dθ  =  2πR, L \;=\; \int_{0}^{2\pi} R\;\dd\theta \;=\; 2\pi R, (3.3.11)

which is Euclid's answer exactly, with no correction of any order.

Straight lines. The curve θ=θ0\theta=\theta_{0} constant has ds=dr\dd s=\dd r, so its length between r=ar=a and r=br=b is bab-a. Radial rays are still straight lines and still have the length they always had.

Angles. The off-diagonal component vanishes, so the coordinate directions r\partial_{r} and θ\partial_{\theta} are perpendicular at every point, as radial and circular directions ought to be.

⚠ The belief this section exists to prevent

Position-dependent metric components do not mean curvature. They mean the coordinate grid is not a Cartesian one. Its lines may not be everywhere parallel, or the coordinate directions may not everywhere have unit length, or both. In (3.3.10) the basis vector θ\partial_{\theta} has length gθθ=r\sqrt{g_{\theta\theta}}=r, so it is short near the origin and long far away, and the metric components are simply reporting that fact.

The sharpest form of the warning is a comparison. The unit sphere carries ds2=dθ2+sin2θdϕ2\dd s^{2}=\dd\theta^{2}+\sin^{2}\theta\,\dd\phi^{2} and the flat plane carries ds2=dr2+r2dθ2\dd s^{2}=\dd r^{2}+r^{2}\dd\theta^{2}. Both have one constant component and one that varies with the other coordinate. One of the two spaces is curved and one is not. Inspection of the components cannot tell them apart, and therefore a genuine test for curvature has to be built. Building it is the whole of Chapter 3.4.

Two more things will be shown about (3.3.10) before this chapter ends. It is worth stating them now so that you know what is coming. In §7 we compute the connection coefficients of the flat plane in polar coordinates and find them non-zero: Γrθθ=r\Gamma^{r}{}_{\theta\theta}=-r and Γθrθ=1/r\Gamma^{\theta}{}_{r\theta}=1/r. In Chapter 3.4 we compute the curvature of the same metric and find it identically zero, component by component.

Non-zero connection, zero curvature, in a space that is flat by construction. That pair of facts is the warning in its final form, and it also explains where Chapter 1.1's centrifugal and Coriolis terms came from.

Those two coefficients are quoted here, five sections early, because the figure below needs them. The comparison the box above asks for cannot be made on the page. It can only be run. One routine transports one vector round one coordinate rectangle, and the only thing that differs between the two panels is which Γ\Gamma it is handed.

1.60
1.10 rad
0.80
1.20 rad
32% round
plane |V| start 0.61846584384265 end 0.61846584384265 drift 2.2e-16
sphere |V| start 0.61638163516818 end 0.61638163516818 drift 3.3e-16
plane marker r=2.200 θ=0.658 V^r=+0.617238 V^θ=-0.017705 Cartesian (+0.512189, +0.346645)
sphere marker θ=1.400 φ=0.686 V^θ=+0.607079 V^φ=+0.108259
closing angle: plane 4.16e-15° sphere 45.059816966° enclosed area/a² 45.059816966° difference 3.9e-12°
One routine, two spaces, two different answers. Both panels integrate the same equation, dVμ/dλ=ΓμνρVνdxρ/dλ\dd V^{\mu}/\dd\lambda = -\Gamma^{\mu}{}_{\nu\rho}V^{\nu}\,\dd x^{\rho}/\dd\lambda, by fourth-order Runge–Kutta with six hundred steps along each of the four legs. Nothing enters but the two coordinates and the connection: no embedding, no rotation of an ambient space, no formula from the text. Left, the flat plane with Γrθθ=r\Gamma^{r}{}_{\theta\theta}=-r and Γθrθ=1/r\Gamma^{\theta}{}_{r\theta}=1/r, drawn in Cartesian coordinates with its polar grid laid over the top. Right, the unit sphere with Γθφφ=sinθcosθ\Gamma^{\theta}{}_{\varphi\varphi}=-\sin\theta\cos\theta and Γφθφ=cotθ\Gamma^{\varphi}{}_{\theta\varphi}=\cot\theta. The green loop is a rectangle in coordinate space — out in the first coordinate, round in the second, back, and round again — and the region it encloses is shaded. Blue arrows are the transported vector at thirteen stations, black is what set out, purple is what came back, and orange is the marker. The length is conserved on both. The readout gives it at the start and at the end to fourteen decimal places, and across the whole family of loops the sliders reach the largest drift between the two is 1.4×10131.4\times10^{-13} — round-off, not physics. That is metric compatibility, which §7 will impose by hand as the first of its two demands, and because it holds on both it is not what tells the two spaces apart. On the sphere the closing angle is the enclosed area, the two agreeing to better than 2×10102\times10^{-10} degrees everywhere the sliders reach — the result Chapter 3.4 will prove, reached here from Γ\Gamma alone. On the plane the closing angle is zero. Over the whole two-dimensional family of rectangles the sliders reach it never exceeds 2×10132\times10^{-13} degrees, which is where double-precision arithmetic stops, not where the geometry does. Now drag the marker along the left-hand loop and read the third line: on the default rectangle VθV^{\theta} swings from +0.250000+0.250000 through +0.068182+0.068182 to 0.777808-0.777808 and back, with Γθrθ\Gamma^{\theta}{}_{r\theta} running from 1.66671.6667 to 0.45450.4545 the whole way — press show the connection at the marker and watch it — while the Cartesian pair printed beside them sits at (0.512189,0.346645)(0.512189,\,0.346645) and does not move in the sixth decimal place, because the arrow on screen is not moving at all. Press shrink both loops: the sphere's angle falls by about a factor of four each time and the plane's does not fall, because it was never anything to fall from.
In plain terms 3.3.2

Here is a habit of thought that has to be broken now rather than later. Write down the rule for measuring distances on a flat page using ordinary square coordinates and it looks utterly plain, with the same two numbers everywhere. Describe the very same page using distance-from-the-centre and angle-around-the-centre instead, and the rule acquires a factor that grows with distance and never stops growing. Nothing was done to the paper. Only the labelling changed.

The factor that appeared reports something real but modest, namely that a step of one unit in the angle coordinate carries you further when you are further out. Coordinate directions in such a description have lengths depending on where you stand, and the rule for measuring distances has to say so. That is bookkeeping about labels rather than a fact about the page.

The reason this matters so much is a coincidence of appearance. The rule for distances on a globe has exactly the same shape as the rule for the flat page in circular labels: one plain term and one term carrying a position-dependent factor. One of those surfaces is curved and one is not, and no amount of staring at the two rules will tell you which is which. A real test has to be built from scratch, and building it is the business of the chapter after this one.

3 · Proper time, and the action of a free particle

Chapter 2.3 §5 defined proper time along a worldline by c2dτ2=ds2c^{2}\dd\tau^{2}=\dd s^{2}, the interval being positive along timelike curves in our signature. That definition referred to ημν\eta_{\mu\nu} only through ds2\dd s^{2}, so replacing the flat metric by the general one carries it across untouched:

c2dτ2  =  gμν(x)  dxμdxν. c^{2}\,\dd\tau^{2} \;=\; g_{\mu\nu}(x)\;\dd x^{\mu}\,\dd x^{\nu}. (3.3.12)

Parametrise a worldline by anything at all, say λ\lambda, write x˙μ=dxμ/dλ\dot x^{\mu}=\dd x^{\mu}/\dd\lambda, and divide (3.3.12) by dλ2\dd\lambda^{2} to get the elapsed proper time as an integral:

τ  =  1cλ1λ2gμν(x)  x˙μx˙ν    dλ. \tau \;=\; \frac{1}{c}\int_{\lambda_{1}}^{\lambda_{2}} \sqrt{g_{\mu\nu}(x)\;\dot x^{\mu}\dot x^{\nu}}\;\;\dd\lambda. (3.3.13)

Two immediate readings of (3.3.12), both of which are collected debts.

A clock standing still. An observer who stays at fixed spatial coordinates has dxi=0\dd x^{i}=0, so only the 0000 term survives. With x0=ctx^{0}=ct as always,

c2dτ2  =  g00(dx0)2  =  g00c2dt2,sodτ  =  g00  dt. c^{2}\dd\tau^{2} \;=\; g_{00}\,\big(\dd x^{0}\big)^{2} \;=\; g_{00}\,c^{2}\,\dd t^{2}, \qquad\text{so}\qquad \dd\tau \;=\; \sqrt{g_{00}}\;\dd t. (3.3.14)

Chapter 3.1 §6.5 wrote exactly this line in advance and used it, together with the redshift derived from the accelerating cabin, to read off g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2} to first order. That argument used no general relativity, and it is now expressed in the notation it was waiting for.

It is also the reason the metric cannot be ημν\eta_{\mu\nu} near a mass. Two static clocks at different heights tick at different rates, and only a position-dependent g00g_{00} can say so.

The action. Chapter 2.5 §5 obtained the action of a free relativistic particle by demanding a Lorentz-invariant number attached to each history and finding that only one was available: S=mc2 ⁣dτS=-mc^{2}\!\int\dd\tau. Nothing in that argument used flatness. So

  S  =  mc2 ⁣dτ  =  mcgμν(x)x˙μx˙ν    dλ.   \boxed{\;S \;=\; -\,mc^{2}\!\int\dd\tau \;=\; -\,mc\int \sqrt{g_{\mu\nu}(x)\,\dot x^{\mu}\dot x^{\nu}}\;\;\dd\lambda.\;} (3.3.15)

Check the units before going on, since the constants are being kept explicit: gμνg_{\mu\nu} is dimensionless, xμx^{\mu} is a length, so the square root has the dimensions of a length per unit λ\lambda, and mc×mc\times length is kgm2s1\mathrm{kg\,m^{2}\,s^{-1}}, which is an action. And setting gμν=ημνg_{\mu\nu}=\eta_{\mu\nu} returns Chapter 2.5's expression character for character.

What has just been claimed, and what has not

Equation (3.3.15) is the entire dynamics of a particle in a gravitational field. There is no force term, no potential, and no coupling constant. The mass appears only as an overall factor and will cancel out of the equations of motion, which is Chapter 3.1's universality wearing new clothes. Everything a freely falling body does is contained in that integral, once gμνg_{\mu\nu} is known.

What has not been claimed is any way of finding gμνg_{\mu\nu}. The metric is at this stage an arbitrary field, and nothing so far says which one nature uses. Determining it from the matter present is Chapter 3.6, and requires curvature, which requires this chapter.

In plain terms 3.3.3

The quantity a wristwatch reads was defined in the previous part using the fixed geometry of empty spacetime, and the definition survives the move to a variable geometry with no change of wording, because it only ever referred to the rule for measuring separations. Feed it the new rule and it computes elapsed time along any history whatever. For somebody sitting still, only one entry in the rule survives, and their watch differs from the coordinate clock by the square root of that single entry.

That last sentence is a promise being collected. The equivalence-principle chapter showed, using only an accelerating cabin and the ordinary Doppler effect, that two stationary clocks at different heights tick at different rates. It then wrote down what the timekeeping entry of the geometry would have to be, and left the notation for later. The notation has arrived, and the conclusion stands unchanged.

The number attached to each possible history is likewise inherited without alteration: elapsed watch time, multiplied by the mass and the square of the speed of light, with a minus sign. The previous part cornered it into that form by demanding that everyone compute the same value. What is startling is how little there is. No force appears and no potential energy appears, and the mass factors out and drops from the equations of motion, which is the universality of falling that started this part.

4 · The coordinate derivative is not a tensor

Here is the obstruction, in full. Chapter 3.2 §5.4 asserted it and deferred the computation to this section. Chapter 3.2 §7 listed it as the single entry of Chapter 2.4's table that goes wrong on a manifold. Everything from §5 onward exists to remove it, so it is shown completely rather than summarised.

Here is the destination, announced before we set out. We differentiate the vector transformation law, find two terms where a tensor would have one, and identify the extra term exactly. It contains a second derivative of the coordinate change, and it vanishes if and only if that change is affine.

The calculation is three lines long, and each line is labelled below.

Line 1 · what we are differentiating. Chapter 3.2 (3.2.16) established

Vν  =  xνxβ  Vβ. V'^{\nu} \;=\; \pdv{x'^{\nu}}{x^{\beta}}\;V^{\beta}. (3.3.16)

To differentiate that, we need to know what the primed derivative operator even means in terms of the unprimed one, so that is the next thing to get.

Line 2 · the derivative operator in the new chart. By the chain rule (Chapter 0.6 §5), differentiating with respect to a primed coordinate means going through the unprimed ones:

xμ  =  xαxμ  xα,that isμ  =  xαxμ  α. \pdv{}{x'^{\mu}} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\;\pdv{}{x^{\alpha}}, \qquad\text{that is}\qquad \partial'_{\mu} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\;\partial_{\alpha}. (3.3.17)

Line 3 · apply it. Act with (3.3.17) on (3.3.16). The object being differentiated is a product of two position-dependent things, the Jacobian factor and the components, so the product rule of Chapter 0.1 §4 applies and produces two terms:

μVν  =  xαxμ  α ⁣(xνxβ  Vβ)  =  xαxμxνxβ  αVβthe tensor part  +  xαxμ2xνxαxβ  Vβthe obstruction. \partial'_{\mu}V'^{\nu} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\;\partial_{\alpha}\!\left(\pdv{x'^{\nu}}{x^{\beta}}\;V^{\beta}\right) \;=\; \underbrace{\pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x'^{\nu}}{x^{\beta}}\;\partial_{\alpha}V^{\beta}}_{\text{the tensor part}} \;+\; \underbrace{\pdv{x^{\alpha}}{x'^{\mu}}\,\frac{\partial^{2}x'^{\nu}}{\partial x^{\alpha}\,\partial x^{\beta}}\;V^{\beta}}_{\text{the obstruction}}. (3.3.18)

Read the two terms carefully, because the whole chapter turns on the difference between them.

The first term is exactly what Chapter 3.2 §7's definition demands of a type (1,1)(1,1) tensor: one factor of x/x\partial x'/\partial x for the upper index, one factor of x/x\partial x/\partial x' for the lower one, contracted against the components in the old chart. Had the calculation stopped there, αVβ\partial_{\alpha}V^{\beta} would be a tensor and this chapter would not exist.

The second term is the obstruction, and note precisely what it is made of. It carries a second derivative of the new coordinates with respect to the old ones, and it is proportional to the components VβV^{\beta} themselves rather than to their derivatives. Nothing about it can be absorbed into a redefinition of αVβ\partial_{\alpha}V^{\beta}, because it does not contain that object at all.

4.1 · When the obstruction vanishes, and why Part II never met it

The extra term dies exactly when 2xν/xαxβ=0\partial^{2}x'^{\nu}/\partial x^{\alpha}\partial x^{\beta}=0 for all indices. That is a differential equation for the coordinate change, and integrating it twice means

xν  =  Aνβxβ  +  bνwith A,b constant x'^{\nu} \;=\; A^{\nu}{}_{\beta}\,x^{\beta} \;+\; b^{\nu} \qquad\text{with } A,b \text{ constant} (3.3.19)

which is an affine change of coordinates. Lorentz transformations and translations are precisely of this form. That is why Chapter 2.4 could list νVμ\partial_{\nu}V^{\mu} in its table of tensors without qualification and never be caught out. The moment arbitrary smooth relabellings are admitted, as Chapter 3.2 §7 said they must be, the second derivative is generically non-zero and the entry is false.

4.2 · A concrete demonstration, in the plane

Abstract failures are easy to nod at and hard to believe. So here is the failure with numbers in it, staged in the flat plane of §2. There is no curvature anywhere in what follows, only a change of chart.

Take the constant vector field V=xV=\partial_{x}: the same arrow at every point of the plane, pointing along the xx-axis, with Cartesian components (Vx,Vy)=(1,0)(V^{x},V^{y})=(1,0). Every derivative of every component vanishes in the Cartesian chart. Now find its components in the polar chart. By (3.3.16), with primed meaning polar,

Vr  =  rxVx  =  cosθ,Vθ  =  θxVx  =  sinθr, V^{r} \;=\; \pdv{r}{x}\,V^{x} \;=\; \cos\theta, \qquad V^{\theta} \;=\; \pdv{\theta}{x}\,V^{x} \;=\; -\,\frac{\sin\theta}{r}, (3.3.20)

using r=x2+y2r=\sqrt{x^{2}+y^{2}} so that r/x=x/r=cosθ\partial r/\partial x = x/r = \cos\theta, and θ=arctan(y/x)\theta=\arctan(y/x) so that θ/x=y/r2=sinθ/r\partial\theta/\partial x = -y/r^{2} = -\sin\theta/r.

Those polar components depend on position. The question we care about is whether their coordinate derivatives come out zero, as they did in the Cartesian chart, so differentiate them and look:

θVr  =  sinθ    0,rVθ  =  +sinθr2    0. \partial_{\theta}V^{r} \;=\; -\sin\theta \;\neq\; 0, \qquad \partial_{r}V^{\theta} \;=\; +\,\frac{\sin\theta}{r^{2}} \;\neq\; 0. (3.3.21)

So the array μVν\partial_{\mu}V^{\nu} is zero in one chart and non-zero in another, for a field that is manifestly the same field. No tensor behaves like that. The transformation law (3.3.3) is linear and homogeneous in the components, so a tensor vanishing in one chart vanishes in all of them (Chapter 2.4 §6). The array μVν\partial_{\mu}V^{\nu} fails that test, and (3.3.18) says why: the polar-to-Cartesian map is not affine.

Recap — what went in, what came out

In: the vector transformation law, the chain rule, and the product rule. Nothing else.

Out: μVν\partial'_{\mu}V'^{\nu} equals the tensor answer plus a term proportional to VβV^{\beta} carrying the second derivative of the coordinate change. The coordinate derivative of a vector field is therefore not a tensor unless the charts are affinely related.

What it costs, and what it buys: the cost is that the most obvious way of differentiating a field is unusable in physics, since its answer depends on a choice nobody made. The purchase is that we now know the exact shape of the disease, which is what makes the cure of §5 constructible rather than guessed.

In plain terms 3.3.4

Differentiating a field means comparing its values at neighbouring points, and the previous chapter established that on a curved arena such a comparison is unavailable. One escape presents itself: work in a set of labels, treat the components as ordinary functions of those labels, and differentiate those. The result is a perfectly good array of numbers. The trouble is that the array describes the labelling as much as the field.

Following the calculation through shows precisely where the damage enters. Relating one labelling to another involves a matrix of rates of change, and taking a derivative makes that matrix get differentiated too, which produces a surplus term that a genuine measuring device would never carry. In the previous part the matrix was the same everywhere, so its derivative vanished and the problem never appeared; the moment arbitrary smooth relabellings are permitted, it does.

The failure is worth seeing with numbers rather than symbols, and the flat plane supplies them. Take a field of identical arrows all pointing the same way, which is as unchanging as a field can be. Describe it in circular labels and the components acquire a plain dependence on position, since the arrow points along the radius here and across it there. Differentiate those components and you get something non-zero, from a field that does not change. The array is reporting the turning of the labels, and nothing else.

5 · The connection, defined as the repair

We do not now guess an object and check that it works. We demand that the disease be cured, and read off what the cure must be.

Here is the destination. Write down the most general first-order correction to the coordinate derivative, require the corrected object to be a tensor, and derive the transformation law that the correction is thereby forced to obey. The correction does not itself come out a tensor, which is the surprise, and that is exactly why it can cancel something which is not a tensor either.

5.1 · The definition

Definition — covariant derivative and connection

Introduce n3n^{3} functions Γνμλ(x)\Gamma^{\nu}{}_{\mu\lambda}(x) per chart, entirely unspecified for the moment, and define the covariant derivative of a vector field by

μVν    μVν  +  ΓνμλVλ. \nabla_{\mu}V^{\nu} \;\equiv\; \partial_{\mu}V^{\nu} \;+\; \Gamma^{\nu}{}_{\mu\lambda}\,V^{\lambda}.

The functions Γνμλ\Gamma^{\nu}{}_{\mu\lambda} are the connection coefficients. The index convention used throughout this book is that the first lower index is the direction of differentiation. Books differ on this, so check before copying a formula.

The correction has been taken linear in VV and free of derivatives of VV. That is the least it can be if it is to cancel the obstruction term of (3.3.18), since that term is itself linear in VV and free of derivatives of VV. Nothing else about Γ\Gamma is assumed.

5.2 · What the connection must do, derived

Everything now follows from one demand, so let's write the demand down first. We want μVν\nabla_{\mu}V^{\nu} to transform as a type (1,1)(1,1) tensor, which says exactly this:

μVν  =!  xαxμ  xνxβ  αVβ. \nabla'_{\mu}V'^{\nu} \;\overset{!}{=}\; \pdv{x^{\alpha}}{x'^{\mu}}\;\pdv{x'^{\nu}}{x^{\beta}}\;\nabla_{\alpha}V^{\beta}. (3.3.22)

The plan is to write each side of that demand out in full and see what is left over. Start with the left side. Use the definition of \nabla, then substitute (3.3.18) for the coordinate-derivative piece:

μVν  =  xαxμxνxβαVβ  +  xαxμ2xνxαxβVβ  +  ΓνμλVλ. \nabla'_{\mu}V'^{\nu} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\pdv{x'^{\nu}}{x^{\beta}}\,\partial_{\alpha}V^{\beta} \;+\; \pdv{x^{\alpha}}{x'^{\mu}}\frac{\partial^{2}x'^{\nu}}{\partial x^{\alpha}\partial x^{\beta}}\,V^{\beta} \;+\; \Gamma'^{\nu}{}_{\mu\lambda}\,V'^{\lambda}. (3.3.23)

Now do the same to the other side, so that the two can be compared term against term. The right side of (3.3.22) carries αVβ\nabla_{\alpha}V^{\beta}, which is a coordinate derivative plus a connection term, so it opens up into two pieces as well:

xαxμxνxβαVβ  =  xαxμxνxβαVβ  +  xαxμxνxβΓβαγVγ. \pdv{x^{\alpha}}{x'^{\mu}}\pdv{x'^{\nu}}{x^{\beta}}\,\nabla_{\alpha}V^{\beta} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\pdv{x'^{\nu}}{x^{\beta}}\,\partial_{\alpha}V^{\beta} \;+\; \pdv{x^{\alpha}}{x'^{\mu}}\pdv{x'^{\nu}}{x^{\beta}}\,\Gamma^{\beta}{}_{\alpha\gamma}V^{\gamma}. (3.3.24)

The first terms of (3.3.23) and (3.3.24) are identical and cancel between the two sides. What is left is a requirement on Γ\Gamma':

ΓνμλVλ  =  xαxμxνxβΓβαγVγ    xαxμ2xνxαxβVβ. \Gamma'^{\nu}{}_{\mu\lambda}\,V'^{\lambda} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\pdv{x'^{\nu}}{x^{\beta}}\,\Gamma^{\beta}{}_{\alpha\gamma}\,V^{\gamma} \;-\; \pdv{x^{\alpha}}{x'^{\mu}}\frac{\partial^{2}x'^{\nu}}{\partial x^{\alpha}\partial x^{\beta}}\,V^{\beta}. (3.3.25)

We want a statement about Γ\Gamma alone, with no VV in it. To strip the vector off, every term has to be expressed through the same VλV'^{\lambda}. So use (3.3.16) in reverse, Vγ=(xγ/xλ)VλV^{\gamma}=(\partial x^{\gamma}/\partial x'^{\lambda})V'^{\lambda}, on both right-hand terms. In the second term the summed index is called β\beta, so relabel it γ\gamma first and the two terms then carry the same letter. Every term is now proportional to VλV'^{\lambda}, and VλV'^{\lambda} is arbitrary, so the coefficients must agree:

  Γνμλ  =  xαxμxγxλxνxβ  Γβαγ    xαxμxγxλ2xνxαxγ.   \boxed{\;\Gamma'^{\nu}{}_{\mu\lambda} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x^{\gamma}}{x'^{\lambda}}\,\pdv{x'^{\nu}}{x^{\beta}}\;\Gamma^{\beta}{}_{\alpha\gamma} \;-\; \pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x^{\gamma}}{x'^{\lambda}}\,\frac{\partial^{2}x'^{\nu}}{\partial x^{\alpha}\partial x^{\gamma}}.\;} (3.3.26)

5.3 · Three consequences, and the surprising one first

(a) The connection is not a tensor. Compare (3.3.26) with the tensor law of Chapter 3.2 §7. The first term is exactly right, with one Jacobian per index, in the correct places, for a type (1,2)(1,2) tensor. The second term has no business being there. A tensor law is homogeneous in the components. This one is inhomogeneous, carrying an additive piece that does not involve Γ\Gamma at all.

The consequence is startling on first meeting, and it is worth stating baldly: the connection can be zero in one chart and non-zero in another. Set Γβαγ=0\Gamma^{\beta}{}_{\alpha\gamma}=0 in (3.3.26) and the first term dies, but the second does not, so Γ0\Gamma'\neq0 in general.

That is not a defect. It is precisely the property required, because Γ\Gamma's job is to cancel a term that is itself not tensorial, and only a non-tensor can do that. Section 7 will exhibit the phenomenon concretely. The flat plane has Γ=0\Gamma=0 in Cartesian coordinates and Γ0\Gamma\neq0 in polar coordinates, and (3.3.26) is what relates the two. Chapter 3.1 §2.1 already met the same phenomenon in physical dress, when a uniform gravitational field was deleted by changing coordinates.

(b) The difference of two connections is a tensor. Suppose Γ\Gamma and Γ~\tilde\Gamma both satisfy (3.3.26). Subtract the two statements. The inhomogeneous term is identical in both and cancels exactly, leaving

(ΓΓ~)νμλ  =  xαxμxγxλxνxβ  (ΓΓ~)βαγ, \big(\Gamma'-\tilde\Gamma'\big)^{\nu}{}_{\mu\lambda} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x^{\gamma}}{x'^{\lambda}}\,\pdv{x'^{\nu}}{x^{\beta}}\;\big(\Gamma-\tilde\Gamma\big)^{\beta}{}_{\alpha\gamma}, (3.3.27)

which is the tensor law exactly. So the set of connections is not a vector space but an affine space: you cannot add two connections, but you can subtract them and get a genuine tensor. This is used twice more in the book, once in §7 and once in Part VI where the same structure reappears with an internal space in place of the tangent space.

(c) The obstruction is now exactly cancelled. By construction, not by luck. That was the demand from which (3.3.26) was derived.

5.4 · Extending \nabla to everything else

We have defined \nabla on vector fields. Two further requirements fix it on everything else, and both are natural rather than arbitrary:

  • On a scalar field, μfμf\nabla_{\mu}f\equiv\partial_{\mu}f. This is forced, because μf\partial_{\mu}f is already a tensor. It is the covector df\dd f of Chapter 3.2 §6, and nothing needs repairing.
  • \nabla obeys the Leibniz rule on products of tensors, as any object calling itself a derivative must.

From those two, the rule for covectors is derived rather than posited. Take a covector ων\omega_{\nu} and a vector VνV^{\nu}. Their contraction ωνVν\omega_{\nu}V^{\nu} is a scalar, so the first bullet says \nabla acts on it as a plain \partial, while the second says it must split across the product:

μ(ωνVν)  =  (μων)Vν  +  ωνμVν. \partial_{\mu}\big(\omega_{\nu}V^{\nu}\big) \;=\; \big(\nabla_{\mu}\omega_{\nu}\big)V^{\nu} \;+\; \omega_{\nu}\,\nabla_{\mu}V^{\nu}. (3.3.28)

The unknown in that line is μων\nabla_{\mu}\omega_{\nu}, so the goal is to isolate it. Expand both sides. On the left, the ordinary product rule gives (μων)Vν+ωνμVν(\partial_{\mu}\omega_{\nu})V^{\nu}+\omega_{\nu}\partial_{\mu}V^{\nu}. On the right, substitute the definition of μVν\nabla_{\mu}V^{\nu}:

(μων)Vν+ωνμVν  =  (μων)Vν  +  ωνμVν  +  ωνΓνμλVλ. \big(\partial_{\mu}\omega_{\nu}\big)V^{\nu} + \omega_{\nu}\partial_{\mu}V^{\nu} \;=\; \big(\nabla_{\mu}\omega_{\nu}\big)V^{\nu} \;+\; \omega_{\nu}\partial_{\mu}V^{\nu} \;+\; \omega_{\nu}\,\Gamma^{\nu}{}_{\mu\lambda}V^{\lambda}. (3.3.29)

The terms ωνμVν\omega_{\nu}\partial_{\mu}V^{\nu} appear on both sides and cancel. In the surviving Γ\Gamma term the summed indices are ν\nu and λ\lambda. Relabel them, swapping the two letters, so that the free vector index reads ν\nu like everywhere else. Both are dummies, so that changes nothing. The term becomes ωλΓλμνVν\omega_{\lambda}\Gamma^{\lambda}{}_{\mu\nu}V^{\nu}, and now every term carries a factor VνV^{\nu} with VV arbitrary. Matching coefficients:

  μων  =  μων    Γλμνωλ.   \boxed{\;\nabla_{\mu}\omega_{\nu} \;=\; \partial_{\mu}\omega_{\nu} \;-\; \Gamma^{\lambda}{}_{\mu\nu}\,\omega_{\lambda}.\;} (3.3.30)

Note the minus sign, and note that it was not chosen. It was forced by the Leibniz rule and the requirement that the pairing be a scalar. Iterating the same argument on a general tensor gives the pattern: one +Γ+\Gamma term per upper index and one Γ-\Gamma term per lower index.

Two particular cases are used later in this chapter and in the next, so let's write them out. First a tensor with two lower indices, which by the pattern picks up two minus terms:

ρTμν  =  ρTμν    ΓλρμTλν    ΓλρνTμλ, \nabla_{\rho}T_{\mu\nu} \;=\; \partial_{\rho}T_{\mu\nu} \;-\; \Gamma^{\lambda}{}_{\rho\mu}\,T_{\lambda\nu} \;-\; \Gamma^{\lambda}{}_{\rho\nu}\,T_{\mu\lambda}, (3.3.31)

and then one index up and one down, where the pattern gives a plus term for the upper index and a minus term for the lower one:

μTρν  =  μTρν  +  ΓρμλTλν    ΓλμνTρλ. \nabla_{\mu}T^{\rho}{}_{\nu} \;=\; \partial_{\mu}T^{\rho}{}_{\nu} \;+\; \Gamma^{\rho}{}_{\mu\lambda}\,T^{\lambda}{}_{\nu} \;-\; \Gamma^{\lambda}{}_{\mu\nu}\,T^{\rho}{}_{\lambda}. (3.3.32)

Equation (3.3.32) is the one the whole of Chapter 3.4 will lean on, so it is worth a second look now. A (1,1)(1,1) tensor gets three terms: a derivative, and one connection term per index, with the sign set by whether the index is up or down.

In plain terms 3.3.5

Rather than inventing a repair and hoping it works, the move is to write down the most general thing that could repair the damage and let the requirement of good behaviour decide what it is. The damage was an unwanted extra term proportional to the field itself, so the repair is an added term proportional to the field itself, carrying unknown coefficients. Call them the comparison coefficients, since their whole business is comparison at a distance; insisting that the sum behave properly under relabelling determines exactly how they must respond to relabelling, and nothing is left to choose.

What emerges has a property that reliably surprises people. They do not respond to relabelling the way an honest measuring device does. They pick up an extra additive piece, so they can be zero in one description of a situation and non-zero in another. That is not a flaw but the entire point, since their assignment is to cancel a quantity that misbehaves in exactly that way, and only something equally badly behaved can do the cancelling.

Anyone who has followed this part should feel a jolt of recognition. A quantity that can be made to vanish everywhere by choosing your description well, but which no choice removes once you look at a large enough region, is precisely how the first chapter of this part described the gravitational field. The comparison coefficients are that field.

6 · Parallel transport

Now that we have a derivative that behaves itself, we can spend it. This section turns it into the one thing Chapter 3.2 said was missing: a way of carrying a vector from one point to another.

The covariant derivative measures how a field changes. Setting it to zero along a curve is therefore the statement that the field does not change along that curve, which is exactly the comparison Chapter 3.2 §5.3 could not supply.

Definition — parallel transport

Let γ\gamma be a curve with tangent uμ=dxμ/dλu^{\mu}=\dd x^{\mu}/\dd\lambda. A vector field VμV^{\mu} defined along γ\gamma is parallel-transported along it if

uννVμ  =  0at every point of γ. u^{\nu}\,\nabla_{\nu}V^{\mu} \;=\; 0 \qquad\text{at every point of } \gamma.

That definition is compact but not yet usable. What we want is an equation we could hand to a computer, so let's unpack it. Expand the covariant derivative first, and deal with the leading term afterwards:

uννVμ  =  uννVμ  +  ΓμνρuνVρ. u^{\nu}\nabla_{\nu}V^{\mu} \;=\; u^{\nu}\partial_{\nu}V^{\mu} \;+\; \Gamma^{\mu}{}_{\nu\rho}\,u^{\nu}V^{\rho}. (3.3.33)

The first term is a directional derivative along the curve, and by the chain rule (Chapter 0.6 §5) that is an ordinary derivative with respect to the parameter:

uννVμ  =  dxνdλVμxν  =  dVμdλ. u^{\nu}\partial_{\nu}V^{\mu} \;=\; \dv{x^{\nu}}{\lambda}\,\pdv{V^{\mu}}{x^{\nu}} \;=\; \dv{V^{\mu}}{\lambda}. (3.3.34)

Put that back into the first equation and set the whole thing to zero, as the definition demands. Parallel transport is then a system of ordinary differential equations for the components:

  dVμdλ  =  Γμνρ(x(λ))  dxνdλ  Vρ.   \boxed{\;\dv{V^{\mu}}{\lambda} \;=\; -\,\Gamma^{\mu}{}_{\nu\rho}\big(x(\lambda)\big)\;\dv{x^{\nu}}{\lambda}\;V^{\rho}.\;} (3.3.35)

Three observations, all of which matter later.

It always has a solution, and only one. Equation (3.3.35) is a linear first-order system in the nn unknowns Vμ(λ)V^{\mu}(\lambda), with coefficients that are smooth functions of λ\lambda along the curve. ⚑ By the standard existence-and-uniqueness theorem for such systems, quoted here and used without proof, a value of VV at one point of the curve determines VV at every other point of it. So the comparison Chapter 3.2 lacked now exists: given a path, a vector at one end can be carried to the other.

The answer depends on the path. Equation (3.3.35) is an equation along a curve, and different curves joining the same two points give different coefficient functions and hence, in general, different answers. Chapter 3.2 §5.3's figure showed exactly this happening on a sphere, before there was any machinery to describe it. Chapter 3.4 takes the path-dependence, makes it quantitative, and calls it curvature.

Nothing yet says lengths are preserved. With Γ\Gamma still arbitrary, a transported vector may grow or shrink. Section 7 imposes the condition that stops it, and the figure below measures the result.

45°
200°
40°
length of the transported vector: start 1.00000000000000 end 1.00000000000001
angle to the direction of travel: start 50.000000° end 268.578636° change -141.421364°
not a great circle: predicted drift −sin(lat)×arc = -141.421356° [tangency 1.6e-15]
A vector carried along a path, held as parallel as the surface permits. The path is a circle of latitude on a sphere, and the blue arrows are one vector transported along it and drawn at fourteen stations. Every arrow is computed from the transport rule itself — the surface is chopped into short great-circle steps and the vector is carried across each by the rotation that takes the start of the step to its end, which is (3.3.35) integrated numerically and uses no formula from the text. The thin grey arrows are the direction of travel. Watch two readouts. The length never changes, to fourteen decimal places, at every latitude and every arc — that is metric compatibility, imposed in §7, showing up as a conserved quantity. The angle to the direction of travel changes only when the path is not a geodesic. Set the latitude to zero, or press the button: the path becomes a great circle, and the angle holds constant to six decimal places all the way round, which is §8's statement that a geodesic transports its own tangent. Raise the latitude and the angle drifts steadily, at a rate that grows as the circle tightens toward the pole — and the third readout gives the drift predicted by integrating (3.3.35) along a circle of latitude, namely minus the sine of the latitude times the arc travelled, against which the measured value may be checked. The second button plants the tangent vector itself at the start and transports that, so you can watch the transported tangent separate from the actual direction of travel — the gap between them is the sense in which a circle of latitude fails to be straight.
In plain terms 3.3.6

Setting the repaired derivative to zero along a path says the field is not changing as you walk that path, and that single sentence supplies the missing comparison. Plant a direction at one end, insist that it stay as unchanged as the arena permits at each step, and you arrive at the other end with a definite answer. The rule is an ordinary equation of the kind the differential-equations chapter solved, so a starting value fixes the result all along the route.

Notice the three words that had to be included. The answer is well defined given a path. Two routes between the same pair of points involve different equations, since the comparison coefficients are evaluated along different tracks, and their answers have no reason to agree. On a globe they visibly do not, and the previous chapter showed as much before any machinery existed to describe it. The disagreement is not a defect and no better construction removes it.

That disagreement is the whole subject of the next chapter. Make the route a small closed loop, carry a direction round it, and compare what returns with what set out. The mismatch shrinks as the loop shrinks, and the rate at which it shrinks belongs to the point rather than to the loop. That quantity is curvature, and by the end of the next chapter it will be the same thing as the tide.

7 · The Christoffel formula

This is the chapter's set piece, so we go slowly and announce everything before doing it.

Where we are going, in three steps

So far Γ\Gamma is any array obeying (3.3.26), and there are infinitely many. Two demands cut the field down to exactly one, and the point of this section is to derive it.

Demand 1 · metric compatibility, ρgμν=0\nabla_{\rho}g_{\mu\nu}=0. Parallel transport should not change the length of a vector or the angle between two. We show below that this is exactly what the equation says.

Demand 2 · vanishing torsion, Γλμν=Γλνμ\Gamma^{\lambda}{}_{\mu\nu}=\Gamma^{\lambda}{}_{\nu\mu}. Second covariant derivatives of a scalar should commute, as ordinary second partial derivatives do.

The derivation, in three steps. Step 1: write Demand 1 out in full, with the connection appearing twice. Step 2: write two more copies of the same equation with the three free indices cyclically permuted. Step 3: add two of the three and subtract the third. Four of the six connection terms then cancel in pairs against each other, leaving one, which is freed by raising an index.

7.1 · Demand 1, and what it actually says

Take two vectors VV and WW, both parallel-transported along a curve with tangent uu, and ask how their inner product changes along the way. The inner product is a scalar, so its rate of change along the curve is the ordinary derivative, which by (3.3.34) is uρρu^{\rho}\nabla_{\rho} acting on it:

ddλ(gμνVμWν)  =  uρρ(gμνVμWν). \dv{}{\lambda}\Big(g_{\mu\nu}V^{\mu}W^{\nu}\Big) \;=\; u^{\rho}\,\nabla_{\rho}\Big(g_{\mu\nu}V^{\mu}W^{\nu}\Big). (3.3.36)

The right-hand side has \nabla acting on a product of three things, so the Leibniz rule of §5.4 applies and gives one term per factor:

uρρ(gμνVμWν)  =  (uρρgμν)VμWν  +  gμν(uρρVμ)Wν  +  gμνVμ(uρρWν). u^{\rho}\nabla_{\rho}\big(g_{\mu\nu}V^{\mu}W^{\nu}\big) \;=\; \big(u^{\rho}\nabla_{\rho}g_{\mu\nu}\big)V^{\mu}W^{\nu} \;+\; g_{\mu\nu}\big(u^{\rho}\nabla_{\rho}V^{\mu}\big)W^{\nu} \;+\; g_{\mu\nu}V^{\mu}\big(u^{\rho}\nabla_{\rho}W^{\nu}\big). (3.3.37)

The last two terms vanish, because VV and WW are parallel-transported and that is precisely the statement uρρVμ=0u^{\rho}\nabla_{\rho}V^{\mu}=0. So

ddλ(gμνVμWν)  =  (uρρgμν)VμWν. \dv{}{\lambda}\Big(g_{\mu\nu}V^{\mu}W^{\nu}\Big) \;=\; \big(u^{\rho}\nabla_{\rho}g_{\mu\nu}\big)\,V^{\mu}W^{\nu}. (3.3.38)

Read (3.3.38) both ways. If ρgμν=0\nabla_{\rho}g_{\mu\nu}=0 then every inner product is constant along every curve, so lengths and angles are preserved by transport. Conversely, if inner products are preserved for all choices of VV, WW and all curves, then the coefficients of VμWνV^{\mu}W^{\nu} must vanish for every uu, giving ρgμν=0\nabla_{\rho}g_{\mu\nu}=0. The condition and the geometric statement are the same thing.

That is why the figure in §6 reports a length constant to fourteen digits. It is not an accident of the numerical scheme. It is Demand 1, which we are about to impose, showing itself.

7.2 · Demand 2, and why it is a real condition

Apply two covariant derivatives to a scalar ff. The first gives the covector νf\partial_{\nu}f. The second then differentiates a covector, so (3.3.30) is the rule that applies:

μνf  =  μνf    Γλμνλf. \nabla_{\mu}\nabla_{\nu}f \;=\; \partial_{\mu}\partial_{\nu}f \;-\; \Gamma^{\lambda}{}_{\mu\nu}\,\partial_{\lambda}f. (3.3.39)

What we want to know is whether the two derivatives commute, so form the difference. Swap μ\mu and ν\nu in that line and subtract it from the original. The ordinary second derivatives cancel by Clairaut's theorem (Chapter 0.6 §6.1), leaving

μνf    νμf  =  (ΓλμνΓλνμ)λf    Tλμνλf. \nabla_{\mu}\nabla_{\nu}f \;-\; \nabla_{\nu}\nabla_{\mu}f \;=\; -\,\Big(\Gamma^{\lambda}{}_{\mu\nu}-\Gamma^{\lambda}{}_{\nu\mu}\Big)\,\partial_{\lambda}f \;\equiv\; -\,T^{\lambda}{}_{\mu\nu}\,\partial_{\lambda}f. (3.3.40)

The array TλμνT^{\lambda}{}_{\mu\nu} is called the torsion, and two facts about it are worth deriving rather than asserting.

Torsion is a tensor, even though Γ\Gamma is not. Look at the inhomogeneous term in (3.3.26). It carries 2xν/xαxγ\partial^{2}x'^{\nu}/\partial x^{\alpha}\partial x^{\gamma}, which is symmetric in α\alpha and γ\gamma by Clairaut, and it is contracted with two Jacobians carrying the lower indices μ\mu and λ\lambda. So the whole inhomogeneous term is symmetric under μλ\mu\leftrightarrow\lambda, and it cancels when the antisymmetric part is taken. What survives transforms homogeneously, which is the tensor law exactly. Setting a tensor to zero is a statement independent of chart, so Demand 2 is a legitimate condition rather than a coordinate convention.

Setting it to zero is a genuine choice, with a reason. A connection with torsion is mathematically consistent, and nothing so far forbids it. We set it to zero because (3.3.40) shows that torsion is precisely the failure of second derivatives of a scalar to commute, and Chapter 3.2 built the entire manifold on smooth functions whose mixed partials do commute. ⚑ That general relativity chooses the torsion-free connection is an assumption about nature. Theories with torsion have been constructed and tested, and the observational situation is that no torsion has been detected. We flag the choice and proceed.

7.3 · Counting, before any algebra

It is worth knowing in advance whether a unique answer should be expected. In four dimensions:

  • Unknowns. Γλμν\Gamma^{\lambda}{}_{\mu\nu} has one free upper index (4 values) and a pair of lower indices which Demand 2 makes symmetric, so 4×5/2=104\times5/2=10 independent pairs. Total 4×10=404\times10=40 functions.
  • Equations. ρgμν=0\nabla_{\rho}g_{\mu\nu}=0 has one free index ρ\rho (4 values) and a symmetric pair μν\mu\nu (10 values). Total 4×10=404\times10=40 equations.

Forty linear equations for forty unknowns. A unique solution is exactly what one should expect, and the algebra below produces it. Had the counting come out unequal we would have known in advance that either no connection or infinitely many satisfy the demands.

7.4 · Step 1 — write the condition out

The metric is a type (0,2)(0,2) tensor, so (3.3.31) gives its covariant derivative. Setting that to zero:

(A)ρgμν    Γλρμgλν    Γλρνgμλ  =  0. \textbf{(A)}\qquad \partial_{\rho}\,g_{\mu\nu} \;-\; \Gamma^{\lambda}{}_{\rho\mu}\,g_{\lambda\nu} \;-\; \Gamma^{\lambda}{}_{\rho\nu}\,g_{\mu\lambda} \;=\; 0. (3.3.41)

7.5 · Step 2 — two cyclic copies

Equation (3.3.41) holds for every assignment of the three free indices, so we may write it again with the letters moved around. Cycle them once, ρμ\rho\to\mu, μν\mu\to\nu, νρ\nu\to\rho, meaning: wherever (A) has ρ\rho write μ\mu, wherever it has μ\mu write ν\nu, wherever it has ν\nu write ρ\rho. The summed index λ\lambda is untouched, being private to each term:

(B)μgνρ    Γλμνgλρ    Γλμρgνλ  =  0. \textbf{(B)}\qquad \partial_{\mu}\,g_{\nu\rho} \;-\; \Gamma^{\lambda}{}_{\mu\nu}\,g_{\lambda\rho} \;-\; \Gamma^{\lambda}{}_{\mu\rho}\,g_{\nu\lambda} \;=\; 0. (3.3.42)

Cycle once more, by the same rule applied to (B), so that overall ρν\rho\to\nu, μρ\mu\to\rho, νμ\nu\to\mu relative to (A):

(C)νgρμ    Γλνρgλμ    Γλνμgρλ  =  0. \textbf{(C)}\qquad \partial_{\nu}\,g_{\rho\mu} \;-\; \Gamma^{\lambda}{}_{\nu\rho}\,g_{\lambda\mu} \;-\; \Gamma^{\lambda}{}_{\nu\mu}\,g_{\rho\lambda} \;=\; 0. (3.3.43)

A third cycle would return (A), which is the check that the permutation was done correctly.

7.6 · Step 3 — the combination (B) + (C) − (A)

Now form that combination. We take it slowly, doing the derivative terms first and then the six connection terms one pair at a time. The derivative terms collect:

μgνρ  +  νgρμ    ρgμν. \partial_{\mu}g_{\nu\rho} \;+\; \partial_{\nu}g_{\rho\mu} \;-\; \partial_{\rho}g_{\mu\nu}. (3.3.44)

The six connection terms are, written out with the signs they carry after (A) has been subtracted:

from (B):Γλμνgλρ        Γλμρgνλfrom (C):Γλνρgλμ        Γλνμgρλfrom (A):+Γλρμgλν    +    Γλρνgμλ. \begin{aligned} \text{from (B):}\quad & -\,\Gamma^{\lambda}{}_{\mu\nu}\,g_{\lambda\rho} \;\;-\;\; \Gamma^{\lambda}{}_{\mu\rho}\,g_{\nu\lambda}\\[3pt] \text{from (C):}\quad & -\,\Gamma^{\lambda}{}_{\nu\rho}\,g_{\lambda\mu} \;\;-\;\; \Gamma^{\lambda}{}_{\nu\mu}\,g_{\rho\lambda}\\[3pt] \text{from }-\text{(A):}\quad & +\,\Gamma^{\lambda}{}_{\rho\mu}\,g_{\lambda\nu} \;\;+\;\; \Gamma^{\lambda}{}_{\rho\nu}\,g_{\mu\lambda}. \end{aligned} (3.3.45)

First cancellation. Compare the second term of the (B) row with the first term of the -(A) row:

Γλμρgνλ  +  Γλρμgλν  =  0. -\,\Gamma^{\lambda}{}_{\mu\rho}\,g_{\nu\lambda} \;+\; \Gamma^{\lambda}{}_{\rho\mu}\,g_{\lambda\nu} \;=\; 0. (3.3.46)

Two separate symmetries were used and both should be named. The connections match because Γλμρ=Γλρμ\Gamma^{\lambda}{}_{\mu\rho}=\Gamma^{\lambda}{}_{\rho\mu}, which is Demand 2. The metrics match because gνλ=gλνg_{\nu\lambda}=g_{\lambda\nu}, which is the symmetry of the metric from §1. With both factors equal and the signs opposite, the pair cancels exactly.

Second cancellation. Compare the first term of the (C) row with the second term of the -(A) row:

Γλνρgλμ  +  Γλρνgμλ  =  0, -\,\Gamma^{\lambda}{}_{\nu\rho}\,g_{\lambda\mu} \;+\; \Gamma^{\lambda}{}_{\rho\nu}\,g_{\mu\lambda} \;=\; 0, (3.3.47)

by exactly the same two symmetries, applied to the index pair νρ\nu\rho this time.

What survives. The two remaining terms are the first of the (B) row and the second of the (C) row:

Γλμνgλρ    Γλνμgρλ  =  2Γλμνgλρ, -\,\Gamma^{\lambda}{}_{\mu\nu}\,g_{\lambda\rho} \;-\; \Gamma^{\lambda}{}_{\nu\mu}\,g_{\rho\lambda} \;=\; -\,2\,\Gamma^{\lambda}{}_{\mu\nu}\,g_{\lambda\rho}, (3.3.48)

where again both symmetries were used: Demand 2 turns Γλνμ\Gamma^{\lambda}{}_{\nu\mu} into Γλμν\Gamma^{\lambda}{}_{\mu\nu}, and the symmetry of the metric turns gρλg_{\rho\lambda} into gλρg_{\lambda\rho}, so the two terms are identical and add.

Putting (3.3.44) and (3.3.48) together, the combination (B) + (C) − (A) reads

μgνρ  +  νgρμ    ρgμν    2Γλμνgλρ  =  0. \partial_{\mu}g_{\nu\rho} \;+\; \partial_{\nu}g_{\rho\mu} \;-\; \partial_{\rho}g_{\mu\nu} \;-\; 2\,\Gamma^{\lambda}{}_{\mu\nu}\,g_{\lambda\rho} \;=\; 0. (3.3.49)

7.7 · Step 4 — free the connection

Only one thing remains: the surviving Γ\Gamma carries a metric factor, and we want it bare. Multiply (3.3.49) through by 12gρσ\half g^{\rho\sigma} and sum over ρ\rho. On the connection term, gρσgλρ=δσλg^{\rho\sigma}g_{\lambda\rho}=\delta^{\sigma}{}_{\lambda} by (3.3.5), and the Kronecker delta then does its only job, which is to rename λ\lambda into σ\sigma and collapse the sum. Rearranging so that the connection is alone on the left, and finally relabelling the free index σλ\sigma\to\lambda and the summed index ρσ\rho\to\sigma so the result reads in the conventional letters:

  Γλμν  =  12gλσ(μgνσ  +  νgσμ    σgμν).   \boxed{\;\Gamma^{\lambda}{}_{\mu\nu} \;=\; \half\,g^{\lambda\sigma}\Big(\partial_{\mu}g_{\nu\sigma} \;+\; \partial_{\nu}g_{\sigma\mu} \;-\; \partial_{\sigma}g_{\mu\nu}\Big).\;} (3.3.50)

These are the Christoffel symbols, and (3.3.50) is the Levi-Civita connection of the metric gg. Two consistency checks, both immediate:

It is symmetric in μν\mu\nu, as Demand 2 requires. Swap the two letters on the right: the first two terms exchange with each other, and the third is already symmetric because gμν=gνμg_{\mu\nu}=g_{\nu\mu}. Nothing changes. Had the formula come out unsymmetric we would have derived a contradiction.

It vanishes when the components are constant. Suppose σgμν=0\partial_{\sigma}g_{\mu\nu}=0 everywhere, as it is in Cartesian coordinates on the plane, or in the Minkowski chart of Part II. Then every term on the right of (3.3.50) is zero, so Γ=0\Gamma=0 and μ=μ\nabla_{\mu}=\partial_{\mu}. All of Part II's calculus is recovered exactly, which it had better be.

Recap — what went in, what came out

In: the requirement that transport preserve lengths and angles, the requirement that second derivatives of a scalar commute, the symmetry of the metric, and one raising of an index.

Out: a unique connection, (3.3.50), built entirely out of the metric and its first derivatives. There is no freedom left. Once you have fixed how lengths are measured, you have fixed how vectors are compared, and you had no say in the matter.

What it cost: two assumptions, both flagged. Metric compatibility is what makes geometry geometric rather than merely differential. Vanishing torsion is a genuine physical choice, marked ⚑ in §7.2. Everything else was algebra.

7.8 · The flat plane in polar coordinates, computed

Section 2 promised this calculation. The metric is (3.3.10), so grr=1g_{rr}=1, gθθ=r2g_{\theta\theta}=r^{2}, grθ=0g_{r\theta}=0, and the inverse is grr=1g^{rr}=1, gθθ=1/r2g^{\theta\theta}=1/r^{2}, grθ=0g^{r\theta}=0. Exactly one derivative of the metric is non-zero:

rgθθ  =  2r,and every other σgμν=0. \partial_{r}g_{\theta\theta} \;=\; 2r, \qquad\text{and every other } \partial_{\sigma}g_{\mu\nu}=0. (3.3.51)

Feed that into (3.3.50), one component at a time. Since the metric is diagonal, the sum over σ\sigma has a single surviving term in each case. Take Γrθθ\Gamma^{r}{}_{\theta\theta} first, where λ=r\lambda=r picks out grrg^{rr} and both lower indices are θ\theta:

Γrθθ  =  12grr(θgθr=0+θgrθ=0rgθθ=2r)  =  121(2r)  =  r. \Gamma^{r}{}_{\theta\theta} \;=\; \half\,g^{rr}\Big(\underbrace{\partial_{\theta}g_{\theta r}}_{=0} + \underbrace{\partial_{\theta}g_{r\theta}}_{=0} - \underbrace{\partial_{r}g_{\theta\theta}}_{=2r}\Big) \;=\; \half\cdot 1\cdot(-2r) \;=\; -\,r. (3.3.52)

Now take Γθrθ\Gamma^{\theta}{}_{r\theta}, where the upper index picks out gθθ=1/r2g^{\theta\theta}=1/r^{2} and the one surviving metric derivative sits in the first slot rather than the third, so this time it arrives with a plus sign:

Γθrθ  =  12gθθ(rgθθ=2r+θgθr=0θgrθ=0)  =  121r22r  =  1r. \Gamma^{\theta}{}_{r\theta} \;=\; \half\,g^{\theta\theta}\Big(\underbrace{\partial_{r}g_{\theta\theta}}_{=2r} + \underbrace{\partial_{\theta}g_{\theta r}}_{=0} - \underbrace{\partial_{\theta}g_{r\theta}}_{=0}\Big) \;=\; \half\cdot\frac{1}{r^{2}}\cdot 2r \;=\; \frac{1}{r}. (3.3.53)

By the symmetry just established, Γθθr=Γθrθ=1/r\Gamma^{\theta}{}_{\theta r}=\Gamma^{\theta}{}_{r\theta}=1/r. The remaining four components all vanish: Γrrr\Gamma^{r}{}_{rr}, Γrrθ\Gamma^{r}{}_{r\theta}, Γθrr\Gamma^{\theta}{}_{rr} and Γθθθ\Gamma^{\theta}{}_{\theta\theta} each involve only derivatives that are zero by (3.3.51), or in the last case θgθθ=θ(r2)=0\partial_{\theta}g_{\theta\theta}=\partial_{\theta}(r^{2})=0. So

Γrθθ  =  r,Γθrθ  =  Γθθr  =  1r,all others zero. \Gamma^{r}{}_{\theta\theta} \;=\; -\,r, \qquad \Gamma^{\theta}{}_{r\theta} \;=\; \Gamma^{\theta}{}_{\theta r} \;=\; \frac{1}{r}, \qquad \text{all others zero.} (3.3.54)
⚠ Read this against §2

The plane is flat. Its connection coefficients in Cartesian coordinates are zero and in polar coordinates are not. Nothing about the plane changed between those two sentences. Only the chart did, and (3.3.26) is the law that relates the two answers. Since the first term of that law is zero here, the entire content of (3.3.54) is its inhomogeneous term.

So a non-zero connection is not evidence of curvature, any more than position-dependent metric components were. In Chapter 3.4 §8 we compute the curvature of exactly this metric and find every component identically zero, which is the third and final form of §2's warning.

Familiar ground — where the centrifugal term went

Chapter 1.1 §4 wrote Newton's second law in plane polar coordinates and found two unexplained terms, attributing them vaguely to "the basis directions turning as you move". Its problem set promised that the residue would be called Γijk\Gamma^{i}{}_{jk} from this chapter onward, and Chapter 3.2's Worked example 2 sharpened the diagnosis without being able to complete it.

Equation (3.3.54) is the residue, and §9's second worked example puts it into the equation of motion and recovers r¨rθ˙2=0\ddot r - r\dot\theta^{2}=0 and θ¨+(2/r)r˙θ˙=0\ddot\theta + (2/r)\dot r\dot\theta=0 exactly. The centrifugal term is Γrθθ\Gamma^{r}{}_{\theta\theta} and the Coriolis term is Γθrθ\Gamma^{\theta}{}_{r\theta}. They were never forces. They were the connection of a flat space in a curvilinear chart.

Chapter 3.1 §2.1 explained in advance why gravity is nonetheless different from the centrifugal term, and the difference is now sayable in one line. Both are connection coefficients. The centrifugal one can be removed everywhere at once by returning to Cartesian coordinates. The gravitational one cannot be removed anywhere except at a point, and the obstruction to removing it is the subject of the next chapter.

In plain terms 3.3.7

Up to this point the comparison coefficients were almost entirely free. Any set of them behaving correctly under relabelling would do, and there were infinitely many. Two requirements cut that down to exactly one, and neither is a technicality. The first is that carrying a direction along a path must not change its length or its angle to a companion, since a rule that quietly stretched things would be measuring nothing. The second is that two derivatives of an ordinary quantity taken in either order must agree, a property the arena was built to have.

Counting first tells you what to expect. In four dimensions the unknown coefficients number forty and the first requirement supplies forty equations, so a single answer is what should come out. The derivation obliges, by a trick worth naming: write the requirement three times with its indices rotated, add two copies and subtract the third, and four of the six unwanted terms annihilate in pairs because the comparison coefficients are symmetric and so is the measuring device.

What survives is a formula giving the comparison coefficients entirely in terms of the rule for measuring lengths and how it varies from place to place. That is a strong statement about the world rather than about the algebra. Fix how distances are measured and you have, without any further choice, fixed how directions at neighbouring points are matched up.

8 · Geodesics, twice

A geodesic is the curved-space replacement for a straight line, and there are two entirely different ways to say what a straight line is. It is the curve that never turns, and it is the shortest route between two points. In Euclidean geometry these agree, and the agreement is so familiar it is invisible. Here they are two separate definitions with two separate derivations, and we do both and compare.

8.1 · First route — the curve that transports its own tangent

"Never turns" means: the direction you are travelling in, carried forward by the transport rule of §6, is still the direction you are travelling in. In symbols, the tangent is parallel-transported along its own curve:

uννuμ  =  0,uμ  =  dxμdλ. u^{\nu}\,\nabla_{\nu}u^{\mu} \;=\; 0, \qquad u^{\mu} \;=\; \dv{x^{\mu}}{\lambda}. (3.3.55)

Expand it exactly as in §6. The first term becomes a derivative with respect to the parameter by (3.3.34), and the second is the connection term with VV replaced by uu:

duμdλ  +  Γμνρuνuρ  =  0. \dv{u^{\mu}}{\lambda} \;+\; \Gamma^{\mu}{}_{\nu\rho}\,u^{\nu}u^{\rho} \;=\; 0. (3.3.56)

Writing uμ=dxμ/dλu^{\mu}=\dd x^{\mu}/\dd\lambda in both places gives the geodesic equation:

  d2xμdλ2  +  Γμνρ  dxνdλdxρdλ  =  0.   \boxed{\;\frac{\dd^{2}x^{\mu}}{\dd\lambda^{2}} \;+\; \Gamma^{\mu}{}_{\nu\rho}\;\dv{x^{\nu}}{\lambda}\,\dv{x^{\rho}}{\lambda} \;=\; 0.\;} (3.3.57)

The parameter is not free. Equation (3.3.57) is not invariant under an arbitrary reparametrisation. It holds for affine parameters, meaning those related to each other by λaλ+b\lambda\to a\lambda+b with constants a,ba,b. Proper time is one of them, and here is the derivation of that fact.

What we want is a quantity that stays fixed along the curve, since that is what will pin the parameter down. Compute the rate of change of gμνuμuνg_{\mu\nu}u^{\mu}u^{\nu} along the curve using (3.3.38) with V=W=uV=W=u. Metric compatibility kills the first factor, and (3.3.55) kills what is left:

ddλ(gμνuμuν)  =  0. \dv{}{\lambda}\Big(g_{\mu\nu}u^{\mu}u^{\nu}\Big) \;=\; 0. (3.3.58)

So the "speed" gμνuμuνg_{\mu\nu}u^{\mu}u^{\nu} is constant along a geodesic. For a timelike geodesic we may therefore normalise it to c2c^{2} once and for all, and by (3.3.12) that choice is exactly λ=τ\lambda=\tau.

8.2 · Second route — extremise the proper time

Now forget parallel transport entirely. Take the action (3.3.15), which came from Chapter 2.5 and knows nothing about connections, and extremise it by Euler–Lagrange. Chapter 1.2 Worked example 2 did precisely this for the plane and the sphere and promised that the general case would be this chapter's business.

Drop the constant mc-mc and work with

F(x,x˙)  =  gαβ(x)x˙αx˙β,x˙μ  =  dxμdλ. F\big(x,\dot x\big) \;=\; \sqrt{g_{\alpha\beta}(x)\,\dot x^{\alpha}\dot x^{\beta}}, \qquad \dot x^{\mu} \;=\; \dv{x^{\mu}}{\lambda}. (3.3.59)

A remark on the parameter, before varying. FF is homogeneous of degree one in x˙\dot x: scale x˙\dot x by kk and FF scales by kk. So under a change of parameter λλ~\lambda\to\tilde\lambda the integrand transforms as F(x,dx/dλ~)dλ~=F(x,dx/dλ)(dλ/dλ~)dλ~=F(x,dx/dλ)dλF(x,\dd x/\dd\tilde\lambda)\,\dd\tilde\lambda = F(x,\dd x/\dd\lambda)\,(\dd\lambda/\dd\tilde\lambda)\,\dd\tilde\lambda = F(x,\dd x/\dd\lambda)\,\dd\lambda. So the integral is the same whichever parameter is used, and therefore so is its Euler–Lagrange equation. That is Chapter 1.2 §7's form-invariance in a particularly clean special case. We may therefore vary with an arbitrary λ\lambda and choose the parameter afterwards, which is what we do. Choosing it beforehand would restrict the competing paths and is not allowed.

The two derivatives. Euler–Lagrange needs FF differentiated with respect to velocity and with respect to position, so take those in turn. The argument of the square root is gαβx˙αx˙βg_{\alpha\beta}\dot x^{\alpha}\dot x^{\beta}. Differentiating it with respect to x˙μ\dot x^{\mu} hits each of the two velocity factors in turn and, since gαβg_{\alpha\beta} is symmetric, the two contributions are equal, giving 2gμβx˙β2g_{\mu\beta}\dot x^{\beta}. The chain rule on the square root supplies 1/(2F)1/(2F):

Fx˙μ  =  gμβx˙βF. \pdv{F}{\dot x^{\mu}} \;=\; \frac{g_{\mu\beta}\,\dot x^{\beta}}{F}. (3.3.60)

Now the position derivative, which is the easier of the two. Inside FF the only thing that depends on xx at all is gαβ(x)g_{\alpha\beta}(x), so the same 1/(2F)1/(2F) from the square root is all that accompanies it:

Fxμ  =  μgαβ  x˙αx˙β2F. \pdv{F}{x^{\mu}} \;=\; \frac{\partial_{\mu}g_{\alpha\beta}\;\dot x^{\alpha}\dot x^{\beta}}{2F}. (3.3.61)

Assemble Euler–Lagrange. The equation is Chapter 1.2 (1.2.22), in the form ddλ(F/x˙μ)F/xμ=0\dv{}{\lambda}\big(\partial F/\partial\dot x^{\mu}\big)-\partial F/\partial x^{\mu}=0, so drop the two derivatives just computed into their slots:

ddλ(gμβx˙βF)    μgαβx˙αx˙β2F  =  0. \dv{}{\lambda}\left(\frac{g_{\mu\beta}\dot x^{\beta}}{F}\right) \;-\; \frac{\partial_{\mu}g_{\alpha\beta}\,\dot x^{\alpha}\dot x^{\beta}}{2F} \;=\; 0. (3.3.62)

Now choose the parameter. Those factors of FF are what make the line awkward, and a good choice of parameter removes them. Take λ=τ\lambda=\tau. Then (3.3.12) gives F=gαβx˙αx˙β=cF=\sqrt{g_{\alpha\beta}\dot x^{\alpha}\dot x^{\beta}}=c, a constant, which therefore passes straight through the derivative. Multiply (3.3.62) through by cc:

ddτ(gμβx˙β)    12μgαβ  x˙αx˙β  =  0. \dv{}{\tau}\Big(g_{\mu\beta}\,\dot x^{\beta}\Big) \;-\; \half\,\partial_{\mu}g_{\alpha\beta}\;\dot x^{\alpha}\dot x^{\beta} \;=\; 0. (3.3.63)

Expand the total derivative. The factor gμβg_{\mu\beta} depends on τ\tau only through the position, so the chain rule gives dgμβ/dτ=αgμβx˙α\dd g_{\mu\beta}/\dd\tau=\partial_{\alpha}g_{\mu\beta}\dot x^{\alpha}:

gμβx¨β  +  αgμβ  x˙αx˙β    12μgαβ  x˙αx˙β  =  0. g_{\mu\beta}\,\ddot x^{\beta} \;+\; \partial_{\alpha}g_{\mu\beta}\;\dot x^{\alpha}\dot x^{\beta} \;-\; \half\,\partial_{\mu}g_{\alpha\beta}\;\dot x^{\alpha}\dot x^{\beta} \;=\; 0. (3.3.64)

Symmetrise the middle term. This is the one move in the derivation that is easy to miss, so here it is in full. The middle term is contracted with x˙αx˙β\dot x^{\alpha}\dot x^{\beta}, which is symmetric under exchanging α\alpha and β\beta. Therefore only the symmetric part of αgμβ\partial_{\alpha}g_{\mu\beta} contributes, and we may replace it by its own average with the αβ\alpha\leftrightarrow\beta relabelled copy:

αgμβ  x˙αx˙β  =  12(αgμβ+βgμα)x˙αx˙β. \partial_{\alpha}g_{\mu\beta}\;\dot x^{\alpha}\dot x^{\beta} \;=\; \half\Big(\partial_{\alpha}g_{\mu\beta} + \partial_{\beta}g_{\mu\alpha}\Big)\dot x^{\alpha}\dot x^{\beta}. (3.3.65)

Substituting (3.3.65) into (3.3.64) collects the three metric-derivative terms under a single factor of 12\half:

gμβx¨β  +  12(αgμβ+βgμαμgαβ)x˙αx˙β  =  0. g_{\mu\beta}\,\ddot x^{\beta} \;+\; \half\Big(\partial_{\alpha}g_{\mu\beta} + \partial_{\beta}g_{\mu\alpha} - \partial_{\mu}g_{\alpha\beta}\Big)\dot x^{\alpha}\dot x^{\beta} \;=\; 0. (3.3.66)

Raise the index. Multiply through by gνμg^{\nu\mu} and sum over μ\mu. On the first term gνμgμβ=δνβg^{\nu\mu}g_{\mu\beta}=\delta^{\nu}{}_{\beta}, which renames β\beta into ν\nu and collapses the sum:

x¨ν  +  12gνμ(αgμβ+βgμαμgαβ)x˙αx˙β  =  0. \ddot x^{\nu} \;+\; \half\,g^{\nu\mu}\Big(\partial_{\alpha}g_{\mu\beta} + \partial_{\beta}g_{\mu\alpha} - \partial_{\mu}g_{\alpha\beta}\Big)\dot x^{\alpha}\dot x^{\beta} \;=\; 0. (3.3.67)

Recognise the bracket. Compare the coefficient in (3.3.67) with the Christoffel formula (3.3.50). Take (3.3.50) and rename its indices λν\lambda\to\nu, μα\mu\to\alpha, νβ\nu\to\beta and σμ\sigma\to\mu, purely to get the two expressions into the same letters. It becomes

Γναβ  =  12gνμ(αgβμ+βgμαμgαβ), \Gamma^{\nu}{}_{\alpha\beta} \;=\; \half\,g^{\nu\mu}\Big(\partial_{\alpha}g_{\beta\mu} + \partial_{\beta}g_{\mu\alpha} - \partial_{\mu}g_{\alpha\beta}\Big), (3.3.68)

which is character for character the coefficient in (3.3.67), using gβμ=gμβg_{\beta\mu}=g_{\mu\beta} in the first term. Hence

d2xνdτ2  +  Γναβ  dxαdτdxβdτ  =  0, \frac{\dd^{2}x^{\nu}}{\dd\tau^{2}} \;+\; \Gamma^{\nu}{}_{\alpha\beta}\;\dv{x^{\alpha}}{\tau}\,\dv{x^{\beta}}{\tau} \;=\; 0, (3.3.69)

which is (3.3.57) exactly.

Why the agreement is worth something

Count the shared assumptions of the two derivations, because there are almost none.

The first route used parallel transport, metric compatibility and vanishing torsion, and never mentioned an action, a variational principle, or the length of anything.

The second route used the action mc2 ⁣dτ-mc^{2}\!\int\dd\tau inherited from Chapter 2.5, the Euler–Lagrange equation from Chapter 1.2, and the symmetry of x˙αx˙β\dot x^{\alpha}\dot x^{\beta}. It never mentioned transport, never imposed ρgμν=0\nabla_{\rho}g_{\mu\nu}=0, and never required the connection to be symmetric. The symmetric combination (3.3.68) appeared anyway, because only the symmetric part of anything survives contraction with x˙αx˙β\dot x^{\alpha}\dot x^{\beta}.

The same equation came out both times. Straightest and shortest are the same curve, and they are the same curve because of the theorem of §7. The connection that preserves lengths is the connection that the length functional produces. That is the content of the name Levi-Civita, and it is the reason this book's two halves, geometry and action principles, meet here.

8.3 · The Newtonian limit, and gravity as geometry

The claim of Part III is that a falling body is following a geodesic. That claim earns nothing until it reproduces Newton, so here is the check, with each approximation named as it is made.

Assumption 1 · slow motion. The body moves slowly compared with light, so dxi/dτcdt/dτ\abs{\dd x^{i}/\dd\tau}\ll c\,\dd t/\dd\tau. Then in the sum Γμνρx˙νx˙ρ\Gamma^{\mu}{}_{\nu\rho}\dot x^{\nu}\dot x^{\rho} the dominant term is by far the one with ν=ρ=0\nu=\rho=0, and we keep only it.

Assumption 2 · a static field. All time derivatives of the metric vanish, 0gμν=0\partial_{0}g_{\mu\nu}=0.

Assumption 3 · a weak field. g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2} with Φ/c21\abs{\Phi}/c^{2}\ll1, which is Chapter 3.1 §6.5's derived result, obtained there from the accelerating cabin and the Doppler effect with no general relativity used. The spatial part of the metric is gij=δij+O(Φ/c2)g_{ij}=-\delta_{ij}+O(\Phi/c^{2}). We shall not need to know that correction, and it is worth watching why.

Compute Γi00\Gamma^{i}{}_{00} from (3.3.50). Two of the three terms carry 0\partial_{0} and die by Assumption 2, leaving

Γi00  =  12giσ(0g0σ=0+0gσ0=0σg00)  =  12gijjg00, \Gamma^{i}{}_{00} \;=\; \half\,g^{i\sigma}\Big(\underbrace{\partial_{0}g_{0\sigma}}_{=0} + \underbrace{\partial_{0}g_{\sigma0}}_{=0} - \partial_{\sigma}g_{00}\Big) \;=\; -\,\half\,g^{ij}\,\partial_{j}g_{00}, (3.3.70)

where the sum over σ\sigma collapsed to spatial values because 0g00=0\partial_{0}g_{00}=0 removes the σ=0\sigma=0 term. Now insert gij=δij+O(Φ/c2)g^{ij}=-\delta^{ij}+O(\Phi/c^{2}) and jg00=2jΦ/c2\partial_{j}g_{00}=2\partial_{j}\Phi/c^{2}. The correction to gijg^{ij} multiplies a quantity that is already first order in Φ\Phi, so it contributes at second order and may be dropped. That is why the spatial metric never had to be specified:

Γi00  =  12(δij)2jΦc2  +  O ⁣(Φ2c4)  =  iΦc2  +  O ⁣(Φ2c4). \Gamma^{i}{}_{00} \;=\; -\,\half\big(-\delta^{ij}\big)\frac{2\,\partial_{j}\Phi}{c^{2}} \;+\; O\!\left(\frac{\Phi^{2}}{c^{4}}\right) \;=\; \frac{\partial_{i}\Phi}{c^{2}} \;+\; O\!\left(\frac{\Phi^{2}}{c^{4}}\right). (3.3.71)

Put that into the geodesic equation (3.3.57) for a spatial component, keeping only the 0000 term by Assumption 1:

d2xidτ2  =  Γi00(dx0dτ)2  =  iΦc2(cdtdτ)2. \frac{\dd^{2}x^{i}}{\dd\tau^{2}} \;=\; -\,\Gamma^{i}{}_{00}\left(\dv{x^{0}}{\tau}\right)^{2} \;=\; -\,\frac{\partial_{i}\Phi}{c^{2}}\left(c\,\dv{t}{\tau}\right)^{2}. (3.3.72)

For slow motion dt/dτ=1+O(v2/c2)\dd t/\dd\tau=1+O(v^{2}/c^{2}), so to leading order τ\tau and tt may be interchanged and the two factors of cc cancel the c2c^{2} in the denominator:

  d2xidt2  =  iΦ  =  gi.   \boxed{\;\frac{\dd^{2}x^{i}}{\dd t^{2}} \;=\; -\,\partial_{i}\Phi \;=\; g^{i}.\;} (3.3.73)

Newton's law of motion under gravity, recovered with no force anywhere in the derivation. The right-hand side of (3.3.73) came out of Γ\Gamma, which came out of the metric, which is the geometry. Chapter 3.1 §2 argued that gravity was eligible to be geometry. Equation (3.3.73) is the first place where it actually is.

One thing is still missing, and it is worth naming so that no more is claimed than has been shown. Nothing here determines Φ\Phi, or equivalently gμνg_{\mu\nu}. We have shown how matter moves in a given geometry. What geometry a given lump of matter produces requires field equations, which require curvature, which is Chapter 3.4 and then Chapter 3.6.

In plain terms 3.3.8

Two different sentences describe a straight line, and on a flat page they agree so quietly that nobody notices there were two. One says a straight line never turns: keep going in the direction you are already going. The other says it is the shortest route between its ends. Each survives the move to a curved arena, each becomes an equation, and the two equations are derived here separately with almost no shared assumptions.

The first derivation says that the direction of travel, carried forward by the comparison coefficients, is still the direction of travel. The second forgets them entirely, takes the number attached to each history from the previous part, and asks which history makes it stationary, using machinery built for pendulums four parts ago. Both give the same equation with the same coefficients, which is no accident: the rule preserving lengths is the one the length functional generates.

Then the test that matters. Feed in a weak, unchanging field, allow only slow motion, and use the timekeeping entry the equivalence-principle chapter obtained from an accelerating cabin and nothing else. Out comes the law of falling bodies, with the familiar acceleration appearing not as a force but as a coefficient of the geometry. Gravity has stopped pushing and become a fact about the shape of the arena. Still entirely missing is any rule saying which shape a given lump of matter produces.

9 · Worked examples

Two examples, each of which runs the whole chapter's machinery from one end to the other. The first takes a genuinely curved surface, the sphere, and goes from its metric to its straight lines. The second takes the flat plane in polar coordinates and settles the debt from Chapter 1.1.

Worked example 1 — the sphere, from metric to great circles

Take the sphere of radius aa with the metric it inherits from ordinary geometry. Compute every Christoffel symbol from (3.3.50), write down the geodesic equations, and show that the solutions are great circles.

The metric. Use the polar angle θ\theta measured from the north pole and the azimuth ϕ\phi. Moving in θ\theta at fixed ϕ\phi traverses a great circle of radius aa, so that displacement has length adθa\,\dd\theta. Moving in ϕ\phi at fixed θ\theta traverses a circle of latitude whose radius is asinθa\sin\theta, so that displacement has length asinθdϕa\sin\theta\,\dd\phi. The two directions are perpendicular. Hence

ds2  =  a2dθ2  +  a2sin2θ  dϕ2,gθθ=a2,gϕϕ=a2sin2θ,gθϕ=0. \dd s^{2} \;=\; a^{2}\,\dd\theta^{2} \;+\; a^{2}\sin^{2}\theta\;\dd\phi^{2}, \qquad g_{\theta\theta}=a^{2},\quad g_{\phi\phi}=a^{2}\sin^{2}\theta,\quad g_{\theta\phi}=0.

The inverse of a diagonal matrix is the diagonal matrix of reciprocals: gθθ=1/a2g^{\theta\theta}=1/a^{2} and gϕϕ=1/(a2sin2θ)g^{\phi\phi}=1/(a^{2}\sin^{2}\theta). Exactly one derivative is non-zero:

θgϕϕ  =  2a2sinθcosθ,all others =0. \partial_{\theta}\,g_{\phi\phi} \;=\; 2a^{2}\sin\theta\cos\theta, \qquad\text{all others } =0.

The Christoffel symbols. Because the metric is diagonal, the sum over σ\sigma in (3.3.50) keeps only σ=λ\sigma=\lambda. Work through the six independent components.

Γθϕϕ  =  12gθθ(ϕgϕθ+ϕgθϕθgϕϕ)  =  121a2(2a2sinθcosθ)  =  sinθcosθ. \Gamma^{\theta}{}_{\phi\phi} \;=\; \half g^{\theta\theta}\Big(\partial_{\phi}g_{\phi\theta}+\partial_{\phi}g_{\theta\phi}-\partial_{\theta}g_{\phi\phi}\Big) \;=\; \half\cdot\frac{1}{a^{2}}\cdot\big(-2a^{2}\sin\theta\cos\theta\big) \;=\; -\sin\theta\cos\theta. Γϕθϕ  =  12gϕϕ(θgϕϕ+ϕgϕθϕgθϕ)  =  121a2sin2θ2a2sinθcosθ  =  cotθ, \Gamma^{\phi}{}_{\theta\phi} \;=\; \half g^{\phi\phi}\Big(\partial_{\theta}g_{\phi\phi}+\partial_{\phi}g_{\phi\theta}-\partial_{\phi}g_{\theta\phi}\Big) \;=\; \half\cdot\frac{1}{a^{2}\sin^{2}\theta}\cdot 2a^{2}\sin\theta\cos\theta \;=\; \cot\theta,

and Γϕϕθ=Γϕθϕ=cotθ\Gamma^{\phi}{}_{\phi\theta}=\Gamma^{\phi}{}_{\theta\phi}=\cot\theta by symmetry. The other four vanish:

Γθθθ=Γθθϕ=Γϕθθ=Γϕϕϕ=0, \Gamma^{\theta}{}_{\theta\theta}=\Gamma^{\theta}{}_{\theta\phi}=\Gamma^{\phi}{}_{\theta\theta}=\Gamma^{\phi}{}_{\phi\phi}=0,

each because every metric derivative appearing in it is zero. Take Γϕϕϕ=12gϕϕϕgϕϕ\Gamma^{\phi}{}_{\phi\phi}=\half g^{\phi\phi}\partial_{\phi}g_{\phi\phi} as the example, and note that gϕϕ=a2sin2θg_{\phi\phi}=a^{2}\sin^{2}\theta has no ϕ\phi in it. Note also that aa has cancelled out of every symbol: the connection of a sphere does not know its radius. (Every symbol above was confirmed independently by computer algebra.)

The geodesic equations. Substitute into (3.3.57), remembering that the sum over νρ\nu\rho counts Γϕθϕ\Gamma^{\phi}{}_{\theta\phi} and Γϕϕθ\Gamma^{\phi}{}_{\phi\theta} separately, which is where the factor of 22 comes from:

θ¨    sinθcosθ  ϕ˙2  =  0,ϕ¨  +  2cotθ  θ˙ϕ˙  =  0. \ddot\theta \;-\; \sin\theta\cos\theta\;\dot\phi^{2} \;=\; 0, \qquad \ddot\phi \;+\; 2\cot\theta\;\dot\theta\,\dot\phi \;=\; 0.

The equator solves them. Put θ=π/2\theta=\pi/2 constant, so θ˙=θ¨=0\dot\theta=\ddot\theta=0. The first equation needs sinθcosθϕ˙2=0\sin\theta\cos\theta\,\dot\phi^{2}=0, and cos(π/2)=0\cos(\pi/2)=0 makes it so for any ϕ˙\dot\phi. The second needs ϕ¨+2cot(π/2)0ϕ˙=0\ddot\phi+2\cot(\pi/2)\cdot 0\cdot\dot\phi=0, and cot(π/2)=0\cot(\pi/2)=0, so ϕ¨=0\ddot\phi=0, meaning ϕ\phi advances at a constant rate. Both are satisfied. The equator, traversed at constant rate, is a geodesic.

Every great circle solves them. The metric is unchanged by any rotation of the sphere. Rotations move points around but preserve all distances, so they carry solutions of the geodesic equation to solutions. Any great circle is the image of the equator under some rotation. Hence every great circle is a geodesic. This is Chapter 1.2's argument, unchanged. What has changed is that we now have the equation rather than only the variational problem. \blacksquare

Two checks. First, the second geodesic equation can be integrated once. Multiply it by sin2θ\sin^{2}\theta and notice that the result is a total derivative:

ddλ(sin2θ  ϕ˙)  =  sin2θϕ¨+2sinθcosθθ˙ϕ˙  =  sin2θ(ϕ¨+2cotθθ˙ϕ˙)  =  0, \dv{}{\lambda}\Big(\sin^{2}\theta\;\dot\phi\Big) \;=\; \sin^{2}\theta\,\ddot\phi + 2\sin\theta\cos\theta\,\dot\theta\dot\phi \;=\; \sin^{2}\theta\Big(\ddot\phi + 2\cot\theta\,\dot\theta\dot\phi\Big) \;=\; 0,

so sin2θϕ˙\sin^{2}\theta\,\dot\phi is constant along every geodesic. That is Clairaut's relation, which Chapter 1.2's grind box obtained from the absence of ϕ\phi in the Lagrangian and identified as Noether's theorem in cartographic disguise. The two routes agree, as §8 said they must.

Second, numerical: integrating the two geodesic equations from a tilted initial direction with a fourth-order Runge–Kutta scheme and fitting a plane to the resulting points gives a plane through the centre of the sphere to a relative precision of 1.8×10131.8\times10^{-13}, while sin2θϕ˙\sin^{2}\theta\,\dot\phi holds constant to 3.3×10133.3\times10^{-13}. A plane through the centre cuts the sphere in a great circle.

Worked example 2 — the free particle in polar coordinates, and a debt from Chapter 1.1

Use (3.3.54) to write the geodesic equation of the flat plane in polar coordinates, and compare with the equations Chapter 1.1 §4 obtained by brute force. Then check the connection against the transformation law (3.3.26).

The equations. With Γrθθ=r\Gamma^{r}{}_{\theta\theta}=-r and Γθrθ=Γθθr=1/r\Gamma^{\theta}{}_{r\theta}=\Gamma^{\theta}{}_{\theta r}=1/r, equation (3.3.57) gives, for μ=r\mu=r and then μ=θ\mu=\theta,

r¨    rθ˙2  =  0,θ¨  +  2rr˙θ˙  =  0. \ddot r \;-\; r\,\dot\theta^{2} \;=\; 0, \qquad \ddot\theta \;+\; \frac{2}{r}\,\dot r\,\dot\theta \;=\; 0.

The factor of 22 in the second equation is again the two orderings Γθrθr˙θ˙\Gamma^{\theta}{}_{r\theta}\dot r\dot\theta and Γθθrθ˙r˙\Gamma^{\theta}{}_{\theta r}\dot\theta\dot r being counted separately by the double sum.

These are exactly Chapter 1.1 §4's equations for a free particle, in which rθ˙2-r\dot\theta^{2} was called the centrifugal term and (2/r)r˙θ˙(2/r)\dot r\dot\theta the Coriolis term. Multiplying the second by r2r^{2} gives ddt(r2θ˙)=0\dv{}{t}(r^{2}\dot\theta)=0, conservation of angular momentum, by exactly the manipulation used on the sphere above.

What that settles. Chapter 1.1's problem set promised these terms would be called Γijk\Gamma^{i}{}_{jk} from this chapter onward, and Chapter 3.2's Worked example 2 diagnosed them as the residue left over from comparing velocity vectors at neighbouring points without a rule for doing so. Both promises are now kept, and with a formula rather than a hand wave: the terms are connection coefficients of a flat space in a curvilinear chart, and (3.3.50) computes them.

The independent check. Section 5.3 said the polar connection ought to be recoverable from the transformation law alone, since the Cartesian connection is zero. With Γ=0\Gamma=0 in Cartesian coordinates, (3.3.26) reduces to its inhomogeneous term:

Γνμλ  =  xαxμxγxλ2xνxαxγ, \Gamma'^{\nu}{}_{\mu\lambda} \;=\; -\,\pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x^{\gamma}}{x'^{\lambda}}\,\frac{\partial^{2}x'^{\nu}}{\partial x^{\alpha}\partial x^{\gamma}},

where primed means polar. Take ν=r\nu=r, μ=λ=θ\mu=\lambda=\theta. The needed Jacobian entries are x/θ=rsinθ\partial x/\partial\theta = -r\sin\theta and y/θ=rcosθ\partial y/\partial\theta = r\cos\theta, and the needed second derivatives of r=x2+y2r=\sqrt{x^{2}+y^{2}} are, by two applications of the quotient rule,

2rx2=y2r3,2rxy=xyr3,2ry2=x2r3. \frac{\partial^{2}r}{\partial x^{2}} = \frac{y^{2}}{r^{3}}, \qquad \frac{\partial^{2}r}{\partial x\,\partial y} = -\frac{xy}{r^{3}}, \qquad \frac{\partial^{2}r}{\partial y^{2}} = \frac{x^{2}}{r^{3}}.

Assemble the four terms of the double sum, writing x=rcosθx=r\cos\theta, y=rsinθy=r\sin\theta:

Γrθθ=[r2sin2θy2r3+2r2sinθcosθxyr3(1)(1)+r2cos2θx2r3], \Gamma^{r}{}_{\theta\theta} = -\Big[ r^{2}\sin^{2}\theta\cdot\frac{y^{2}}{r^{3}} + 2r^{2}\sin\theta\cos\theta\cdot\frac{xy}{r^{3}}\cdot(-1)\cdot(-1) + r^{2}\cos^{2}\theta\cdot\frac{x^{2}}{r^{3}} \Big],

which on substituting x,yx,y becomes [rsin4θ+2rsin2θcos2θ+rcos4θ]=r(sin2θ+cos2θ)2=r-\big[r\sin^{4}\theta + 2r\sin^{2}\theta\cos^{2}\theta + r\cos^{4}\theta\big] = -r\big(\sin^{2}\theta+\cos^{2}\theta\big)^{2} = -r. That is (3.3.52), obtained by a completely different route: the first calculation used the metric and the Christoffel formula, this one used only the fact that Γ=0\Gamma=0 in Cartesian coordinates and the way Γ\Gamma transforms. (Both routes were also confirmed symbolically, and the other components agree in the same way.)

10 · Your turn

Problem 1 — every surface of revolution at once

A large family of two-dimensional metrics can be written ds2=du2+f(u)2dv2\dd s^{2}=\dd u^{2}+f(u)^{2}\,\dd v^{2} for some positive function ff. (a) Compute all Christoffel symbols in terms of ff and ff'. (b) Check your answer against §7.8 by taking f(u)=uf(u)=u, and against Worked example 1 by taking f(u)=asin(u/a)f(u)=a\sin(u/a) with u=aθu=a\theta. (c) Take f(u)=kuf(u)=ku with 0<k<10<k<1. Show that the geodesic equations are those of the plane with θ\theta replaced by kvkv, so this surface is locally indistinguishable from the flat plane. Then say what is nevertheless different about it globally.

Solution

(a) guu=1g_{uu}=1, gvv=f2g_{vv}=f^{2}, guu=1g^{uu}=1, gvv=1/f2g^{vv}=1/f^{2}, and the only non-zero metric derivative is ugvv=2ff\partial_{u}g_{vv}=2ff'. From (3.3.50),

Γuvv=12guu(ugvv)=ff,Γvuv=Γvvu=12gvvugvv=ff, \Gamma^{u}{}_{vv} = \half g^{uu}\big(-\partial_{u}g_{vv}\big) = -f f', \qquad \Gamma^{v}{}_{uv} = \Gamma^{v}{}_{vu} = \half g^{vv}\,\partial_{u}g_{vv} = \frac{f'}{f},

and all others vanish, every term in them being a derivative that is zero.

(b) With f=uf=u: Γuvv=u\Gamma^{u}{}_{vv}=-u and Γvuv=1/u\Gamma^{v}{}_{uv}=1/u, which is (3.3.54) with u=ru=r, v=θv=\theta. With u=aθu=a\theta and f=asinθf=a\sin\theta: f=df/du=cosθf'=\dd f/\dd u=\cos\theta, so Γuvv=asinθcosθ\Gamma^{u}{}_{vv}=-a\sin\theta\cos\theta and Γvuv=cosθ/(asinθ)\Gamma^{v}{}_{uv}=\cos\theta/(a\sin\theta). Converting from uu back to θ\theta divides the first by aa and multiplies the second by aa, giving sinθcosθ-\sin\theta\cos\theta and cotθ\cot\theta: Worked example 1 exactly.

(c) With f=kuf=ku, Γuvv=k2u\Gamma^{u}{}_{vv}=-k^{2}u and Γvuv=1/u\Gamma^{v}{}_{uv}=1/u. The geodesic equations are u¨k2uv˙2=0\ddot u - k^{2}u\dot v^{2}=0 and v¨+(2/u)u˙v˙=0\ddot v + (2/u)\dot u\dot v=0. Substituting ϑ=kv\vartheta=kv turns them into u¨uϑ˙2=0\ddot u - u\dot\vartheta^{2}=0 and ϑ¨+(2/u)u˙ϑ˙=0\ddot\vartheta+(2/u)\dot u\dot\vartheta=0, which are the flat plane's equations. So locally this is a plane: the map (u,v)(u,kv)(u,v)\mapsto(u,kv) is an isometry onto a piece of the plane.

Globally it is not. If vv runs over [0,2π)[0,2\pi) then ϑ=kv\vartheta=kv runs only over [0,2πk)[0,2\pi k), so the surface is a plane with a wedge of angle 2π(1k)2\pi(1-k) cut out and the edges glued, which is a cone. Walking a circle of radius RR around the apex covers a length 2πkR2\pi kR, not 2πR2\pi R. The local geometry is flat everywhere away from the apex, and the defect is invisible to any purely local measurement. This is precisely the loophole that Chapter 3.4 §8 has to be careful about when it asks whether vanishing curvature implies flatness.

Problem 2 — the inverse metric is compatible too

(a) Starting from ρgμν=0\nabla_{\rho}g_{\mu\nu}=0 and the definition (3.3.5), prove that ρgμν=0\nabla_{\rho}g^{\mu\nu}=0 as well. (b) Deduce that raising and lowering indices commutes with covariant differentiation, so that ρ(gμνVν)=gμνρVν\nabla_{\rho}\big(g_{\mu\nu}V^{\nu}\big)=g_{\mu\nu}\nabla_{\rho}V^{\nu} and it is unambiguous to write ρVμ\nabla_{\rho}V_{\mu}. (c) Show that ρδμν=0\nabla_{\rho}\delta^{\mu}{}_{\nu}=0 directly from (3.3.32), and say why this had to be true.

Solution

(a) Apply ρ\nabla_{\rho} to (3.3.5) and use the Leibniz rule:

0  =  ρδμν  =  ρ(gμλgλν)  =  (ρgμλ)gλν  +  gμλρgλν=0, 0 \;=\; \nabla_{\rho}\delta^{\mu}{}_{\nu} \;=\; \nabla_{\rho}\big(g^{\mu\lambda}g_{\lambda\nu}\big) \;=\; \big(\nabla_{\rho}g^{\mu\lambda}\big)g_{\lambda\nu} \;+\; g^{\mu\lambda}\underbrace{\nabla_{\rho}g_{\lambda\nu}}_{=\,0},

using part (c) for the first equality. So (ρgμλ)gλν=0(\nabla_{\rho}g^{\mu\lambda})g_{\lambda\nu}=0. To free the inverse metric from that trailing factor, contract with gνσg^{\nu\sigma}, which turns gλνgνσg_{\lambda\nu}g^{\nu\sigma} into δλσ\delta_{\lambda}{}^{\sigma} and renames λ\lambda to σ\sigma, giving ρgμσ=0\nabla_{\rho}g^{\mu\sigma}=0.

(b) By Leibniz, ρ(gμνVν)=(ρgμν)Vν+gμνρVν\nabla_{\rho}(g_{\mu\nu}V^{\nu})=(\nabla_{\rho}g_{\mu\nu})V^{\nu}+g_{\mu\nu}\nabla_{\rho}V^{\nu}, and the first term is zero. So it does not matter whether one lowers then differentiates or differentiates then lowers, and ρVμ\nabla_{\rho}V_{\mu} has one meaning. This is used constantly from Chapter 3.4 onward and is worth noticing as a convenience that had to be earned.

(c) From (3.3.32) with Tρν=δρνT^{\rho}{}_{\nu}=\delta^{\rho}{}_{\nu}: the ordinary derivative of a constant array is zero, and the two connection terms are +ΓρμλδλνΓλμνδρλ=ΓρμνΓρμν=0+\Gamma^{\rho}{}_{\mu\lambda}\delta^{\lambda}{}_{\nu}-\Gamma^{\lambda}{}_{\mu\nu}\delta^{\rho}{}_{\lambda} = \Gamma^{\rho}{}_{\mu\nu}-\Gamma^{\rho}{}_{\mu\nu}=0, each delta having renamed one index. It had to be true because δμν\delta^{\mu}{}_{\nu} is the identity map on each tangent space, and that map is the same object in every chart. A tensor whose components are the same numbers in every chart cannot have a non-zero derivative.

Problem 3 — what the geodesic equation cannot see

Let Γ\Gamma and Γ~\tilde\Gamma be two connections with Γ~λμν=Γλμν+Sλμν\tilde\Gamma^{\lambda}{}_{\mu\nu}=\Gamma^{\lambda}{}_{\mu\nu}+S^{\lambda}{}_{\mu\nu}. (a) Show SS is a tensor. (b) Show that the two connections have the same geodesics if and only if SλμνS^{\lambda}{}_{\mu\nu} is antisymmetric in μν\mu\nu. (c) Conclude that the geodesic equation alone cannot detect torsion, and say which of §7's two demands is therefore doing the work of singling out the Levi-Civita connection when we require the two derivations of §8 to agree.

Solution

(a) This is (3.3.27): both connections obey (3.3.26) with the same inhomogeneous term, which cancels on subtraction.

(b) The geodesic equations differ by Sμνρx˙νx˙ρS^{\mu}{}_{\nu\rho}\dot x^{\nu}\dot x^{\rho}. The product x˙νx˙ρ\dot x^{\nu}\dot x^{\rho} is symmetric in νρ\nu\rho, so contracting it with SS picks out only the symmetric part of SS, by exactly the argument of (3.3.65). The difference therefore vanishes for all curves precisely when Sμ(νρ)=0S^{\mu}{}_{(\nu\rho)}=0, that is when SS is antisymmetric in its lower indices.

(c) The torsion is Tλμν=ΓλμνΓλνμT^{\lambda}{}_{\mu\nu}=\Gamma^{\lambda}{}_{\mu\nu}-\Gamma^{\lambda}{}_{\nu\mu}, which is antisymmetric in μν\mu\nu by construction. By (b), adding any antisymmetric SS changes the torsion but not a single geodesic. So no observation of freely falling bodies can determine the torsion, and §8's agreement between the two routes is therefore secured by metric compatibility alone. The variational derivation produced a symmetric coefficient because only the symmetric part survived, not because torsion was ruled out. Demand 2 remains a separate assumption, which is why §7.2 flagged it.

Problem 4 — standing still is an accelerated motion

In a static metric, consider an observer who stays at fixed spatial coordinates. (a) Show that their four-velocity is uμ=(c/g00,0,0,0)u^{\mu}=\big(c/\sqrt{g_{00}},\,0,\,0,\,0\big). (b) Show that their four-acceleration aμuννuμa^{\mu}\equiv u^{\nu}\nabla_{\nu}u^{\mu} has spatial components ai=c2Γi00/g00a^{i}=c^{2}\Gamma^{i}{}_{00}/g_{00}. (c) Evaluate it in the weak field of §8.3 and interpret the answer.

Solution

(a) Fixed spatial coordinates means ui=0u^{i}=0. Normalisation gμνuμuν=c2g_{\mu\nu}u^{\mu}u^{\nu}=c^{2} then reads g00(u0)2=c2g_{00}(u^{0})^{2}=c^{2}, so u0=c/g00u^{0}=c/\sqrt{g_{00}}, taking the root that points to the future.

(b) Expand aμ=duμ/dτ+Γμνρuνuρa^{\mu}=\dd u^{\mu}/\dd\tau+\Gamma^{\mu}{}_{\nu\rho}u^{\nu}u^{\rho}. The metric is static and the observer does not move, so uμu^{\mu} is the same at every point of the worldline and the first term vanishes. Only ν=ρ=0\nu=\rho=0 contributes to the second, giving aμ=Γμ00(u0)2=c2Γμ00/g00a^{\mu}=\Gamma^{\mu}{}_{00}(u^{0})^{2}=c^{2}\Gamma^{\mu}{}_{00}/g_{00}.

(c) With (3.3.71), Γi00=iΦ/c2\Gamma^{i}{}_{00}=\partial_{i}\Phi/c^{2}, and g001g_{00}\approx1, so ai=iΦa^{i}=\partial_{i}\Phi. For Φ=GM/r\Phi=-GM/r this is +GMxi/r3+GMx^{i}/r^{3}, pointing away from the mass with magnitude GM/r2GM/r^{2}.

The interpretation is the one Chapter 3.1 kept insisting on. Free fall is unaccelerated. It is standing on the floor that requires a four-acceleration, supplied by the floor pushing up, of exactly the magnitude Newton would have called the gravitational field. The accelerometer in your pocket agrees: it reads 9.8 ms29.8\ \mathrm{m\,s^{-2}} on the table and zero in free fall. Nothing in this problem used curvature, only the connection.

The brick you just laid

The arena now measures and compares. The metric gμν(x)g_{\mu\nu}(x) is an inner product on each tangent space, with signature (+,,,)(+,-,-,-) everywhere, and it supplies lengths, angles, the classification of vectors into timelike, spacelike and null, the raising and lowering of indices, and the gradient, which settles a debt from Chapter 0.6. Proper time and the free particle's action S=mc2 ⁣dτS=-mc^{2}\!\int\dd\tau transfer from Part II with η\eta replaced by gg and no other change.

The connection was not guessed. Section 4 showed in full that μVν\partial_{\mu}V^{\nu} fails to be a tensor, and it exhibited the offending second-derivative term. Section 5 defined Γ\Gamma as whatever cancels that term and derived the inhomogeneous law (3.3.26) it must therefore obey, from which it follows that Γ\Gamma is not a tensor and can be zero in one chart and non-zero in another. Section 7 imposed metric compatibility and vanishing torsion, counted forty equations against forty unknowns, and derived the unique answer (3.3.50) by the cyclic-permutation trick. Section 8 derived the geodesic equation twice, once as a curve that transports its own tangent and once by extremising proper time, and the two agreed, which collects Chapter 1.2's promise in full. It then recovered Newton's law of falling bodies in the weak, slow, static limit, using nothing beyond Chapter 3.1's independently derived g00g_{00}.

The warning you should be carrying. Position-dependent metric components do not mean curvature: the flat plane in polar coordinates has them. Non-zero connection coefficients do not mean curvature either: the same plane has Γrθθ=r\Gamma^{r}{}_{\theta\theta}=-r. Both facts are about the chart. A genuine test has not yet been built.

Where this gets spent. Chapter 3.4 builds the test, by carrying a vector around a closed loop and asking whether it comes back unchanged. The answer is the Riemann tensor, and the difference between two geodesics governed by it is Chapter 3.1's tidal equation. Chapter 3.5 generalises §6's transport into the Lie derivative and finds the conserved quantities that make Chapter 3.7 solvable. Chapter 3.6 writes field equations for gμνg_{\mu\nu}, which is the one thing this chapter conspicuously did not supply. And in Part VI the whole construction is run again with an internal space in place of the tangent space, at which point (3.3.26) becomes the transformation law of a gauge field.