Part I · The Action Principle — Chapter 1.4

Noether's Theorem

Every continuous symmetry of the action is a conservation law. It is the most important theorem in this book, and the proof is half a page.

Where we are

Conservation laws keep turning up, and every time they have been earned by a different trick. In Chapter 1.1 energy appeared because someone thought to dot F=maF=ma with v\vv v. Momentum appeared because Newton's third law made the internal forces cancel in pairs. In Chapter 0.8 the oscillator's energy appeared as a "first integral" found by inspection, and a three-mass chain turned out to have a zero-frequency mode that was really momentum conservation in disguise. In Chapter 1.2 the Beltrami identity produced a conserved EE out of the assumption L/t=0\partial L/\partial t=0, and the sphere's geodesics produced Clairaut's relation out of the absence of ϕ\phi from the integrand. Five separate conjuring tricks.

They are one theorem. Emmy Noether proved it in 1918, and its content is not a computational convenience. It is this: conservation laws are not facts about forces. They are facts about symmetry. Momentum is conserved because space is the same here as over there. Energy is conserved because time is the same now as it was yesterday. Neither statement mentions what the matter is made of, and neither can be broken by discovering a new force. The only way to break one is to discover that space or time is not uniform.

That is worth the price of admission on its own. But it is not why this chapter is the hinge of the book. Read the theorem backwards and it stops being a way of finding conservation laws and becomes a way of building theories. If you know in advance what is conserved, you know what symmetry the action must have. And demanding that symmetry is often enough to determine the action almost uniquely.

From Part II onward that is the entire method. We will not discover the Dirac equation or the Standard Model by looking at data and guessing. We will write down a symmetry group, ask for the most general action invariant under it, and read off the physics. Chapter 1.2 gave us the machine that turns an action into equations of motion. This chapter gives us the only known reason to prefer one action over another.

It also closes Part I. Four chapters ago mechanics was a list of forces.

Tools you'll need  — Chapter 0.6: the multivariable chain rule, and the equality of mixed partial derivatives (the proof turns on it twice). Chapter 0.7: the continuity equation tρ+J=0\partial_{t}\rho+\nabla\cdot\vv J=0 and the divergence theorem, since §6 shows that the field version of this theorem produces exactly that equation. Chapter 1.2: the action S=LdtS=\int L\,\dd t, the Euler–Lagrange equation, and above all its Problem 4, which established that adding a total time derivative dFdt\dv{F}{t} to LL changes nothing physical. That freedom is not a footnote here. It is part of the hypothesis of the theorem, and §1 shows a symmetry that fails without it. Chapter 1.3: the Poisson bracket and the fact that a phase-space function generates an infinitesimal transformation. §7 needs it, and states what it needs.

1 · What a symmetry is, precisely

The word "symmetry" in ordinary speech means a shape looks the same after you flip it. The word in physics means something narrower and more useful, and getting the definition exactly right is most of the work. Once the hypothesis is stated properly, the proof in §2 runs to four lines. So this section is all definition, and everything afterwards is spending it.

1.1 · A continuous family of transformations

Start with a system described by coordinates q1(t),,qn(t)q_{1}(t),\dots,q_{n}(t). A continuous transformation is a recipe that takes any path and deforms it, smoothly, by an amount controlled by a single real parameter ϵ\epsilon:

qi(t)    qi(t,ϵ),withqi(t,0)=qi(t). q_{i}(t) \;\longmapsto\; q_{i}(t,\epsilon), \qquad\text{with}\qquad q_{i}(t,0) = q_{i}(t). (1.4.1)

The condition qi(t,0)=qi(t)q_{i}(t,0)=q_{i}(t) says that ϵ=0\epsilon=0 is the do-nothing transformation. That is what makes the family continuous in the sense we need: we can turn the transformation on gradually, and therefore we can differentiate with respect to how much of it we have turned on.

Only the derivative at ϵ=0\epsilon=0 will matter. Define

Ki(q,q˙,t)    qi(t,ϵ)ϵϵ=0,so thatδqi    qi(t,ϵ)qi(t)  =  ϵKi  +  O(ϵ2). \begin{aligned} K_{i}\big(q,\dot q,t\big) \;&\equiv\; \left.\pdv{q_{i}(t,\epsilon)}{\epsilon}\right|_{\epsilon=0},\\[5pt] \text{so that}\qquad \delta q_{i} \;\equiv\; q_{i}(t,\epsilon)-q_{i}(t) \;&=\; \epsilon\,K_{i} \;+\; O(\epsilon^{2}). \end{aligned} (1.4.2)

KiK_{i} is called the generator of the transformation. It is a vector field on configuration space, telling you which way each coordinate moves when you nudge ϵ\epsilon. Here are three examples, all of which we will use:

TransformationActing onGenerator KK
Translate by ϵ\epsilon along n\vv nxx+ϵn\vv x \mapsto \vv x + \epsilon\,\vv nK=n\vv K = \vv n
Rotate by ϵ\epsilon about n\vv nxx+ϵn×x+O(ϵ2)\vv x \mapsto \vv x + \epsilon\,\vv n\times\vv x + O(\epsilon^{2})K=n×x\vv K = \vv n\times\vv x
Boost by ϵu\epsilon\vv uxx+ϵut\vv x \mapsto \vv x + \epsilon\,\vv u\,tK=ut\vv K = \vv u\,t

The rotation entry is the only one that needs an argument, so let's give it one. A rotation by a small angle ϵ\epsilon about the axis n\vv n moves the tip of x\vv x along a circle of radius xsinθ\abs{\vv x}\sin\theta. That motion is perpendicular to both n\vv n and x\vv x, and it covers an arc length ϵxsinθ\epsilon\abs{\vv x}\sin\theta. The vector with exactly that direction and that magnitude is ϵn×x\epsilon\,\vv n\times\vv x. So the infinitesimal version of a rotation is a cross product, which is why angular momentum will turn out to be one.

1.2 · Invariance of the action — with one crucial loosening

Now for the question the rest of the chapter rests on. What should it mean for the transformation to be a symmetry? The naive answer is "LL doesn't change". That answer is too strong. Adopt it and you will conclude, wrongly, that Newtonian mechanics is not invariant under a change of reference velocity.

The right condition comes from remembering what LL is for. Nothing physical depends on LL directly. The physics is in the stationary points of S=LdtS=\int L\,\dd t. And Chapter 1.2's Problem 4 established that two Lagrangians differing by a total time derivative,

L  =  L  +  dF(q,t)dt, L' \;=\; L \;+\; \dv{F(q,t)}{t}, (1.4.3)

have identical Euler–Lagrange equations. The reason is that the extra term integrates to F(q(t2),t2)F(q(t1),t1)F(q(t_{2}),t_{2})-F(q(t_{1}),t_{1}), which is the same number for every path in the competition and therefore invisible to the variation. So a transformation that changes LL by a total derivative has not changed the physics at all, and refusing to call it a symmetry would be a mistake of bookkeeping.

Definition — a symmetry of the action

The family (1.4.1) is a symmetry of the Lagrangian LL if there exists a function F(q,t)F(q,t) such that

δL    L(q+δq,  q˙+δq˙,  t)L(q,q˙,t)  =  ϵdFdt  +  O(ϵ2) \delta L \;\equiv\; L\big(q+\delta q,\;\dot q+\delta\dot q,\;t\big) - L\big(q,\dot q,t\big) \;=\; \epsilon\,\dv{F}{t} \;+\; O(\epsilon^{2})

for every path q(t)q(t), not merely for solutions of the equations of motion. The case F=0F=0 ("LL is strictly invariant") is allowed and common. The general case is called invariance up to a total derivative, or quasi-invariance.

Two features of this definition will do real work later. First, it is a condition on the action, not on the equations of motion. §7 exhibits a transformation that preserves the equations of motion but not the action, and produces no conserved quantity. Second, the hypothesis is off shell, meaning it is required for all paths, while the conclusion of §2 will hold only on shell, along actual trajectories. Getting those two quantifiers the wrong way round is the single most common way to misstate this theorem.

1.3 · Why the total derivative is not a technicality: Galilean boosts

Here is the case that forces the loosening. Take NN particles interacting through a potential that depends only on their separations. That covers every non-relativistic system of mutually interacting bodies you have ever seen:

L  =  a=1N12max˙ax˙a    V({xaxb}). L \;=\; \sum_{a=1}^{N}\half m_{a}\,\dot{\vv x}_{a}\cdot\dot{\vv x}_{a} \;-\; V\big(\{\vv x_{a}-\vv x_{b}\}\big). (1.4.4)

Now change to a reference frame moving at a small constant velocity ϵu-\epsilon\vv u. In the new frame every particle's position is shifted by ϵut\epsilon\vv u\,t:

xa    xa+ϵut,x˙a    x˙a+ϵu. \vv x_{a} \;\longmapsto\; \vv x_{a} + \epsilon\,\vv u\,t, \qquad\qquad \dot{\vv x}_{a} \;\longmapsto\; \dot{\vv x}_{a} + \epsilon\,\vv u. (1.4.5)

Galilean relativity says this had better be a symmetry. That principle is the statement that you cannot detect uniform motion by any mechanical experiment inside a sealed cabin. So let's check whether the boost qualifies. The potential is untouched, since every separation xaxb\vv x_{a}-\vv x_{b} is unchanged by shifting all positions equally. But the kinetic energy is not:

12max˙a+ϵu2=12max˙a2+ϵmax˙au+O(ϵ2),δL=ϵuamax˙a  +  O(ϵ2). \begin{aligned} \half m_{a}\abs{\dot{\vv x}_{a}+\epsilon\vv u}^{2} &= \half m_{a}\abs{\dot{\vv x}_{a}}^{2} + \epsilon\,m_{a}\,\dot{\vv x}_{a}\cdot\vv u + O(\epsilon^{2}),\\[3pt] \delta L &= \epsilon\,\vv u\cdot\sum_{a}m_{a}\dot{\vv x}_{a} \;+\; O(\epsilon^{2}). \end{aligned} (1.4.6)

So δL0\delta L\neq0: the Lagrangian is not invariant under a boost. If "symmetry" meant "LL unchanged" we would be forced to say that Newtonian mechanics violates Galilean relativity, which is absurd.

Look again at what δL\delta L actually is. The masses are constants, so

uamax˙a  =  ddt(uamaxa)  =  dFdt,F  =  MuRcm, \vv u\cdot\sum_{a}m_{a}\dot{\vv x}_{a} \;=\; \dv{}{t}\left(\vv u\cdot\sum_{a}m_{a}\vv x_{a}\right) \;=\; \dv{F}{t}, \qquad F \;=\; M\,\vv u\cdot\vv R_{\rm cm}, (1.4.7)

where M=amaM=\sum_{a}m_{a} and Rcm=1Mamaxa\vv R_{\rm cm}=\frac{1}{M}\sum_{a}m_{a}\vv x_{a} is the centre of mass. So the change in LL is a total time derivative, which is exactly the case the definition was widened to admit. A boost is a symmetry, in the only sense that matters, and it is not a symmetry in the naive sense.

The moral is worth stating flatly, because it will save you confusion in three later chapters: if you ever find that a symmetry you believe in appears to fail, check whether δL\delta L is a total derivative before concluding anything. Worked example 1 extracts the conserved quantity that this particular FF produces, and it is a conservation law that nobody thinks to name.

In plain terms 1.4.1

Getting the definition right is most of the work, and the theorem itself is four lines long because of it. A symmetry in the sense required is a family of deformations of the path controlled by one dial, with the do-nothing deformation at zero on the dial, so the transformation can be turned on gradually and differentiated with respect to how much of it is on. Only that first derivative ever matters, and it is called the generator.

The second half of the definition is a loosening, and refusing it would produce nonsense. The natural demand is that the scalar be unchanged, and under a change to a frame moving at a small steady speed it is not, which would force the conclusion that Newtonian mechanics can detect uniform motion inside a sealed cabin. What the scalar changes by is the rate of change of something, and anything of that shape contributes the same fixed amount to every competing path, so it is invisible to the comparison.

One quantifier in the definition does real work later and is the commonest place to get the theorem wrong. The requirement is that the scalar behave this way along every path you could draw, including the wild ones that solve nothing. A transformation leaving it alone only along the paths the system actually takes describes solutions you already have, and yields nothing.

2 · The theorem, proved

Everything is now in place. The proof is short, and each step uses exactly one hypothesis, so it is worth watching which.

Noether's theorem (transformations that do not move the clock)

Let δqi=ϵKi\delta q_{i}=\epsilon K_{i} be a symmetry of LL in the sense of §1.2, with δL=ϵdFdt\delta L = \epsilon\dv{F}{t}. Then along every solution of the Euler–Lagrange equations the quantity

Q  =  iLq˙iKi    F Q \;=\; \sum_{i}\pdv{L}{\dot q_{i}}\,K_{i} \;-\; F

is constant in time.

2.1 · Step 1 — the two derivatives commute

We will need to know how the velocities respond, and there is a small fact to establish first. The transformed velocity is q˙i(t,ϵ)=tqi(t,ϵ)\dot q_{i}(t,\epsilon)=\partial_{t}q_{i}(t,\epsilon), so

δq˙i  =  ϵϵqitϵ=0  =  ϵtqiϵϵ=0  =  ϵdKidt  =  ddt(δqi). \delta\dot q_{i} \;=\; \epsilon\left.\pdv{}{\epsilon}\pdv{q_{i}}{t}\right|_{\epsilon=0} \;=\; \epsilon\left.\pdv{}{t}\pdv{q_{i}}{\epsilon}\right|_{\epsilon=0} \;=\; \epsilon\,\dv{K_{i}}{t} \;=\; \dv{}{t}\big(\delta q_{i}\big). (1.4.8)

The middle equality is the equality of mixed partial derivatives (Chapter 0.6). In words: varying and then differentiating in time gives the same thing as differentiating and then varying. That is only true because δ\delta is taken at fixed tt. It is precisely why §2.5, where the transformation moves tt as well, needs extra care rather than more of the same.

2.2 · Step 2 — expand δL\delta L by the chain rule

We want to know how LL itself responds to the nudge. It is a function of 2n+12n+1 slots, so nudge the first 2n2n of them and apply the multivariable chain rule of Chapter 0.6:

δL  =  i[Lqiδqi  +  Lq˙iδq˙i]  =  ϵi[LqiKi  +  Lq˙idKidt], \delta L \;=\; \sum_{i}\left[\pdv{L}{q_{i}}\,\delta q_{i} \;+\; \pdv{L}{\dot q_{i}}\,\delta\dot q_{i}\right] \;=\; \epsilon\sum_{i}\left[\pdv{L}{q_{i}}\,K_{i} \;+\; \pdv{L}{\dot q_{i}}\,\dv{K_{i}}{t}\right], (1.4.9)

using Step 1 in the second equality. Note that tt itself has not moved, so there is no L/t\partial L/\partial t term. Nothing has been assumed yet except smoothness. Equation (1.4.9) is true for every path.

2.3 · Step 3 — use the equations of motion

Now, and only now, restrict attention to a path that actually solves the Euler–Lagrange equation,

ddt ⁣(Lq˙i)  =  Lqi. \dv{}{t}\!\left(\pdv{L}{\dot q_{i}}\right) \;=\; \pdv{L}{q_{i}}. (1.4.10)

Substitute the left side of (1.4.10) in place of L/qi\partial L/\partial q_{i} in (1.4.9). The bracket becomes something you have seen before:

δL=ϵi[ddt ⁣(Lq˙i)Ki  +  Lq˙idKidt]=ϵddt ⁣(iLq˙iKi). \begin{aligned} \delta L &= \epsilon\sum_{i}\left[\dv{}{t}\!\left(\pdv{L}{\dot q_{i}}\right)K_{i} \;+\; \pdv{L}{\dot q_{i}}\,\dv{K_{i}}{t}\right]\\[4pt] &= \epsilon\,\dv{}{t}\!\left(\sum_{i}\pdv{L}{\dot q_{i}}\,K_{i}\right). \end{aligned} (1.4.11)

2.4 · Step 4 — recognise a total derivative and compare

The second line of (1.4.11) is the product rule read backwards. That is the whole trick: the equation of motion is exactly what is needed to turn δL\delta L into a total time derivative of something.

But the symmetry hypothesis says δL\delta L is also a total time derivative, namely ϵdFdt\epsilon\dv{F}{t}. Two expressions for the same quantity:

ϵddt ⁣(iLq˙iKi)  =  ϵdFdtddt ⁣(iLq˙iKiF)=0. \epsilon\,\dv{}{t}\!\left(\sum_{i}\pdv{L}{\dot q_{i}}K_{i}\right) \;=\; \epsilon\,\dv{F}{t} \qquad\Longrightarrow\qquad \dv{}{t}\!\left(\sum_{i}\pdv{L}{\dot q_{i}}K_{i} - F\right) = 0. (1.4.12)

Both sides are the rate of change of something, so the difference of those two somethings has a rate of change of zero. That is what the right-hand half of (1.4.12) says. Give the constant quantity a name and the theorem is proved:

  Q  =  ipiKi    F,dQdt  =  0   \boxed{\;Q \;=\; \sum_{i}p_{i}K_{i} \;-\; F, \qquad \dv{Q}{t} \;=\; 0\;} (1.4.13)

with pi=L/q˙ip_{i}=\partial L/\partial\dot q_{i} the canonical momentum of Chapter 1.2. QQ is the Noether charge. \blacksquare

⚠ Why this isn't obvious

Let's count the hypotheses and see where each one was spent. The popular one-line version of this theorem, "symmetry implies conservation", is missing one of them.

The symmetry hypothesis was used off shell. Step 2 computed δL\delta L for an arbitrary path. The definition in §1.2 demanded δL=ϵdFdt\delta L=\epsilon\dv{F}{t} for every path, and that is what licenses comparing the two expressions in (1.4.12). A transformation that happens to leave LL alone only on solutions tells you nothing. It is a statement about the solutions you already have, not an independent input.

The equations of motion were used, so the conclusion is on shell. Step 3 is where (1.4.10) entered, and it is indispensable. Without it δL\delta L is not a total derivative of anything in particular. So QQ is not constant along an arbitrary path. It is constant along paths that solve the equations of motion. "Energy is conserved" is a statement about trajectories the system actually takes, and it is silent about the wiggly competitor paths that Chapter 1.2 varied over. In the path integral of Chapter 5.6, where the competitor paths are physically present and contribute, this distinction stops being pedantic. There, conservation laws hold as statements about expectation values (Ehrenfest, Ward identities), not path by path.

Sharpen it into a slogan you can check: the hypothesis is off shell, the conclusion is on shell. Swap them and the theorem is either vacuous or false.

2.5 · Transformations that also move time

One case is not yet covered, and it happens to be the most important one. Time translation is the statement that the laws are the same today as they were yesterday, and it moves tt itself. The derivation above assumed tt stood still. Most treatments wave at this step. It is worth doing properly, because the extra terms are exactly where the Hamiltonian comes from.

Allow the general case: a transformation that shifts the clock and the coordinates,

t    t=t+ϵτ(t),qi(t)    qi(t)=qi(t)+ϵKi. t \;\longmapsto\; t' = t + \epsilon\,\tau(t), \qquad\qquad q_{i}(t) \;\longmapsto\; q'_{i}(t') = q_{i}(t) + \epsilon\,K_{i}. (1.4.14)

Read the second equation carefully. It compares the new path at the new time with the old path at the old time, which is what a transformation of spacetime does. But Step 1 above needed a variation at fixed tt. The grind box reconciles the two and grinds out the general result. Here is the answer.

  Q  =  ipiKi    Hτ    F  H    ipiq˙iL. \boxed{\;Q \;=\; \sum_{i}p_{i}K_{i} \;-\; H\,\tau \;-\; F\;} \qquad H \;\equiv\; \sum_{i}p_{i}\dot q_{i} - L. (1.4.15)

The object HH that appears attached to the time-shift τ\tau is the combination Chapter 1.2 found by the Beltrami identity and Chapter 1.3 calls the Hamiltonian. It did not have to be put in by hand: it fell out of the bookkeeping of moving the clock. Setting τ=0\tau=0 recovers (1.4.13), as it must.

Grind box — the general variation, when the clock moves too

The two kinds of variation. Define the total variation, comparing new-at-new with old-at-old, and the form variation, comparing new-at-tt with old-at-tt:

Δqiqi(t)qi(t)=ϵKi,δqiqi(t)qi(t). \Delta q_{i} \equiv q'_{i}(t') - q_{i}(t) = \epsilon K_{i}, \qquad \delta q_{i} \equiv q'_{i}(t) - q_{i}(t).

Only δ\delta commutes with d/dt\dd/\dd t, so only δ\delta can be fed into §2.1–§2.4. Relate them by Taylor expansion (Chapter 0.3), keeping first order in ϵ\epsilon:

qi(t)=qi(t+ϵτ)=qi(t)+ϵτq˙i(t)+O(ϵ2)=qi(t)+ϵτq˙i(t)+O(ϵ2), q'_{i}(t') = q'_{i}(t+\epsilon\tau) = q'_{i}(t) + \epsilon\tau\,\dot q'_{i}(t) + O(\epsilon^{2}) = q'_{i}(t) + \epsilon\tau\,\dot q_{i}(t) + O(\epsilon^{2}),

where in the last step q˙=q˙+O(ϵ)\dot q'=\dot q+O(\epsilon) was used, legitimate because it is already multiplied by ϵ\epsilon. Rearranging,

δqi  =  ϵ(Kiq˙iτ)    ϵKˉi. \delta q_{i} \;=\; \epsilon\big(K_{i} - \dot q_{i}\,\tau\big) \;\equiv\; \epsilon\,\bar K_{i}.

The generator to use in the fixed-tt derivation is not KK but Kˉ=Kq˙τ\bar K = K-\dot q\,\tau. That single substitution is the entire content of this grind box. The rest is showing what happens to the measure dt\dd t.

The action, with limits and measure transformed. The transformed action is an integral over the transformed interval:

S  =  t1t2L(q(s),  dqds(s),  s)ds. S' \;=\; \int_{t'_{1}}^{t'_{2}} L\Big(q'(s),\;\tfrac{\dd q'}{\dd s}(s),\;s\Big)\,\dd s.

Substitute s=t+ϵτ(t)s=t+\epsilon\tau(t), so that ds=(1+ϵτ˙)dt\dd s=(1+\epsilon\dot\tau)\,\dd t and the limits become t1,t2t_{1},t_{2} again. Now expand the integrand about s=ts=t exactly as above, using G(t+ϵτ)=G(t)+ϵτG˙(t)+O(ϵ2)G(t+\epsilon\tau)=G(t)+\epsilon\tau\dot G(t)+O(\epsilon^{2}) for any function GG. That gives

S=t1t2[L(q(t),q˙(t),t)+ϵτdLdt](1+ϵτ˙)dt+O(ϵ2)=S+t1t2[δL+ϵτdLdt+ϵτ˙L]dt+O(ϵ2)=S+t1t2[δL+ϵddt(τL)]dt+O(ϵ2), \begin{aligned} S' &= \int_{t_{1}}^{t_{2}}\Big[L\big(q'(t),\dot q'(t),t\big) + \epsilon\tau\,\dv{L}{t}\Big]\big(1+\epsilon\dot\tau\big)\,\dd t + O(\epsilon^{2})\\[4pt] &= S + \int_{t_{1}}^{t_{2}}\Big[\delta L + \epsilon\,\tau\,\dv{L}{t} + \epsilon\,\dot\tau\,L\Big]\dd t + O(\epsilon^{2})\\[4pt] &= S + \int_{t_{1}}^{t_{2}}\Big[\delta L + \epsilon\,\dv{}{t}\big(\tau L\big)\Big]\dd t + O(\epsilon^{2}), \end{aligned}

where δL\delta L is the fixed-tt variation driven by Kˉ\bar K, and the last line collected two terms into one by the product rule. So the natural statement of the symmetry hypothesis, that the action changes at most by a boundary term, reads

δL+ϵddt(τL)  =  ϵdFdt. \delta L + \epsilon\,\dv{}{t}\big(\tau L\big) \;=\; \epsilon\,\dv{F}{t}.

Finish. Steps 2 and 3 of §2 applied verbatim with generator Kˉ\bar K give, on shell, δL=ϵddt(ipiKˉi)\delta L = \epsilon\,\dv{}{t}\big(\sum_{i}p_{i}\bar K_{i}\big). Substituting,

ddt(ipiKˉi+τLF)=0. \dv{}{t}\left(\sum_{i}p_{i}\bar K_{i} + \tau L - F\right) = 0.

Finally unpack Kˉ\bar K and recognise the Hamiltonian:

ipiKˉi+τL=ipiKiτ(ipiq˙iL)=ipiKiτH, \sum_{i}p_{i}\bar K_{i} + \tau L = \sum_{i}p_{i}K_{i} - \tau\left(\sum_{i}p_{i}\dot q_{i} - L\right) = \sum_{i}p_{i}K_{i} - \tau H,

which is (1.4.15). \blacksquare

Why the sign of HH looks backwards. In (1.4.15) the Hamiltonian enters with a minus sign, so pure time translation (τ=1\tau=1, K=0K=0, F=0F=0) gives the conserved charge Q=HQ=-H. Since a constant multiple of a conserved quantity is conserved, that is the same statement as "HH is conserved". But the minus is not an accident, and §7 shows it is exactly the sign needed for HH to generate forward time evolution.

In plain terms 1.4.2

No result in this book matters more than this one, and the whole of its proof is four lines. Expand the change in the scalar under the deformation by the chain rule, which is true along any path whatever, and then restrict attention to a path that solves the equation of motion. What the equation of motion does, and this is the entire trick, is turn that expression into the rate of change of one particular quantity.

The hypothesis has already said that the change is the rate of change of something else. Two expressions for one thing, so their difference has a rate of change of zero, and a quantity whose rate of change is zero is conserved. The two halves of that argument sit on opposite sides of a distinction worth keeping as a slogan: the hypothesis is required along every path, the conclusion holds only along paths that solve the equations, and swapping the two makes the theorem either vacuous or false.

One further case is usually waved at and repays being done properly, namely a transformation moving the clock as well as the coordinates. The bookkeeping then produces an extra term attached to the size of the time shift, and the quantity sitting in that term is the energy. It was not put in by hand. It arrived because the clock moved.

3 · The three classics, each from one formula

Everything in this section is (1.4.15) with different entries. Nothing new gets proved here. The point is that three conservation laws, each with its own unrelated textbook derivation, are three substitutions into one equation.

3.1 · Time translation ⇒ energy

Take τ=1\tau=1 (shift every clock reading by the same amount) and Ki=0K_{i}=0 (do not touch the coordinates). By the grind box the form variation is Kˉi=q˙i\bar K_{i}=-\dot q_{i}, so

δL  =  ϵi[Lqiq˙i+Lq˙iq¨i]. \delta L \;=\; -\epsilon\sum_{i}\left[\pdv{L}{q_{i}}\dot q_{i} + \pdv{L}{\dot q_{i}}\ddot q_{i}\right]. (1.4.16)

We need one more ingredient before we can test the symmetry condition, and it is the rate at which LL changes along the path all by itself. By the chain rule that is

dLdt  =  i[Lqiq˙i+Lq˙iq¨i]+Lt. \dv{L}{t} \;=\; \sum_{i}\left[\pdv{L}{q_{i}}\dot q_{i} + \pdv{L}{\dot q_{i}}\ddot q_{i}\right] + \pdv{L}{t}. (1.4.17)

Add them as the symmetry condition of the grind box requires, δL+ϵddt(τL)=δL+ϵdLdt\delta L + \epsilon\dv{}{t}(\tau L)=\delta L+\epsilon\dv{L}{t}, and the bracketed sums cancel exactly:

δL+ϵdLdt  =  ϵLt. \delta L + \epsilon\,\dv{L}{t} \;=\; \epsilon\,\pdv{L}{t}. (1.4.18)

So time translation is a symmetry, with F=0F=0, if and only if LL has no explicit tt-dependence. Feeding τ=1\tau=1, K=0K=0, F=0F=0 into (1.4.15):

Q=HE    ipiq˙iL  =  constant. Q = -H \qquad\Longrightarrow\qquad E \;\equiv\; \sum_{i}p_{i}\dot q_{i} - L \;=\; \text{constant}. (1.4.19)

That is the Beltrami identity of Chapter 1.2, Problem 2, which was obtained there by differentiating a lucky combination and watching terms cancel. It is also the Hamiltonian of Chapter 1.3, which was obtained there by a Legendre transform. Neither derivation said why that particular combination of pp's and q˙\dot q's should be the conserved one. This one does. It is the charge conjugate to moving the clock, and HH appears multiplying τ\tau in (1.4.15) for the same reason that pp appears multiplying KK.

For L=a12max˙a2V({x})L=\sum_{a}\half m_{a}\abs{\dot{\vv x}_{a}}^{2}-V(\{\vv x\}) the momentum is p=mx˙p=m\dot x and E=amax˙a2(TV)=T+VE=\sum_{a}m_{a}\abs{\dot{\vv x}_{a}}^{2}-\big(T-V\big)=T+V, the familiar total energy. But note carefully what the theorem conserves. It conserves pq˙L\sum p\dot q-L. Whether that equals T+VT+V is a separate question about the coordinates, and Chapter 1.2's Problem 3 (the bead on the motor-driven hoop) supplied a system where EE is conserved and T+VT+V is not.

3.2 · Space translation ⇒ momentum

Translate every particle by the same displacement, xaxa+ϵn\vv x_{a}\mapsto\vv x_{a}+\epsilon\vv n, so Ka=n\vv K_{a}=\vv n for all aa and τ=0\tau=0. Since the velocities are untouched, only VV can respond:

δL  =  ϵnaVxa. \delta L \;=\; -\epsilon\,\vv n\cdot\sum_{a}\pdv{V}{\vv x_{a}}. (1.4.20)

This vanishes for every n\vv n precisely when aV/xa=0\sum_{a}\partial V/\partial\vv x_{a}=0, which is the statement that VV is unchanged by moving everything together. In other words, VV depends only on differences xaxb\vv x_{a}-\vv x_{b}. Then F=0F=0, and (1.4.15) gives

Q  =  apan  =  Pn,P  =  amax˙a. Q \;=\; \sum_{a}\vv p_{a}\cdot\vv n \;=\; \vv P\cdot\vv n, \qquad \vv P \;=\; \sum_{a}m_{a}\dot{\vv x}_{a}. (1.4.21)

Since n\vv n is arbitrary, all three components of P\vv P are conserved. Compare this with the Newtonian derivation, which invokes the third law and cancels internal forces in pairs. The third law is nowhere in sight here. What replaced it is the observation that a two-body potential can only depend on the separation. That is the same physical input, now recognisable as a symmetry statement rather than an axiom about forces.

The one-coordinate version is worth naming, because it is how you will use the theorem in practice. Suppose some coordinate qjq_{j} is entirely absent from LL, making it a cyclic or ignorable coordinate. Then δqj=ϵ\delta q_{j}=\epsilon is a symmetry with F=0F=0, and Q=pjQ=p_{j}. That is the remark Chapter 1.2 made in passing after deriving p˙i=L/qi\dot p_{i}=\partial L/\partial q_{i}. It is now a special case of a theorem.

3.3 · Rotation ⇒ angular momentum

Rotate everything by ϵ\epsilon about the fixed axis n\vv n, so that from the table in §1.1

δxa=ϵn×xa,δx˙a=ϵn×x˙a,τ=0. \delta\vv x_{a} = \epsilon\,\vv n\times\vv x_{a}, \qquad \delta\dot{\vv x}_{a} = \epsilon\,\vv n\times\dot{\vv x}_{a}, \qquad \tau = 0. (1.4.22)

Before reading off a charge we have to check that this really is a symmetry, so take the two pieces of LL in turn. The kinetic term responds as

δ ⁣(12mx˙x˙)=ϵmx˙(n×x˙)=0, \delta\!\left(\half m\,\dot{\vv x}\cdot\dot{\vv x}\right) = \epsilon\, m\,\dot{\vv x}\cdot\big(\vv n\times\dot{\vv x}\big) = 0, (1.4.23)

because a cross product is perpendicular to both its factors. The potential responds through the separations, and

δxaxb2=2ϵ(xaxb)(n×(xaxb))=0 \delta\abs{\vv x_{a}-\vv x_{b}}^{2} = 2\,\epsilon\,\big(\vv x_{a}-\vv x_{b}\big)\cdot\Big(\vv n\times\big(\vv x_{a}-\vv x_{b}\big)\Big) = 0 (1.4.24)

for the same reason. So if VV depends only on the lengths of the separations, which is the case of central forces, then δL=0\delta L=0 exactly, with F=0F=0. Rotational symmetry is strict here, not merely up to a total derivative. Now read off the charge, using the cyclic property of the scalar triple product a(b×c)=b(c×a)\vv a\cdot(\vv b\times\vv c)=\vv b\cdot(\vv c\times\vv a):

Q=apa(n×xa)=na(xa×pa)=nL. Q = \sum_{a}\vv p_{a}\cdot\big(\vv n\times\vv x_{a}\big) = \vv n\cdot\sum_{a}\big(\vv x_{a}\times\vv p_{a}\big) = \vv n\cdot\vv L. (1.4.25)

Arbitrary n\vv n again, so the whole vector L=axa×pa\vv L=\sum_{a}\vv x_{a}\times\vv p_{a} is conserved. And now the definition of angular momentum stops being an arbitrary-looking cross product that someone decided to write down: x×p\vv x\times\vv p is what you get when you contract the momentum with the generator of rotations. The cross product is there because the infinitesimal rotation was.

Grind box — rotations in index notation, and where εijk\varepsilon_{ijk} comes from

The vector notation above hides the structure that generalises. Write the rotation generator in components using the Levi-Civita symbol εijk\varepsilon_{ijk} of Chapter 0.7 (antisymmetric in every pair, ε123=+1\varepsilon_{123}=+1):

δxi=ϵεijknjxk,δx˙i=ϵεijknjx˙k, \delta x_{i} = \epsilon\,\varepsilon_{ijk}\,n_{j}x_{k}, \qquad \delta \dot x_{i} = \epsilon\,\varepsilon_{ijk}\,n_{j}\dot x_{k},

with repeated indices summed. The kinetic term:

δ ⁣(12mx˙ix˙i)=ϵmx˙iεijknjx˙k=0, \delta\!\left(\half m\dot x_{i}\dot x_{i}\right) = \epsilon\,m\,\dot x_{i}\varepsilon_{ijk}n_{j}\dot x_{k} = 0,

because εijk\varepsilon_{ijk} is antisymmetric under iki\leftrightarrow k while x˙ix˙k\dot x_{i}\dot x_{k} is symmetric, and the full contraction of a symmetric with an antisymmetric pair vanishes identically. That two-line argument is worth internalising. It is the reason rotations preserve lengths, and in Chapter 2.3 the same manoeuvre with a different symbol shows that Lorentz transformations preserve the spacetime interval.

The charge:

Q=piεijknjxk=njεjkixkpi=njLj,Ljεjkixkpi, Q = p_{i}\,\varepsilon_{ijk}n_{j}x_{k} = n_{j}\,\varepsilon_{jki}\,x_{k}p_{i} = n_{j}L_{j}, \qquad L_{j} \equiv \varepsilon_{jki}x_{k}p_{i},

where the middle step used εijk=εjki\varepsilon_{ijk}=\varepsilon_{jki} (a cyclic permutation of three indices is even). Choosing n=z^\vv n=\hat{\vv z} picks out

L3=ε312x1p2+ε321x2p1=xpyypx, L_{3} = \varepsilon_{312}x_{1}p_{2} + \varepsilon_{321}x_{2}p_{1} = xp_{y} - yp_{x},

the familiar LzL_{z}.

The three rotations do not commute, and that matters. Rotating about xx then yy is not the same as yy then xx. The discrepancy is, at second order, a rotation about zz. Since the generators are what survive at first order, that failure to commute must show up as an algebraic relation among Lx,Ly,LzL_{x},L_{y},L_{z}, and §7 computes it to be {Li,Lj}=εijkLk\{L_{i},L_{j}\}=\varepsilon_{ijk}L_{k}. Translations, by contrast, commute with each other in any order, and correspondingly {Pi,Pj}=0\{P_{i},P_{j}\}=0. The conserved charges inherit the group structure of the symmetries that produced them, which is the subject of Chapter 6.1 and the reason quantum angular momentum is quantised in Chapter 4.11.

3.4 · The table, and what it means

Symmetry of the actionGeneratorConserved charge
Time translation, tt+ϵt\to t+\epsilonτ=1\tau=1Energy H=piq˙iLH=\sum p_{i}\dot q_{i}-L
Space translation along n\vv nK=n\vv K=\vv nMomentum nP\vv n\cdot\vv P
Rotation about n\vv nK=n×x\vv K=\vv n\times\vv xAngular momentum nL\vv n\cdot\vv L
Galilean boost by u\vv uK=ut\vv K=\vv u\,t, F=MuRcmF=M\vv u\cdot\vv R_{\rm cm}PtMRcm\vv P t-M\vv R_{\rm cm} (Worked Ex. 1)
Phase rotation ϕeiαϕ\phi\to\ee^{\ii\alpha}\phiK=iϕK=\ii\phiElectric charge (§6)

Read the first three rows as physics rather than as bookkeeping. The hypothesis in row one is that the laws are the same at every moment. The hypothesis in row two is that they are the same at every place. The hypothesis in row three is that they are the same in every direction. So:

What the classical conservation laws are actually about

Momentum is conserved because space is uniform. Energy is conserved because time is uniform. Angular momentum is conserved because space is isotropic.

None of these is a fact about matter. You cannot invent a new particle, or a new force, or a new material, that violates them. So long as it lives in a uniform, isotropic space and an unchanging time, its action has those symmetries and the charges are conserved whatever else it does. Conversely, the only way to break one is to break the corresponding symmetry, and §4 shows three situations where that genuinely happens. This is why conservation laws survived the transitions to relativity and quantum mechanics untouched while nearly everything else about mechanics was rewritten. They were never claims about mechanics in the first place.

In plain terms 1.4.3

Energy, momentum and angular momentum arrive here as one calculation performed three times with different entries. Shift every clock by the same amount and, provided the scalar has no explicit dependence on the time, the conserved quantity is the energy. Shift every particle by the same displacement and, provided the potential depends only on separations, it is the total momentum. Turn everything through one angle about an axis and, provided the potential depends only on the lengths of those separations, it is the angular momentum about it.

Read as physics rather than bookkeeping, the three hypotheses say the laws are the same at every moment, the same at every place, and the same in every direction. Three separate empirical laws have become one statement about the uniformity of the arena. None is a fact about matter, and no particle, force or material can be invented that violates them, because so long as it lives in a uniform isotropic space and an unchanging time its scalar carries those symmetries and the quantities are conserved whatever else it does.

Notice what has replaced the third law. Momentum conservation followed there from an assumption that internal forces cancel in pairs; it follows here from the observation that a potential between two bodies can depend only on how far apart they are — the same physical content, no longer an axiom about forces.

4 · When symmetry fails, conservation fails

A theorem earns trust by being falsifiable. If symmetry really is the reason for conservation, then removing the symmetry must remove the conservation law. It must also remove that one only, leaving the others intact. This section takes the theorem at its word.

4.1 · The rate at which conservation fails

Start with the quantitative version, which is stronger than "the theorem no longer applies". Differentiate E=ipiq˙iLE=\sum_{i}p_{i}\dot q_{i}-L along a solution, with LL now allowed an explicit tt-dependence:

dEdt=i[p˙iq˙i+piq¨i]i[Lqiq˙i+Lq˙iq¨i]Lt=iq˙i[p˙iLqi]=0 on shell+iq¨i[piLq˙i]=0 by definitionLt, \begin{aligned} \dv{E}{t} &= \sum_{i}\Big[\dot p_{i}\dot q_{i} + p_{i}\ddot q_{i}\Big] - \sum_{i}\left[\pdv{L}{q_{i}}\dot q_{i} + \pdv{L}{\dot q_{i}}\ddot q_{i}\right] - \pdv{L}{t}\\[4pt] &= \sum_{i}\dot q_{i}\underbrace{\left[\dot p_{i} - \pdv{L}{q_{i}}\right]}_{=\,0\ \text{on shell}} + \sum_{i}\ddot q_{i}\underbrace{\left[p_{i}-\pdv{L}{\dot q_{i}}\right]}_{=\,0\ \text{by definition}} - \pdv{L}{t}, \end{aligned} (1.4.26)

Look at the two braced brackets before reading on. The first one vanishes because the path solves the Euler–Lagrange equation, and the second vanishes because that is how pip_{i} was defined in the first place. Only the last term survives:

  dEdt  =  Lt.   \boxed{\;\dv{E}{t} \;=\; -\pdv{L}{t}.\;} (1.4.27)

Read that as an accounting identity. The left side is the rate at which the conservation law is violated. The right side is the rate at which the symmetry is violated. They are equal. Noether's theorem is the special case in which both sides are zero.

Example. A mass on a spring being driven by an external agent, L=12mx˙212kx2+xf(t)L=\half m\dot x^{2}-\half kx^{2}+x\,f(t), has L/t=xf˙\partial L/\partial t = x\dot f, so dEdt=xf˙(t)\dv{E}{t}=-x\dot f(t). The energy of the oscillator is not conserved, and it should not be, because something outside the system is doing work on it. The Lagrangian's explicit tt-dependence is the mathematical signature of "there is an outside". Enlarge the system to include whatever is driving ff, and the enlarged LL has no explicit tt, and total energy is conserved again. Every apparent violation of energy conservation in ordinary physics is of that kind. It is an incomplete system.

4.2 · A crystal: continuous symmetry broken to discrete

Now a case where the symmetry is not destroyed but reduced. An electron moving through a crystal lattice sees a potential that repeats:

V(x+a)  =  V(x)for every lattice vector a, V(\vv x + \vv a) \;=\; V(\vv x) \qquad\text{for every lattice vector }\vv a, (1.4.28)

but V(x+δ)V(x)V(\vv x+\vv\delta)\neq V(\vv x) for a general small displacement δ\vv\delta. Translation invariance survives only for a discrete set of displacements.

Noether's theorem has nothing to say here. Its hypothesis was a family q(t,ϵ)q(t,\epsilon) continuous in ϵ\epsilon, so that /ϵ\partial/\partial\epsilon exists at ϵ=0\epsilon=0. A discrete symmetry has no ϵ\epsilon to differentiate with respect to, so there is no generator KK to put into (1.4.15). And indeed momentum is not conserved. The lattice exerts forces on the electron, and p\vv p changes.

⚑ What survives is quoted rather than derived here, and it belongs to solid-state physics. Because the symmetry group is still nontrivial, the eigenstates can be labelled by a quantity k\hbar\vv k called the crystal momentum, and it is conserved modulo a reciprocal lattice vector G\vv G. Collisions can change k\vv k by any G\vv G at no cost.

The pattern is general and worth carrying. A continuous symmetry gives an additively conserved quantity taking values in R\R. Break it down to a discrete subgroup and you are left with a quantity conserved only up to the discrete "dual" of that subgroup. The residual freedom is exactly the amount of symmetry you threw away.

4.3 · An expanding universe: no conserved total energy

The striking case, and the one where the honest answer is more interesting than the tidy one.

⚑ Here is one fact quoted forward from Chapter 3.9. On the largest scales the universe is well described by a spacetime whose spatial distances all grow by a common factor a(t)a(t), the scale factor. Matter fields living in that spacetime have an action whose Lagrangian contains a(t)a(t) explicitly. So L/t0\partial L/\partial t\neq0, and not because of some external driving agent that could be absorbed into a larger system. It is because the geometry itself is time-dependent.

By (1.4.27) there is then no conserved energy, and this is not a technicality that a cleverer definition repairs. The most visible consequence is one you have heard described in a misleading way. Light from a distant galaxy arrives redshifted. Its wavelength has been stretched by the factor aa, so each photon's energy has fallen by 1/a1/a. The energy is not transferred to the expansion, and it is not stored anywhere. There is no conserved global energy in this spacetime for it to be a part of, so there is nothing that needs to balance. Asking where it went presupposes the very conservation law that the absence of time-translation symmetry has removed.

⚠ Say exactly this much and no more

Three claims, in decreasing order of confidence.

What is certainly true. ⚑ In a general curved spacetime there is no time-translation symmetry, hence no Noether charge for it, hence no meaningful "total energy of the universe". Chapter 3.5 makes this precise: the symmetries of a spacetime are its Killing vector fields, a conserved energy requires a timelike one, and the expanding solutions of Chapter 3.9 do not have one. Killing vectors turn out to be Noether's theorem written geometrically, which is why that chapter is where this loose end gets tied.

What survives. ⚑ The local statement does hold, always and exactly. It is μTμν=0\nabla_{\mu}T^{\mu\nu}=0, the covariant version of Chapter 0.7's continuity equation for the energy–momentum tensor, and it says energy and momentum are not created or destroyed at any point. But covariant conservation cannot be integrated over a large region to give a constant the way μJμ=0\partial_{\mu}J^{\mu}=0 can, and two separate things block it. The extra connection terms of Chapter 3.3 obstruct exactly that step. And ν\nu is a free index, so the object you would like to add up is a vector at each point, and there is no coordinate-free way to add vectors at different points on a curved manifold (Chapter 3.2). Local conservation is real. Global conservation is a boundary condition, and here the boundary conditions are not available.

What is true only sometimes. ⚑ Some spacetimes are asymptotically flat, meaning they settle back to flatness far from everything, as around an isolated star or a black hole in an otherwise empty universe. There the symmetry is restored far away, and a conserved total energy can be defined (the ADM mass). The expanding universe is not of this type. So "energy conservation fails in general relativity" is too strong, and "energy is always conserved, you just have to look harder" is false. The correct statement is the one Noether's theorem hands you directly: you get a conserved energy exactly when you have a time-translation symmetry, and not otherwise.

4.4 · Breaking one symmetry at a time

The figure below is the theorem's own falsification test. A particle orbits in a central potential, and two dials sit next to it, each of which breaks exactly one symmetry.

The system is a particle of mass mm in the plane, with

V(r,θ,t)  =  12mω2r2(1+λcos2θ)(1+μsinΩt). V(r,\theta,t) \;=\; \half m\omega^{2}r^{2}\,\big(1+\lambda\cos 2\theta\big)\,\big(1+\mu\sin\Omega t\big). (1.4.29)

At λ=μ=0\lambda=\mu=0 this is the isotropic harmonic oscillator. It is rotationally symmetric, so LzL_{z} is conserved, and it is time-independent, so EE is conserved. Each dial spoils one of those two properties and leaves the other alone.

  • The dial λ\lambda squashes the potential into an ellipse. Rotational symmetry gone, time-translation symmetry untouched.
  • The dial μ\mu makes the whole well breathe in time. Time-translation symmetry gone, rotational symmetry untouched.

Before looking, extract a prediction. The angular equation of motion is ddt(mr2θ˙)=V/θ\dv{}{t}\big(mr^{2}\dot\theta\big)=-\partial V/\partial\theta, and mr2θ˙=Lzmr^{2}\dot\theta = L_{z}, so

dLzdt  =  Vθ  =  λ  mω2r2sin2θ  (1+μsinΩt), \dv{L_{z}}{t} \;=\; -\pdv{V}{\theta} \;=\; \lambda\;m\omega^{2}r^{2}\sin 2\theta\;\big(1+\mu\sin\Omega t\big), (1.4.30)

That is the rate at which the rotational conservation law fails. Now do the same for the energy, where (1.4.27) has already done the work for us and gives

dEdt  =  Vt  =  μΩcosΩt  12mω2r2(1+λcos2θ). \dv{E}{t} \;=\; \pdv{V}{t} \;=\; \mu\,\Omega\cos\Omega t\;\cdot\half m\omega^{2}r^{2}\big(1+\lambda\cos2\theta\big). (1.4.31)

Let's look at what those two lines are saying. Each rate is exactly proportional to its own dial and completely independent of the other. The theorem does not merely predict that conservation fails. It predicts which one fails, and how fast.

λ = 0.000
μ = 0.000
max |ΔLz| = 2.3e-13
max |ΔE| = 2.9e-13
One dial, one conservation law. Units m=ω=1m=\omega=1; the particle starts at (x,y)=(1,0)(x,y)=(1,0) with velocity (0,0.5)(0,\,0.5) and is integrated by fourth-order Runge–Kutta for 3232 time units (about five orbits) at step 0.0040.004. Top: the trajectory, with one level set of VV dashed — normalised to enclose a fixed area, so that only its shape responds to the dial: a circle when λ=0\lambda=0, an ellipse otherwise. That is the broken symmetry made visible. Bottom: Lz(t)L_{z}(t) in blue and E(t)E(t) in orange, against dashed lines at their initial values. With both dials at zero, both traces are flat to 3×10133\times10^{-13} — that residue is double-precision round-off in the integrator, not physics. Push λ\lambda to 0.0050.005 — a half-percent squash — and LzL_{z} departs by 1.2%1.2\% while EE is still flat to 3×10133\times10^{-13}: ten orders of magnitude separate the broken law from the intact one. Push μ\mu to 0.020.02 instead and the roles swap exactly, EE wandering by 2%2\% with LzL_{z} flat to 2×10132\times10^{-13}. Turn both up and both fail. There is no setting in which the wrong one fails.

One detail in the figure repays attention. At λ=0\lambda=0 the orbit is a closed ellipse, and the particle retraces its own path exactly. Turn λ\lambda up and the ellipse begins to precess, so the path never closes. Closure of orbits is not a generic property. It is a symptom of extra conserved quantities, and Worked example 2 is about the most famous case of that.

In plain terms 1.4.4

A theorem earns trust by being falsifiable, and this one is easy to put on trial. Differentiate the energy along a solution while assuming nothing, and out comes the statement that its rate of change is exactly minus the rate at which the scalar depends on the time. The left side is how fast the conservation law fails, the right side how fast the symmetry fails, and they are the same number.

An oscillator driven from outside has a scalar depending on the time, and its energy is not conserved, because something outside is doing work on it. Explicit time dependence is the mathematical signature of there being an outside, and enlarging the system to include whatever drives it restores the symmetry and the law together. Every ordinary violation of energy conservation is an incomplete system.

One case is not of that kind, and there the honest answer beats the tidy one. On the largest scales every distance grows by a common factor depending on the time, so the geometry itself is time-dependent and there is nowhere larger to escape into. Light from a distant galaxy arrives stretched, each photon carrying less energy than it set out with, and that energy has not been transferred or stored anywhere, because no conserved global energy exists for it to belong to. Asking where it went assumes the law the missing symmetry removed.

5 · Discrete symmetries — an honest exception

Nature has symmetries that are not continuous, and the theorem simply does not reach them. The three that matter are parity PP (reflect all spatial coordinates, xx\vv x\to-\vv x), charge conjugation CC (exchange every particle for its antiparticle), and time reversal TT (ttt\to-t). None of them can be applied a little bit.

Parity is the cleanest case to see why. The matrix I-I has determinant 1-1 in three dimensions, and the identity has determinant +1+1. The determinant of a rotation-or-reflection is a continuous function taking only those two values, so no continuous path of transformations connects PP to the identity without leaving the group. There is therefore no one-parameter family to differentiate. No ϵ\epsilon, no generator KK, no (1.4.15), and no conserved current.

What a discrete symmetry gives instead is a multiplicative quantum number, typically ±1\pm1, whose product over the particles present is the same before and after a reaction. It gives that only in quantum mechanics, where states can be eigenvectors of the operation. Additive charges come from continuous symmetries and multiplicative ones from discrete symmetries. They are different objects with different arithmetic.

⚑ Here is a fact quoted now and developed in Chapter 6.6. Both PP and CPCP are violated in nature. Parity was found to fail in the weak interaction in 1957, and the combination CPCP was found to fail in 1964. This was a genuine surprise, and it is instructive about how much of physics is convention and how much is fact. Nobody had ever suspected that the laws could distinguish left from right, because there is no reason within mechanics or electromagnetism why they should.

Notice what does not happen when they fail. There is no Noether charge to lose, and correspondingly no continuity equation quietly breaks. The symmetry just is not there. The one combination that appears to survive is CPTCPT, and ⚑ that one is not an accident. It is a theorem, forced by Lorentz invariance and locality, which is why it has a different status from the others.

In plain terms 1.4.5

Some of nature's symmetries cannot be applied a little at a time, and the theorem does not reach them. Reflecting every spatial direction is the clean case: that operation and the do-nothing operation differ by the sign of a determinant, the sign takes only two values, and no continuous path of transformations gets from one to the other without leaving the set of allowed transformations. There is no dial to turn, so there is no derivative at zero, no generator, and no conserved current.

What such a symmetry offers instead exists only once states can carry labels, and it is a multiplicative label, usually plus or minus one, whose product over everything present is the same before a reaction as after. Additive quantities come from symmetries you can apply by degrees and multiplicative ones from symmetries you cannot, and they are different objects obeying different arithmetic.

Nature turns out to violate two of the three, which was a genuine surprise, since nothing in mechanics or electromagnetism gives any reason for the laws to distinguish left from right. Nothing quietly breaks when they fail, because there was no charge to lose and no continuity equation to spoil; the symmetry is merely absent. The one combination that does appear to survive is not luck but a theorem, forced by requirements the later parts of this book impose.

6 · Fields: symmetry gives a conserved current

Chapter 1.2 §8 showed that the Euler–Lagrange derivation goes through verbatim for fields, at the cost of one substitution. Replace L(q,q˙,t)L(q,\dot q,t) by a Lagrangian density L(ϕ,μϕ)\mathcal{L}(\phi,\partial_{\mu}\phi), replace dt\int\dd t by d4x\int\dd^{4}x, and replace ddt\dv{}{t} by μ\partial_{\mu}. The same substitution turns §2 into the field version of Noether's theorem, and the result is more useful than the mechanics version.

Take an internal symmetry, meaning one that transforms the field at each point without moving the point:

δϕ  =  ϵK(ϕ),δ(μϕ)  =  μ(δϕ),δL  =  ϵμFμ. \delta\phi \;=\; \epsilon\,K(\phi), \qquad \delta\big(\partial_{\mu}\phi\big) \;=\; \partial_{\mu}\big(\delta\phi\big), \qquad \delta\mathcal{L} \;=\; \epsilon\,\partial_{\mu}F^{\mu}. (1.4.32)

The middle equation is Step 1 again, the statement that variations commute with derivatives. The right-hand one is the symmetry hypothesis, with the total time derivative promoted to a four-divergence, since that is now what integrates to a boundary term. Steps 2 and 3 are unchanged. Expand by the chain rule, then use the field Euler–Lagrange equation μ(L/(μϕ))=L/ϕ\partial_{\mu}\big(\partial\mathcal{L}/\partial(\partial_{\mu}\phi)\big)=\partial\mathcal{L}/\partial\phi:

δL=Lϕδϕ+L(μϕ)μ(δϕ)=ϵ[μ ⁣(L(μϕ))K+L(μϕ)μK]  =  ϵμ ⁣[L(μϕ)K]. \begin{aligned} \delta\mathcal{L} &= \pdv{\mathcal{L}}{\phi}\,\delta\phi + \pdv{\mathcal{L}}{(\partial_{\mu}\phi)}\,\partial_{\mu}\big(\delta\phi\big)\\[4pt] &= \epsilon\left[\partial_{\mu}\!\left(\pdv{\mathcal{L}}{(\partial_{\mu}\phi)}\right)K + \pdv{\mathcal{L}}{(\partial_{\mu}\phi)}\,\partial_{\mu}K\right] \;=\; \epsilon\,\partial_{\mu}\!\left[\pdv{\mathcal{L}}{(\partial_{\mu}\phi)}\,K\right]. \end{aligned} (1.4.33)

Now compare that with the symmetry hypothesis. Once again we have two expressions for the same quantity, so the four-divergence of their difference vanishes, and that difference is the object we were after:

  jμ  =  L(μϕ)K    Fμ,μjμ  =  0.   \boxed{\;j^{\mu} \;=\; \pdv{\mathcal{L}}{(\partial_{\mu}\phi)}\,K \;-\; F^{\mu}, \qquad \partial_{\mu}j^{\mu} \;=\; 0.\;} (1.4.34)

6.1 · This is the continuity equation, and that is a promotion

Split jμ=(j0,j)j^{\mu}=(j^{0},\vv j) into its time and space parts. Then μjμ=tj0+j\partial_{\mu}j^{\mu}=\partial_{t}j^{0}+\nabla\cdot\vv j, and (1.4.34) is

j0t  +  j  =  0, \pdv{j^{0}}{t} \;+\; \nabla\cdot\vv j \;=\; 0, (1.4.35)

which is precisely the continuity equation of Chapter 0.7, with j0j^{0} the density of the conserved stuff and j\vv j its current. Chapter 0.7 promised that Noether's theorem would deliver one of these, and here it is. Integrating over a region VV and applying the divergence theorem, exactly as in that chapter,

Q  =  Vj0d3xdQdt  =  VjdA. Q \;=\; \int_{V} j^{0}\,\dd^{3}x \qquad\Longrightarrow\qquad \dv{Q}{t} \;=\; -\oint_{\partial V}\vv j\cdot\dd\vv A. (1.4.36)

Now notice how much stronger this is than the mechanics version. In §2 we proved that a number Q(t)Q(t) does not change. Here we have proved something local. The charge in any region changes only by flux through that region's boundary. Charge cannot vanish in London and reappear in Sydney with the books still balancing, because to get from one to the other it must cross every surface in between. Field conservation laws forbid teleportation, and mechanical ones do not. That is not a stylistic upgrade. It is the difference between a conservation law compatible with special relativity and one that is not, since a global bookkeeping rule requires a notion of "at the same instant everywhere" that Chapter 2.2 will destroy.

Grind box — the current is not unique, and the ambiguity is useful

Suppose jμj^{\mu} is conserved and let Σνμ\Sigma^{\nu\mu} be any field antisymmetric in its two indices, Σνμ=Σμν\Sigma^{\nu\mu}=-\Sigma^{\mu\nu}. Define

j~μ  =  jμ+νΣνμ. \tilde j^{\mu} \;=\; j^{\mu} + \partial_{\nu}\Sigma^{\nu\mu}.

Then μj~μ=μjμ+μνΣνμ=0\partial_{\mu}\tilde j^{\mu} = \partial_{\mu}j^{\mu} + \partial_{\mu}\partial_{\nu}\Sigma^{\nu\mu} = 0, because μν\partial_{\mu}\partial_{\nu} is symmetric in μν\mu\nu (equality of mixed partials, Chapter 0.6) while Σνμ\Sigma^{\nu\mu} is antisymmetric, and a symmetric object fully contracted with an antisymmetric one vanishes. That is the identical argument used in §3.3's grind box. The new current is conserved too.

Does it carry the same charge? The added piece contributes νΣν0d3x\int\partial_{\nu}\Sigma^{\nu0}\dd^{3}x. The ν=0\nu=0 term vanishes by antisymmetry (Σ00=0\Sigma^{00}=0), leaving iΣi0d3x\int\partial_{i}\Sigma^{i0}\dd^{3}x, a total spatial divergence, which by the divergence theorem is a surface integral at infinity. If the fields fall off, it is zero and Q~=Q\tilde Q = Q. So the current is ambiguous but the charge is not.

This is not idle. The Noether procedure applied to spacetime translations produces an energy–momentum tensor TμνT^{\mu\nu} that is in general not symmetric in its indices, which is a disaster for general relativity, where TμνT^{\mu\nu} sits on the right-hand side of a manifestly symmetric field equation (Chapter 3.6). The repair is exactly this freedom. One adds a suitable νΣνμρ\partial_{\nu}\Sigma^{\nu\mu\rho}, called an "improvement term", to symmetrise TμνT^{\mu\nu} without changing a single conserved charge. Chapter 5.2 does it.

6.2 · The example that sets up Part VI: a global phase

Take a complex scalar field ϕ\phi, with ϕ\phi and its conjugate ϕ\phi^{*} treated as independent. That is legitimate for the same reason xx and x˙\dot x were in Chapter 1.2, namely that they are independent slots of the function, and any complex ϕ\phi can be traded for the real pair (Reϕ,Imϕ)(\operatorname{Re}\phi,\operatorname{Im}\phi) and back:

L  =  μϕμϕ    m2ϕϕ. \mathcal{L} \;=\; \partial_{\mu}\phi^{*}\,\partial^{\mu}\phi \;-\; m^{2}\phi^{*}\phi. (1.4.37)

Now rotate the phase of the field by the same angle everywhere:

ϕ    eiαϕ,ϕ    eiαϕ,α a real constant. \phi \;\longmapsto\; \ee^{\ii\alpha}\phi, \qquad \phi^{*} \;\longmapsto\; \ee^{-\ii\alpha}\phi^{*}, \qquad \alpha \text{ a real constant}. (1.4.38)

Every term in (1.4.37) pairs a ϕ\phi with a ϕ\phi^{*}, so the phases cancel exactly. That makes L\mathcal{L} strictly invariant, with Fμ=0F^{\mu}=0. This is a global U(1)U(1) symmetry. It is U(1)U(1) because the transformations are the unit complex numbers eiα\ee^{\ii\alpha}, which form a circle group under multiplication, and global because α\alpha is the same at every point of spacetime.

Infinitesimally, ϵ=α\epsilon=\alpha and Kϕ=iϕK_{\phi}=\ii\phi, Kϕ=iϕK_{\phi^{*}}=-\ii\phi^{*}. The two derivatives needed are L/(μϕ)=μϕ\partial\mathcal{L}/\partial(\partial_{\mu}\phi)=\partial^{\mu}\phi^{*} and L/(μϕ)=μϕ\partial\mathcal{L}/\partial(\partial_{\mu}\phi^{*})=\partial^{\mu}\phi, so (1.4.34) summed over both fields gives

jμ  =  i(ϕμϕ    ϕμϕ). j^{\mu} \;=\; \ii\Big(\phi\,\partial^{\mu}\phi^{*} \;-\; \phi^{*}\partial^{\mu}\phi\Big). (1.4.39)

Two checks are worth making. First, the current is real, because complex-conjugating swaps the two terms and the i\ii out front, giving back the same expression. Second, it is conserved directly, without invoking the theorem at all. Take the divergence, note that the terms with two derivatives on different factors cancel in pairs, and use the field equation ϕ=m2ϕ\Box\phi=-m^{2}\phi that follows from (1.4.37):

μjμ  =  i(ϕϕϕϕ)  =  i(m2ϕϕ+m2ϕϕ)  =  0. \partial_{\mu}j^{\mu} \;=\; \ii\Big(\phi\,\Box\phi^{*} - \phi^{*}\Box\phi\Big) \;=\; \ii\Big(-m^{2}\phi\phi^{*} + m^{2}\phi^{*}\phi\Big) \;=\; 0. (1.4.40)

Multiplying a conserved current by any constant leaves it conserved, so the normalisation and overall sign are conventions. Fix them by demanding that a positive-energy mode carry positive charge, and the physical electromagnetic current of a field of charge qq is qjμ-q\,j^{\mu}. The conserved quantity Q=j0d3xQ=\int j^{0}\dd^{3}x is, up to that constant, electric charge.

The question you should be holding

The symmetry above required α\alpha to be the same everywhere. That is a peculiar demand. It says that if you rotate the phase of the field in this room you must simultaneously rotate it, by exactly the same angle, in a galaxy ten billion light-years away. And Chapter 2.2 is about to make "simultaneously" a suspect word. A relativist's instinct is that a symmetry requiring instantaneous agreement across the universe is not a symmetry that should be taken seriously in that form.

So: what happens if you insist that α\alpha may depend on position, αα(x)\alpha\to\alpha(x)?

The answer is that (1.4.37) is no longer invariant, because the derivative μ\partial_{\mu} hits α(x)\alpha(x) and leaves a term behind. The only way to repair it is to introduce a new field whose transformation law is designed to cancel that term. The new field turns out to be the electromagnetic potential AμA_{\mu}, and its dynamics turn out to be Maxwell's equations.

Demanding that a global symmetry be made local is called the gauge principle, and it does not merely accommodate electromagnetism. It generates it, and with a bigger group it generates the strong and weak forces as well. That is Chapter 6.3, and it is the single largest cheque this book writes. Nothing more will be said about it here. The point is only that you now know exactly which question leads there.

⚑ One warning to file, since it explains why Chapter 6.3 cannot simply run §6 again. Noether proved two theorems in 1918. The first is the one above, for symmetries with finitely many parameters. The second concerns symmetries depending on arbitrary functions, which is what gauge symmetries are, and its conclusion is different in kind. Instead of conserved charges it yields identities among the equations of motion, which is why gauge theories have constraints, why the Bianchi identity of Chapter 3.4 exists, and why μTμν=0\nabla_{\mu}T^{\mu\nu}=0 is automatic in general relativity rather than imposed.

Familiar ground — the indicator-dilution equation is a conservation law

You have taken a cardiac output off an indicator-dilution curve. Inject a known mass mm of dye or a known thermal deficit upstream, sample downstream, and

Q  =  m0c(t)dt, Q \;=\; \frac{m}{\displaystyle\int_{0}^{\infty} c(t)\,\dd t},

which is the Stewart–Hamilton equation. It is (1.4.36) with a different name on the conserved stuff, and its derivation is the one just given.

Take the conserved density to be the indicator concentration, j0=cj^{0}=c, and the current to be j=cv\vv j = c\,\vv v. Indestructibility of the indicator is exactly (1.4.35). Now integrate over the vascular volume lying between the injection site and the sampling site. Immediately after the injection that region holds all of mm. Long afterwards it holds none. And the only boundary through which anything leaves is the sampling plane. So (1.4.36), integrated over all time, reads 0 ⁣VjdA  dt=m\int_{0}^{\infty}\!\oint_{\partial V}\vv j\cdot\dd\vv A\;\dd t = m, and with flow QQ across a plane of uniform concentration the flux is Qc(t)Q\,c(t). Hence Qcdt=mQ\int c\,\dd t = m. That is the whole derivation.

The local statement is the one doing the work. A purely global rule, of the form "all the dye that goes in comes out eventually", is perfectly compatible with dye vanishing at the injection site and reappearing at the catheter tip. Under that hypothesis the area under the curve tells you nothing whatever about a flow. It is §6.1's stronger claim, that the amount inside any region changes only by crossing that region's boundary, which turns an area under a curve into millilitres per minute.

Both standard corrections to the method are repairs to a violated conservation law rather than to the arithmetic.

  • Recirculation means indicator crosses the sampling plane twice, so the integral double-counts and QQ comes out too low. That is why the downslope is fitted with an exponential and extrapolated.
  • Thermal loss to the catheter and the vessel wall is a genuine sink term on the right of (1.4.35), removing indicator without transporting it, and QQ comes out too high.

Where the correspondence stops, and it stops sharply. Noether's theorem does not supply this conservation law. Nobody derives the indestructibility of indocyanine green from an invariance of an action. It is assumed, and the assumption is precisely what those two corrections are defending. What §6 delivers is the arrow running the other way. Hand it a symmetry and a conserved current appears whether or not anyone was expecting one, together with an explicit formula (1.4.34) for what the conserved stuff is. That is a stronger kind of statement than a mass balance, and §6.2 spends it on a quantity nobody would have guessed.

In plain terms 1.4.6

Repeating the proof with a field in place of a coordinate costs nothing and returns much more. The same steps, with a four-dimensional divergence in place of a time derivative, deliver not a number that fails to change but a density and a current, tied by the statement that the density's rate of increase at a point equals what flows out. That is the equation promised when local conservation was first set out.

The gain deserves stating exactly. A global law says a grand total, added over everything, is the same at one moment as another. A local law says the amount inside any region changes only by crossing its boundary, so nothing may leave one place and arrive at another without traversing the ground between. The difference is not stylistic: a global ledger needs a notion of the same moment everywhere, and the next part of this book dismantles it.

The example to keep is a field of complex values whose scalar is unchanged when the phase is turned by one angle everywhere, and whose conserved quantity is, up to a constant, electric charge. That the angle must be the same everywhere should look suspicious for the reason already given, demanding as it does instantaneous agreement across the universe. Asking what happens when the angle may vary from place to place is the question that produces every force in nature.

7 · The other direction: charges generate their own symmetries

So far the arrow has run one way, from symmetry to conserved charge. This section runs it backwards, and the result is the real destination of Part I.

Chapter 1.3 supplies the tool. On phase space, coordinates qiq_{i} and momenta pip_{i} are independent, and any two functions of them have a Poisson bracket

{f,g}  =  i(fqigpifpigqi), \{f,g\} \;=\; \sum_{i}\left(\pdv{f}{q_{i}}\pdv{g}{p_{i}} - \pdv{f}{p_{i}}\pdv{g}{q_{i}}\right), (1.4.41)

with the two facts we need: time evolution is dfdt={f,H}+f/t\dv{f}{t}=\{f,H\}+\partial f/\partial t, and a function GG generates an infinitesimal transformation of phase space through

δf  =  ϵ{f,G}. \delta f \;=\; \epsilon\,\{f,G\}. (1.4.42)

7.1 · The charge is the generator — in one line

Take a Noether charge of the standard form (1.4.15) with τ=0\tau=0, namely Q=ipiKi(q,t)F(q,t)Q=\sum_{i}p_{i}K_{i}(q,t)-F(q,t), and ask what transformation it generates. Only the first term contains any pp, and it is linear in them. That is exactly where the restriction to Ki(q,t)K_{i}(q,t) is spent. §1.1 allowed the wider form Ki(q,q˙,t)K_{i}(q,\dot q,t), and a velocity-dependent generator makes QQ quadratic in the momenta, so the bracket below would return something else. For generators of the narrower kind,

δqj  =  ϵ{qj,Q}  =  ϵQpj  =  ϵKj. \delta q_{j} \;=\; \epsilon\,\{q_{j},Q\} \;=\; \epsilon\,\pdv{Q}{p_{j}} \;=\; \epsilon\,K_{j}. (1.4.43)

That is not an analogy or a coincidence. It is an identity, and it says that the Noether charge of a symmetry, used as a generator, reproduces exactly the symmetry it came from. The theorem and its converse are the same equation read in opposite directions. Here are three instances, worth doing explicitly because the brackets are how you will compute in Part IV.

Momentum generates translations. With G=pxG=p_{x}, equation (1.4.41) gives {x,px}=1\{x,p_{x}\}=1 and {y,px}={pi,px}=0\{y,p_{x}\}=\{p_{i},p_{x}\}=0, so

δx=ϵ,δy=δz=0,δp=0. \delta x = \epsilon, \qquad \delta y = \delta z = 0, \qquad \delta\vv p = 0. (1.4.44)

A rigid translation along xx, exactly as advertised.

Angular momentum generates rotations. With G=Lz=xpyypxG=L_{z}=xp_{y}-yp_{x}, the derivatives are Lz/px=y\partial L_{z}/\partial p_{x}=-y, Lz/py=x\partial L_{z}/\partial p_{y}=x, Lz/x=py\partial L_{z}/\partial x=p_{y}, Lz/y=px\partial L_{z}/\partial y=-p_{x}, so

δx=ϵ{x,Lz}=ϵy,δy=ϵ{y,Lz}=+ϵx,δpx=ϵ{px,Lz}=ϵpy,δpy=ϵ{py,Lz}=+ϵpx. \begin{aligned} \delta x &= \epsilon\{x,L_{z}\} = -\epsilon y, & \delta y &= \epsilon\{y,L_{z}\} = +\epsilon x,\\[3pt] \delta p_{x} &= \epsilon\{p_{x},L_{z}\} = -\epsilon p_{y}, & \delta p_{y} &= \epsilon\{p_{y},L_{z}\} = +\epsilon p_{x}. \end{aligned} (1.4.45)

The position rotates by ϵ\epsilon about the zz-axis, and so does the momentum. That second part is a bonus the Lagrangian derivation never showed us, and it is exactly right, since momentum is a vector and vectors rotate.

The Hamiltonian generates time evolution. With G=HG=H, equation (1.4.42) reads δf=ϵ{f,H}=ϵdfdt\delta f=\epsilon\{f,H\}=\epsilon\,\dv{f}{t} for any ff without explicit time dependence, which is the statement that the transformation generated by HH is advance everything by ϵ\epsilon. Time evolution is not a different kind of process from a symmetry transformation. It is the symmetry transformation generated by the energy.

This is also where the minus sign in (1.4.15) pays. The Noether charge of time translation is H-H, which generates δf=ϵdfdt\delta f=-\epsilon\dv{f}{t}. That is the shift of the graph of q(t)q(t) forward by ϵ\epsilon, which at fixed tt shows you the value the old path had at tϵt-\epsilon. The two signs describe the same motion from opposite sides.

7.2 · The converse, and why it is not decorative

Now the reverse implication. Suppose QQ is conserved and has no explicit time dependence. Then 0=dQdt={Q,H}0=\dv{Q}{t}=\{Q,H\}, and therefore the transformation QQ generates does this to the Hamiltonian:

δH  =  ϵ{H,Q}  =  ϵ{Q,H}  =  0. \delta H \;=\; \epsilon\,\{H,Q\} \;=\; -\epsilon\,\{Q,H\} \;=\; 0. (1.4.46)

The Hamiltonian is invariant under the transformation generated by QQ. So QQ is not merely produced by a symmetry. Viewed as a generator, QQ is a symmetry. There is nothing left to prove in either direction.

The destination of Part I

Symmetries and conservation laws are not cause and effect. They are the same object, viewed twice. Given a symmetry, (1.4.15) hands you a function on phase space. Given that function, (1.4.42) hands you back the symmetry. The Noether charge and the generator are one thing wearing two names, and which name you use depends only on whether you are asking "what stays the same?" or "what moves?"

This has structural consequences that outrun mechanics. The charges close under the bracket. §3.3's grind box promised, and §7.3 verifies, that {Li,Lj}=εijkLk\{L_{i},L_{j}\}=\varepsilon_{ijk}L_{k}, so the bracket of two conserved charges is again a conserved charge, with coefficients fixed by how the symmetries themselves compose. That structure is a Lie algebra, and it is the algebra of the symmetry group. Chapter 6.1 studies it systematically. Chapter 4.11 finds the identical relation with commutators in place of brackets, [L^i,L^j]=iεijkL^k[\hat L_{i},\hat L_{j}]=\ii\hbar\varepsilon_{ijk}\hat L_{k}, and with Chapter 4.12 derives the entire quantum theory of angular momentum from nothing but that algebra, including half-integer spin, which has no classical counterpart.

This is what licenses the sentence this book keeps repeating. Pick a symmetry, write the most general action invariant under it, quantise. Step one is now a precise instruction. Name a Lie group. Step two is Chapter 1.2. Step three replaces {,}\{\,,\} by 1i[,]\tfrac{1}{\ii\hbar}[\,,] and is Chapter 5.3. Every remaining part of this book is that recipe with a different group in step one.

7.3 · The algebra of the rotation charges

Grind box — computing {Li,Lj}\{L_{i},L_{j}\} by hand

Take Lx=ypzzpyL_{x}=yp_{z}-zp_{y} and Ly=zpxxpzL_{y}=zp_{x}-xp_{z} and grind (1.4.41) over the three coordinate pairs. Only terms in which the same variable is differentiated in both factors survive. The nonzero partials are

Lxy=pz,Lxz=py,Lxpy=z,Lxpz=y,Lyz=px,Lyx=pz,Lypz=x,Lypx=z. \begin{aligned} &\pdv{L_{x}}{y}=p_{z}, && \pdv{L_{x}}{z}=-p_{y}, && \pdv{L_{x}}{p_{y}}=-z, && \pdv{L_{x}}{p_{z}}=y,\\[3pt] &\pdv{L_{y}}{z}=p_{x}, && \pdv{L_{y}}{x}=-p_{z}, && \pdv{L_{y}}{p_{z}}=-x, && \pdv{L_{y}}{p_{x}}=z. \end{aligned}

Assemble the bracket term by term. From the xx-pair: xLxpxLypxLxxLy=0z0(pz)=0\partial_{x}L_{x}\,\partial_{p_{x}}L_{y}-\partial_{p_{x}}L_{x}\,\partial_{x}L_{y} = 0\cdot z - 0\cdot(-p_{z}) = 0. From the yy-pair: pz0(z)(0)=0p_{z}\cdot 0 - (-z)(0) = 0. From the zz-pair: (py)(x)(y)(px)=xpyypx(-p_{y})(-x) - (y)(p_{x}) = xp_{y}-yp_{x}. Hence

{Lx,Ly}  =  xpyypx  =  Lz, \{L_{x},L_{y}\} \;=\; xp_{y}-yp_{x} \;=\; L_{z},

and by relabelling xyzxx\to y\to z\to x, which permutes the three definitions cyclically,

{Li,Lj}  =  εijkLk. \{L_{i},L_{j}\} \;=\; \varepsilon_{ijk}L_{k}.

(Verified symbolically for all nine pairs.) Two remarks. First, the right-hand side is not zero, and that non-vanishing is the algebraic residue of the geometric fact that rotations about different axes do not commute. The bracket of generators measures the failure of the transformations to commute. Second, the same computation for translations gives {Pi,Pj}=0\{P_{i},P_{j}\}=0, because translations do commute. The bracket is a faithful record of the group.

Here is one consequence you can already read off, and Chapter 4.12 turns it into the whole theory of spin. Since {Lx,Ly}0\{L_{x},L_{y}\}\neq0, the three components of angular momentum are not simultaneously "diagonalisable" in the sense that will matter after quantisation. You will be able to know L\abs{\vv L} and one component, and no more.

⚠ Why this isn't obvious

A symmetry of the equations of motion is not the same thing as a symmetry of the action, and only the second kind gives a Noether charge. This is the second hypothesis that popular statements of the theorem drop.

Take the free particle, L=12mx˙2L=\half m\dot x^{2}, whose equation of motion is mx¨=0m\ddot x=0. Now scale the coordinate, x(1+ϵ)xx\mapsto(1+\epsilon)x, leaving tt alone. The equation of motion is completely unbothered. It is linear and homogeneous, so if x(t)x(t) is a solution then so is any multiple of it. Every solution maps to a solution. Yet the action is not invariant:

L    12m(1+ϵ)2x˙2  =  L+2ϵL+O(ϵ2), L \;\longmapsto\; \half m(1+\epsilon)^{2}\dot x^{2} \;=\; L + 2\epsilon L + O(\epsilon^{2}),

so δL=2ϵL\delta L = 2\epsilon L, which is not a total time derivative for a general path (if it were, LL would contribute nothing to the equations of motion, and it plainly does). So this is not a symmetry in the sense of §1.2, and the theorem does not apply.

Test the conclusion rather than trusting it. If the theorem did apply, the generator is K=xK=x and the charge would be Q=px=mxx˙Q=px=m x\dot x. Differentiate along a solution: dQdt=mx˙2+mxx¨=mx˙20\dv{Q}{t} = m\dot x^{2}+mx\ddot x = m\dot x^{2} \neq 0. Not conserved. The theorem was right to refuse.

What such a transformation gives instead is a map from solutions to other solutions. That is useful information, but of a different type. Problem 4 develops the most famous example. It is the scaling that maps Kepler orbits to other Kepler orbits, which is not a symmetry of the action, has no conserved charge, and is nevertheless the origin of Kepler's third law.

In plain terms 1.4.7

The last step reverses the arrow, and shows there was never an arrow to reverse. Take the conserved quantity a symmetry produced, feed it back as a generator of motion, and what it generates is the symmetry it came from. Not an analogy but an identity. A conserved quantity and a symmetry are one object under two names, and which you reach for depends on whether you ask what stays the same or what moves.

These quantities close among themselves: combining two returns a third, with coefficients fixed by how the transformations compose. Rotations about different axes do not commute and their combination is not zero; translations do commute and theirs is. That structure is what a symmetry group concretely amounts to, and it makes the book's recurring instruction precise: name a group, write the most general scalar invariant under it, quantise.

Part I is finished. Forces are gone, replaced by one number attached to each history, nature selecting the one where that number stops changing; a single point now fixes a whole future; and every continuous symmetry hands over a conserved quantity that is the symmetry itself. One thing was left broken, when the momentum of two moving charges refused to add up. Repairing it means surrendering something so ordinary nobody has troubled to state it, that everybody shares one clock; the next part opens on the collision between electromagnetism and that assumption.

8 · Worked examples

Worked example 1 — the Galilean boost, and a conservation law nobody names

§1.3 established that the boost (1.4.5) is a symmetry of the NN-body Lagrangian (1.4.4), but only in the loosened sense. It changes LL by the total derivative of F=MuRcmF=M\vv u\cdot\vv R_{\rm cm}. Now collect the charge. The ingredients are

Ka=ut,τ=0,F=uamaxa, \vv K_{a} = \vv u\,t, \qquad \tau = 0, \qquad F = \vv u\cdot\sum_{a}m_{a}\vv x_{a},

and (1.4.15) gives

Q=apa(ut)uamaxa=u(PtMRcm). Q = \sum_{a}\vv p_{a}\cdot\big(\vv u\,t\big) - \vv u\cdot\sum_{a}m_{a}\vv x_{a} = \vv u\cdot\Big(\vv P\,t - M\vv R_{\rm cm}\Big).

Since u\vv u was an arbitrary boost direction, all three components are conserved:

  N  =  Pt    MRcm  =  constant.   \boxed{\;\vv N \;=\; \vv P\,t \;-\; M\vv R_{\rm cm} \;=\; \text{constant.}\;}

Check it directly. dNdt=P+tdPdtMR˙cm\dv{\vv N}{t} = \vv P + t\,\dv{\vv P}{t} - M\dot{\vv R}_{\rm cm}. The middle term vanishes because P\vv P is conserved, by §3.2 and the same potential-depends-on-differences hypothesis. And by definition MR˙cm=amax˙a=PM\dot{\vv R}_{\rm cm}=\sum_{a}m_{a}\dot{\vv x}_{a}=\vv P. So dNdt=PP=0\dv{\vv N}{t}=\vv P-\vv P=0. ✓

What it says. Rearrange:

Rcm(t)  =  PMt    NM. \vv R_{\rm cm}(t) \;=\; \frac{\vv P}{M}\,t \;-\; \frac{\vv N}{M}.

The centre of mass moves in a straight line at constant velocity, forever, whatever the particles inside are doing to each other. That is a statement everybody knows and almost nobody files under "conservation law", because the conserved quantity has no name and no symbol in the standard curriculum. It has one here. It is the Noether charge of the boost, and it sits in the table of §3.4 alongside energy and momentum with exactly the same status.

Two things worth noticing. First, N\vv N depends explicitly on tt. Nothing in the theorem forbade that, since KK was allowed to depend on tt, and here K=ut\vv K=\vv u t does. A conserved quantity need not be a function of the instantaneous state alone. It is a combination of the state and the clock that happens to be constant.

Second, this is the charge that has no analogue after Chapter 2. The Galilean group's boosts get replaced by Lorentz boosts, whose Noether charges are the components of a relativistic object (Chapter 2.5) with a rather different physical reading. The centre-of-mass theorem is a specifically Newtonian statement, and it is a good illustration of the general rule that changing the symmetry group changes the list of conserved quantities, which is the whole reason the rest of this book is organised by symmetry group.

Worked example 2 — the Laplace–Runge–Lenz vector, and a hidden symmetry

Take the Kepler problem: one particle of mass mm in the attractive potential

V(r)=kr,so thatp˙=kr2r^,r^=rr. V(r) = -\frac{k}{r}, \qquad\text{so that}\qquad \dot{\vv p} = -\frac{k}{r^{2}}\,\hat{\vv r}, \qquad \hat{\vv r} = \frac{\vv r}{r}.

Rotational symmetry gives L=r×p\vv L=\vv r\times\vv p conserved (§3.3), and time-translation symmetry gives EE conserved (§3.1). That is four constants, and a generic central force gives no more. But the 1/r1/r potential has a fifth, discovered and rediscovered so often it carries three names:

  A  =  p×L    mkr^   \boxed{\;\vv A \;=\; \vv p\times\vv L \;-\; mk\,\hat{\vv r}\;}

It is conserved, and the proof is three lines of vector algebra done in the grind box below. The key move is that the awkward-looking r^×(r×r˙)\hat{\vv r}\times(\vv r\times\dot{\vv r}) collapses, by the a×(b×c)\vv a\times(\vv b\times\vv c) expansion, into exactly r2ddt(r/r)r^{2}\dv{}{t}(\vv r/r). So the rate of change of p×L\vv p\times\vv L is mkdr^dtmk\,\dv{\hat{\vv r}}{t}, and subtracting mkr^mk\hat{\vv r} cancels it.

Grind box — proving dAdt=0\dv{\vv A}{t}=0, and two identities that come free

Conservation. L\vv L is constant (§3.3), so only the first factor moves:

ddt(p×L)=p˙×L=kr2r^×(mr×r˙). \dv{}{t}\big(\vv p\times\vv L\big) = \dot{\vv p}\times\vv L = -\frac{k}{r^{2}}\,\hat{\vv r}\times\big(m\,\vv r\times\dot{\vv r}\big).

Apply a×(b×c)=b(ac)c(ab)\vv a\times(\vv b\times\vv c)=\vv b(\vv a\cdot\vv c)-\vv c(\vv a\cdot\vv b) with a=r^\vv a=\hat{\vv r}, b=r\vv b=\vv r, c=r˙\vv c=\dot{\vv r}, and use r^r˙=r˙\hat{\vv r}\cdot\dot{\vv r}=\dot r together with r^r=r\hat{\vv r}\cdot\vv r=r:

p˙×L=mkr2[rr˙r˙r]=mk[r˙rrr˙r2]=mkddt ⁣(rr)=mkdr^dt, \begin{aligned} \dot{\vv p}\times\vv L &= -\frac{mk}{r^{2}}\Big[\vv r\,\dot r - \dot{\vv r}\,r\Big] = mk\left[\frac{\dot{\vv r}}{r} - \frac{\vv r\,\dot r}{r^{2}}\right]\\[4pt] &= mk\,\dv{}{t}\!\left(\frac{\vv r}{r}\right) = mk\,\dv{\hat{\vv r}}{t}, \end{aligned}

the last line by the quotient rule. Hence ddt(p×Lmkr^)=0\dv{}{t}\big(\vv p\times\vv L - mk\hat{\vv r}\big)=0. \blacksquare (Verified symbolically in three dimensions, componentwise, on the equations of motion.)

Identity one: A\vv A lies in the orbital plane. Dot with L\vv L. The first term gives (p×L)L=0(\vv p\times\vv L)\cdot\vv L=0 because a cross product is perpendicular to its factors. The second gives mkr^L=0-mk\,\hat{\vv r}\cdot\vv L = 0 because L=r×p\vv L=\vv r\times\vv p is perpendicular to r\vv r. So AL=0\vv A\cdot\vv L=0, meaning the vector lies in the plane of the orbit, which is why it can point at the perihelion.

Identity two: its length is fixed by EE and LL. Square it, using pL\vv p\perp\vv L (so p×L=pL\abs{\vv p\times\vv L}=pL) and (p×L)r^=L(r×p)/r=L2/r(\vv p\times\vv L)\cdot\hat{\vv r}=\vv L\cdot(\vv r\times\vv p)/r=L^{2}/r:

A2=p2L2+m2k22mkL2r=L2 ⁣(p22mkr)+m2k2=2mEL2+m2k2, \begin{aligned} A^{2} &= p^{2}L^{2} + m^{2}k^{2} - 2mk\,\frac{L^{2}}{r} = L^{2}\!\left(p^{2}-\frac{2mk}{r}\right) + m^{2}k^{2}\\[4pt] &= 2mE\,L^{2} + m^{2}k^{2}, \end{aligned}

since E=p2/2mk/rE=p^{2}/2m-k/r gives p22mk/r=2mEp^{2}-2mk/r=2mE. So A\vv A is not a sixth independent constant. Its direction in the plane is one new number, and its length is determined. That leaves five independent constants rather than six, which is exactly the count needed to fix an orbit in space completely apart from where along it the particle currently sits.

What it is. Dot A\vv A into r\vv r, using the cyclic property of the scalar triple product:

Ar=(p×L)rmkr=L(r×p)mkr=L2mkr. \vv A\cdot\vv r = (\vv p\times\vv L)\cdot\vv r - mkr = \vv L\cdot(\vv r\times\vv p) - mkr = L^{2} - mkr.

Writing Ar=Arcosθ\vv A\cdot\vv r = Ar\cos\theta with θ\theta measured from the fixed direction of A\vv A and solving for rr:

r(θ)  =  L2/mk1+ecosθ,e    Amk. r(\theta) \;=\; \frac{L^{2}/mk}{1 + e\cos\theta}, \qquad e \;\equiv\; \frac{A}{mk}.

That is the polar equation of a conic of eccentricity ee, an ellipse for e<1e\lt1, with rr smallest at θ=0\theta=0. So A\vv A points from the focus to the perihelion, and its length is the eccentricity times mkmk. Its conservation is the statement that the orbit's axis never moves, which is to say that Kepler ellipses close and do not precess. Adding any deviation from exact 1/r1/r destroys A\vv A and the orbit precesses. That is exactly what the λ\lambda dial did to the closed orbit in §4.4's figure, and exactly what the Sun's other planets, and then general relativity, do to Mercury (Chapter 3.7).

Where is the symmetry? By §7.2 there must be one. A\vv A has no explicit time dependence and is conserved, so {A,H}=0\{\vv A,H\}=0, so the transformation A\vv A generates leaves HH invariant. But look at the list of §3.4. Nothing there produces A\vv A. The reason is that the transformation generated by A\vv A is not a motion of space at all. It mixes positions with momenta, deforming an orbit into a different orbit of the same energy, turning a circle into a thin ellipse, say. No rearrangement of the coordinates does that, which is why it is called a hidden symmetry.

The Lagrangian derivation of §2 assumed the transformation acted on qq alone and could never have found it. The phase-space picture of §7 finds it immediately, because on phase space qq and pp are on the same footing and there are transformations that mix them.

Now the structure. The brackets of the six conserved quantities close on themselves. This part is a computation of exactly the kind §7.3 did by hand, only longer:

{Li,Lj}=εijkLk,{Li,Aj}=εijkAk,{Ai,Aj}=2mHεijkLk, \{L_{i},L_{j}\}=\varepsilon_{ijk}L_{k}, \qquad \{L_{i},A_{j}\}=\varepsilon_{ijk}A_{k}, \qquad \{A_{i},A_{j}\}=-2mH\,\varepsilon_{ijk}L_{k},

(all three verified symbolically here.) ⚑ The identification that follows is quoted, since naming an abstract Lie algebra is Chapter 6.1's business. For bound orbits E<0E\lt0 the rescaled vector D=A/2mE\vv D=\vv A/\sqrt{-2mE} makes the last relation match the first. The resulting six-generator Lie algebra is that of SO(4)SO(4), the rotation group of four-dimensional space. The Kepler problem in three dimensions has the symmetry of rotations in four. That is the hidden symmetry, and it is why there are five constants instead of four and why the orbits close.

⚑ And the payoff arrives in Chapter 4.14. The hydrogen atom is the quantum 1/r1/r problem, so it inherits this same SO(4)SO(4). Its energy levels depend only on the principal quantum number nn and not on the orbital angular momentum \ell, so the 2s2s and 2p2p states are degenerate. That looks like a numerical accident and is presented as one in most first courses. It is not an accident. It is the extra symmetry, showing up as a degeneracy, exactly as the λ=0\lambda=0 orbit in §4.4's figure showed its extra symmetry by closing.

Which is the moral, and it is worth carrying out of this chapter as a working heuristic. An unexplained coincidence in physics is usually an unrecognised symmetry. A degeneracy that "just happens", a quantity that "happens" to stay constant, a term that "happens" to cancel: each is a symmetry that has not yet been named. Looking for it is one of the most reliably productive things a physicist can do, and Chapter 6.6 is what happens when the search succeeds and Chapter 5.11 is what happens when it fails.

9 · Your turn

Problem 1 — a scaling that is a Noether symmetry

Consider a particle on the half-line x>0x\gt0 with

L  =  12mx˙2    kx2,k>0. L \;=\; \half m\dot x^{2} \;-\; \frac{k}{x^{2}}, \qquad k\gt0.

(a) Show that the combined scaling xxx\mapsto\ell x, t2tt\mapsto\ell^{2}t leaves the action exactly invariant. (b) Identify KK and τ\tau and use (1.4.15) to find the conserved charge. (c) Verify directly that it is conserved. (d) Why does this scaling work when §7's free-particle scaling did not?

Solution

(a) Under x(t)=x(t)x'(t')=\ell x(t) with t=2tt'=\ell^{2}t, the velocity scales as dx/dt=/2x˙=1x˙\dd x'/\dd t' = \ell/\ell^{2}\cdot\dot x=\ell^{-1}\dot x, so the kinetic term scales by 2\ell^{-2}. The potential k/x2k/x'^{2} scales by 2\ell^{-2} as well. Both terms of LL therefore carry the same factor 2\ell^{-2}, while the measure dt=2dt\dd t'=\ell^{2}\dd t carries +2\ell^{+2}:

S=Ldt=2L2dt=S. S' = \int L'\,\dd t' = \int \ell^{-2}L\cdot\ell^{2}\,\dd t = S.

Exact invariance, F=0F=0. (Confirmed symbolically.) The reason it works is that the two terms of LL have the same scaling weight, and 1/x21/x^{2} is the unique power for which they do. This is the simplest system in physics with a scale symmetry.

(b) Set =1+ϵ\ell=1+\epsilon. Then δx=ϵx\delta x=\epsilon x and δt=2ϵt\delta t=2\epsilon t, so K=xK=x and τ=2t\tau=2t, and (1.4.15) gives

Q  =  px2tH  =  mxx˙    2t(12mx˙2+kx2). Q \;=\; p\,x - 2t\,H \;=\; m x\dot x \;-\; 2t\left(\half m\dot x^{2} + \frac{k}{x^{2}}\right).

(c) The equation of motion is mx¨=ddx ⁣(kx2)=2kx3m\ddot x = -\dv{}{x}\!\left(\dfrac{k}{x^{2}}\right) = \dfrac{2k}{x^{3}}. Differentiate QQ, remembering that H=EH=E is itself conserved because LL has no explicit tt:

dQdt=mx˙2+mxx¨2E=mx˙2+x2kx32(12mx˙2+kx2)=0.   \dv{Q}{t} = m\dot x^{2} + mx\ddot x - 2E = m\dot x^{2} + x\cdot\frac{2k}{x^{3}} - 2\left(\half m\dot x^{2} + \frac{k}{x^{2}}\right) = 0. \;\checkmark

Note the structure. The charge is explicitly time-dependent, like the boost charge of Worked Example 1, and for the same reason, which is that τ\tau and KK contained tt.

(d) Because here the time was scaled too, and by exactly the amount needed to make the measure compensate the integrand. In §7's free particle only xx was scaled, giving S2SS\mapsto\ell^{2}S. The equations of motion survive that, since they are unchanged by an overall constant factor on SS, but the action does not, and the theorem's hypothesis fails. The lesson is that whether a scaling is a Noether symmetry depends on how it acts on time, and finding the right relative weight is exactly what one does when constructing a scale-invariant theory.

⚑ Here is a forward pointer. This Lagrangian is the one-dimensional prototype of a conformally invariant theory. The full symmetry is larger than the single scaling, since there are three charges, closing on the algebra of SL(2,R)SL(2,\R). The two-dimensional version of the same structure, with infinitely many charges instead of three, is what makes string theory calculable in Chapter 7.3.

Problem 2 — a particle on a surface of revolution

A particle of mass mm slides without friction on the surface obtained by rotating the curve ρ=R(z)\rho=R(z) about the zz-axis, with no potential (V=0V=0). Using cylindrical coordinates (ρ,φ,z)(\rho,\varphi,z):

(a) Write LL on the surface. (b) Identify the symmetry and its Noether charge. (c) Use it to reduce the motion to a single first-order equation for z(t)z(t), i.e. to a quadrature. (d) Show that the charge is equivalent to Clairaut's relation: RsinψR\sin\psi is constant, where ψ\psi is the angle between the trajectory and the local meridian.

Solution

(a) In cylindrical coordinates the squared speed is ρ˙2+ρ2φ˙2+z˙2\dot\rho^{2}+\rho^{2}\dot\varphi^{2}+\dot z^{2}. On the surface ρ=R(z)\rho=R(z), so ρ˙=R(z)z˙\dot\rho=R'(z)\dot z and

L  =  12m[(1+R(z)2)z˙2  +  R(z)2φ˙2]. L \;=\; \half m\Big[\big(1+R'(z)^{2}\big)\dot z^{2} \;+\; R(z)^{2}\dot\varphi^{2}\Big].

Two coordinates, because imposing the constraint by choosing coordinates is exactly what Chapter 1.2 §6 recommended.

(b) φ\varphi does not appear in LL, because the surface is unchanged by rotating about its axis. That is the symmetry δφ=ϵ\delta\varphi=\epsilon, Kφ=1K_{\varphi}=1, τ=0\tau=0, F=0F=0. So

Q  =  Lφ˙  =  mR(z)2φ˙      =  constant, Q \;=\; \pdv{L}{\dot\varphi} \;=\; m\,R(z)^{2}\,\dot\varphi \;\equiv\; \ell \;=\; \text{constant,}

the angular momentum about the axis. (This is §3.2's cyclic-coordinate case, and it is what Chapter 1.2 saw in polar coordinates without a name for it.)

(c) LL has no explicit tt, so E=12m[(1+R2)z˙2+R2φ˙2]E=\half m[(1+R'^{2})\dot z^{2}+R^{2}\dot\varphi^{2}] is conserved as well. Eliminate φ˙=/(mR2)\dot\varphi=\ell/(mR^{2}):

E  =  12m(1+R2)z˙2  +  22mR(z)2Veff(z). E \;=\; \half m\big(1+R'^{2}\big)\dot z^{2} \;+\; \underbrace{\frac{\ell^{2}}{2mR(z)^{2}}}_{V_{\rm eff}(z)}.

Two conserved quantities have reduced two coupled second-order equations to one first-order equation, which separates:

tt0  =  m(1+R(z)2)2(E2/2mR(z)2)  dz, t - t_{0} \;=\; \int \sqrt{\frac{m\big(1+R'(z)^{2}\big)}{2\big(E - \ell^{2}/2mR(z)^{2}\big)}}\;\dd z,

and then φ(t)\varphi(t) follows by integrating φ˙=/mR2\dot\varphi=\ell/mR^{2}. The problem is solved by quadrature. Every conserved quantity buys one integration, and two of them turn a two-degree-of-freedom problem into arithmetic.

Note also that Veff1/R2V_{\rm eff}\propto1/R^{2} acts as a barrier wherever the surface narrows. A particle with 0\ell\neq0 cannot reach a neck of radius smaller than Rmin=/2mER_{\min}=\ell/\sqrt{2mE}. That is the centrifugal barrier, arriving here as a consequence of a symmetry rather than as a fictitious force.

(d) With V=0V=0 the speed vv is constant (energy conservation). Decompose the velocity into a component along the local circle of latitude, vφ=Rφ˙v_{\varphi}=R\dot\varphi, and one along the meridian. If ψ\psi is the angle between the trajectory and the meridian, then sinψ=vφ/v\sin\psi = v_{\varphi}/v, so

=mR2φ˙=mRvφ=mvRsinψRsinψ=constant. \ell = mR^{2}\dot\varphi = mR\,v_{\varphi} = m\,v\,R\sin\psi \qquad\Longrightarrow\qquad R\sin\psi = \text{constant}.

This is Clairaut's relation, and it is what Chapter 1.2's grind box on spherical geodesics found by integrating the Euler–Lagrange equation and then remarked was "Noether's theorem arriving three chapters early disguised as cartography". Here is the general version. On any surface of revolution, the geodesics obey Rsinψ=R\sin\psi= constant because the surface has an axial symmetry. It is also the reason a great-circle route from London to Tokyo goes over the Arctic, since the path must turn toward the axis as RR decreases to keep the product fixed.

Problem 3 — the current of a complex field, and what the charge is

For the field theory of (1.4.37), derive the Noether current for the global phase symmetry from scratch (do not quote §6.2), and evaluate the charge density j0j^{0} on the plane-wave solution ϕ=Nei(Etpx)\phi=N\ee^{-\ii(Et-\vv p\cdot\vv x)} with NN real. Then evaluate it on the complex-conjugate solution ϕ\phi^{*} and say what the sign means.

Solution

Setup. Treat ϕ\phi and ϕ\phi^{*} as two independent fields. Writing L=gμνμϕνϕm2ϕϕ\mathcal{L}=g^{\mu\nu}\partial_{\mu}\phi^{*}\partial_{\nu}\phi - m^{2}\phi^{*}\phi,

L(μϕ)=μϕ,L(μϕ)=μϕ. \pdv{\mathcal{L}}{(\partial_{\mu}\phi)} = \partial^{\mu}\phi^{*}, \qquad \pdv{\mathcal{L}}{(\partial_{\mu}\phi^{*})} = \partial^{\mu}\phi.

The transformation. ϕeiαϕ\phi\to\ee^{\ii\alpha}\phi gives, to first order in ϵ=α\epsilon=\alpha,

δϕ=iαϕ    Kϕ=iϕ,δϕ=iαϕ    Kϕ=iϕ. \delta\phi = \ii\alpha\,\phi \;\Rightarrow\; K_{\phi}=\ii\phi, \qquad \delta\phi^{*} = -\ii\alpha\,\phi^{*} \;\Rightarrow\; K_{\phi^{*}}=-\ii\phi^{*}.

L\mathcal{L} pairs each ϕ\phi with a ϕ\phi^{*}, so the phases cancel and δL=0\delta\mathcal{L}=0 exactly: Fμ=0F^{\mu}=0.

The current. Sum (1.4.34) over the two fields:

jμ=μϕ(iϕ)+μϕ(iϕ)=i(ϕμϕϕμϕ). j^{\mu} = \partial^{\mu}\phi^{*}\,(\ii\phi) + \partial^{\mu}\phi\,(-\ii\phi^{*}) = \ii\big(\phi\,\partial^{\mu}\phi^{*} - \phi^{*}\partial^{\mu}\phi\big).

On a plane wave. With ϕ=Nei(Etpx)\phi = N\ee^{-\ii(Et-\vv p\cdot\vv x)} we have 0ϕ=tϕ=iEϕ\partial^{0}\phi=\partial_{t}\phi=-\ii E\phi and 0ϕ=+iEϕ\partial^{0}\phi^{*}=+\ii E\phi^{*}, so

j0=i(ϕ(iEϕ)ϕ(iEϕ))=i(2iEϕ2)=2EN2. j^{0} = \ii\Big(\phi\,(\ii E\phi^{*}) - \phi^{*}(-\ii E\phi)\Big) = \ii\big(2\ii E\abs{\phi}^{2}\big) = -2EN^{2}.

The conjugate solution. ϕ=Ne+i(Etpx)\phi^{*}=N\ee^{+\ii(Et-\vv p\cdot\vv x)} is also a solution of ϕ=m2ϕ\Box\phi=-m^{2}\phi, and running the same computation gives j0=+2EN2j^{0}=+2EN^{2}: the same magnitude, opposite sign.

What it means. A complex field carries a conserved charge, and the field has two families of solutions carrying that charge with opposite signs and equal magnitude. A real field has no such symmetry, since there is no phase to rotate, and therefore no conserved charge at all. So a field carries charge if and only if it is complex. Multiplying jμj^{\mu} by the constant q-q (allowed, since any constant multiple of a conserved current is conserved) makes the first solution carry charge +q+q per unit 2EN22EN^{2} and the second q-q.

⚑ Looking forward: after quantisation in Chapter 5.3 the two families become particles and antiparticles. Antimatter is not an extra postulate. It is the second sign of the Noether charge of a U(1)U(1) symmetry, and it exists for exactly the same reason the charge does. A neutral particle like the photon is described by a real field precisely because it has no U(1)U(1) charge to carry, and is therefore its own antiparticle.

Problem 4 — a symmetry with no charge, and Kepler's third law

For the Kepler Lagrangian L=12mr˙2+k/rL=\half m\abs{\dot{\vv r}}^{2}+k/r, consider the scaling

r    r,t    3/2t. \vv r \;\longmapsto\; \ell\,\vv r, \qquad t \;\longmapsto\; \ell^{3/2}\,t.

(a) Show that it maps solutions to solutions. (b) Show that it is not a symmetry of the action, and that the would-be Noether charge is not conserved. (c) What does the transformation give you instead? (d) Contrast with Problem 1.

Solution

(a) The equation of motion is mr¨=kr/r3m\ddot{\vv r} = -k\vv r/r^{3}. Under the scaling, the left side picks up /3=2\ell/\ell^{3}=\ell^{-2} and the right side picks up /3=2\ell/\ell^{3}=\ell^{-2} as well. Both sides scale identically, so the equation is unchanged and every solution maps to a solution.

(b) Now scale the action. The velocity scales by /3/2=1/2\ell/\ell^{3/2}=\ell^{-1/2}, so T1TT\mapsto\ell^{-1}T. The potential goes as k/r1k/rk/r\mapsto\ell^{-1}k/r. Both terms of LL carry 1\ell^{-1}, while dt\dd t carries 3/2\ell^{3/2}, so

S    13/2S  =  1/2S    S. S \;\longmapsto\; \ell^{-1}\cdot\ell^{3/2}\,S \;=\; \ell^{1/2}\,S \;\neq\; S.

Not invariant, and not invariant up to a boundary term either. The change is proportional to SS itself, which is not the integral of a total derivative for a general path. The hypothesis of §1.2 fails.

Test the conclusion. With =1+ϵ\ell=1+\epsilon the would-be generators are K=r\vv K=\vv r and τ=32t\tau=\tfrac32t, so (1.4.15) would offer

Q  =?  pr32tE. Q \;\overset{?}{=}\; \vv p\cdot\vv r - \tfrac32 tE.

Differentiate along a solution, using ddt(pr)=p2/m+rp˙=2Tk/r\dv{}{t}(\vv p\cdot\vv r)=\abs{\vv p}^{2}/m+\vv r\cdot\dot{\vv p} = 2T - k/r (the last step because rp˙=k/r\vv r\cdot\dot{\vv p}=-k/r):

dQdt=2Tkr32E=2Tkr32(Tkr)=12(T+kr)  =  12L    0. \dv{Q}{t} = 2T - \frac{k}{r} - \tfrac32E = 2T - \frac{k}{r} - \tfrac32\left(T - \frac{k}{r}\right) = \half\left(T + \frac{k}{r}\right) \;=\; \half\,L \;\neq\; 0.

Not conserved, and not even conserved on average. Over one closed orbit pr\vv p\cdot\vv r returns to its starting value, so dQdt=32E0\avg{\dv{Q}{t}} = -\tfrac32 E \neq 0 for a bound orbit. What does average to zero is the simpler combination G=prG=\vv p\cdot\vv r on its own, whose rate is dGdt=2Tk/r\dv{G}{t}=2T-k/r. Setting its average to zero gives

2T  =  k/r, \avg{2T} \;=\; \avg{k/r},

the virial theorem. That is the residue a non-symmetry leaves behind. Not a constant, but a relation between time-averages.

(c) The transformation relates different solutions rather than constraining a single one. Take any orbit, scale all lengths by \ell and all times by 3/2\ell^{3/2}, and you get another legitimate orbit. Since the semi-major axis scales as aaa\mapsto\ell a and the period as T3/2TT\mapsto\ell^{3/2}T, the combination

T2a3    3T23a3  =  T2a3 \frac{T^{2}}{a^{3}} \;\longmapsto\; \frac{\ell^{3}T^{2}}{\ell^{3}a^{3}} \;=\; \frac{T^{2}}{a^{3}}

is the same for every orbit. That is Kepler's third law, obtained without solving the equations of motion at all, and from a symmetry that has no Noether charge. Symmetries of the equations of motion are not useless. They produce a different kind of information, namely relations between solutions rather than constants along one.

(d) The two problems differ in exactly one respect, which is whether the time-scaling exponent is the one that makes Ldt\int L\,\dd t balance. In Problem 1, LL scaled by 2\ell^{-2} and dt\dd t by +2\ell^{+2}, so they cancel and a Noether charge exists. Here LL scales by 1\ell^{-1} and dt\dd t by 3/2\ell^{3/2}, so they do not, and there is none. Same kind of transformation, same kind of system, and the presence or absence of a conservation law turns on a single exponent. That is a useful reminder that "is this a symmetry?" is a question about the action, and can only be answered by computing.

The brick you just laid

You can now convert symmetry into conservation, in either direction, by formula. Given a transformation δqi=ϵKi\delta q_{i}=\epsilon K_{i}, δt=ϵτ\delta t=\epsilon\tau under which the action is invariant up to a boundary term ϵF\epsilon F, the quantity Q=ipiKiHτFQ=\sum_{i}p_{i}K_{i}-H\tau-F is constant along solutions (1.4.15). And given a conserved QQ, the bracket δf=ϵ{f,Q}\delta f=\epsilon\{f,Q\} hands the symmetry back. The proof was four steps: variations commute with d/dt\dd/\dd t, expand by the chain rule, substitute the Euler–Lagrange equation, recognise a total derivative. The only subtlety worth remembering is that the hypothesis is off shell while the conclusion is on shell.

From the one formula came energy, momentum, angular momentum, the uniform motion of the centre of mass, and electric charge. You saw conservation fail, on screen, in exactly the component whose symmetry you broke, at exactly the predicted rate. You saw that discrete symmetries give no current, that in fields the theorem upgrades to a local statement, which is Chapter 0.7's continuity equation, and that an unexplained degeneracy is usually a symmetry nobody has named yet.

Where this gets spent. Chapter 2.5 applies it to spacetime translations in Minkowski space and gets four-momentum conservation as a single relation. Chapter 3.5 geometrises it: Killing vectors are Noether generators written as vector fields on a manifold, and Chapter 3.6 finds μTμν=0\nabla_{\mu}T^{\mu\nu}=0 falling out of the invariance of the Einstein–Hilbert action under coordinate changes. Chapter 5.2 runs §6 in earnest for relativistic fields and constructs the energy–momentum tensor. Chapter 6.3 takes §6.2's U(1)U(1), makes α\alpha depend on position, and out comes electromagnetism. Chapter 6.8 does it with the full Standard Model group. And Chapter 4.11 finds §7's bracket algebra again with commutators, which is where Chapter 4.12's spin comes from.

Part I, closed. Four chapters ago mechanics was a list of forces, one per interaction, each resolved into components in a coordinate system chosen for convenience. It is now four objects.

  • A single scalar LL.
  • The condition δS=0\delta S=0 that turns it into equations of motion.
  • A bracket {,}\{\,,\} that turns phase-space functions into motions.
  • A theorem converting symmetry into conservation.

Nothing in that list mentions gravity, or electrons, or fields, or any particular physics at all. That is the point, because every remaining part of this book is that same package applied to a different symmetry group. Lorentz transformations in Part II. Diffeomorphisms, which are arbitrary smooth coordinate changes, in Part III. Unitary phase rotations, first global and then local, in Parts IV to VI. Conformal transformations of a two-dimensional worldsheet in Part VII. You now have the whole method. What remains is to learn the groups.