Part I · The Action Principle — Chapter 1.3

Hamilton and Phase Space

Trade nn second-order equations for 2n2n first-order ones, and the geometry starts doing the work.

Where we are

Chapter 1.2 handed you a machine. Feed it one scalar function L(q,q˙,t)L(q,\dot q,t) and it returns nn second-order differential equations, one per degree of freedom, in a form that survives any change of coordinates you care to make. That machine is not going away. It is the formulation that generalises to special relativity (Chapter 2.5), to fields (Chapter 5.2), and to curved spacetime (Chapter 3.6), and the reason is that the action is a scalar. Scalars do not care what you call your coordinates.

This chapter builds a second machine out of the same materials. It costs one change of variables. We swap the velocities q˙i\dot q^{i} for the momenta pi=L/q˙ip_{i}=\partial L/\partial\dot q^{i}, and in exchange we get three things the Lagrangian formulation does not give us.

The first is 2n2n first-order equations in place of nn second-order ones. That means the state of the system is a single point, and the dynamics is a flow.

The second is a theorem, Liouville's, saying that flow is incompressible. It is the foundation of statistical mechanics. Read correctly, it is also the classical ancestor of the statement that quantum evolution is unitary.

The third is the one to keep your eye on. It is an operation on functions called the Poisson bracket, and it has the property that

replacing it by 1i[  ,  ]\dfrac{1}{\ii\hbar}\big[\;,\;\big] and changing nothing else turns classical mechanics into quantum mechanics.

That is what Part I is for. Chapter 4.9 will spend one line doing the substitution and the rest of the chapter reaping it, and the reason that line can be short is that this chapter is long.

Here is the summary you should finish with. The Lagrangian formulation is the one that generalises; the Hamiltonian formulation is the one that quantises. You need both, for different reasons, and mistaking one for the other is a standard way to get stuck.

Tools you'll need  — Chapter 0.4: the determinant as a volume factor, the trace, and det(I+ϵA)=1+ϵtrA+O(ϵ2)\det(I+\epsilon A)=1+\epsilon\,\mathrm{tr}A+O(\epsilon^{2}). Section 5 is that identity. Chapter 0.5: quadratic forms and positive definiteness, which is exactly the condition §1 needs. Chapter 0.6: the Hessian, convexity, the multivariable chain rule, and Clairaut's theorem on mixed partials. Sections 5, 6 and 7 each turn on Clairaut and nothing else. Chapter 0.7: divergence as the trace of the Jacobian, hence as the fractional rate of change of volume carried by a flow. Section 5 is one substitution into that result. Chapter 0.8: Picard–Lindelöf and the uniqueness of solutions, the harmonic oscillator, the phase-space ellipse, and the pendulum's separatrix. Chapter 1.2: Euler–Lagrange, the canonical momentum pi=L/q˙ip_{i}=\partial L/\partial\dot q^{i}, and Problem 2's Beltrami identity, which, you are about to discover, was the Hamiltonian all along.

1 · The Legendre transform, properly

Chapter 1.2's Problem 2 asked you to show that when LL has no explicit time dependence the quantity

E  =  q˙Lq˙    L E \;=\; \dot q\,\pdv{L}{\dot q} \;-\; L (1.3.1)

is conserved. We are about to promote that combination from a conserved quantity to the central object of a whole formulation of mechanics. It will be a great deal less mysterious if you first see what kind of operation it is.

It is not an ad hoc grouping of terms. It is an instance of a standard construction with a clean geometric meaning, and the same construction is what turns the internal energy of a gas into its free energy. So here is the plan for this section: §1 is mathematics, and there is no mechanics in it at all.

1.1 · Two ways to describe a curve

Take a smooth, strictly convex function f:RRf:\R\to\R. For now just picture a parabola. There are two ways to say which function it is.

By points. List the pairs (x,f(x))\big(x,f(x)\big). This is the description you grew up with.

By tangent lines. For each slope pp, there is exactly one line of that slope tangent to the graph. Give me the family of all those lines, meaning for each pp its intercept, and I can reconstruct the curve. The reason is that a convex curve is the upper envelope of its own tangent lines. Every tangent line lies below the graph and touches it once, and the graph is what you get by taking, at each xx, the highest of them.

Both descriptions contain the same information. The Legendre transform is the dictionary between them. That is the whole idea, and the formula is going to be an immediate consequence of it rather than a definition to memorise.

Let's set it up. A line of slope pp has the form y=pxcy = px - c, where c-c is its yy-intercept. We carry the minus sign because it makes every formula below come out with a plus.

Now push this line down from above, increasing cc, until it first touches the graph. Touching means the vertical gap between line and graph has shrunk to zero somewhere. Before that moment the line sits above the graph everywhere. So the critical value of cc is the largest gap there ever was, which we write as

g(p)    supx[px    f(x)]. g(p) \;\equiv\; \sup_{x}\Big[\,p\,x \;-\; f(x)\,\Big]. (1.3.2)

Read the bracket as: how far the line y=pxy=px rises above the graph at the point xx. Take the largest such gap. That is the amount you must lower y=pxy=px to make it tangent, and it is therefore precisely (intercept of the tangent line of slope p)-(\text{intercept of the tangent line of slope }p).

What gg is, in one sentence

g(p)g(p) is minus the yy-intercept of the tangent line to ff whose slope is pp. The function gg is the "by tangent lines" description of ff, packaged so that its argument is the slope.

Everything else in §1 is a consequence of that sentence: the formula p=f(x)p=f'(x), the involution g=fg^{**}=f, the thermodynamic potentials, and H=pq˙LH=p\dot q-L.

1.2 · Where the supremum sits

Now let's do the calculus. The bracket in (1.3.2) is, for fixed pp, an ordinary function of the single variable xx. So we want its stationary point, and Chapter 0.6's rule for finding one applies without change:

x[pxf(x)]  =  pf(x)  =  0  p=f(x).   \pdv{}{x}\Big[px-f(x)\Big] \;=\; p - f'(x) \;=\; 0 \qquad\Longrightarrow\qquad \boxed{\;p = f'(x).\;} (1.3.3)

So the maximising xx is the point where the graph's own slope equals pp, which is exactly what "tangent line of slope pp" meant. Nothing has been assumed yet except differentiability. Two questions remain, and they are what convexity is for.

Is the stationary point a maximum? The second derivative of the bracket is f(x)-f''(x), which is negative wherever f>0f''\gt0. So yes, provided ff is strictly convex.

Is there exactly one? The equation f(x)=pf'(x)=p must have a unique solution for each pp in range. Since f>0f''\gt0 means ff' is strictly increasing, ff' is injective, and so it is invertible on its image. Again convexity, and again it is doing genuine work rather than decorating the theorem.

That unique maximiser is a function of pp, so give it a name and write x(p)x(p). Substituting it back into the definition turns the supremum into an ordinary formula:

g(p)  =  px(p)    f(x(p)),where f(x(p))=p. g(p) \;=\; p\,x(p) \;-\; f\big(x(p)\big), \qquad\text{where } f'\big(x(p)\big)=p. (1.3.4)

Put (1.3.4) beside (1.3.1) and the shape is already the same: slope times variable, minus the function. We will make that identification precise in §2.

1.3 · The transform is its own inverse

Here is the property that makes the construction worth having, and getting it takes one differentiation. We want to know what gg does when its own argument pp changes, so differentiate (1.3.4) with respect to pp, using the product rule and the chain rule:

dgdp  =  x(p)  +  pdxdp    f(x(p))dxdp  =  x(p), \dv{g}{p} \;=\; x(p) \;+\; p\,\dv{x}{p} \;-\; f'\big(x(p)\big)\,\dv{x}{p} \;=\; x(p), (1.3.5)

The last two terms cancel exactly, because f(x(p))=pf'(x(p))=p. That cancellation is not a coincidence, and it will happen again in §3 for the same reason: we are differentiating at a stationary point, so the motion of the stationary point itself contributes nothing at first order.

Now set the result beside the equation that defined the maximiser in the first place. Reading (1.3.5) next to (1.3.3) gives

f(x)=p,g(p)=x. f'(x) = p, \qquad\qquad g'(p) = x. (1.3.6)

Perfectly symmetric. The transform swaps the roles of the variable and the slope. And since g(p)=dx/dp=1/f(x)>0g''(p) = \dd x/\dd p = 1/f''(x) \gt 0, the transform of a strictly convex function is strictly convex, so we are entitled to transform a second time. Let's do exactly that and see where it lands:

g(x)  =  supp[xpg(p)], g^{*}(x) \;=\; \sup_{p}\big[xp - g(p)\big], (1.3.7)

Its supremum sits where g(p)=xg'(p)=x, which by (1.3.6) is the pp with x(p)=xx(p)=x, which is p=f(x)p=f'(x). So evaluate the bracket there and see what number comes out:

g(x)  =  xf(x)[f(x)xf(x)]  =  f(x). g^{*}(x) \;=\; x\,f'(x) - \Big[f'(x)\,x - f(x)\Big] \;=\; f(x). (1.3.8)

The Legendre transform is an involution on strictly convex functions. Do it twice and you are back where you started. That is what "dictionary" meant. No information is lost going from the point description to the tangent-line description, because you can always go back.

⚑ One statement here is quoted rather than proved. For a function that is not convex, the transform still exists, since the supremum in (1.3.2) is well defined whenever it is finite. What fails is the involution. The double transform gg^{*} returns the convex hull of ff, meaning the largest convex function lying below it, and any non-convex wiggles are erased permanently. This is the Fenchel–Moreau theorem, and it is the precise sense in which "the tangent lines determine the curve" needs convexity. It matters because it tells us which Lagrangians the Hamiltonian formulation applies to, and we will need it for exactly that in §2.3.

1.4 · A worked scalar example

Abstract dictionaries are easier to trust once you have run one. Take f(x)=14x4f(x)=\tfrac14 x^{4} on x0x\ge0, which is strictly convex there (f=3x2>0f''=3x^{2}\gt0 for x0x\neq0). Then p=f(x)=x3p = f'(x) = x^{3}, so x=p1/3x = p^{1/3}, and putting that into (1.3.4) gives

g(p)  =  pp1/314(p1/3)4  =  p4/314p4/3  =  34p4/3. g(p) \;=\; p\cdot p^{1/3} - \tfrac14\big(p^{1/3}\big)^{4} \;=\; p^{4/3} - \tfrac14 p^{4/3} \;=\; \tfrac34\,p^{4/3}. (1.3.9)

Now check the involution on this example, without appealing to the theorem. First the inverse relation: g(p)=3443p1/3=p1/3=xg'(p) = \tfrac34\cdot\tfrac43 p^{1/3} = p^{1/3} = x ✓, matching (1.3.6). Then transform a second time and see whether ff comes back:

g(x)  =  supp[xp34p4/3]atp=x3x434x4=14x4=f(x).   g^{*}(x) \;=\; \sup_{p}\Big[xp - \tfrac34 p^{4/3}\Big] \quad\text{at}\quad p = x^{3} \quad\Longrightarrow\quad x^{4} - \tfrac34 x^{4} = \tfrac14 x^{4} = f(x). \;\checkmark (1.3.10)

It does. And the supremum is a real number you can measure rather than a formal device. Numerically, at p=2p=2 the supremum of 2x14x42x-\tfrac14x^{4} over a fine grid is 1.88988161.8898816, and 3424/3=1.8898816\tfrac34\cdot2^{4/3} = 1.8898816. The formula is measuring something real.

Familiar ground

You have used this transform for years under a different name, and knowing that it is the same operation makes the zoo of thermodynamic potentials collapse into one idea.

Start from the internal energy U(S,V)U(S,V) of a simple system, whose differential is the first law,

dU  =  TdS    PdV,soT=USV,P=UVS. \dd U \;=\; T\,\dd S \;-\; P\,\dd V, \qquad\text{so}\qquad T = \pdv{U}{S}\bigg|_{V}, \quad P = -\pdv{U}{V}\bigg|_{S}.

Look at what that says about temperature. Temperature is the slope of UU with respect to entropy. Now suppose you would rather label states by TT than by SS, which is what a laboratory forces on you, since thermometers exist and entropy meters do not. That is exactly the swap (1.3.6) performs. So define

F(T,V)  =  U(S,V)TSevaluated at the S where US=T, F(T,V) \;=\; U(S,V) - T S \qquad\text{evaluated at the } S \text{ where } \pdv{U}{S}=T,

and the Helmholtz free energy is 1-1 times the Legendre transform of UU in the variable SS. (Physicists carry the extra minus sign so that FF is minimised at equilibrium. The mathematical content is identical.) To see which variables FF naturally belongs to, take its differential, which is one line:

dF=dUTdSSdT=SdTPdV, \dd F = \dd U - T\,\dd S - S\,\dd T = -S\,\dd T - P\,\dd V,

The TdST\,\dd S terms cancel, and it is the same cancellation as (1.3.5), for the same reason. So the natural variables of FF are (T,V)(T,V), and S=F/TS=-\partial F/\partial T is the inverse relation (1.3.6).

Every other potential is the same move on a different slot. Transform VPV\to P instead and you get the enthalpy H=U+PV\mathcal{H}=U+PV with dH=TdS+VdP\dd\mathcal{H}=T\,\dd S+V\,\dd P. Transform both and you get the Gibbs free energy G=HTS=UTS+PVG=\mathcal{H}-TS=U-TS+PV with dG=SdT+VdP\dd G=-S\,\dd T+V\,\dd P, which is why GG is the one a chemist uses. Its natural variables are the two a bench can control.

Convexity is not decoration here either. The statement 2U/S2=T/S=T/CV>0\partial^{2}U/\partial S^{2} = \partial T/\partial S = T/C_{V} \gt 0 is thermodynamic stability, and it is exactly the condition (1.3.3) needed for the transform to be invertible. So a system with negative heat capacity has no well-defined free energy. That is not a mathematical curiosity. Self-gravitating systems have exactly that pathology, and it is why a star heats up when it loses energy.

In §2 we do the identical operation on L(q,q˙)L(q,\dot q), swapping q˙\dot q for its slope p=L/q˙p=\partial L/\partial\dot q. Same transform, same cancellation, same inverse relation. Only the letters change.

Grind box — Young's inequality, and what the supremum is really enforcing

Definition (1.3.2) says g(p)pxf(x)g(p)\ge px-f(x) for every xx, with equality only at the maximiser. Rearranged, that is

px    f(x)+g(p)for all x,p, p\,x \;\le\; f(x) + g(p) \qquad\text{for all } x,p,

with equality precisely when p=f(x)p=f'(x). This is Young's inequality, and it is not a separate theorem. It is the definition of the transform, read sideways.

The classic case is f(x)=xa/af(x)=x^{a}/a with a>1a\gt1 on x0x\ge0. Then p=xa1p=x^{a-1}, so x=p1/(a1)x=p^{1/(a-1)} and

g(p)=pa/(a1)1apa/(a1)=a1apa/(a1)=pbb,1a+1b=1, g(p) = p^{a/(a-1)} - \frac{1}{a}p^{a/(a-1)} = \frac{a-1}{a}\,p^{a/(a-1)} = \frac{p^{b}}{b}, \qquad \frac1a+\frac1b=1,

since ba/(a1)b\equiv a/(a-1) satisfies 1/b=(a1)/a=11/a1/b = (a-1)/a = 1-1/a. So Young's inequality reads pxxa/a+pb/bpx\le x^{a}/a + p^{b}/b, the inequality from which Hölder's and Minkowski's inequalities are built, and the reason conjugate exponents come in pairs summing to 11 in their reciprocals. Our worked example was a=4a=4, b=4/3b=4/3. ✓

What breaks without convexity. Take f(x)=14x412x2f(x)=\tfrac14x^{4}-\tfrac12x^{2}, the double well, which has f=3x21<0f''=3x^{2}-1\lt0 on x<1/3\abs{x}\lt1/\sqrt3. For each slope pp there can now be three points with f(x)=pf'(x)=p, and (1.3.2) selects whichever gives the largest gap, the global one. The information about the other two is discarded, and the double transform gg^{*} comes back as the convex hull: the double well with its non-convex middle replaced by a straight segment. This is not a defect. It is exactly the Maxwell construction of a first-order phase transition: the flat segment is the coexistence region, the tie-line joining liquid and gas, and its slope is the common tangent, which is the shared value of the chemical potential. Legendre-transforming a non-convex free energy and getting a flat piece is the phase transition. Chapter 6.6 meets the same picture again with the Higgs potential.

The condition in nn dimensions. With several variables, pi=f/xip_{i}=\partial f/\partial x^{i} is invertible for xx precisely when the Jacobian of that map, the Hessian 2f/xixj\partial^{2}f/\partial x^{i}\partial x^{j}, is nonsingular, by the inverse function theorem of Chapter 0.6. Strict convexity in the multivariable sense means the Hessian is positive definite (Chapter 0.5), which is more than enough. Hold that thought. In §2 the Hessian in question is 2L/q˙iq˙j\partial^{2}L/\partial\dot q^{i}\partial\dot q^{j}, which Chapter 1.2's grind box already showed is the positive-definite mass matrix whenever the kinetic energy is an honest quadratic form.

In plain terms 1.3.1

A curve can be given by listing its points, or by listing all the straight lines that graze it without crossing, and the two lists carry the same information, because a curve bending only one way is the highest of its own tangent lines. Swapping between the two descriptions is a standard operation, and performing it twice returns you exactly to where you began.

Which description you want is decided by what you are able to set. A laboratory has thermometers and no entropy meters, so a chemist labels a state by its temperature, which is the slope of the internal energy against entropy rather than the entropy itself, and the free energies are that one swap performed on one slot or another. What makes it reversible is that the curve bends one way only. Where it does not, the swap flattens the offending stretch permanently, and that flat stretch is a first-order phase transition rather than a defect in the mathematics.

In mechanics the variable traded away is the velocity and the slope accepted in its place is the momentum, and the reversibility condition becomes the statement that the kinetic energy is an honest positive quantity in the velocities. What the trade produces has been met before, as the combination a problem with no explicit clock in it happens to conserve.

2 · Momentum and the Hamiltonian

2.1 · The canonical momentum is not mvm\vv v

Before we transform anything, we need to be clear about the variable we are trading for. Chapter 1.2 defined, for any Lagrangian and any coordinates,

pi    Lq˙i, p_{i} \;\equiv\; \pdv{L}{\dot q^{i}}, (1.3.11)

the canonical momentum conjugate to qiq^{i}, and noted that this makes the Euler–Lagrange equation read p˙i=L/qi\dot p_{i}=\partial L/\partial q^{i}. In Cartesian coordinates with L=12mx˙2V(x)L=\half m\dot{\vv x}^{2}-V(\vv x), (1.3.11) returns mx˙im\dot x^{i}, which is why the name is not absurd. But that agreement is a special case, and it fails in two different ways that both matter.

It need not have the dimensions of momentum. For the pendulum with coordinate θ\theta, L=12m2θ˙2+mgcosθL=\half m\ell^{2}\dot\theta^{2}+mg\ell\cos\theta gives pθ=m2θ˙p_{\theta}=m\ell^{2}\dot\theta. That is an angular momentum, with units of Js\mathrm{J\,s} rather than kgms1\mathrm{kg\,m\,s^{-1}}. The word "conjugate" is the point. Whatever pip_{i} turns out to be, it is the thing that pairs with qiq^{i} so that pidqip_{i}\,\dd q^{i} has the dimensions of action. Chapter 4.9's uncertainty relation ΔqΔp/2\Delta q\,\Delta p\ge\hbar/2 is dimensionally consistent for every conjugate pair for exactly this reason.

It need not be mvm\vv v even in Cartesian coordinates. This is the one worth being careful about, and seeing it costs one example. Chapter 2.6 will derive, from relativity, that a particle of charge ee in an electromagnetic field described by the scalar potential ϕ\phi and vector potential A\vv A has the Lagrangian

L  =  12mvv    eϕ(x,t)  +  evA(x,t). L \;=\; \half m\,\vv v\cdot\vv v \;-\; e\,\phi(\vv x,t) \;+\; e\,\vv v\cdot\vv A(\vv x,t). (1.3.12)

Worked example 1 checks that this reproduces the Lorentz force, so it is the right Lagrangian. Our question is what its canonical momentum is, so apply (1.3.11) to it, component by component, remembering that A\vv A depends on position but not on velocity:

pi  =  vi[12mvjvj+evjAj]  =  mvi+eAip  =  mv+eA. p_{i} \;=\; \pdv{}{v^{i}}\left[\half m\,v^{j}v^{j} + e\,v^{j}A_{j}\right] \;=\; m v^{i} + e A_{i} \qquad\Longleftrightarrow\qquad \vv p \;=\; m\vv v + e\vv A. (1.3.13)

The canonical momentum and the mechanical momentum mvm\vv v differ by eAe\vv A. This is the first time the vector potential has entered mechanics in this book, and it will not leave. Chapter 2.6 shows that A\vv A is the spatial part of a four-vector AμA^{\mu} and that Maxwell's equations are its equations of motion. Chapter 6.3 shows that the entire structure of the Standard Model follows from demanding that the substitution ppeA\vv p\to\vv p - e\vv A be forced on you by a symmetry principle. The name for that substitution is minimal coupling, and you have just watched it appear, three parts early, as an unremarkable partial derivative.

⚠ Why this isn't obvious

It is tempting to file (1.3.13) under "notation", on the grounds that mvm\vv v must be the real momentum and p\vv p a bookkeeping device. That is exactly backwards, and the difference is measurable. Here is why it matters in practice.

The gauge problem. Chapter 1.2's Problem 4 showed that AA+χ\vv A\to\vv A+\nabla\chi, ϕϕtχ\phi\to\phi-\partial_{t}\chi changes LL by a total time derivative and therefore changes no physics. But it does change p\vv p, by eχe\nabla\chi. So the canonical momentum is gauge-dependent, and by itself it is not an observable. Meanwhile mv=peAm\vv v = \vv p - e\vv A is gauge-invariant, because A\vv A shifts by exactly the same amount. The mechanical momentum is what you measure. The canonical momentum is what the formalism runs on. Both are needed and they are not the same object.

And yet it is p\vv p that gets quantised. In Chapter 4.6 the operator conjugate to position is p^=i\hat{\vv p}=-\ii\hbar\nabla, and the rule for which classical quantity it represents is fixed by the bracket {qi,pj}=δji\{q^{i},p_{j}\}=\delta^{i}_{j} of §6, which is a statement about the canonical momentum. Feed (1.3.13) into a Hamiltonian and the kinetic term becomes (p^eA)2/2m(\hat{\vv p}-e\vv A)^{2}/2m, so that i-\ii\hbar\nabla is replaced by ieA-\ii\hbar\nabla-e\vv A acting on the wavefunction. Every magnetic effect in quantum mechanics comes out of that one replacement: Landau levels, the Aharonov–Bohm phase, the Zeeman effect, superconducting flux quantisation. None of them would, if mvm\vv v were the thing you quantised.

⚑ The sharpest experimental statement is quoted forward to Chapter 6.3. In the Aharonov–Bohm effect an electron beam is split around a solenoid and recombined. Outside the solenoid the magnetic field is exactly zero, so a particle described by mvm\vv v alone can feel nothing. The interference fringes nonetheless shift, by an amount eAd/e\oint\vv A\cdot\dd\vv\ell/\hbar, which is the enclosed flux in units of /e\hbar/e. The potential is not a computational convenience. It is where the physics lives, and p=mv+eA\vv p = m\vv v + e\vv A is the first hint.

2.2 · The Hamiltonian

Now we have both halves of the trade: a function LL that depends on velocity, and a slope p=L/q˙p=\partial L/\partial\dot q we would rather use instead. So perform §1's transform on LL, taking q˙i\dot q^{i} as the variable and holding qq and tt fixed. The conjugate slope is pi=L/q˙ip_{i}=\partial L/\partial\dot q^{i} by (1.3.11), and the transform (1.3.4) reads

  H(q,p,t)    ipiq˙i    L(q,q˙,t),   \boxed{\;H(q,p,t) \;\equiv\; \sum_{i}p_{i}\,\dot q^{i} \;-\; L\big(q,\dot q,t\big),\;} (1.3.14)

Every q˙i\dot q^{i} on the right is to be eliminated in favour of the pip_{i} by inverting (1.3.11), and that instruction is the entire content of the definition. The result is a function of qq, pp and tt only. This object is the Hamiltonian. Compare it with (1.3.1) and you can see that Chapter 1.2's Problem 2 was computing it: the conserved quantity of a time-independent Lagrangian is the Hamiltonian.

The inversion step is not free, and §1 told us exactly what it needs. Solving pi=L/q˙ip_{i}=\partial L/\partial\dot q^{i} for q˙\dot q requires that the map from velocities to momenta be invertible, and by the inverse function theorem that means asking about the Hessian

Wij    2Lq˙iq˙j W_{ij} \;\equiv\; \frac{\partial^{2}L}{\partial\dot q^{i}\,\partial\dot q^{j}} (1.3.15)

which must be nonsingular. Chapter 1.2's grind box on the double pendulum established that when the coordinate change from Cartesians is time-independent, T=12ijMij(q)q˙iq˙jT=\half\sum_{ij}M_{ij}(q)\dot q^{i}\dot q^{j} with MM symmetric positive definite, so W=MW=M and the condition holds automatically.

⚑ It does not always hold, and the failure has a name. When LL is linear in some velocity, the corresponding row of WW vanishes, the transform is not invertible, and the system is called constrained. Dirac built a whole formalism for that case. It is quoted here and not developed, but you should know where it lives, because it is not an exotic corner. Gauge theories are there (Chapter 6.3: the Lagrangian of electromagnetism has no A˙0\dot A_{0} at all), and so is the reparametrisation-invariant string action (Chapter 7.2). Every fundamental theory in Parts VI and VII is a constrained system. So the honest statement is this. The clean Legendre transform of this section is the easy case, and it is the one that teaches you what the hard case is deforming.

2.3 · When H=T+VH=T+V, and when it isn't

You have surely been told that the Hamiltonian is the total energy. It often is. But that is a theorem rather than a definition, and the theorem has two hypotheses that fail in ordinary situations. Let's state it, prove it, and then break it.

Theorem

Suppose (i) VV does not depend on the velocities, and (ii) TT is a homogeneous quadratic form in the generalised velocities, T=12ijMij(q)q˙iq˙jT=\half\sum_{ij}M_{ij}(q)\dot q^{i}\dot q^{j}. Then H=T+VH = T+V.

Proof. With L=TVL=T-V and hypothesis (i), pi=L/q˙i=T/q˙ip_{i}=\partial L/\partial\dot q^{i} = \partial T/\partial\dot q^{i}. Hypothesis (ii) says TT is homogeneous of degree 22: replacing every q˙\dot q by λq˙\lambda\dot q multiplies TT by λ2\lambda^{2}. What we want is the combination piq˙i\sum p_{i}\dot q^{i}, so differentiate that homogeneity identity T(q,λq˙)=λ2T(q,q˙)T(q,\lambda\dot q) = \lambda^{2}T(q,\dot q) with respect to λ\lambda and set λ=1\lambda=1. The chain rule gives

iq˙iTq˙i  =  2T \sum_{i}\dot q^{i}\,\pdv{T}{\dot q^{i}} \;=\; 2\,T (1.3.16)

which is Euler's theorem on homogeneous functions, derived rather than quoted. (The grind box does the general degree.) That is the whole input. Feed it into the definition of HH and every term is now known:

H  =  ipiq˙iL  =  iq˙iTq˙i(TV)  =  2TT+V  =  T+V. H \;=\; \sum_{i}p_{i}\dot q^{i} - L \;=\; \sum_{i}\dot q^{i}\pdv{T}{\dot q^{i}} - (T-V) \;=\; 2T - T + V \;=\; T+V. (1.3.17)

\blacksquare

Now let's break the hypotheses, using a system Chapter 1.2 already built. Problem 3 there put a bead on a hoop of radius RR spun about its vertical diameter at a rate Ω\Omega fixed by a motor, and found

L  =  12mR2θ˙2  +  12mR2Ω2sin2θ  +  mgRcosθ. L \;=\; \half mR^{2}\dot\theta^{2} \;+\; \half mR^{2}\Omega^{2}\sin^{2}\theta \;+\; mgR\cos\theta. (1.3.18)

The middle term is part of the kinetic energy, since it is the speed the bead has purely because the hoop is carrying it around. But it contains no θ˙\dot\theta. So TT is not homogeneous of degree 22 in the velocity, and hypothesis (ii) fails. The theorem is therefore unavailable, and we must compute honestly instead. Since pθ=mR2θ˙p_{\theta}=mR^{2}\dot\theta, we have θ˙=pθ/mR2\dot\theta = p_{\theta}/mR^{2}, and putting that into the definition gives HH alongside the true energy for comparison:

H=pθθ˙L  =  pθ22mR2    12mR2Ω2sin2θ    mgRcosθ,T+V=pθ22mR2  +  12mR2Ω2sin2θ    mgRcosθ. \begin{aligned} H &= p_{\theta}\dot\theta - L \;=\; \frac{p_{\theta}^{2}}{2mR^{2}} \;-\; \half mR^{2}\Omega^{2}\sin^{2}\theta \;-\; mgR\cos\theta,\\[4pt] T+V &= \frac{p_{\theta}^{2}}{2mR^{2}} \;+\; \half mR^{2}\Omega^{2}\sin^{2}\theta \;-\; mgR\cos\theta. \end{aligned} (1.3.19)

They differ by mR2Ω2sin2θmR^{2}\Omega^{2}\sin^{2}\theta, which changes as the bead slides. HH is conserved, since it has no explicit tt. T+VT+V is not, because the motor is doing work to hold Ω\Omega fixed. Conserved and "the energy" are different predicates, and HH is the first one.

Notice also what HH is here. It is the kinetic energy of the θ\theta motion plus the effective potential VeffV_{\text{eff}} that Chapter 1.2 found. The centrifugal term has flipped sign on its way through the transform and become part of the potential. That is not sleight of hand. It is (1.3.16) refusing to apply to a degree-zero term.

Grind box — Euler's theorem in general, the three-piece decomposition, and the relativistic particle

Euler's theorem. Let FF be homogeneous of degree kk in the variables uiu^{i}, meaning F(λu)=λkF(u)F(\lambda u) = \lambda^{k}F(u) for all λ>0\lambda \gt 0. Differentiate both sides with respect to λ\lambda using the chain rule:

iuiFuiλu=kλk1F(u), \sum_{i}u^{i}\,\pdv{F}{u^{i}}\bigg|_{\lambda u} = k\,\lambda^{k-1}F(u),

and set λ=1\lambda=1 to get iuiF/ui=kF\sum_{i}u^{i}\partial F/\partial u^{i} = kF. That is the whole proof. Degree 22 gives (1.3.16), degree 11 gives uiiF=F\sum u^{i}\partial_{i}F = F, and degree 00 gives uiiF=0\sum u^{i}\partial_{i}F = 0.

The general Lagrangian. Split L=L2+L1+L0L = L_{2}+L_{1}+L_{0} by degree in q˙\dot q. Then piq˙i=2L2+L1p_{i}\dot q^{i} = 2L_{2}+L_{1} by the three cases above, so

H  =  (2L2+L1)(L2+L1+L0)  =  L2L0. H \;=\; \big(2L_{2}+L_{1}\big) - \big(L_{2}+L_{1}+L_{0}\big) \;=\; L_{2}-L_{0}.

Every degree-one piece drops out of HH entirely. Three consequences, all checkable against what we have already done:

  • Ordinary mechanics: L2=TL_{2}=T, L1=0L_{1}=0, L0=VL_{0}=-V, giving H=T+VH=T+V. ✓
  • Rotating hoop: L2=12mR2θ˙2L_{2}=\half mR^{2}\dot\theta^{2}, L1=0L_{1}=0, L0=12mR2Ω2sin2θ+mgRcosθL_{0}=\half mR^{2}\Omega^{2}\sin^{2}\theta+mgR\cos\theta, giving exactly (1.3.19). ✓ The centrifugal term's sign flip is the minus in L2L0L_{2}-L_{0}.
  • Charged particle: L2=12mv2L_{2}=\half mv^{2}, L1=evAL_{1}=e\,\vv v\cdot\vv A, L0=eϕL_{0}=-e\phi, giving H=12mv2+eϕH = \half mv^{2}+e\phi. The magnetic term vanishes from HH, which is the statement that magnetic forces do no work, derived rather than asserted. It reappears the moment you express HH in terms of p\vv p rather than v\vv v, because 12mv2=(peA)2/2m\half mv^{2}=(\vv p-e\vv A)^{2}/2m. Worked example 1 does this properly.

The relativistic free particle. Chapter 1.2's table listed S=mc2 ⁣dτS=-mc^{2}\!\int\dd\tau, i.e. in one dimension L=mc21x˙2/c2L = -mc^{2}\sqrt{1-\dot x^{2}/c^{2}}, which is not of the form TVT-V and is not homogeneous of any degree. Turn the crank anyway:

p=Lx˙=mc2x˙/c21x˙2/c2=mx˙1x˙2/c2=γmx˙. p = \pdv{L}{\dot x} = -mc^{2}\cdot\frac{-\dot x/c^{2}}{\sqrt{1-\dot x^{2}/c^{2}}} = \frac{m\dot x}{\sqrt{1-\dot x^{2}/c^{2}}} = \gamma m\dot x.

Then, using 1/γ=1x˙2/c21/\gamma = \sqrt{1-\dot x^{2}/c^{2}},

H=px˙L=γmx˙2+mc2γ=m(γ2x˙2+c2)γ=mc2γ2γ=γmc2, H = p\dot x - L = \gamma m\dot x^{2} + \frac{mc^{2}}{\gamma} = \frac{m\big(\gamma^{2}\dot x^{2}+c^{2}\big)}{\gamma} = \frac{mc^{2}\gamma^{2}}{\gamma} = \gamma mc^{2},

where the middle step used γ2x˙2+c2=c2γ2(x˙2/c2+1/γ2)=c2γ2\gamma^{2}\dot x^{2}+c^{2} = c^{2}\gamma^{2}\big(\dot x^{2}/c^{2}+1/\gamma^{2}\big) = c^{2}\gamma^{2}. So H=γmc2H=\gamma mc^{2}, the relativistic energy including the rest energy, produced by a Legendre transform and nothing else. Expressed in the right variable, which is what HH demands, we use p2c2=γ2m2x˙2c2p^{2}c^{2}=\gamma^{2}m^{2}\dot x^{2}c^{2} and the same identity to get

H(p)=p2c2+m2c4. H(p) = \sqrt{p^{2}c^{2}+m^{2}c^{4}}.

That is the relation Chapter 2.5 derives from four-vectors and Chapter 5.5 quantises into the Dirac equation. It fell out of §1's transform applied to a square root. Note in passing that 2L/x˙2=mγ3>0\partial^{2}L/\partial\dot x^{2} = m\gamma^{3} \gt 0, so the Hessian condition (1.3.15) holds and the transform is legitimate. It is the same positivity Chapter 1.2's second-variation grind box used to show that a relativistic worldline maximises proper time.

In plain terms 1.3.2

The quantity paired with a coordinate here is not always mass times velocity, and the two ways it can differ both matter. For a pendulum described by an angle it is an angular momentum, with the wrong units for a momentum entirely, because what the pairing requires is that coordinate and partner multiply to something with the units of action. For a charge in a magnetic field it is not the mechanical momentum even in Cartesian coordinates, the field's potential having been added to it.

That second case is not notation. The mechanical momentum is what an instrument reads, while the conjugate one shifts when the potential is rewritten in a way that changes no physics, so the conjugate one is not by itself observable. Yet the conjugate one is what quantum mechanics turns into an operator, which is why every magnetic effect in that subject, including the interference shift around a solenoid whose field the electrons never enter, follows from a single substitution.

The other object built here is usually introduced as the total energy. It is the total energy under two hypotheses, both of which fail in ordinary situations, and a bead on a motor-driven hoop is enough to break them. The conserved quantity and the energy come apart there, and it is the first that the formulation runs on.

3 · Hamilton's equations, derived

We have a new function H(q,p,t)H(q,p,t). We do not yet have equations of motion in terms of it. There are two ways to get them, and they are worth seeing both, because they emphasise different things.

3.1 · First route: take the differential

Chapter 0.6's total differential is the statement that for any smooth function of several variables, the first-order change is the sum of the partial derivatives times the changes in their arguments. Our aim is to find out which variables HH genuinely depends on, so apply that rule to the right-hand side of (1.3.14), treating qq, q˙\dot q and tt as the independent variables there:

dH=i(q˙idpi+pidq˙i)    i(Lqidqi+Lq˙idq˙i)Ltdt. \begin{aligned} \dd H &= \sum_{i}\Big(\dot q^{i}\,\dd p_{i} + p_{i}\,\dd\dot q^{i}\Big) \;-\; \sum_{i}\left(\pdv{L}{q^{i}}\dd q^{i} + \pdv{L}{\dot q^{i}}\,\dd\dot q^{i}\right) - \pdv{L}{t}\,\dd t. \end{aligned} (1.3.20)

Now look at the two terms containing dq˙i\dd\dot q^{i}. Their coefficients are pip_{i} and L/q˙i-\partial L/\partial\dot q^{i}, and by the definition (1.3.11) those are equal and opposite. They cancel identically.

That cancellation is the entire trick, and it deserves a sentence of its own: the Legendre transform is engineered so that the variable you are trying to get rid of drops out of the differential. It is the same cancellation as (1.3.5), and the same one as dF=dUTdSSdT\dd F = \dd U - T\dd S - S\dd T in the thermodynamics callout. Not an analogy. One theorem, appearing for the third time.

With those two terms gone, what survives is

dH  =  iq˙idpi    iLqidqi    Ltdt, \dd H \;=\; \sum_{i}\dot q^{i}\,\dd p_{i} \;-\; \sum_{i}\pdv{L}{q^{i}}\,\dd q^{i} \;-\; \pdv{L}{t}\,\dd t, (1.3.21)

which contains only dq\dd q, dp\dd p and dt\dd t. That confirms HH really is a function of (q,p,t)(q,p,t) and not secretly of the velocities. And once we know that, we may write its differential a second way, straight from the definition of a partial derivative:

dH  =  iHqidqi  +  iHpidpi  +  Htdt. \dd H \;=\; \sum_{i}\pdv{H}{q^{i}}\,\dd q^{i} \;+\; \sum_{i}\pdv{H}{p_{i}}\,\dd p_{i} \;+\; \pdv{H}{t}\,\dd t. (1.3.22)

Two expressions for the same differential, in terms of the same independent increments. Their coefficients must therefore match one by one. Comparing dpi\dd p_{i}, then dqi\dd q^{i}, then dt\dd t gives three identities:

Hpi=q˙i,Hqi=Lqi,Ht=Lt. \pdv{H}{p_{i}} = \dot q^{i}, \qquad \pdv{H}{q^{i}} = -\pdv{L}{q^{i}}, \qquad \pdv{H}{t} = -\pdv{L}{t}. (1.3.23)

Only the middle one still mentions LL, and we want equations in HH alone, so kill it with the Euler–Lagrange equation, which says L/qi=ddt(L/q˙i)=p˙i\partial L/\partial q^{i} = \dv{}{t}\big(\partial L/\partial\dot q^{i}\big) = \dot p_{i}. And there they are:

  q˙i  =  Hpi,p˙i  =  Hqi,Ht  =  Lt.   \boxed{\;\dot q^{i} \;=\; \pdv{H}{p_{i}}, \qquad\qquad \dot p_{i} \;=\; -\pdv{H}{q^{i}}, \qquad\qquad \pdv{H}{t} \;=\; -\pdv{L}{t}.\;} (1.3.24)

Hamilton's equations. Count what happened. Our nn second-order equations became 2n2n first-order ones. Nothing was gained or lost in information, because a second-order equation always splits into two first-order ones, and the standard way to do that is to call q˙\dot q a new variable. What is special here is which new variable was chosen. Choosing p=L/q˙p=\partial L/\partial\dot q rather than q˙\dot q itself is what produces the near-symmetry of (1.3.24), and that symmetry is worth everything that follows.

One immediate dividend, free of charge. Suppose HH has no explicit time dependence, and ask how HH changes along a trajectory. The chain rule plus (1.3.24) answers in one line:

dHdt=i(Hqiq˙i+Hpip˙i)+Ht=i(p˙iq˙i+q˙ip˙i)+0  =  0. \dv{H}{t} = \sum_{i}\left(\pdv{H}{q^{i}}\dot q^{i} + \pdv{H}{p_{i}}\dot p_{i}\right) + \pdv{H}{t} = \sum_{i}\Big(-\dot p_{i}\dot q^{i} + \dot q^{i}\dot p_{i}\Big) + 0 \;=\; 0. (1.3.25)

HH is conserved, by cancellation between the two halves of (1.3.24). That is Chapter 1.2's Beltrami identity again, now in three lines instead of five. Notice which feature did the work. The minus sign that made the two equations not symmetric is precisely what made the cancellation happen. Remember that, because §5 and §6 are both that minus sign.

3.2 · Second route: vary the action in phase space

The first derivation used the Euler–Lagrange equation as an input. Here is a derivation that does not, and that treats qq and pp on an equal footing from the start.

Rearrange (1.3.14) as L=ipiq˙iHL = \sum_i p_{i}\dot q^{i}-H and substitute it into the action. Then declare qi(t)q^{i}(t) and pi(t)p_{i}(t) to be 2n2n independent functions, with no relation between them assumed. The object to be made stationary is

S[q,p]  =  t1t2[ipiq˙i    H(q,p,t)]dt. S[q,p] \;=\; \int_{t_{1}}^{t_{2}}\left[\sum_{i}p_{i}\,\dot q^{i} \;-\; H(q,p,t)\right]\dd t. (1.3.26)

This is a functional on a space of paths in the 2n2n-dimensional space of (q,p)(q,p), and we ask for its stationary points exactly as Chapter 1.2 did. So vary both: qq+ϵηq\to q+\epsilon\eta, pp+ϵξp\to p+\epsilon\xi. Note the asymmetry that is about to matter. In (1.3.26) the velocity q˙\dot q appears and p˙\dot p does not. Collecting the terms linear in ϵ\epsilon,

δS=t1t2i[ξiq˙i+piη˙iHqiηiHpiξi]dt. \delta S = \int_{t_{1}}^{t_{2}}\sum_{i}\left[\xi_{i}\dot q^{i} + p_{i}\dot\eta^{i} - \pdv{H}{q^{i}}\eta^{i} - \pdv{H}{p_{i}}\xi_{i}\right]\dd t. (1.3.27)

Only one term carries a derivative of a variation, and Chapter 1.2's move three handles it. We integrate piη˙idt\int p_{i}\dot\eta^{i}\,\dd t by parts to get [piηi]t1t2p˙iηidt\big[p_{i}\eta^{i}\big]_{t_{1}}^{t_{2}} - \int\dot p_{i}\eta^{i}\,\dd t, so that every remaining term carries η\eta or ξ\xi as a plain factor. Collecting those coefficients separately,

δS=[ipiηi]t1t2+t1t2i[(q˙iHpi)ξi    (p˙i+Hqi)ηi]dt. \delta S = \Big[\textstyle\sum_i p_{i}\eta^{i}\Big]_{t_{1}}^{t_{2}} + \int_{t_{1}}^{t_{2}}\sum_{i}\left[\left(\dot q^{i}-\pdv{H}{p_{i}}\right)\xi_{i} \;-\; \left(\dot p_{i}+\pdv{H}{q^{i}}\right)\eta^{i}\right]\dd t. (1.3.28)

Kill the boundary term by requiring ηi(t1)=ηi(t2)=0\eta^{i}(t_{1})=\eta^{i}(t_{2})=0, which fixes the positions at both ends, as before. Note carefully that no condition on ξ\xi is needed. Since pp never appears differentiated, it generates no boundary term and is free at the endpoints.

That is not a technicality, and it is worth saying what it buys. It says the natural boundary-value problem in phase space is "specify where you start and where you finish", not "specify the whole initial state". Chapter 5.6's path integral sums over paths with fixed endpoints, and it inherits exactly this asymmetry.

Now apply the fundamental lemma (Chapter 1.2 §3.4) twice, once with ξ\xi arbitrary and η=0\eta=0, and once the other way round. Each bracket must vanish separately, which gives

q˙iHpi=0,p˙i+Hqi=0, \dot q^{i} - \pdv{H}{p_{i}} = 0, \qquad\qquad \dot p_{i} + \pdv{H}{q^{i}} = 0, (1.3.29)

and that is (1.3.24) again. The two derivations agree, as they must, and each tells you something the other does not. Route one tells you the equations are the Legendre transform of Euler–Lagrange. Route two tells you they are themselves a variational principle, one in which position and momentum are independent variables of equal status.

3.3 · The minus sign, and what it is the seed of

Stare at (1.3.24) until the near-symmetry bothers you. The equations for qq and pp are the same equation with qpq\leftrightarrow p, except for one minus sign. That asymmetry is irreducible. You cannot rescale it away, because rescaling ppp\to-p just moves the sign to the other equation.

Since we cannot remove it, the next best thing is to give it a home of its own where we can see exactly what it does. Start by stacking the coordinates into one list of 2n2n numbers,

z  =  (q1,,qn,p1,,pn), z \;=\; \big(q^{1},\dots,q^{n},\,p_{1},\dots,p_{n}\big), (1.3.30)

so that a single letter zz now names the whole state. The minus sign is a relation between the first half of that list and the second half, so it should live in a 2n×2n2n\times2n matrix. Define

Ω  =  (0InIn0),ΩT=Ω,Ω2=I2n,detΩ=1. \Omega \;=\; \begin{pmatrix} 0 & I_{n}\\ -I_{n} & 0\end{pmatrix}, \qquad \Omega^{\mathsf T} = -\Omega, \qquad \Omega^{2} = -I_{2n}, \qquad \det\Omega = 1. (1.3.31)

with the three properties on the right all read straight off the block form. With that one matrix in hand, both halves of (1.3.24) collapse into a single line:

z˙  =  ΩH,(H)a=Hza. \dot z \;=\; \Omega\,\nabla H, \qquad \big(\nabla H\big)_{a} = \pdv{H}{z^{a}}. (1.3.32)

Check it: the top block of ΩH\Omega\nabla H is +H/p+\partial H/\partial p and the bottom block is H/q-\partial H/\partial q. ✓

Now read that line geometrically, because it explains a fact we have already proved algebraically. The gradient of HH points uphill, at right angles to the level sets. Ω\Omega rotates that direction by a quarter turn. So the system moves along the level sets, which is why HH is conserved. The flow is everywhere perpendicular to the gradient, by construction.

Ω\Omega is called the symplectic form, and the antisymmetry ΩT=Ω\Omega^{\mathsf T}=-\Omega is the minus sign, isolated. Compare it with the object that runs Chapter 3.3, the metric gμνg_{\mu\nu}, which is symmetric and measures lengths. A symplectic form measures oriented areas instead, and it is the reason phase space has a geometry at all.

Every structural fact in the rest of this chapter is a consequence of the antisymmetry of Ω\Omega: Liouville's theorem, the Poisson bracket, and the [q^,p^]=i[\hat q,\hat p]=\ii\hbar of quantum mechanics. So is the fact that phase space is always even-dimensional. An antisymmetric matrix in odd dimensions has detΩ=detΩT=det(Ω)=(1)2n+1detΩ\det\Omega = \det\Omega^{\mathsf T} = \det(-\Omega) = (-1)^{2n+1}\det\Omega, forcing detΩ=0\det\Omega=0, so no invertible Ω\Omega exists there.

In plain terms 1.3.3

Trading one second-order equation for two first-order ones is available to anybody and costs nothing, since calling the velocity a new variable does it. All of the content here is in which new variable was chosen. Taking the conjugate partner rather than the velocity produces two equations that are the same equation with the roles of coordinate and momentum exchanged, apart from a single minus sign sitting in one of them.

That sign cannot be removed. Reversing the sign of the momentum only moves it to the other equation, and it is what makes the two halves of the calculation cancel when you ask whether the energy changes, so that the energy does not change. Written compactly the sign becomes an antisymmetric object standing between the two halves of the space, whose effect is to turn a gradient through a quarter turn, so the motion runs along the level surfaces of the energy rather than across them.

Everything structural in the rest of the chapter is that antisymmetry and nothing else. A symmetric object of the same general kind measures lengths and will run the chapters on gravity; this one measures oriented areas instead, which is a different geometry and the reason the space must have an even number of dimensions, an antisymmetric table in an odd number of them being necessarily degenerate.

4 · Phase space

The space whose points are the 2n2n numbers z=(q,p)z=(q,p) is phase space. Chapter 0.8 introduced it for the oscillator and promised it would come back. This is where it becomes the setting rather than a picture.

4.1 · One point fixes everything, and trajectories cannot cross

Hamilton's equations (1.3.32) are first order. That is the structural difference from the Lagrangian formulation, and it changes what a "state" is. Chapter 0.8's Picard–Lindelöf theorem says that a first-order system z˙=F(z)\dot z = F(z) with FF Lipschitz has exactly one solution through each initial point. Applied here, that gives us the following.

The defining property of phase space

A single point of phase space determines the entire past and future of the system. Not a point plus a velocity. A point.

Configuration space does not have this property. A pendulum passing through θ=0\theta=0 might be moving left or right, fast or slow, so "θ=0\theta=0" is not a state. Phase space is precisely the space of things that are states, which is why Chapter 0.8 said the list the existence theorem asks for is the list phase space is made of.

There is an immediate geometric consequence, and it is what makes phase portraits legible.

Claim. If HH has no explicit time dependence, two distinct trajectories can never cross.

Proof. Suppose the trajectory z1z_{1} passes through the point zz_{\ast} at time t1t_{1} and the trajectory z2z_{2} passes through the same zz_{\ast} at time t2t_{2}. Because HH does not depend on tt, the right-hand side of (1.3.32) depends only on the point, so shifting a solution in time gives another solution. Define z~2(t)=z2(t+t2t1)\tilde z_{2}(t) = z_{2}(t+t_{2}-t_{1}), which also solves the equations and satisfies z~2(t1)=z\tilde z_{2}(t_{1}) = z_{\ast}. Now z1z_{1} and z~2\tilde z_{2} are two solutions with the same value at t1t_{1}, so by uniqueness they are the same solution. Hence z2z_{2} is z1z_{1} retimed, and it is the same curve. \blacksquare

So through every point of phase space there passes exactly one curve, and phase space is filled by these curves like the streamlines of a steady fluid, with no branching, no merging and no intersections. The only points where curves appear to meet are fixed points, where H=0\nabla H=0 and the "trajectory" is a single stationary point. We will meet one shortly.

4.2 · The harmonic oscillator, again

The quickest way to get a feel for a flow is to draw one we already know. Take H=p22m+12mω2q2H = \dfrac{p^{2}}{2m} + \half m\omega^{2}q^{2}, which is T+VT+V, since both hypotheses of §2.3 hold. Differentiating it in pp and in qq gives Hamilton's equations for the oscillator:

q˙=Hp=pm,p˙=Hq=mω2q. \dot q = \pdv{H}{p} = \frac{p}{m}, \qquad\qquad \dot p = -\pdv{H}{q} = -m\omega^{2}q. (1.3.33)

The first says p=mq˙p=m\dot q, recovering the mechanical momentum. Substituting it into the second gives mq¨=mω2qm\ddot q = -m\omega^{2}q, the oscillator equation. Nothing new so far. But now read the pair as a flow instead of as a route back to a second-order equation.

The level sets of HH are the ellipses of Chapter 0.8 §4.4, and by §3.3 the motion runs along them. To see what the motion actually is, we would like the ellipses to be circles, so rescale the two axes by equal and opposite factors:

Q  =  mω  q,P  =  pmω, Q \;=\; \sqrt{m\omega}\;q, \qquad\qquad P \;=\; \frac{p}{\sqrt{m\omega}}, (1.3.34)

which has the virtue of preserving areas, since dQdP=dqdp\dd Q\,\dd P = \dd q\,\dd p. That is a property §8 will name. In these variables H=ω2(P2+Q2)H = \tfrac{\omega}{2}\big(P^{2}+Q^{2}\big), and rewriting (1.3.33) in them gives

Q˙  =  ωP,P˙  =  ωQ. \dot Q \;=\; \omega P, \qquad\qquad \dot P \;=\; -\omega Q. (1.3.35)

That is a rigid rotation of the whole plane at angular velocity ω\omega, clockwise. Every point, whatever its energy, goes round at the same rate, which is the phase-space statement of the fact that a harmonic oscillator's period is independent of amplitude. Chapter 0.1 observed that i\ii is a rotation by 90°. Here the rotation is generated by Ω\Omega, which squares to 1-1 for the same reason. The oscillator flow is multiplication by eiωt\ee^{-\ii\omega t}, drawn.

4.3 · The pendulum, and the separatrix

Now a system where the flow is not uniform. Take the pendulum of Chapter 1.2 §6.2, with q=θq=\theta and p=m2θ˙p=m\ell^{2}\dot\theta. Both hypotheses of §2.3 hold, so its Hamiltonian is the total energy:

H  =  p22m2  +  mg(1cosq), H \;=\; \frac{p^{2}}{2m\ell^{2}} \;+\; mg\ell\big(1-\cos q\big), (1.3.36)

where the constant has been chosen so that H=0H=0 at the bottom. Hamilton's equations are q˙=p/m2\dot q = p/m\ell^{2} and p˙=mgsinq\dot p = -mg\ell\sin q, reproducing q¨=(g/)sinq\ddot q = -(g/\ell)\sin q. Write ω02=g/\omega_{0}^{2}=g/\ell.

Since HH is conserved, trajectories are its level sets, and there are three kinds.

Libration. For 0<H<2mg0 \lt H \lt 2mg\ell the bob cannot reach the top. Setting p=0p=0 in (1.3.36) gives a turning point at cosq=1H/mg\cos q = 1 - H/mg\ell, which has a solution only when H2mgH\le2mg\ell. The level set is a closed curve encircling the origin, and the motion is back-and-forth swinging. Near the origin, expanding 1cosqq2/21-\cos q\approx q^{2}/2 recovers the oscillator of §4.2 and the level curves are Chapter 0.8's ellipses.

Rotation. For H>2mgH \gt 2mg\ell there is no turning point. Here pp never vanishes, so the bob goes over the top and keeps going. The level set is an open curve running across the whole range of qq, and since qq and q+2πq+2\pi are the same physical configuration, phase space here is really a cylinder and the curve is a loop around it.

The separatrix. Between those two behaviours sits one energy at which the bob arrives at the inverted position q=πq=\pi with exactly zero momentum. Putting q=πq=\pi, p=0p=0 into (1.3.36) gives its value:

Hsep  =  2mg, H_{\text{sep}} \;=\; 2mg\ell, (1.3.37)

We would like the curve itself, not only its energy, so solve (1.3.36) for pp at this energy, using 1cosq=2sin2(q/2)1-\cos q = 2\sin^{2}(q/2) to make the square root come out cleanly:

p22m2=mg(1+cosq)=2mgcos2q2    p=±2m2ω0cosq2, \frac{p^{2}}{2m\ell^{2}} = mg\ell\big(1+\cos q\big) = 2mg\ell\cos^{2}\frac{q}{2} \;\Longrightarrow\; p = \pm\,2m\ell^{2}\omega_{0}\cos\frac{q}{2}, (1.3.38)

These are the two arcs Chapter 0.8 §8.2 found. They meet at q=±πq=\pm\pi, p=0p=0, and that meeting point is a fixed point, since H/q=mgsinπ=0\partial H/\partial q = mg\ell\sin\pi = 0 and H/p=0\partial H/\partial p=0. So §4.1's no-crossing theorem is not violated. To see what the flow does near it, linearise by writing q=π+uq=\pi+u, so that sinq=sinuu\sin q = -\sin u \approx -u:

u˙=pm2,p˙=+mguu¨=+ω02u, \dot u = \frac{p}{m\ell^{2}}, \qquad \dot p = +mg\ell\,u \qquad\Longrightarrow\qquad \ddot u = +\omega_{0}^{2}\,u, (1.3.39)

whose solutions are e±ω0t\ee^{\pm\omega_{0}t}, one growing and one decaying. This is a hyperbolic fixed point, or saddle. The level curves near it are hyperbolas, and the separatrix is the pair of straight lines they asymptote to. The decaying solution is the trajectory that creeps toward the inverted position and takes infinite time to arrive. The growing one is the same trajectory run backwards.

Compare the origin, where the same calculation gives u¨=ω02u\ddot u = -\omega_{0}^{2}u and the level curves are closed. That is an elliptic fixed point, or centre. Two fixed points, two signs, two completely different local pictures, and the entire qualitative structure of the pendulum is the statement that they are joined by the separatrix.

4.4 · What the enclosed area means

For a closed orbit, meaning a libration, the natural number to attach to it is the area it encloses:

A(E)  =  H=Epdq, \mathcal{A}(E) \;=\; \oint_{H=E} p\,\dd q, (1.3.40)

where the line integral around the closed curve equals the enclosed area by Green's theorem (Chapter 0.7 §5.1). That is what a line integral of pdqp\,\dd q around a loop is. Two facts make this the right quantity to care about.

First: its derivative is the period. To see it, ask how the area responds when the energy changes. Solve H(q,p)=EH(q,p)=E for pp on the upper branch and differentiate (1.3.40) under the integral sign. Since p/E\partial p/\partial E at fixed qq is the reciprocal of H/p\partial H/\partial p at fixed qq, and H/p=q˙\partial H/\partial p = \dot q, the integrand turns into dt\dd t:

dAdE  =  pEdq  =  dqH/p  =  dqq˙  =  dt  =  T(E). \dv{\mathcal{A}}{E} \;=\; \oint\pdv{p}{E}\,\dd q \;=\; \oint\frac{\dd q}{\partial H/\partial p} \;=\; \oint\frac{\dd q}{\dot q} \;=\; \oint\dd t \;=\; T(E). (1.3.41)

The area is the antiderivative of the period. For the oscillator, Chapter 0.8 computed A=2πE/ω\mathcal{A}=2\pi E/\omega, so dA/dE=2π/ω=T\dd\mathcal{A}/\dd E = 2\pi/\omega = T ✓, and since TT does not depend on EE there, the area is exactly linear in the energy. For the pendulum it is not. Numerical integration of (1.3.40) at ω0=1\omega_{0}=1, m2=1m\ell^{2}=1 gives dA/dE=6.449765\dd\mathcal{A}/\dd E = 6.449765 at E=0.2E=0.2 and 11.63334911.633349 at E=1.9E=1.9, against periods 4K(E/2)/ω04K(E/2)/\omega_{0} of 6.4497656.449765 and 11.63334911.633349 from Chapter 0.8's elliptic-integral formula. Agreement to ten digits, which is the sort of thing worth checking once.

Second: it is what quantum mechanics rations. ⚑ Quoted forward to Chapters 4.8 and 4.10. The Bohr–Sommerfeld condition states that in the semiclassical regime the allowed orbits are those with

pdq  =  (n+12)h,n=0,1,2, \oint p\,\dd q \;=\; \left(n+\tfrac12\right)h, \qquad n=0,1,2,\dots (1.3.42)

with h=2πh=2\pi\hbar Planck's constant. For the oscillator this reads 2πEn/ω=(n+12)h2\pi E_{n}/\omega = (n+\tfrac12)h, that is En=(n+12)ωE_{n}=(n+\tfrac12)\hbar\omega, which Chapter 4.8 will derive exactly, with ladder operators and no semiclassical approximation, and get precisely this answer. That is the one case in which the condition is exact rather than semiclassical, and it is the only part of this flag 4.8 discharges. The general statement, and the measurement of how far it is from the truth when the potential is not a parabola, is Chapter 4.10 §6.

So the classical phase-space ellipse is not deleted by quantum mechanics. It is quantised. The continuum of orbits becomes a discrete stack of them, one per unit of hh of enclosed area, and the ground state is the one with half a unit. Phase-space area is measured in units of Planck's constant. That sentence is the reason hh has the units it does, namely Js\mathrm{J\,s}, the units of pdqp\,\dd q, and it is the single most useful thing to remember about phase space.

Grind box — the area–period relation done carefully, and the adiabatic invariant

The derivation, without hand-waving. Take a one-dimensional system with H(q,p)H(q,p) and a closed orbit at energy EE. On the upper half of the orbit write p=p+(q,E)p=p_{+}(q,E). The enclosed area is A(E)=pdq\mathcal{A}(E)=\oint p\,\dd q, traversed so the integral is positive. Differentiating with respect to EE moves the curve. The endpoints, which are the turning points, move too, but p=0p=0 there, so the boundary contribution vanishes and we may differentiate under the integral:

dAdE=(pE) ⁣qdq. \dv{\mathcal{A}}{E} = \oint\left(\pdv{p}{E}\right)_{\!q}\dd q.

Now, differentiating the identity H(q,p(q,E))=EH\big(q,p(q,E)\big) = E with respect to EE at fixed qq gives (H/p)(p/E)=1\big(\partial H/\partial p\big)\big(\partial p/\partial E\big) = 1, so p/E=1/q˙\partial p/\partial E = 1/\dot q. Hence the integrand is dq/q˙=dt\dd q/\dot q = \dd t, and integrating once round the orbit gives one period. \blacksquare

The action variable. Define J=A/2πJ = \mathcal{A}/2\pi. Then dE/dJ=2π/T=ω(E)\dd E/\dd J = 2\pi/T = \omega(E), the angular frequency. Problem 3 constructs a change of variables in which JJ is literally the new momentum and the new coordinate is an angle advancing uniformly at rate ω\omega. These are action–angle variables, and they turn any one-dimensional bound motion into the trivial system ϑ˙=ω\dot\vartheta=\omega, J˙=0\dot J = 0.

⚑ One quoted fact, because it explains why (1.3.42) was ever a sensible guess. If a parameter of the system (a pendulum's length, say) is changed slowly compared with the period, then EE and TT both change but JJ does not, to all orders in the slowness. JJ is an adiabatic invariant. Ehrenfest's argument was that only an adiabatic invariant can sensibly be quantised, since if you could change JJ by gently stretching a wire you could move a system off its allowed levels without a transition, and that is how Bohr and Sommerfeld chose which quantity to set equal to nhnh. The modern statement is Chapter 4.17's adiabatic theorem: a quantum system stays in the nn-th eigenstate under slow change, which is the same sentence with JJ replaced by nn.

Numbers. For the pendulum at m2=1m\ell^{2}=1, ω0=1\omega_{0}=1, quadrature of (1.3.40) against the elliptic-integral period gives

EEdA/dE\dd\mathcal{A}/\dd E (numerical)T=4K(E/2)T=4K(E/2)relative difference
0.20.26.4497656.4497656.4497656.4497652×10112\times10^{-11}
0.80.87.1100777.1100777.1100777.1100771×10121\times10^{-12}
1.41.48.3014538.3014538.3014538.3014534×10114\times10^{-11}
1.91.911.63334911.63334911.63334911.6333493×10103\times10^{-10}

The period diverges as E2E\to2, since the separatrix takes infinite time, and (1.3.41) says the area therefore has infinite slope there while remaining finite. Both statements are visible in the interactive below. Near the separatrix, neighbouring orbits have wildly different periods, so a blob of initial conditions is sheared apart. That shearing is the whole show.

In plain terms 1.3.4

A single point of this space fixes the entire past and future of the system, which is the whole reason for building it. A configuration by itself is not a state, since a pendulum passing through the bottom of its swing might be going either way at any speed. A configuration together with its conjugate partner is a state — this is the space of states the existence theorem always asked for.

Two things follow. Exactly one solution passes through each point, so trajectories never cross and the space fills with curves like the streamlines of a steady fluid, the only apparent meetings being where nothing moves at all. And the energy is unchanging along each curve, so a pendulum's whole qualitative behaviour reduces to three kinds of curve: closed loops for swinging, open ones for going over the top, and the single curve dividing them, along which the bob takes forever to arrive upside down.

The number worth attaching to a closed loop is the area it encloses, whose rate of change with energy is the period. A promise made when the oscillator was first drawn is collected here: quantum mechanics does not delete these loops, it rations them, permitting only those enclosing a whole number of units of Planck's constant, with half a unit for the lowest. That is why the constant has the units it does.

5 · Liouville's theorem

This is the centrepiece, and it costs almost nothing, because Chapter 0.7 already did the work.

Forget for a moment that (1.3.24) describes mechanics and read it as what it literally is: a formula assigning a velocity vector to every point of a 2n2n-dimensional space. That is a vector field in the sense of Chapter 0.7 §1, and writing out its components gives

v(z)  =  (q˙1,,q˙n,p˙1,,p˙n)  =  (Hp1,,Hpn,Hq1,,Hqn). \vv v(z) \;=\; \big(\dot q^{1},\dots,\dot q^{n},\,\dot p_{1},\dots,\dot p_{n}\big) \;=\; \left(\pdv{H}{p_{1}},\dots,\pdv{H}{p_{n}},\,-\pdv{H}{q^{1}},\dots,-\pdv{H}{q^{n}}\right). (1.3.43)

Chapter 0.7 spent a section establishing what the divergence of such a field means, and the answer was not "something to do with flux" but something much more useful. The divergence is the trace of the Jacobian, hence the fractional rate at which a blob carried by the flow changes its volume:

1VdVdt  =  trJ  =  v. \frac{1}{V}\,\dv{V}{t} \;=\; \mathrm{tr}\,J \;=\; \nabla\cdot\vv v. (1.3.44)

So the question "does a Hamiltonian flow squeeze or spread the blobs it carries" is now a question about one derivative. Compute the divergence of (1.3.43), which is the sum of (component)/(corresponding coordinate)\partial(\text{component})/\partial(\text{corresponding coordinate}) over all 2n2n coordinates:

v  =  i[qi ⁣(Hpi)+pi ⁣(Hqi)]=  i[2Hqipi2Hpiqi]  =  0 \begin{aligned} \nabla\cdot\vv v \;&=\; \sum_{i}\left[\pdv{}{q^{i}}\!\left(\pdv{H}{p_{i}}\right) + \pdv{}{p_{i}}\!\left(-\pdv{H}{q^{i}}\right)\right]\\[3pt] &=\; \sum_{i}\left[\frac{\partial^{2}H}{\partial q^{i}\,\partial p_{i}} - \frac{\partial^{2}H}{\partial p_{i}\,\partial q^{i}}\right] \;=\; 0 \end{aligned} (1.3.45)

by the equality of mixed partial derivatives, which is Clairaut's theorem, Chapter 0.6 §6.1. Every term cancels against its partner, and the cancellation happens because of the minus sign of §3.3 and nothing else. Now substitute that zero into (1.3.44) and read off what it says about volume:

  dVdt  =  0.   \boxed{\;\dv{V}{t} \;=\; 0.\;} (1.3.46)

Liouville's theorem: Hamiltonian flow preserves phase-space volume exactly. Not approximately, not on average, and not for special systems. For every Hamiltonian, including time-dependent ones, since nothing in the calculation used H/t=0\partial H/\partial t=0, including chaotic ones, and including the entire universe if you are willing to write down its Hamiltonian.

What has actually been proved

Take any region R0R_{0} of phase space, that is, a set of possible initial conditions. Let every point of it evolve for a time tt, and call the resulting set RtR_{t}. Then vol(Rt)=vol(R0)\mathrm{vol}(R_{t}) = \mathrm{vol}(R_{0}).

The shape of RtR_{t} is entirely unconstrained. It may be stretched, sheared, folded, wound into a spiral of a million turns, or drawn out into a filament thinner than any resolution you can afford. The theorem says nothing whatever about shape. It says the measure is exactly invariant. Arbitrary distortion together with exactly conserved measure is the combination that makes the theorem worth having.

Three payoffs, in increasing order of importance.

Statistical mechanics has a measure. To do statistical mechanics you must say what "equally likely" means for a continuous system, and that requires a measure on phase space that the dynamics does not distort. Liouville supplies exactly one, namely dnqdnp\dd^{n}q\,\dd^{n}p. The microcanonical ensemble, uniform on a surface of constant energy, is well defined because the flow is incompressible. Had the flow contracted volumes, an initially uniform ensemble would pile up and "uniform" would not be a stable notion. Every ensemble average you have ever seen rests on (1.3.46).

There is no classical attractor. A damped oscillator spirals into the origin, so a blob of initial conditions shrinks to a point and volume is destroyed. Liouville therefore says damping is never fundamental. It is what a Hamiltonian system looks like when you have thrown away the degrees of freedom that the energy leaked into: the air, the wire, the thermal bath. The full system, including the bath, conserves volume. This is the reason a fundamental theory can never contain a friction term, and why every dissipative equation in physics is an effective description.

Information is not destroyed, only stirred. This is the one to keep. Suppose you know the state to within a small blob of volume V0V_{0}, which is your measurement precision. After evolution the blob has some horrible filamentary shape, but its volume is still V0V_{0}. In principle, running the equations backwards recovers the initial blob exactly, so Hamiltonian evolution is reversible and loses nothing. Two forward pointers, both of which are this sentence in different clothes.

  • Chapter 4.6: quantum time evolution is generated by eiH^t/\ee^{-\ii\hat Ht/\hbar}, which is a unitary operator. It preserves inner products, hence probabilities, hence distinguishability, so two distinct initial states stay distinct forever. That is the quantum statement of exactly this theorem, with "volume in phase space" replaced by "overlap of state vectors" and the symplectic form Ω\Omega replaced by the factor of i\ii.
  • Chapters 3.8 and 7.9: the black-hole information problem is what happens when the two statements appear to disagree. Hawking's calculation says a black hole evaporates into radiation that depends only on its mass, charge and spin, so the volume of states that could have formed it collapses to a point and information is destroyed. Something in that argument must be wrong, because unitarity is not a convenience you can trade away, and the search for what is wrong produced holography. The problem is a problem precisely because (1.3.46) and its quantum successor are load-bearing.
p₀ = 1.55 (H at blob top = 1.428)
t = 0.00
boundary vertices = 400
area now = 0.0615727 initial = 0.0615727
drift = +0.0 ppm — exactly zero is the theorem
Liouville, made visible. The pendulum (1.3.36) in units m2=1m\ell^{2}=1, ω0=1\omega_{0}=1, so the separatrix (orange) sits at H=2H=2 and the grey curves are other level sets of HH. A disc of radius 0.140.14 is released centred on the pp-axis, drawn as 240240 interior points plus a 400400-vertex boundary polygon, and every point is advanced by fourth-order Runge–Kutta with step 0.010.01. What to do. Press play with the slider near the bottom of its range. The blob is deep inside the libration region, where the flow is nearly the rigid rotation of §4.2, and it goes round almost undeformed. Now slide it up toward 1.821.82, where the top of the blob has energy 1.9211.921 against a separatrix at 22. Different parts of the blob are now on orbits with substantially different periods, because (1.3.41) says the period diverges at the separatrix, so the blob is sheared. Within one circuit it is a crescent. Within four it is a filament wrapped twice around the level curves and threaded along the separatrix. The dashed grey circle shows where it started. The number that refuses to move. The readout is the shoelace area of the boundary polygon, which is exact for a polygon, so it is a measurement and not a formula. The boundary is adaptively refined, meaning points are inserted whenever neighbours drift more than 0.030.03 apart and the count is reported, because a fixed polygon would eventually cut corners off a curve that is growing without bound in length. Watch the area sit there while the shape becomes unrecognisable. Measured drift over 100100 time units, which is ten to sixteen pendulum periods depending on the setting, is 1818 parts per million at the gentlest end and never worse than ±300\pm300 ppm anywhere on the slider. That is integrator and polygon-resolution error rather than physics, since the exact statement (1.3.46) has no error at all. What no static figure shows. That the distortion is unbounded and the measure is exactly fixed, simultaneously, in the same object.
⚠ Why this isn't obvious

Liouville's theorem seems to forbid the second law of thermodynamics. Entropy is supposed to increase. Entropy is a measure of the phase-space volume compatible with what you know. But the volume is exactly constant. Something has to give.

Look again at the figure. After a few periods the blob is a filament winding through a large region. Its area is what it was. But ask a different question: how much of the plane is within ε\varepsilon of some point of the blob, for a small but nonzero ε\varepsilon? That number grows enormously, because a long thin filament has a huge neighbourhood. Any description of the system with finite resolution sees the coarse-grained blob rather than the blob, and any thermodynamic description is of that kind, since you cannot record 102310^{23} coordinates. The coarse-grained volume is not conserved. It increases, monotonically in practice, and its logarithm is the entropy.

So the arrow of time is not in the dynamics. The dynamics is exactly reversible and exactly volume-preserving, forwards and backwards, and the figure runs equally happily either way. The arrow is in the coarse-graining. It is in the fact that stirring makes a distribution look uniform at any finite resolution, and that we are obliged to describe systems at finite resolution. Gibbs's own image was a drop of ink stirred into water. The ink occupies the same volume it always did, but you will never unstir it, because unstirring requires knowing where every filament went.

Two consequences worth having. First, "entropy increases" is a statement about our description rather than a term in any equation of motion. There is no +σ+\sigma anywhere in (1.3.24). Second, ⚑ quoted: Poincaré's recurrence theorem says that a Hamiltonian system in a bounded phase-space region must return arbitrarily close to its initial state, and the proof is Liouville. If it never returned, the images of a small blob would be disjoint forever and their total volume would exceed the finite volume available. The recurrence times are absurd, being exponentially large in 102310^{23} for a mole of gas, which is why this does not overturn thermodynamics. But it is not a loophole. It is a theorem, and it exists because of the three lines above.

In plain terms 1.3.5

The centrepiece of the chapter costs three lines and is hard to credit at first hearing. Take any region of this space, regard it as a collection of possible starting conditions, let every point evolve for as long as you like, and the volume occupied afterwards is exactly what it was. Not nearly, not on average, and not for well-behaved systems only.

The shape is unconstrained while the measure is exact, and that combination is what makes the theorem worth having. A blob released near a pendulum's dividing curve is drawn into a filament wound many times round, thinner than any resolution you can afford, and its area does not move. Damping can therefore never be fundamental, since a system spiralling into rest destroys volume; friction is a larger volume-preserving system with some coordinates thrown away. And information is stirred rather than destroyed, because running the equations backwards recovers the original blob exactly.

That last sentence travels furthest. Its quantum successor says evolution preserves overlaps, so two states that begin distinguishable stay distinguishable forever, and an evaporating black hole is a crisis rather than a curiosity because the usual calculation appears to contradict it. The arrow of time, meanwhile, is not in the dynamics, which runs equally happily either way, but in the fact that a filament with an enormous neighbourhood looks like a filled region at finite resolution.

6 · Poisson brackets

We now extract from Hamilton's equations the structure that will survive quantisation. Everything in this section is elementary. Its importance is entirely in what it becomes.

6.1 · The equation of motion for any observable whatsoever

An observable is any function f(q,p,t)f(q,p,t) on phase space: the energy, a component of angular momentum, the distance between two particles, the electric dipole moment, anything you could in principle measure. Ask how it changes along a trajectory. By the multivariable chain rule (Chapter 0.6),

dfdt  =  i(fqiq˙i+fpip˙i)+ft. \dv{f}{t} \;=\; \sum_{i}\left(\pdv{f}{q^{i}}\,\dot q^{i} + \pdv{f}{p_{i}}\,\dot p_{i}\right) + \pdv{f}{t}. (1.3.47)

which is true of any function on any space and says nothing about mechanics yet. The mechanics enters when we say what q˙\dot q and p˙\dot p actually are, so substitute Hamilton's equations (1.3.24) for both of them:

dfdt  =  i(fqiHpifpiHqi)+ft. \dv{f}{t} \;=\; \sum_{i}\left(\pdv{f}{q^{i}}\,\pdv{H}{p_{i}} - \pdv{f}{p_{i}}\,\pdv{H}{q^{i}}\right) + \pdv{f}{t}. (1.3.48)

The bracketed combination has appeared without being asked for, so let's give it a name. For any two functions f,gf,g on phase space define the Poisson bracket

  {f,g}    i(fqigpi    fpigqi)  =  (f)TΩ(g),   \boxed{\;\{f,g\} \;\equiv\; \sum_{i}\left(\pdv{f}{q^{i}}\,\pdv{g}{p_{i}} \;-\; \pdv{f}{p_{i}}\,\pdv{g}{q^{i}}\right) \;=\; \big(\nabla f\big)^{\mathsf T}\,\Omega\,\big(\nabla g\big),\;} (1.3.49)

where the second form uses the symplectic matrix of (1.3.31). Check the block structure and you will find it reproduces the sum exactly. With that abbreviation, (1.3.48) shortens to

  dfdt  =  {f,H}  +  ft.   \boxed{\;\dv{f}{t} \;=\; \{f,H\} \;+\; \pdv{f}{t}.\;} (1.3.50)

This is not a formula about a particular system. It is the equation of motion for every observable of every Hamiltonian system, and it contains Hamilton's equations as the special cases f=qif=q^{i} and f=pif=p_{i} (verify: {qi,H}=H/pi\{q^{i},H\}=\partial H/\partial p_{i} and {pi,H}=H/qi\{p_{i},H\}=-\partial H/\partial q^{i} ✓).

It also contains the master conservation law, immediately. Set f/t=0\partial f/\partial t=0 in (1.3.50) and the left-hand side vanishes exactly when the bracket does:

f is conserved{f,H}=0. f \text{ is conserved} \qquad\Longleftrightarrow\qquad \{f,H\} = 0. (1.3.51)

Conservation has become an algebraic condition. You no longer solve the equations of motion to find out whether something is conserved. You compute one bracket. Taking f=Hf=H gives {H,H}=0\{H,H\}=0 by antisymmetry alone, recovering (1.3.25).

6.2 · The fundamental brackets

The first brackets to compute are the ones between the coordinates themselves, since everything else is built from those. For {qi,pj}\{q^{i},p_{j}\}, note that qi/qk=δki\partial q^{i}/\partial q^{k}=\delta^{i}_{k}, qi/pk=0\partial q^{i}/\partial p_{k}=0, pj/pk=δjk\partial p_{j}/\partial p_{k}=\delta_{jk} and pj/qk=0\partial p_{j}/\partial q^{k}=0, so only the first term of (1.3.49) survives:

{qi,pj}  =  kδkiδjk  =  δji,{qi,qj}=0,{pi,pj}=0. \{q^{i},p_{j}\} \;=\; \sum_{k}\delta^{i}_{k}\,\delta_{jk} \;=\; \delta^{i}_{j}, \qquad\qquad \{q^{i},q^{j}\} = 0, \qquad\qquad \{p_{i},p_{j}\} = 0. (1.3.52)

The last two vanish because in each case one of the two factors in every term is a derivative of a qq with respect to a pp, or the other way round. These three lines are the canonical commutation relations of classical mechanics. Every coordinate commutes with every other coordinate. Every momentum commutes with every other momentum. And a coordinate fails to commute with its own conjugate momentum, by exactly 11.

6.3 · The algebra

The bracket has four properties, and they are the properties rather than incidental features.

Bilinearity. {af+bg,h}=a{f,h}+b{g,h}\{af+bg,h\} = a\{f,h\}+b\{g,h\} for constants a,ba,b, and likewise in the second slot. Immediate from (1.3.49), since differentiation is linear.

Antisymmetry. {f,g}={g,f}\{f,g\} = -\{g,f\}, and hence {f,f}=0\{f,f\}=0. Immediate from ΩT=Ω\Omega^{\mathsf T}=-\Omega. Swapping ff and gg in (f)TΩg(\nabla f)^{\mathsf T}\Omega\,\nabla g transposes a number, which does nothing, and transposing the middle factor flips the sign.

Leibniz. {f,gh}={f,g}h+g{f,h}\{f,gh\} = \{f,g\}h + g\{f,h\}. Immediate from the product rule, since every term of (1.3.49) differentiates gg exactly once. This says {f,}\{f,\,\cdot\,\} acts as a derivation, meaning it behaves like a directional derivative, and §7 will explain that by exhibiting the direction.

The Jacobi identity. The fourth is the only one that is not a line of inspection, and it says that the three ways of nesting a bracket inside a bracket sum to nothing:

{f,{g,h}}  +  {g,{h,f}}  +  {h,{f,g}}  =  0. \big\{f,\{g,h\}\big\} \;+\; \big\{g,\{h,f\}\big\} \;+\; \big\{h,\{f,g\}\big\} \;=\; 0. (1.3.53)

This one is not obvious and is proved in the grind box. It is the compatibility condition that makes everything hang together. Bilinear, antisymmetric, and Jacobi is the definition of a Lie algebra, the structure Chapter 6.1 builds a whole part of the book on.

One concrete consequence is available right now: if ff and gg are both conserved, so is {f,g}\{f,g\}. To see it, put h=Hh=H in (1.3.53) and use {f,H}={g,H}=0\{f,H\}=\{g,H\}=0. The first two terms die, leaving {H,{f,g}}=0\{H,\{f,g\}\}=0. This is Poisson's theorem, and it is a machine for manufacturing conservation laws from ones you already have. It occasionally produces something new and usually produces something you knew, and either way it does not cost you an integration.

Grind box — proving the Jacobi identity, and why no second derivatives can survive

Write everything in the compact notation of (1.3.49). Let zaz^{a}, a=1,,2na=1,\dots,2n, be the phase-space coordinates, a=/za\partial_{a}=\partial/\partial z^{a}, and Ωab\Omega^{ab} the constant antisymmetric matrix of (1.3.31), so that

{f,g}=Ωabafbg \{f,g\} = \Omega^{ab}\,\partial_{a}f\,\partial_{b}g

with repeated indices summed. The two facts we will use are that Ωab\Omega^{ab} is constant, so it passes through derivatives, and antisymmetric.

The counting argument first, because it tells you why the identity is true. Each term of (1.3.53) is a bracket of a bracket, so it contains exactly two derivatives distributed over three functions. Every term therefore contains a second derivative of exactly one of f,g,hf,g,h and first derivatives of the other two. The claim is that the second-derivative terms cancel among themselves, separately for ff, for gg and for hh, and that once they are gone nothing is left, since every term has a second derivative somewhere. So it suffices to check one of the three cancellations, the other two following by relabelling.

The cancellation. Expand the first and third terms by the product rule:

{f,{g,h}}=Ωabafb(Ωcdcgdh)=ΩabΩcd[afbcgdh  +  afcgbdh],{h,{f,g}}=ΩabΩcd[ahbcfdg  +  ahcfbdg]. \begin{aligned} \big\{f,\{g,h\}\big\} &= \Omega^{ab}\,\partial_{a}f\,\partial_{b}\big(\Omega^{cd}\partial_{c}g\,\partial_{d}h\big)\\[2pt] &= \Omega^{ab}\Omega^{cd}\Big[\partial_{a}f\,\partial_{b}\partial_{c}g\,\partial_{d}h \;+\; \partial_{a}f\,\partial_{c}g\,\partial_{b}\partial_{d}h\Big],\\[6pt] \big\{h,\{f,g\}\big\} &= \Omega^{ab}\Omega^{cd}\Big[\partial_{a}h\,\partial_{b}\partial_{c}f\,\partial_{d}g \;+\; \partial_{a}h\,\partial_{c}f\,\partial_{b}\partial_{d}g\Big]. \end{aligned}

Collect the two terms carrying second derivatives of gg: they are

T1=ΩabΩcdafdhbcg,T2=ΩabΩcdahcfbdg. T_{1} = \Omega^{ab}\Omega^{cd}\,\partial_{a}f\,\partial_{d}h\,\partial_{b}\partial_{c}g, \qquad T_{2} = \Omega^{ab}\Omega^{cd}\,\partial_{a}h\,\partial_{c}f\,\partial_{b}\partial_{d}g.

(The middle term of (1.3.53), {g,{h,f}}\{g,\{h,f\}\}, contains second derivatives of hh and of ff only, so it contributes nothing here.) Now relabel the dummy indices in T2T_{2} by the substitution ada\to d, bcb\to c, cac\to a, dbd\to b:

T2=ΩdcΩabdhafcbg=ΩcdΩabafdhbcg=T1, T_{2} = \Omega^{dc}\Omega^{ab}\,\partial_{d}h\,\partial_{a}f\,\partial_{c}\partial_{b}g = -\,\Omega^{cd}\Omega^{ab}\,\partial_{a}f\,\partial_{d}h\,\partial_{b}\partial_{c}g = -\,T_{1},

using Ωdc=Ωcd\Omega^{dc}=-\Omega^{cd} once and the symmetry cbg=bcg\partial_{c}\partial_{b}g=\partial_{b}\partial_{c}g (Clairaut) once. So T1+T2=0T_{1}+T_{2}=0. By the cyclic symmetry of (1.3.53), the second-derivative-of-ff terms and the second-derivative-of-hh terms cancel by the identical relabelling. Nothing remains. \blacksquare

Note what the proof used: antisymmetry of Ω\Omega, constancy of Ω\Omega, and equality of mixed partials. Nothing else, and in particular nothing about mechanics. The identity is a fact about any constant antisymmetric bilinear pairing of gradients.

Verification, since a proof by index gymnastics deserves a check. Taking the deliberately unstructured functions f=x2py+zpxpy+pzsinyf = x^{2}p_{y} + z\,p_{x}p_{y} + p_{z}\sin y, g=ypx2+xzpz+pycoszg = y\,p_{x}^{2} + xz\,p_{z} + p_{y}\cos z, h=z3px+ypypz+xyzh = z^{3}p_{x} + y\,p_{y}p_{z} + xyz on a three-degree-of-freedom phase space, symbolic differentiation gives a Jacobi residual of exactly 00, and the Leibniz residual {f,gh}{f,g}hg{f,h}\{f,gh\}-\{f,g\}h-g\{f,h\} is exactly 00 as well.

6.4 · The payoff, and the reason Part I exists

Here is what the last three pages were for.

Canonical quantisation, in one line

Chapter 4.2 §8 will take the classical structure you now own, replace observables f(q,p)f(q,p) by self-adjoint operators f^\hat f on a Hilbert space, and make the single substitution

{f,g}    1i[f^,g^],[f^,g^]f^g^g^f^. \{f,g\} \;\longmapsto\; \frac{1}{\ii\hbar}\big[\hat f,\hat g\big], \qquad \big[\hat f,\hat g\big] \equiv \hat f\hat g - \hat g\hat f.

Nothing else changes. Every equation of this section survives verbatim.

Watch what that does to the results we have. The fundamental brackets (1.3.52) become

[q^i,p^j]=iδji,[q^i,q^j]=0,[p^i,p^j]=0, \big[\hat q^{i},\hat p_{j}\big] = \ii\hbar\,\delta^{i}_{j}, \qquad \big[\hat q^{i},\hat q^{j}\big] = 0, \qquad \big[\hat p_{i},\hat p_{j}\big] = 0, (1.3.54)

which is the foundational relation of quantum mechanics, from which the uncertainty principle follows in three lines (Chapter 4.9) and for which p^=i/q\hat p = -\ii\hbar\,\partial/\partial q is the standard realisation (Chapter 4.6). Do the same to the equation of motion (1.3.50) and it becomes

df^dt  =  1i[f^,H^]  +  f^t, \dv{\hat f}{t} \;=\; \frac{1}{\ii\hbar}\big[\hat f,\hat H\big] \;+\; \pdv{\hat f}{t}, (1.3.55)

the Heisenberg equation of motion, which is the whole of quantum dynamics in the Heisenberg picture. And the conservation criterion (1.3.51) becomes "f^\hat f is conserved if and only if it commutes with the Hamiltonian", the statement that gets used, without comment, in every quantum mechanics course from the second week onward.

The bracket properties go across too, and they are exactly the properties a commutator has. Bilinearity: immediate. Antisymmetry: [f^,g^]=[g^,f^][\hat f,\hat g]=-[\hat g,\hat f] by inspection. Leibniz: [f^,g^h^]=[f^,g^]h^+g^[f^,h^][\hat f,\hat g\hat h] = [\hat f,\hat g]\hat h + \hat g[\hat f,\hat h], which you can verify in one line by inserting and cancelling g^f^h^\hat g\hat f\hat h. Jacobi: holds for commutators identically, by expanding all twelve products and watching them cancel in pairs. The classical Poisson algebra and the quantum commutator algebra are the same Lie algebra. That is the sense in which the substitution "changes nothing else".

So the answer to "why did we spend a chapter on a change of variables" is that the change of variables is the thing that quantises. You cannot quantise the Lagrangian formulation directly, because there is no bracket there, only a variational principle. That is why Feynman had to invent an entirely different route, the path integral of Chapter 5.6, to quantise from an action. Canonical quantisation, by contrast, is a one-line substitution into a structure you now have.

⚠ Why this isn't obvious

"Replace brackets by commutators" is a correspondence rather than a theorem, and it is not quite consistent. ⚑ Quoted: the Groenewold–van Hove theorem proves that there is no map from classical observables to operators that sends {f,g}\{f,g\} to [f^,g^]/i[\hat f,\hat g]/\ii\hbar for all f,gf,g while also sending qq^q\mapsto\hat q and pp^p\mapsto\hat p. The obstruction shows up as soon as you go beyond quadratic. The bracket {q2,p2}\{q^{2},p^{2}\} and the corresponding commutator disagree once you decide how to order q^\hat q's and p^\hat p's. So the substitution is exact for observables at most quadratic in qq and pp, which covers the harmonic oscillator, free particles, angular momentum and therefore most of Chapter 4, and beyond that it needs a convention, an ordering prescription.

This is not a scandal. It is a statement that classical mechanics does not uniquely determine its quantum parent, which we should have expected, since the arrow of explanation runs the other way. Quantum mechanics is the theory. Classical mechanics is its 0\hbar\to0 shadow, and shadows do not determine the objects that cast them. Keep the correspondence as what it is: an extremely good guess, exact where it matters most, and the reason the structure of this chapter is worth owning.

In plain terms 1.3.6

Out of the equations of motion falls a way of combining two quantities on this space to produce a third, and its importance lies almost entirely in what it later becomes. The immediate use is already substantial. The rate of change of any measurable quantity whatever is its combination with the energy, so asking whether something is conserved stops being a matter of solving the motion and becomes a matter of computing one expression and seeing whether it is zero.

Among the coordinates themselves the answers are as simple as they could be. Two coordinates give nothing, two momenta give nothing, and a coordinate paired with its own conjugate momentum gives exactly one. Those three lines and the four properties the operation possesses are the entire structure, and the fourth property is what makes the whole an algebra of the kind the last parts of this book are built from.

Here is what the chapter was for. Replace this operation by the difference between doing two things in one order and in the other, divided by the imaginary unit and Planck's constant, alter nothing else, and classical mechanics becomes quantum mechanics. The relation between a coordinate and its momentum becomes the founding relation of the subject, and the conservation criterion becomes the familiar remark that a quantity is conserved when it commutes with the energy.

7 · Symmetries generate motion

This is the deepest section in the chapter, and it is short. It answers a question you may not have thought to ask: why should conserved quantities and symmetries have anything to do with one another?

Chapter 1.4 will prove Noether's theorem, which says that every continuous symmetry of the action implies a conserved quantity. That is a one-way street, and it leaves the conserved quantity looking like a by-product, something the symmetry emits. The Hamiltonian picture reveals that the relationship is far tighter than that.

7.1 · Every observable generates a flow

The idea is to notice that Hamilton's equations never used any special property of HH. So take any function G(q,p)G(q,p) on phase space and put it where HH used to be. That is, define an infinitesimal change of every observable ff by

δf  =  ϵ{f,G},ϵ infinitesimal. \delta f \;=\; \epsilon\,\{f,G\}, \qquad \epsilon \text{ infinitesimal}. (1.3.56)

Applied to the coordinates themselves this reads δqi=ϵG/pi\delta q^{i} = \epsilon\,\partial G/\partial p_{i} and δpi=ϵG/qi\delta p_{i} = -\epsilon\,\partial G/\partial q^{i}, which is the same structure as (1.3.24) with ϵ\epsilon playing the role of a small time and GG the role of the Hamiltonian. So GG defines a flow on phase space, and (1.3.56) is how any observable changes along it. We say GG generates the transformation.

Before doing examples, we should check that these flows respect the structure of §3.3, and they do. Compute the fundamental bracket of the shifted variables, to first order in ϵ\epsilon, using bilinearity and then Clairaut:

{q+δq,  p+δp}={q,p}+ϵ{Gp,p}ϵ{q,Gq}+O(ϵ2)=1+ϵ2Gqpϵ2Gpq+O(ϵ2)  =  1+O(ϵ2). \begin{aligned} \big\{q+\delta q,\;p+\delta p\big\} &= \{q,p\} + \epsilon\left\{\pdv{G}{p},\,p\right\} - \epsilon\left\{q,\,\pdv{G}{q}\right\} + O(\epsilon^{2})\\[3pt] &= 1 + \epsilon\,\frac{\partial^{2}G}{\partial q\,\partial p} - \epsilon\,\frac{\partial^{2}G}{\partial p\,\partial q} + O(\epsilon^{2}) \;=\; 1 + O(\epsilon^{2}). \end{aligned} (1.3.57)

So the transformation is automatically canonical, meaning it preserves the fundamental brackets. Mixed partials again, for the third time in this chapter, after (1.3.45) and the Jacobi identity. Whatever GG you pick, the flow it generates is a symmetry of the symplectic structure. Now the three cases that matter.

7.2 · The three canonical examples

The Hamiltonian generates time translation. Take G=HG=H and ϵ=dt\epsilon=\dd t. Then (1.3.56) is δf=dt{f,H}=dtdfdt\delta f = \dd t\,\{f,H\} = \dd t\,\dv{f}{t} for any ff without explicit time dependence, by (1.3.50). So the flow generated by HH is the dynamics itself. Evolving the system forward by dt\dd t and applying the transformation generated by HH are the same operation. The Hamiltonian is not merely the energy. It is the generator of translations in time.

Momentum generates spatial translation. Take G=pxG=p_{x}, the momentum conjugate to xx, and work out its bracket with an arbitrary observable. Every term of (1.3.49) except one vanishes, because pxp_{x} depends on nothing else:

{f,px}  =  fxpxpx  =  fxδf=ϵfx, \{f,p_{x}\} \;=\; \pdv{f}{x}\pdv{p_{x}}{p_{x}} \;=\; \pdv{f}{x} \qquad\Longrightarrow\qquad \delta f = \epsilon\,\pdv{f}{x}, (1.3.58)

And ϵf/x\epsilon\,\partial f/\partial x is precisely the first-order change in ff when you move the system a distance ϵ\epsilon in the xx-direction, which is Chapter 0.1's linear approximation. So pxp_{x} generates translations along xx. In particular δx=ϵ\delta x = \epsilon and δp=0\delta p = 0, so the momentum shifts nothing and the position shifts by ϵ\epsilon, which is exactly what a translation does.

Angular momentum generates rotation. Take G=Lz=xpyypxG = L_{z} = xp_{y}-yp_{x}. This one has four relevant coordinates rather than one, so compute the brackets with all four directly from (1.3.49):

{x,Lz}=Lzpx=y,{y,Lz}=Lzpy=+x,{px,Lz}=Lzx=py,{py,Lz}=Lzy=+px. \begin{aligned} \{x,L_{z}\} &= \pdv{L_{z}}{p_{x}} = -y, & \{y,L_{z}\} &= \pdv{L_{z}}{p_{y}} = +x,\\[3pt] \{p_{x},L_{z}\} &= -\pdv{L_{z}}{x} = -p_{y}, & \{p_{y},L_{z}\} &= -\pdv{L_{z}}{y} = +p_{x}. \end{aligned} (1.3.59)

Feeding those into (1.3.56) gives δx=ϵy\delta x = -\epsilon y, δy=+ϵx\delta y = +\epsilon x, and identically δpx=ϵpy\delta p_{x} = -\epsilon p_{y}, δpy=+ϵpx\delta p_{y} = +\epsilon p_{x}. Written out as a matrix acting on the position, the first pair is

(xy)(1ϵϵ1)(xy), \begin{pmatrix}x\\y\end{pmatrix} \longmapsto \begin{pmatrix}1 & -\epsilon\\ \epsilon & 1\end{pmatrix}\begin{pmatrix}x\\y\end{pmatrix}, (1.3.60)

which is a rotation by the angle ϵ\epsilon about the zz-axis, to first order, and the momenta rotate the same way, as vectors must. So LzL_{z} generates rotations about zz. Nothing was assumed about the system. This is a statement about the function xpyypxxp_{y}-yp_{x} and the bracket, valid for any Hamiltonian whatsoever.

7.3 · The two-way street

Now put the pieces together, and note that everything needed is antisymmetry.

Symmetry and conservation are the same statement, read in two directions

Let GG have no explicit time dependence. Then

dGdt={G,H}G is conserved iff this vanishesandδH=ϵ{H,G}=ϵ{G,H}H is invariant under G’s flow iff this vanishes \underbrace{\dv{G}{t} = \{G,H\}}_{\text{$G$ is conserved iff this vanishes}} \qquad\text{and}\qquad \underbrace{\delta H = \epsilon\{H,G\} = -\epsilon\{G,H\}}_{\text{$H$ is invariant under $G$'s flow iff this vanishes}}

are the same quantity up to sign. Therefore

G is conserved    {G,H}=0    H is invariant under the flow generated by G. G \text{ is conserved} \iff \{G,H\}=0 \iff H \text{ is invariant under the flow generated by } G.

Read it left to right and it is the converse of Noether's theorem. Read it right to left and it is Noether's theorem. And read as a single line it says something stronger than either. The conserved quantity is not merely associated with the symmetry, and it is not a by-product of it. The conserved quantity is the generator of the symmetry. They are one object. Momentum is spatial translation. Angular momentum is rotation. Energy is time translation. The bracket {G,H}\{G,H\} is simultaneously "the rate of change of GG" and "the change in HH under GG", and there is nothing else to say.

This is the single most portable idea in the book, and here is where it goes.

  • Chapter 1.4 proves Noether's theorem for a general Lagrangian, including the field-theory version where the conserved quantity becomes a conserved current. This section is the Hamiltonian shadow of it, and it is the direction Noether does not give you.
  • Chapter 4.2 repeats every word with operators. There, an observable G^\hat G generates the one-parameter family of unitary transformations eiϵG^/\ee^{-\ii\epsilon\hat G/\hbar}, and differentiating that at ϵ=0\epsilon=0 gives δf^=iϵ[G^,f^]\delta\hat f = \tfrac{\ii\epsilon}{\hbar}[\hat G,\hat f], which is (1.3.56) under §6.4's substitution. "Observables generate unitary transformations" is this section, quantised. It is why p^=ix\hat p=-\ii\hbar\partial_{x}, the generator of translations, and why H^\hat H appears in the exponent of the time-evolution operator.
  • Chapters 6.1 and 6.2 take the set of generators seriously as an object in its own right. The brackets of generators with each other close into an algebra. Problem 2 computes {Li,Lj}=ϵijkLk\{L_{i},L_{j}\}=\epsilon_{ijk}L_{k} and finds su(2)\mathfrak{su}(2) arriving classically, and a Lie algebra with its representations is what a "symmetry group" concretely is. Chapter 6.3 then makes the parameter ϵ\epsilon depend on position, discovers that this forces a new field into existence, and that field is the photon. The entire gauge principle is this section with ϵϵ(x)\epsilon\to \epsilon(x).
In plain terms 1.3.7

Any function on this space, not only the energy, can be put into the equations of motion in the energy's place, and doing so makes it push everything around in a definite way. The energy pushes the system forward in time, which is what the equations of motion said all along, so it is not merely the energy: it generates a translation of the clock.

Do the same with the momentum along a direction and everything slides a little that way. Do it with the angular momentum about an axis and everything turns about that axis, positions and momenta alike, as vectors should. Nothing was assumed about the system anywhere in this, and it collects a promise made when exponentials of self-partnered maps first appeared: a generator and the family it builds are two aspects of one object.

Read the combination of a quantity with the energy in both directions. Left to right it is the rate at which that quantity changes, so its vanishing means the quantity is conserved. Right to left it is the change in the energy under the flow that quantity generates, so its vanishing means the energy is untouched by it. They are one expression up to a sign — so a conserved quantity is not accompanied by a symmetry and not emitted by one, but is the symmetry, seen as something that moves things.

8 · Canonical transformations and Hamilton–Jacobi

This section is a signpost rather than a development. Two ideas, stated precisely enough to be usable, with the machinery left for the chapters that need it.

8.1 · Which changes of variable are allowed

Chapter 1.2 §7 proved that the Euler–Lagrange equations keep their form under any change of coordinates. Hamilton's equations are more demanding, because they involve the momenta, and the momenta are not free to be relabelled independently. They were defined by (1.3.11).

So we need a test for which relabellings are legitimate, and §6.2 has already supplied the thing to test. Call a transformation (q,p)(Q,P)(q,p)\to(Q,P) canonical if it preserves the fundamental brackets,

{Qi,Pj}=δji,{Qi,Qj}=0,{Pi,Pj}=0, \{Q^{i},P_{j}\} = \delta^{i}_{j}, \qquad \{Q^{i},Q^{j}\} = 0, \qquad \{P_{i},P_{j}\} = 0, (1.3.61)

with the brackets still computed with respect to the old variables. There is a second way to say the same thing, in the notation of §3.3. If Mab=za/zbM^{a}{}_{b} = \partial z'^{a}/\partial z^{b} is the Jacobian matrix of the transformation, then (1.3.61) is precisely the statement

MΩMT  =  Ω, M\,\Omega\,M^{\mathsf T} \;=\; \Omega, (1.3.62)

which is the definition of a symplectic matrix. Such transformations preserve Hamilton's equations, the new equations being Q˙=K/P\dot Q = \partial K/\partial P, P˙=K/Q\dot P = -\partial K/\partial Q for a suitable KK. They preserve every Poisson bracket whatsoever, since (1.3.49) is built from Ω\Omega. And taking determinants of (1.3.62) gives (detM)2=1(\det M)^{2}=1. ⚑ The sharper statement that detM=+1\det M = +1 always, never 1-1, is quoted here. It follows from a Pfaffian argument.

So canonical transformations preserve phase-space volume, and Liouville's theorem is the special case in which the transformation is "evolve for a time tt". The flow map is canonical, which Problem 3 asks you to check.

Generating functions, in one paragraph. Knowing which transformations are allowed is not the same as knowing how to build one, and there is a practical recipe. Notice that the phase-space action (1.3.26) may differ between old and new variables by a total time derivative without changing the equations of motion (Chapter 1.2, Problem 4). So demand

ipiq˙iH  =  iPiQ˙iK  +  dFdt, \sum_i p_{i}\dot q^{i} - H \;=\; \sum_i P_{i}\dot Q^{i} - K \;+\; \dv{F}{t}, (1.3.63)

and let F=F1(q,Q,t)F=F_{1}(q,Q,t). Expanding dF1/dt\dd F_{1}/\dd t by the chain rule and matching the coefficients of the independent quantities q˙i\dot q^{i}, Q˙i\dot Q^{i} and 11 gives

pi=F1qi,Pi=F1Qi,K=H+F1t. p_{i} = \pdv{F_{1}}{q^{i}}, \qquad P_{i} = -\pdv{F_{1}}{Q^{i}}, \qquad K = H + \pdv{F_{1}}{t}. (1.3.64)

Any F1F_{1} you write down generates a canonical transformation. Three cousins F2(q,P,t)F_{2}(q,P,t), F3(p,Q,t)F_{3}(p,Q,t) and F4(p,P,t)F_{4}(p,P,t) follow by Legendre-transforming F1F_{1} in one or both arguments, which is §1 again in yet another costume. The grind box works the most useful one.

8.2 · Hamilton–Jacobi

Push the idea to its limit. Suppose you could find a canonical transformation making the new Hamiltonian KK identically zero. Then Q˙=K/P=0\dot Q = \partial K/\partial P = 0 and P˙=K/Q=0\dot P = -\partial K/\partial Q = 0, so all 2n2n new variables are constants and the system is completely solved.

The price is one equation for the generating function that does it. From (1.3.64), with the generating function called SS instead of F1F_{1}, K=0K=0 means H+S/t=0H + \partial S/\partial t = 0 with pi=S/qip_{i}=\partial S/\partial q^{i} substituted into HH:

  St  +  H ⁣(q1,,qn,  Sq1,,Sqn,  t)  =  0.   \boxed{\;\pdv{S}{t} \;+\; H\!\left(q^{1},\dots,q^{n},\;\pdv{S}{q^{1}},\dots,\pdv{S}{q^{n}},\;t\right) \;=\; 0.\;} (1.3.65)

This is the Hamilton–Jacobi equation: one first-order partial differential equation for one function S(q,t)S(q,t), entirely equivalent to the 2n2n ordinary differential equations we started with. It is rarely the easiest way to solve a mechanics problem. It is on this page for two reasons.

First, SS is the action. Not a function that resembles the action. The action. To see it, compute its total time derivative along a trajectory, using S/qi=pi\partial S/\partial q^{i}=p_{i} and (1.3.65):

dSdt  =  St+iSqiq˙i  =  H+ipiq˙i  =  L, \dv{S}{t} \;=\; \pdv{S}{t} + \sum_{i}\pdv{S}{q^{i}}\,\dot q^{i} \;=\; -H + \sum_{i}p_{i}\dot q^{i} \;=\; L, (1.3.66)

so S=LdtS=\int L\,\dd t along the path. The generating function that trivialises the dynamics is Chapter 1.2's action regarded as a function of the endpoint. That is why the letter is the same, and it closes a loop opened two chapters ago.

Second, it is the classical limit of the Schrödinger equation. ⚑ Quoted, with the derivation deferred to Chapter 4.10. The way to see it is to feed the wave equation a wave whose phase is S/S/\hbar, so substitute the ansatz

ψ(x,t)  =  A(x,t)eiS(x,t)/ \psi(\vv x,t) \;=\; A(\vv x,t)\,\ee^{\,\ii S(\vv x,t)/\hbar} (1.3.67)

into itψ=22m2ψ+Vψ\ii\hbar\,\partial_{t}\psi = -\tfrac{\hbar^{2}}{2m}\nabla^{2}\psi + V\psi and separate the real and imaginary parts. The real part, at leading order in \hbar, is

St+S22m+V  =  0, \pdv{S}{t} + \frac{\abs{\nabla S}^{2}}{2m} + V \;=\; 0, (1.3.68)

which is exactly (1.3.65) for H=p2/2m+VH=\vv p^{2}/2m+V with p=S\vv p=\nabla S, and the terms that were dropped are down by one power of \hbar. (The imaginary part gives tA2+(A2S/m)=0\partial_{t}\abs{A}^{2} + \nabla\cdot\big(\abs{A}^{2}\nabla S/m\big)=0, which is Chapter 0.7's continuity equation for probability, with velocity S/m\nabla S/m.)

So the phase of a quantum wavefunction is the classical action divided by \hbar, which is also what Chapter 1.2 §5 quoted about the path integral weight eiS/\ee^{\ii S/\hbar}, and what Chapter 5.6 will prove. Classical mechanics is the geometry of surfaces of constant phase, in exactly the way that ray optics is the geometry of surfaces of constant phase of a light wave. Hamilton knew this in 1834, which is ninety-two years before anyone wrote down a wave equation for matter.

Grind box — the F2F_{2} generating function, and the oscillator solved by Hamilton–Jacobi

The transform between generating functions. F1(q,Q)F_{1}(q,Q) is awkward because the identity transformation is not expressible in it. Legendre-transform the QQ argument: set F2(q,P,t)=F1(q,Q,t)+iQiPiF_{2}(q,P,t) = F_{1}(q,Q,t) + \sum_i Q^{i}P_{i}, treating Pi=F1/QiP_{i}=-\partial F_{1}/\partial Q^{i} as the conjugate slope. Then dF2=ipidqi+iQidPi+(F1/t)dt\dd F_{2} = \sum_i p_{i}\dd q^{i} + \sum_i Q^{i}\dd P_{i} + (\partial F_{1}/\partial t)\dd t, giving

pi=F2qi,Qi=F2Pi,K=H+F2t. p_{i} = \pdv{F_{2}}{q^{i}}, \qquad Q^{i} = \pdv{F_{2}}{P_{i}}, \qquad K = H + \pdv{F_{2}}{t}.

Now F2=iqiPiF_{2}=\sum_i q^{i}P_{i} gives pi=Pip_{i}=P_{i} and Qi=qiQ^{i}=q^{i}, the identity, as required. This is the version used in practice, and the one Hamilton–Jacobi uses.

The oscillator, solved from scratch. Take H=p22m+12mω2q2H = \dfrac{p^{2}}{2m}+\half m\omega^{2}q^{2}. Since HH has no explicit tt, look for a separated solution S(q,t)=W(q)EtS(q,t) = W(q) - Et, so that S/t=E\partial S/\partial t = -E and (1.3.65) becomes

12m(dWdq)2+12mω2q2=E    dWdq=2mEm2ω2q2. \frac{1}{2m}\left(\dv{W}{q}\right)^{2} + \half m\omega^{2}q^{2} = E \;\Longrightarrow\; \dv{W}{q} = \sqrt{2mE - m^{2}\omega^{2}q^{2}}.

So W(q)=2mEm2ω2q2  dqW(q) = \int\sqrt{2mE-m^{2}\omega^{2}q^{2}}\;\dd q, and S=W(q)EtS = W(q)-Et. Take the constant EE to be the new momentum PP. Then the new coordinate is

Q=SE=mdq2mEm2ω2q2t=1ωarcsin ⁣(qωm2E)t, Q = \pdv{S}{E} = \int\frac{m\,\dd q}{\sqrt{2mE-m^{2}\omega^{2}q^{2}}} - t = \frac{1}{\omega}\arcsin\!\left(q\,\omega\sqrt{\frac{m}{2E}}\right) - t,

and QQ is constant because K=H+S/t=EE=0K=H+\partial S/\partial t = E - E = 0. Setting Q=t0Q=-t_{0} and solving for qq:

q(t)=2Emω2  sin(ω(tt0)). q(t) = \sqrt{\frac{2E}{m\omega^{2}}}\;\sin\big(\omega(t-t_{0})\big).

The oscillator solution, with the amplitude expressed through the energy exactly as Chapter 0.8 had it. Note what the method did. It turned solving a differential equation into evaluating an integral, and the integral was the one whose derivative is the period, (1.3.41) again, since dW=pdq=A\oint\dd W = \oint p\,\dd q = \mathcal{A}.

And the action variable. Instead of EE, take the new momentum to be J=A/2π=E/ωJ = \mathcal{A}/2\pi = E/\omega. Then K=ωJK = \omega J and Hamilton's equations give ϑ˙=K/J=ω\dot\vartheta = \partial K/\partial J = \omega and J˙=0\dot J = 0: an angle advancing uniformly and a constant. Every one-dimensional bound system can be brought to this form, which is the content of "action–angle variables" and the starting point for perturbation theory in celestial mechanics. It is also, ⚑ quoted forward, the starting point for the KAM theorem, which says what survives when you switch a small perturbation on and is the reason the solar system's stability is a hard question rather than an obvious one.

In plain terms 1.3.8

Not every relabelling is permitted here, and the restriction is real. The equations survived any smooth change of coordinates because the momenta were free to follow; once the momenta are independent variables, only changes preserving the pairing between a coordinate and its partner leave the equations alone. Those changes preserve areas and volumes, and evolving the system forward is itself one of them.

Push the freedom to its limit and ask for a relabelling making the new energy identically zero. Every new variable is then constant and the system solved, the price being one partial differential equation for the function generating the change. That function does not merely resemble the number attached to each history. It is that number, regarded as depending on where the path ends, which closes a loop opened two chapters ago.

One remark makes this part look different in retrospect. Feed a wave whose phase is that number divided by Planck's constant into the equation of quantum mechanics, and the leading term is that same partial differential equation, the discarded terms smaller by one factor of the constant. Nearly everything is an approximation, and this is one: an expansion in a named small quantity, with a known first discarded term. Classical mechanics is the geometry of surfaces of constant phase, as ray optics is, and Hamilton had that structure ninety-two years before anyone wrote a wave equation for matter.

9 · Worked examples

Worked example 1 — the charged particle, and minimal coupling three parts early

Take the Lagrangian (1.3.12) for a particle of mass mm and charge ee in given potentials ϕ(x,t)\phi(\vv x,t) and A(x,t)\vv A(\vv x,t):

L  =  12mvv    eϕ  +  evA. L \;=\; \half m\,\vv v\cdot\vv v \;-\; e\phi \;+\; e\,\vv v\cdot\vv A.

(a) It is the right Lagrangian. Before transforming it, check it. The canonical momentum is (1.3.13), pi=mvi+eAip_{i}=mv_{i}+eA_{i}, so the Euler–Lagrange equation p˙i=L/xi\dot p_{i}=\partial L/\partial x^{i} reads

ddt(mvi+eAi)  =  eϕxi+evjAjxi. \dv{}{t}\big(mv_{i}+eA_{i}\big) \;=\; -e\,\pdv{\phi}{x^{i}} + e\,v^{j}\pdv{A_{j}}{x^{i}}.

We want an equation for v˙\dot v alone, and the left-hand side is carrying an extra term. The total time derivative of AiA_{i} along the trajectory is dAi/dt=tAi+vjjAi\dd A_{i}/\dd t = \partial_{t}A_{i} + v^{j}\partial_{j}A_{i} (chain rule, Chapter 0.6). Move it to the right:

mv˙i=e(ϕxiAit)=  Ei+evj(AjxiAixj). m\dot v_{i} = e\underbrace{\left(-\pdv{\phi}{x^{i}} - \pdv{A_{i}}{t}\right)}_{=\;E_{i}} + e\,v^{j}\left(\pdv{A_{j}}{x^{i}} - \pdv{A_{i}}{x^{j}}\right).

The first bracket is the electric field. The second is the antisymmetric part of the Jacobian of A\vv A, which Chapter 0.7 §4.3 identified with the curl. Writing B=×A\vv B=\nabla\times\vv A, the identity vj(iAjjAi)=(v×B)iv^{j}(\partial_{i}A_{j}-\partial_{j}A_{i}) = (\vv v\times\vv B)_{i} holds component by component, so the whole right-hand side collapses to

mv˙  =  e(E+v×B), m\dot{\vv v} \;=\; e\big(\vv E + \vv v\times\vv B\big),

the Lorentz force. ✓ (Verified symbolically for all three components.) The Lagrangian is correct, and note the structure: the vA\vv v\cdot\vv A term is linear in velocity, so it produces a force linear in velocity, which is what a magnetic force is.

(b) The Hamiltonian. By §2.3's grind box, H=L2L0=12mv2+eϕH = L_{2}-L_{0} = \half mv^{2} + e\phi. But HH must be expressed in the momenta, and from (1.3.13) we have mv=peAm\vv v = \vv p - e\vv A. Substituting that in,

  H  =  (peA)22m  +  eϕ.   \boxed{\;H \;=\; \frac{\big(\vv p - e\vv A\big)^{2}}{2m} \;+\; e\phi.\;}

Look at what happened. The vector potential vanished from HH when written in velocities and reappeared inside a square when written in momenta. The magnetic field does no work, which is the first statement, and yet it is not absent from the dynamics, which is the second. Both are true, and the Hamiltonian formalism is what makes them compatible.

(c) It reproduces the same physics. Hamilton's equations give x˙i=H/pi=(pieAi)/m\dot x^{i} = \partial H/\partial p_{i} = (p^{i}-eA^{i})/m, which is mv=peAm\vv v = \vv p - e\vv A again, and p˙i=H/xi\dot p_{i} = -\partial H/\partial x^{i}, which after substituting the first equation reduces to the Lorentz force. (Verified symbolically. The residual is exactly zero once x˙=(peA)/m\dot{\vv x}=(\vv p-e\vv A)/m is used.)

(d) Why this box exists. The boxed Hamiltonian is the object Chapter 6.3 quantises. The rule there is called minimal coupling, and it will be stated like this: to couple a charged particle to electromagnetism, replace

p    peA,equivalentlyi    ieA, \vv p \;\longrightarrow\; \vv p - e\vv A, \qquad\text{equivalently}\qquad -\ii\hbar\nabla \;\longrightarrow\; -\ii\hbar\nabla - e\vv A,

everywhere in the free Hamiltonian, and nothing else. Chapter 6.3 derives that rule from the requirement that the theory be invariant under a phase rotation ψeieχ/ψ\psi\to\ee^{\ii e\chi/\hbar}\psi whose parameter varies from point to point, and then repeats the derivation with the phase replaced by an SU(3)\mathrm{SU}(3) matrix to obtain the strong force. All of it is the substitution above. You have now derived that substitution, from a Legendre transform, in classical mechanics, with no quantum mechanics and no gauge theory in sight. When Chapter 6.3 tells you that the covariant derivative Dμ=μieAμD_{\mu}=\partial_{\mu}-\ii eA_{\mu} is forced, you will recognise it as this worked example wearing indices.

Worked example 2 — the pendulum phase portrait, with every number computed

Take (1.3.36) with m==1m=\ell=1 and g=ω02g=\omega_{0}^{2}, so H=12p2+ω02(1cosq)H = \half p^{2} + \omega_{0}^{2}(1-\cos q). We work out the complete structure of the phase portrait, which is the thing the interactive draws underneath the blob.

(a) Fixed points. q˙=p=0\dot q = p = 0 and p˙=ω02sinq=0\dot p = -\omega_{0}^{2}\sin q = 0 require p=0p=0 and q=0q = 0 or q=πq=\pi (modulo 2π2\pi). Two fixed points per period, and no others.

(b) Their character. To classify a fixed point we linearise the flow at it. Writing the state as z=(q,p)z=(q,p) and z˙=v(z)\dot z = \vv v(z), the Jacobian of v\vv v is

Dv=(qq˙pq˙qp˙pp˙)=(01ω02cosq0). \mathrm{D}\vv v = \begin{pmatrix} \partial_{q}\dot q & \partial_{p}\dot q\\ \partial_{q}\dot p & \partial_{p}\dot p\end{pmatrix} = \begin{pmatrix}0 & 1\\ -\omega_{0}^{2}\cos q & 0\end{pmatrix}.

Its trace is zero everywhere, which is Liouville's theorem (1.3.45) seen locally, since the trace of the Jacobian is the divergence. Its determinant is ω02cosq\omega_{0}^{2}\cos q, so the eigenvalues satisfy λ2=ω02cosq\lambda^{2} = -\omega_{0}^{2}\cos q:

Fixed pointλ2\lambda^{2}EigenvaluesType
q=0q=0 (hanging)ω02-\omega_{0}^{2}±iω0\pm\ii\omega_{0}centre — closed orbits
q=πq=\pi (inverted)+ω02+\omega_{0}^{2}±ω0\pm\omega_{0}saddle — hyperbolic

Because the trace vanishes, the eigenvalues always come in ±\pm pairs. So a Hamiltonian system in one degree of freedom can only have centres and saddles, never spirals or nodes. Spiralling in would contract volume, and Liouville forbids it. Chapter 0.8's classification of damped oscillators had four cases, and here two of them are illegal.

(c) The separatrix energy, exactly. The saddle sits at (q,p)=(π,0)(q,p)=(\pi,0), so its energy is H(π,0)=0+ω02(1cosπ)=2ω02H(\pi,0) = 0 + \omega_{0}^{2}(1-\cos\pi) = 2\omega_{0}^{2}. Restoring units, Hsep=2mgH_{\text{sep}}=2mg\ell, which is (1.3.37), and which is exactly the potential energy of a bob lifted from the bottom to the top, namely mg2mg\cdot2\ell. The separatrix is the level set through the saddle, given by (1.3.38) as p=±2ω0cos(q/2)p = \pm2\omega_{0}\cos(q/2), which is the orange curve in the figure. Its maximum height is p=2ω0p=2\omega_{0} at q=0q=0, so the pendulum needs exactly the angular velocity 2ω02\omega_{0} at the bottom to just reach the top.

(d) Time on the separatrix. On it, q˙=p=2ω0cos(q/2)\dot q = p = 2\omega_{0}\cos(q/2), which separates and integrates:

ω0t=dq2cos(q/2)=ln[tan ⁣(q4+π4)]+const, \omega_{0}\,t = \int\frac{\dd q}{2\cos(q/2)} = \ln\left[\tan\!\left(\frac{q}{4}+\frac{\pi}{4}\right)\right] + \text{const},

which diverges logarithmically as qπq\to\pi. Inverting, q(t)=4arctan(eω0t)πq(t) = 4\arctan(\ee^{\omega_{0}t})-\pi, the famous homoclinic orbit. It leaves the inverted position at t=t=-\infty and returns to it at t=+t=+\infty, taking infinite time at both ends and a finite time in between. It is a single trajectory whose past and future limit points are the same point, which is why it can appear to "cross itself" at the saddle without violating §4.1.

(e) Why this makes the interactive violent. Just inside the separatrix, the period is T(E)=4K(E/2ω02)/ω0T(E)=4K(E/2\omega_{0}^{2})/\omega_{0} in the elliptic-integral notation of Chapter 0.8, and KK diverges logarithmically as its argument approaches 11. At ω0=1\omega_{0}=1: T=6.4498T=6.4498 at E=0.2E=0.2, T=8.3015T=8.3015 at E=1.4E=1.4, T=11.6333T=11.6333 at E=1.9E=1.9, and TT\to\infty at E=2E=2. A blob of radius 0.140.14 centred at p=1.82p=1.82 spans momenta from 1.681.68 to 1.961.96, hence energies from 1.41121.4112 to 1.92081.9208, hence periods from 8.33488.3348 to 12.083912.0839, a spread of 45%. After ten circuits the fast edge has lapped the slow edge several times, and the blob is a filament. Its area, by (1.3.46), has not moved.

10 · Your turn

Problem 1 — Legendre-transform a Lagrangian, and check the involution

A particle of mass mm moves in a plane under a central potential V(r)V(r). In polar coordinates, L=12m(r˙2+r2θ˙2)V(r)L = \half m\big(\dot r^{2}+r^{2}\dot\theta^{2}\big) - V(r) (Chapter 1.2 §7.1). (a) Find both canonical momenta and construct HH. (b) Write down Hamilton's equations and identify the conserved quantities without solving anything. (c) Verify explicitly that Legendre-transforming HH back with respect to the momenta returns LL.

Solution

(a) The momenta are

pr=Lr˙=mr˙,pθ=Lθ˙=mr2θ˙. p_{r} = \pdv{L}{\dot r} = m\dot r, \qquad\qquad p_{\theta} = \pdv{L}{\dot\theta} = mr^{2}\dot\theta.

Note that pθp_{\theta} is the angular momentum, and that it is not m×(a velocity)m\times(\text{a velocity}). It carries an extra factor of r2r^{2}, which is §2.1's warning made concrete. Inverting, r˙=pr/m\dot r = p_{r}/m and θ˙=pθ/mr2\dot\theta = p_{\theta}/mr^{2}. Both hypotheses of §2.3 hold, since TT is a homogeneous quadratic and VV is velocity-independent, so H=T+VH=T+V. Doing it by the definition anyway,

H=prr˙+pθθ˙L=pr2m+pθ2mr2[pr22m+pθ22mr2V(r)]=pr22m+pθ22mr2+V(r). \begin{aligned} H &= p_{r}\dot r + p_{\theta}\dot\theta - L\\[2pt] &= \frac{p_{r}^{2}}{m} + \frac{p_{\theta}^{2}}{mr^{2}} - \left[\frac{p_{r}^{2}}{2m} + \frac{p_{\theta}^{2}}{2mr^{2}} - V(r)\right] = \frac{p_{r}^{2}}{2m} + \frac{p_{\theta}^{2}}{2mr^{2}} + V(r). \end{aligned}

(b) Differentiating that HH in each of its four slots gives Hamilton's equations:

r˙=prm,θ˙=pθmr2,p˙r=Hr=pθ2mr3V(r),p˙θ=Hθ=0. \dot r = \frac{p_{r}}{m}, \quad \dot\theta = \frac{p_{\theta}}{mr^{2}}, \quad \dot p_{r} = -\pdv{H}{r} = \frac{p_{\theta}^{2}}{mr^{3}} - V'(r), \quad \dot p_{\theta} = -\pdv{H}{\theta} = 0.

Two conservation laws can be read straight off. The coordinate θ\theta does not appear in HH, so pθp_{\theta} is constant, which is angular momentum. And HH has no explicit tt, so HH is constant, which is energy. That is two of the four constants of the motion, obtained by inspection rather than integration, which is the practical advantage of the Hamiltonian form.

Substituting the constant pθ=p_{\theta}=\ell into the rr equations leaves a one-degree-of-freedom problem with effective potential Veff(r)=2/2mr2+V(r)V_{\text{eff}}(r) = \ell^{2}/2mr^{2} + V(r). That is the centrifugal barrier, appearing here as an honest term in HH rather than a fictitious force. With V=k/rV=-k/r this is the Kepler problem, and the phase portrait in (r,pr)(r,p_{r}) has exactly the structure of §4.3: a centre at the circular orbit, closed curves around it for bound elliptical orbits, and open curves for hyperbolic ones.

(c) The involution. Transform HH back with respect to (pr,pθ)(p_{r},p_{\theta}). The conjugate slopes are what we differentiate for, so take them first:

Hpr=prm=r˙,Hpθ=pθmr2=θ˙, \pdv{H}{p_{r}} = \frac{p_{r}}{m} = \dot r, \qquad \pdv{H}{p_{\theta}} = \frac{p_{\theta}}{mr^{2}} = \dot\theta,

reproducing (1.3.6), in that the transform of HH has the velocities as its slopes, exactly mirroring p=L/q˙p=\partial L/\partial\dot q. With the slopes in hand the transform is one substitution:

r˙pr+θ˙pθH=mr˙2+mr2θ˙2[12mr˙2+12mr2θ˙2+V]=12m(r˙2+r2θ˙2)V=L.   \dot r\,p_{r} + \dot\theta\,p_{\theta} - H = m\dot r^{2} + mr^{2}\dot\theta^{2} - \left[\half m\dot r^{2} + \half mr^{2}\dot\theta^{2} + V\right] = \half m\big(\dot r^{2}+r^{2}\dot\theta^{2}\big) - V = L. \;\checkmark

Which is (1.3.8), verified on a real example. The formula H=pq˙LH = \sum p\dot q - L and the formula L=pq˙HL = \sum p\dot q - H are the same formula, and the apparent asymmetry between them is entirely in which variables you regard as independent.

Problem 2 — the angular momentum algebra, arriving three parts early

With L=r×p\vv L = \vv r\times\vv p, so that Lx=ypzzpyL_{x}=yp_{z}-zp_{y} and cyclically, show that

{Li,Lj}  =  kϵijkLk, \{L_{i},L_{j}\} \;=\; \sum_{k}\epsilon_{ijk}L_{k},

where ϵijk\epsilon_{ijk} is +1+1 for cyclic index orders, 1-1 for anticyclic, and 00 if any two indices agree. Then show {L2,Lz}=0\{\vv L^{2},L_{z}\}=0, and say what both results become in Chapter 4.11.

Solution

The computation. Take i=xi=x, j=yj=y. The rest follow by cycling xyzxx\to y\to z\to x, which leaves the definitions invariant. The bracket needs gradients, so write down the four of them for Lx=ypzzpyL_{x}=yp_{z}-zp_{y} and Ly=zpxxpzL_{y}=zp_{x}-xp_{z}:

qLx=(0,  pz,  py),pLx=(0,  z,  y),qLy=(pz,  0,  px),pLy=(z,  0,  x). \begin{aligned} \nabla_{q}L_{x} &= (0,\;p_{z},\;-p_{y}), & \nabla_{p}L_{x} &= (0,\;-z,\;y),\\ \nabla_{q}L_{y} &= (-p_{z},\;0,\;p_{x}), & \nabla_{p}L_{y} &= (z,\;0,\;-x). \end{aligned}

Now assemble (1.3.49) term by term over the three coordinates:

{Lx,Ly}=[0z0(pz)]x term+[pz0(z)0]y term+[(py)(x)ypx]z term=xpyypx  =  Lz. \begin{aligned} \{L_{x},L_{y}\} &= \underbrace{\big[0\cdot z - 0\cdot(-p_{z})\big]}_{x\text{ term}} + \underbrace{\big[p_{z}\cdot 0 - (-z)\cdot 0\big]}_{y\text{ term}} + \underbrace{\big[(-p_{y})(-x) - y\,p_{x}\big]}_{z\text{ term}}\\[3pt] &= xp_{y} - yp_{x} \;=\; L_{z}. \end{aligned}

Only the zz column survived, which is worth noticing. The bracket of two components of angular momentum "about" xx and yy is built entirely from the third direction. Cycling gives {Ly,Lz}=Lx\{L_{y},L_{z}\}=L_{x} and {Lz,Lx}=Ly\{L_{z},L_{x}\}=L_{y}, and antisymmetry supplies the rest, so {Li,Lj}=ϵijkLk\{L_{i},L_{j}\}=\epsilon_{ijk}L_{k} in every case, including {Li,Li}=0\{L_{i},L_{i}\}=0. (Verified symbolically.)

The Casimir. For the second part we need no new brackets, only the ones just computed. Use bilinearity and Leibniz on L2=Lx2+Ly2+Lz2\vv L^{2}=L_{x}^{2}+L_{y}^{2}+L_{z}^{2}:

{L2,Lz}=2Lx{Lx,Lz}+2Ly{Ly,Lz}+2Lz{Lz,Lz}=0=2Lx(Ly)+2Ly(Lx)=0.   \begin{aligned} \{\vv L^{2},L_{z}\} &= 2L_{x}\{L_{x},L_{z}\} + 2L_{y}\{L_{y},L_{z}\} + 2L_{z}\underbrace{\{L_{z},L_{z}\}}_{=0}\\[2pt] &= 2L_{x}(-L_{y}) + 2L_{y}(L_{x}) = 0. \;\checkmark \end{aligned}

The magnitude of the angular momentum commutes with every component, while no two components commute with each other.

What this is. The three functions Lx,Ly,LzL_{x},L_{y},L_{z} close under the bracket, meaning the bracket of any two is a linear combination of the three. A set of generators with that property is a Lie algebra, and this particular one, three generators with structure constants ϵijk\epsilon_{ijk}, is su(2)\mathfrak{su}(2), the same algebra as so(3)\mathfrak{so}(3). You have just derived it inside classical mechanics, with no rotation matrices and no quantum mechanics.

By §6.4's substitution it becomes [L^i,L^j]=iϵijkL^k[\hat L_{i},\hat L_{j}] = \ii\hbar\,\epsilon_{ijk}\hat L_{k} and [L^2,L^z]=0[\hat{\vv L}^{2},\hat L_{z}]=0, the two relations from which Chapters 4.11 and 4.12 derive, with no further input, that angular momentum is quantised, that its magnitude takes the values j(j+1)\sqrt{j(j+1)}\,\hbar, that jj can be a half-integer, and hence that spin exists. Note what §7.2 already told you about these generators: they generate rotations. So "angular momentum is quantised" and "the rotation group has discrete representations" are the same fact, which is what Chapter 6.2 is about. Everything in that sentence was visible here, classically, as one bracket.

Problem 3 — canonical transformations have unit Jacobian

(a) For one degree of freedom, show that a transformation (q,p)(Q,P)(q,p)\to(Q,P) satisfies {Q,P}=1\{Q,P\}=1 if and only if its Jacobian determinant is 11, and explain how this makes Liouville's theorem a corollary. (b) Verify it on a concrete case: show that

q=2PmωsinQ,p=2PmωcosQ q = \sqrt{\frac{2P}{m\omega}}\,\sin Q, \qquad p = \sqrt{2Pm\omega}\,\cos Q

is canonical, and find what it does to the oscillator Hamiltonian.

Solution

(a) With one degree of freedom the Jacobian matrix of the transformation is

M=(Q/qQ/pP/qP/p),detM=QqPpQpPq. M = \begin{pmatrix} \partial Q/\partial q & \partial Q/\partial p\\[2pt] \partial P/\partial q & \partial P/\partial p\end{pmatrix}, \qquad \det M = \pdv{Q}{q}\pdv{P}{p} - \pdv{Q}{p}\pdv{P}{q}.

Now write out the bracket the definition (1.3.49) asks for, and compare the two expressions:

{Q,P}=QqPpQpPq. \{Q,P\} = \pdv{Q}{q}\pdv{P}{p} - \pdv{Q}{p}\pdv{P}{q}.

They are literally the same expression. So {Q,P}=1    detM=1\{Q,P\}=1 \iff \det M = 1, with no computation at all. In nn degrees of freedom the statement is (1.3.62), MΩMT=ΩM\Omega M^{\mathsf T}=\Omega. Taking determinants and using detΩ=10\det\Omega=1\neq0 gives (detM)2=1(\det M)^{2}=1, and the sharper detM=+1\det M=+1 is the fact quoted in §8.1.

Liouville as a corollary. Fix a time tt and consider the map Φt\Phi_{t} that sends each initial condition (q0,p0)(q_{0},p_{0}) to its state at time tt. This map is canonical. The brackets {q(t),p(t)}\{q(t),p(t)\} computed with respect to (q0,p0)(q_{0},p_{0}) are preserved by the flow, because ddt{q(t),p(t)}={{q,H},p}+{q,{p,H}}\dv{}{t}\{q(t),p(t)\} = \{\{q,H\},p\} + \{q,\{p,H\}\}, which is zero by the Jacobi identity (1.3.53) applied to qq, pp and HH together with {q,p}=\{q,p\}= constant. Hence detDΦt=1\det \mathrm{D}\Phi_{t} = 1, and by Chapter 0.6's change-of-variables formula the map preserves areas. That is Liouville's theorem again, derived algebraically rather than through the divergence. The two proofs are §5 and this one, and it is worth having both, because §5's generalises to any divergence-free flow while this one explains why the phase-space structure, and not merely the volume, is what is being conserved.

(b) Compute the four partials, writing c=cosQc=\cos Q, s=sinQs=\sin Q and α=2P/mω\alpha = \sqrt{2P/m\omega}, β=2Pmω\beta=\sqrt{2Pm\omega}, so q=αsq=\alpha s, p=βcp=\beta c, and αβ=2P\alpha\beta = 2P:

qQ=αc,qP=sβ,pQ=βs,pP=cmωβ=cα, \pdv{q}{Q} = \alpha c, \quad \pdv{q}{P} = \frac{s}{\beta}, \quad \pdv{p}{Q} = -\beta s, \quad \pdv{p}{P} = \frac{c\,m\omega}{\beta} = \frac{c}{\alpha},

using α/P=α/2P=1/β\partial\alpha/\partial P = \alpha/2P = 1/\beta and β/P=β/2P=mω/β\partial\beta/\partial P = \beta/2P = m\omega/\beta. Assembling them into the determinant,

qQpPqPpQ=αccαsβ(βs)=c2+s2=1.   \pdv{q}{Q}\pdv{p}{P} - \pdv{q}{P}\pdv{p}{Q} = \alpha c\cdot\frac{c}{\alpha} - \frac{s}{\beta}\cdot(-\beta s) = c^{2} + s^{2} = 1. \;\checkmark

(This is the Jacobian of the inverse map. Since det(M1)=1/detM\det(M^{-1})=1/\det M, the forward Jacobian is 11 too.) The payoff. Substitute into the oscillator Hamiltonian:

H=p22m+12mω2q2=2Pmωcos2Q2m+12mω22Pmωsin2Q=ωP(cos2Q+sin2Q)=ωP. H = \frac{p^{2}}{2m} + \half m\omega^{2}q^{2} = \frac{2Pm\omega\cos^{2}Q}{2m} + \half m\omega^{2}\cdot\frac{2P}{m\omega}\sin^{2}Q = \omega P\big(\cos^{2}Q+\sin^{2}Q\big) = \omega P.

The Hamiltonian has become a single term with no QQ in it. Hamilton's equations are therefore Q˙=H/P=ω\dot Q = \partial H/\partial P = \omega and P˙=H/Q=0\dot P = -\partial H/\partial Q = 0, whose solution is Q=ωt+φQ=\omega t+\varphi and P=P= constant =E/ω=E/\omega. Substituting back gives q(t)=2E/mω2sin(ωt+φ)q(t) = \sqrt{2E/m\omega^{2}}\,\sin(\omega t+\varphi). These are the action–angle variables of §4.4's grind box, with PP the action A/2π\mathcal{A}/2\pi and QQ the angle around the ellipse. The whole solution of the oscillator has been reduced to choosing good coordinates, and the reason we were allowed to choose them is detM=1\det M = 1.

Problem 4 — the area of an oscillator orbit, and what quantum mechanics does to it

For H=p2/2m+12mω2q2H = p^{2}/2m + \half m\omega^{2}q^{2}, compute the phase-space area enclosed by the orbit of energy EE directly from the integral pdq\oint p\,\dd q, and show it equals 2πE/ω2\pi E/\omega. Confirm it against (1.3.41). Then apply the Bohr–Sommerfeld condition and say precisely what it predicts.

Solution

The integral. Solving H=EH=E for pp gives p=±2mEm2ω2q2p = \pm\sqrt{2mE-m^{2}\omega^{2}q^{2}}, with turning points at q=±aq=\pm a, a=2E/mω2a=\sqrt{2E/m\omega^{2}}. Going round the loop, the upper branch contributes aa\int_{-a}^{a} and the lower contributes the same again with both signs reversed, so

A=pdq=2aa2mEm2ω2q2  dq=2mωaaa2q2  dq, \mathcal{A} = \oint p\,\dd q = 2\int_{-a}^{a}\sqrt{2mE - m^{2}\omega^{2}q^{2}}\;\dd q = 2m\omega\int_{-a}^{a}\sqrt{a^{2}-q^{2}}\;\dd q,

using 2mE=m2ω2a22mE = m^{2}\omega^{2}a^{2}. The remaining integral is the area of a half-disc of radius aa, namely πa2/2\pi a^{2}/2 (substitute q=asinuq=a\sin u if you want it by hand: π/2π/2a2cos2udu=πa2/2\int_{-\pi/2}^{\pi/2}a^{2}\cos^{2}u\,\dd u = \pi a^{2}/2). Putting that value in,

A=2mωπa22=πmω2Emω2=2πEω. \mathcal{A} = 2m\omega\cdot\frac{\pi a^{2}}{2} = \pi m\omega\,\frac{2E}{m\omega^{2}} = \frac{2\pi E}{\omega}.

Same answer as the ellipse-area route of Chapter 0.8 §4.4, which is π\pi times the product of the semi-axes aa and mωam\omega a, as it must be. Numerically, at m=1.7m=1.7, ω=2.3\omega=2.3, E=3E=3: the semi-axis product route gives 8.195459108.19545910 and 2πE/ω2\pi E/\omega gives 8.195459108.19545910.

The check. dA/dE=2π/ω\dd\mathcal{A}/\dd E = 2\pi/\omega, which is the period T=2π/ωT=2\pi/\omega. ✓ (1.3.41) holds, and the fact that A\mathcal{A} is exactly linear in EE is the statement that the oscillator's period does not depend on its amplitude, the property that makes it the one system everything is expanded around.

Bohr–Sommerfeld. Impose (1.3.42), A=(n+12)h\mathcal{A}=(n+\tfrac12)h, on the area we have just computed:

2πEnω=(n+12)h=(n+12)2π  En=(n+12)ω.   \frac{2\pi E_{n}}{\omega} = \left(n+\half\right)h = \left(n+\half\right)2\pi\hbar \qquad\Longrightarrow\qquad \boxed{\;E_{n} = \left(n+\half\right)\hbar\omega.\;}

Three things to notice, all of which Chapter 4.8 will confirm by an exact operator calculation that uses none of this reasoning.

The levels are equally spaced, by ω\hbar\omega. That follows from (1.3.41), since equal increments of area correspond to equal increments of energy exactly when the period is energy-independent. A pendulum at large amplitude, whose period grows with energy, has levels that converge. And indeed real molecular vibrations, which are anharmonic, have converging levels that eventually run into the dissociation limit. You can read that off the phase portrait without solving anything.

The ground state is not zero. E0=12ωE_{0}=\half\hbar\omega, the zero-point energy, enclosing half a unit of hh. A classical oscillator can sit at the origin of phase space, a single point of zero area. A quantum one cannot, because the smallest allowed area is h/2h/2. Chapter 4.9 says the same thing as ΔqΔp/2\Delta q\,\Delta p\ge\hbar/2. The uncertainty principle is the statement that a state occupies a minimum area of phase space, and that minimum is Planck's constant. Every "why can't the electron just fall into the nucleus" question is answered by this picture.

The classical limit is visible. For large nn the spacing ω\hbar\omega is negligible compared with EnE_{n}, the rungs of the ladder crowd together, and the discrete set of ellipses is indistinguishable from the continuum. Numerically, an oscillator with ω=2π s1\omega=2\pi\ \mathrm{s^{-1}} carrying 1 J1\ \mathrm{J} sits at n1.5×1033n\approx1.5\times10^{33}. This is why nobody noticed.

The brick you just laid

You traded nn second-order equations for 2n2n first-order ones, by a Legendre transform you now understand geometrically. It is the dictionary between describing a convex curve by its points and describing it by its tangent lines, the same operation that turns internal energy into free energy. Out of it came three things. First, the canonical momentum, which is not mvm\vv v in general: in a magnetic field it is mv+eAm\vv v+e\vv A, and it is that object rather than the mechanical momentum that quantum mechanics turns into an operator. Second, the Hamiltonian, which equals T+VT+V only under two hypotheses you can state. Third, Hamilton's equations, derived twice.

You then met phase space, in which one point fixes the whole future, trajectories cannot cross, and the pendulum's entire qualitative behaviour is a centre, a saddle, and the separatrix joining them. You proved Liouville's theorem in three lines from Chapter 0.7's identity of divergence with the trace of the Jacobian. Hamiltonian flow is exactly incompressible, so information is stirred and never destroyed. You saw it in the figure, where a blob became unrecognisable while its area sat there to within a few parts in ten thousand of integrator error. And you saw that the arrow of time is therefore in the coarse-graining rather than in the dynamics.

And you built the Poisson bracket: bilinear, antisymmetric, Leibniz, Jacobi. With it came dfdt={f,H}+ft\dv{f}{t}=\{f,H\}+\pdv{f}{t} for every observable there is, the fundamental relation {qi,pj}=δji\{q^{i},p_{j}\}=\delta^{i}_{j}, and the two-way street of §7, in which a conserved quantity does not merely accompany a symmetry but is the generator of it.

Where this gets spent. The bracket goes to Chapter 4.9, which replaces it by 1i[  ,  ]\tfrac{1}{\ii\hbar}[\;,\;] and changes nothing else, so that {q,p}=1\{q,p\}=1 becomes [q^,p^]=i[\hat q,\hat p]=\ii\hbar and (1.3.50) becomes the Heisenberg equation. It goes also to Chapter 5.3, which does the same substitution for a field and gets particles. Phase space and Liouville go to Chapter 4.6, where unitarity is the same statement in a Hilbert space, and to statistical mechanics, which cannot define an ensemble without them. The generators of §7 go to Chapter 1.4 (Noether, from the other side), Chapter 4.2 (observables generate unitaries), and Chapter 6.1, where the brackets of generators with each other become a Lie algebra and the whole of Part VI. The canonical momentum goes to Chapter 6.3, where ppeA\vv p\to\vv p-e\vv A is derived from a symmetry principle and named minimal coupling. And Hamilton–Jacobi goes to Chapter 4.10 as the classical limit of the Schrödinger equation, the last stop before the wavefunction.

Chapter 1.4 now closes Part I by proving Noether's theorem in the Lagrangian language, where it applies to fields and to relativity, and where it will be the tool that constructs the conserved currents of every theory in this book.