Part I · The Action Principle — Chapter 1.3
Hamilton and Phase Space
Trade second-order equations for first-order ones, and the geometry starts doing the work.
Chapter 1.2 handed you a machine. Feed it one scalar function and it returns second-order differential equations, one per degree of freedom, in a form that survives any change of coordinates you care to make. That machine is not going away. It is the formulation that generalises to special relativity (Chapter 2.5), to fields (Chapter 5.2), and to curved spacetime (Chapter 3.6), and the reason is that the action is a scalar. Scalars do not care what you call your coordinates.
This chapter builds a second machine out of the same materials. It costs one change of variables. We swap the velocities for the momenta , and in exchange we get three things the Lagrangian formulation does not give us.
The first is first-order equations in place of second-order ones. That means the state of the system is a single point, and the dynamics is a flow.
The second is a theorem, Liouville's, saying that flow is incompressible. It is the foundation of statistical mechanics. Read correctly, it is also the classical ancestor of the statement that quantum evolution is unitary.
The third is the one to keep your eye on. It is an operation on functions called the Poisson bracket, and it has the property that
replacing it by and changing nothing else turns classical mechanics into quantum mechanics.
That is what Part I is for. Chapter 4.9 will spend one line doing the substitution and the rest of the chapter reaping it, and the reason that line can be short is that this chapter is long.
Here is the summary you should finish with. The Lagrangian formulation is the one that generalises; the Hamiltonian formulation is the one that quantises. You need both, for different reasons, and mistaking one for the other is a standard way to get stuck.
Tools you'll need — Chapter 0.4: the determinant as a volume factor, the trace, and . Section 5 is that identity. Chapter 0.5: quadratic forms and positive definiteness, which is exactly the condition §1 needs. Chapter 0.6: the Hessian, convexity, the multivariable chain rule, and Clairaut's theorem on mixed partials. Sections 5, 6 and 7 each turn on Clairaut and nothing else. Chapter 0.7: divergence as the trace of the Jacobian, hence as the fractional rate of change of volume carried by a flow. Section 5 is one substitution into that result. Chapter 0.8: Picard–Lindelöf and the uniqueness of solutions, the harmonic oscillator, the phase-space ellipse, and the pendulum's separatrix. Chapter 1.2: Euler–Lagrange, the canonical momentum , and Problem 2's Beltrami identity, which, you are about to discover, was the Hamiltonian all along.
1 · The Legendre transform, properly
Chapter 1.2's Problem 2 asked you to show that when has no explicit time dependence the quantity
is conserved. We are about to promote that combination from a conserved quantity to the central object of a whole formulation of mechanics. It will be a great deal less mysterious if you first see what kind of operation it is.
It is not an ad hoc grouping of terms. It is an instance of a standard construction with a clean geometric meaning, and the same construction is what turns the internal energy of a gas into its free energy. So here is the plan for this section: §1 is mathematics, and there is no mechanics in it at all.
1.1 · Two ways to describe a curve
Take a smooth, strictly convex function . For now just picture a parabola. There are two ways to say which function it is.
By points. List the pairs . This is the description you grew up with.
By tangent lines. For each slope , there is exactly one line of that slope tangent to the graph. Give me the family of all those lines, meaning for each its intercept, and I can reconstruct the curve. The reason is that a convex curve is the upper envelope of its own tangent lines. Every tangent line lies below the graph and touches it once, and the graph is what you get by taking, at each , the highest of them.
Both descriptions contain the same information. The Legendre transform is the dictionary between them. That is the whole idea, and the formula is going to be an immediate consequence of it rather than a definition to memorise.
Let's set it up. A line of slope has the form , where is its -intercept. We carry the minus sign because it makes every formula below come out with a plus.
Now push this line down from above, increasing , until it first touches the graph. Touching means the vertical gap between line and graph has shrunk to zero somewhere. Before that moment the line sits above the graph everywhere. So the critical value of is the largest gap there ever was, which we write as
Read the bracket as: how far the line rises above the graph at the point . Take the largest such gap. That is the amount you must lower to make it tangent, and it is therefore precisely .
is minus the -intercept of the tangent line to whose slope is . The function is the "by tangent lines" description of , packaged so that its argument is the slope.
Everything else in §1 is a consequence of that sentence: the formula , the involution , the thermodynamic potentials, and .
1.2 · Where the supremum sits
Now let's do the calculus. The bracket in (1.3.2) is, for fixed , an ordinary function of the single variable . So we want its stationary point, and Chapter 0.6's rule for finding one applies without change:
So the maximising is the point where the graph's own slope equals , which is exactly what "tangent line of slope " meant. Nothing has been assumed yet except differentiability. Two questions remain, and they are what convexity is for.
Is the stationary point a maximum? The second derivative of the bracket is , which is negative wherever . So yes, provided is strictly convex.
Is there exactly one? The equation must have a unique solution for each in range. Since means is strictly increasing, is injective, and so it is invertible on its image. Again convexity, and again it is doing genuine work rather than decorating the theorem.
That unique maximiser is a function of , so give it a name and write . Substituting it back into the definition turns the supremum into an ordinary formula:
Put (1.3.4) beside (1.3.1) and the shape is already the same: slope times variable, minus the function. We will make that identification precise in §2.
1.3 · The transform is its own inverse
Here is the property that makes the construction worth having, and getting it takes one differentiation. We want to know what does when its own argument changes, so differentiate (1.3.4) with respect to , using the product rule and the chain rule:
The last two terms cancel exactly, because . That cancellation is not a coincidence, and it will happen again in §3 for the same reason: we are differentiating at a stationary point, so the motion of the stationary point itself contributes nothing at first order.
Now set the result beside the equation that defined the maximiser in the first place. Reading (1.3.5) next to (1.3.3) gives
Perfectly symmetric. The transform swaps the roles of the variable and the slope. And since , the transform of a strictly convex function is strictly convex, so we are entitled to transform a second time. Let's do exactly that and see where it lands:
Its supremum sits where , which by (1.3.6) is the with , which is . So evaluate the bracket there and see what number comes out:
The Legendre transform is an involution on strictly convex functions. Do it twice and you are back where you started. That is what "dictionary" meant. No information is lost going from the point description to the tangent-line description, because you can always go back.
⚑ One statement here is quoted rather than proved. For a function that is not convex, the transform still exists, since the supremum in (1.3.2) is well defined whenever it is finite. What fails is the involution. The double transform returns the convex hull of , meaning the largest convex function lying below it, and any non-convex wiggles are erased permanently. This is the Fenchel–Moreau theorem, and it is the precise sense in which "the tangent lines determine the curve" needs convexity. It matters because it tells us which Lagrangians the Hamiltonian formulation applies to, and we will need it for exactly that in §2.3.
1.4 · A worked scalar example
Abstract dictionaries are easier to trust once you have run one. Take on , which is strictly convex there ( for ). Then , so , and putting that into (1.3.4) gives
Now check the involution on this example, without appealing to the theorem. First the inverse relation: ✓, matching (1.3.6). Then transform a second time and see whether comes back:
It does. And the supremum is a real number you can measure rather than a formal device. Numerically, at the supremum of over a fine grid is , and . The formula is measuring something real.
You have used this transform for years under a different name, and knowing that it is the same operation makes the zoo of thermodynamic potentials collapse into one idea.
Start from the internal energy of a simple system, whose differential is the first law,
Look at what that says about temperature. Temperature is the slope of with respect to entropy. Now suppose you would rather label states by than by , which is what a laboratory forces on you, since thermometers exist and entropy meters do not. That is exactly the swap (1.3.6) performs. So define
and the Helmholtz free energy is times the Legendre transform of in the variable . (Physicists carry the extra minus sign so that is minimised at equilibrium. The mathematical content is identical.) To see which variables naturally belongs to, take its differential, which is one line:
The terms cancel, and it is the same cancellation as (1.3.5), for the same reason. So the natural variables of are , and is the inverse relation (1.3.6).
Every other potential is the same move on a different slot. Transform instead and you get the enthalpy with . Transform both and you get the Gibbs free energy with , which is why is the one a chemist uses. Its natural variables are the two a bench can control.
Convexity is not decoration here either. The statement is thermodynamic stability, and it is exactly the condition (1.3.3) needed for the transform to be invertible. So a system with negative heat capacity has no well-defined free energy. That is not a mathematical curiosity. Self-gravitating systems have exactly that pathology, and it is why a star heats up when it loses energy.
In §2 we do the identical operation on , swapping for its slope . Same transform, same cancellation, same inverse relation. Only the letters change.
Grind box — Young's inequality, and what the supremum is really enforcing
Definition (1.3.2) says for every , with equality only at the maximiser. Rearranged, that is
with equality precisely when . This is Young's inequality, and it is not a separate theorem. It is the definition of the transform, read sideways.
The classic case is with on . Then , so and
since satisfies . So Young's inequality reads , the inequality from which Hölder's and Minkowski's inequalities are built, and the reason conjugate exponents come in pairs summing to in their reciprocals. Our worked example was , . ✓
What breaks without convexity. Take , the double well, which has on . For each slope there can now be three points with , and (1.3.2) selects whichever gives the largest gap, the global one. The information about the other two is discarded, and the double transform comes back as the convex hull: the double well with its non-convex middle replaced by a straight segment. This is not a defect. It is exactly the Maxwell construction of a first-order phase transition: the flat segment is the coexistence region, the tie-line joining liquid and gas, and its slope is the common tangent, which is the shared value of the chemical potential. Legendre-transforming a non-convex free energy and getting a flat piece is the phase transition. Chapter 6.6 meets the same picture again with the Higgs potential.
The condition in dimensions. With several variables, is invertible for precisely when the Jacobian of that map, the Hessian , is nonsingular, by the inverse function theorem of Chapter 0.6. Strict convexity in the multivariable sense means the Hessian is positive definite (Chapter 0.5), which is more than enough. Hold that thought. In §2 the Hessian in question is , which Chapter 1.2's grind box already showed is the positive-definite mass matrix whenever the kinetic energy is an honest quadratic form.
A curve can be given by listing its points, or by listing all the straight lines that graze it without crossing, and the two lists carry the same information, because a curve bending only one way is the highest of its own tangent lines. Swapping between the two descriptions is a standard operation, and performing it twice returns you exactly to where you began.
Which description you want is decided by what you are able to set. A laboratory has thermometers and no entropy meters, so a chemist labels a state by its temperature, which is the slope of the internal energy against entropy rather than the entropy itself, and the free energies are that one swap performed on one slot or another. What makes it reversible is that the curve bends one way only. Where it does not, the swap flattens the offending stretch permanently, and that flat stretch is a first-order phase transition rather than a defect in the mathematics.
In mechanics the variable traded away is the velocity and the slope accepted in its place is the momentum, and the reversibility condition becomes the statement that the kinetic energy is an honest positive quantity in the velocities. What the trade produces has been met before, as the combination a problem with no explicit clock in it happens to conserve.
2 · Momentum and the Hamiltonian
2.1 · The canonical momentum is not
Before we transform anything, we need to be clear about the variable we are trading for. Chapter 1.2 defined, for any Lagrangian and any coordinates,
the canonical momentum conjugate to , and noted that this makes the Euler–Lagrange equation read . In Cartesian coordinates with , (1.3.11) returns , which is why the name is not absurd. But that agreement is a special case, and it fails in two different ways that both matter.
It need not have the dimensions of momentum. For the pendulum with coordinate , gives . That is an angular momentum, with units of rather than . The word "conjugate" is the point. Whatever turns out to be, it is the thing that pairs with so that has the dimensions of action. Chapter 4.9's uncertainty relation is dimensionally consistent for every conjugate pair for exactly this reason.
It need not be even in Cartesian coordinates. This is the one worth being careful about, and seeing it costs one example. Chapter 2.6 will derive, from relativity, that a particle of charge in an electromagnetic field described by the scalar potential and vector potential has the Lagrangian
Worked example 1 checks that this reproduces the Lorentz force, so it is the right Lagrangian. Our question is what its canonical momentum is, so apply (1.3.11) to it, component by component, remembering that depends on position but not on velocity:
The canonical momentum and the mechanical momentum differ by . This is the first time the vector potential has entered mechanics in this book, and it will not leave. Chapter 2.6 shows that is the spatial part of a four-vector and that Maxwell's equations are its equations of motion. Chapter 6.3 shows that the entire structure of the Standard Model follows from demanding that the substitution be forced on you by a symmetry principle. The name for that substitution is minimal coupling, and you have just watched it appear, three parts early, as an unremarkable partial derivative.
It is tempting to file (1.3.13) under "notation", on the grounds that must be the real momentum and a bookkeeping device. That is exactly backwards, and the difference is measurable. Here is why it matters in practice.
The gauge problem. Chapter 1.2's Problem 4 showed that , changes by a total time derivative and therefore changes no physics. But it does change , by . So the canonical momentum is gauge-dependent, and by itself it is not an observable. Meanwhile is gauge-invariant, because shifts by exactly the same amount. The mechanical momentum is what you measure. The canonical momentum is what the formalism runs on. Both are needed and they are not the same object.
And yet it is that gets quantised. In Chapter 4.6 the operator conjugate to position is , and the rule for which classical quantity it represents is fixed by the bracket of §6, which is a statement about the canonical momentum. Feed (1.3.13) into a Hamiltonian and the kinetic term becomes , so that is replaced by acting on the wavefunction. Every magnetic effect in quantum mechanics comes out of that one replacement: Landau levels, the Aharonov–Bohm phase, the Zeeman effect, superconducting flux quantisation. None of them would, if were the thing you quantised.
⚑ The sharpest experimental statement is quoted forward to Chapter 6.3. In the Aharonov–Bohm effect an electron beam is split around a solenoid and recombined. Outside the solenoid the magnetic field is exactly zero, so a particle described by alone can feel nothing. The interference fringes nonetheless shift, by an amount , which is the enclosed flux in units of . The potential is not a computational convenience. It is where the physics lives, and is the first hint.
2.2 · The Hamiltonian
Now we have both halves of the trade: a function that depends on velocity, and a slope we would rather use instead. So perform §1's transform on , taking as the variable and holding and fixed. The conjugate slope is by (1.3.11), and the transform (1.3.4) reads
Every on the right is to be eliminated in favour of the by inverting (1.3.11), and that instruction is the entire content of the definition. The result is a function of , and only. This object is the Hamiltonian. Compare it with (1.3.1) and you can see that Chapter 1.2's Problem 2 was computing it: the conserved quantity of a time-independent Lagrangian is the Hamiltonian.
The inversion step is not free, and §1 told us exactly what it needs. Solving for requires that the map from velocities to momenta be invertible, and by the inverse function theorem that means asking about the Hessian
which must be nonsingular. Chapter 1.2's grind box on the double pendulum established that when the coordinate change from Cartesians is time-independent, with symmetric positive definite, so and the condition holds automatically.
⚑ It does not always hold, and the failure has a name. When is linear in some velocity, the corresponding row of vanishes, the transform is not invertible, and the system is called constrained. Dirac built a whole formalism for that case. It is quoted here and not developed, but you should know where it lives, because it is not an exotic corner. Gauge theories are there (Chapter 6.3: the Lagrangian of electromagnetism has no at all), and so is the reparametrisation-invariant string action (Chapter 7.2). Every fundamental theory in Parts VI and VII is a constrained system. So the honest statement is this. The clean Legendre transform of this section is the easy case, and it is the one that teaches you what the hard case is deforming.
2.3 · When , and when it isn't
You have surely been told that the Hamiltonian is the total energy. It often is. But that is a theorem rather than a definition, and the theorem has two hypotheses that fail in ordinary situations. Let's state it, prove it, and then break it.
Suppose (i) does not depend on the velocities, and (ii) is a homogeneous quadratic form in the generalised velocities, . Then .
Proof. With and hypothesis (i), . Hypothesis (ii) says is homogeneous of degree : replacing every by multiplies by . What we want is the combination , so differentiate that homogeneity identity with respect to and set . The chain rule gives
which is Euler's theorem on homogeneous functions, derived rather than quoted. (The grind box does the general degree.) That is the whole input. Feed it into the definition of and every term is now known:
Now let's break the hypotheses, using a system Chapter 1.2 already built. Problem 3 there put a bead on a hoop of radius spun about its vertical diameter at a rate fixed by a motor, and found
The middle term is part of the kinetic energy, since it is the speed the bead has purely because the hoop is carrying it around. But it contains no . So is not homogeneous of degree in the velocity, and hypothesis (ii) fails. The theorem is therefore unavailable, and we must compute honestly instead. Since , we have , and putting that into the definition gives alongside the true energy for comparison:
They differ by , which changes as the bead slides. is conserved, since it has no explicit . is not, because the motor is doing work to hold fixed. Conserved and "the energy" are different predicates, and is the first one.
Notice also what is here. It is the kinetic energy of the motion plus the effective potential that Chapter 1.2 found. The centrifugal term has flipped sign on its way through the transform and become part of the potential. That is not sleight of hand. It is (1.3.16) refusing to apply to a degree-zero term.
Grind box — Euler's theorem in general, the three-piece decomposition, and the relativistic particle
Euler's theorem. Let be homogeneous of degree in the variables , meaning for all . Differentiate both sides with respect to using the chain rule:
and set to get . That is the whole proof. Degree gives (1.3.16), degree gives , and degree gives .
The general Lagrangian. Split by degree in . Then by the three cases above, so
Every degree-one piece drops out of entirely. Three consequences, all checkable against what we have already done:
- Ordinary mechanics: , , , giving . ✓
- Rotating hoop: , , , giving exactly (1.3.19). ✓ The centrifugal term's sign flip is the minus in .
- Charged particle: , , , giving . The magnetic term vanishes from , which is the statement that magnetic forces do no work, derived rather than asserted. It reappears the moment you express in terms of rather than , because . Worked example 1 does this properly.
The relativistic free particle. Chapter 1.2's table listed , i.e. in one dimension , which is not of the form and is not homogeneous of any degree. Turn the crank anyway:
Then, using ,
where the middle step used . So , the relativistic energy including the rest energy, produced by a Legendre transform and nothing else. Expressed in the right variable, which is what demands, we use and the same identity to get
That is the relation Chapter 2.5 derives from four-vectors and Chapter 5.5 quantises into the Dirac equation. It fell out of §1's transform applied to a square root. Note in passing that , so the Hessian condition (1.3.15) holds and the transform is legitimate. It is the same positivity Chapter 1.2's second-variation grind box used to show that a relativistic worldline maximises proper time.
The quantity paired with a coordinate here is not always mass times velocity, and the two ways it can differ both matter. For a pendulum described by an angle it is an angular momentum, with the wrong units for a momentum entirely, because what the pairing requires is that coordinate and partner multiply to something with the units of action. For a charge in a magnetic field it is not the mechanical momentum even in Cartesian coordinates, the field's potential having been added to it.
That second case is not notation. The mechanical momentum is what an instrument reads, while the conjugate one shifts when the potential is rewritten in a way that changes no physics, so the conjugate one is not by itself observable. Yet the conjugate one is what quantum mechanics turns into an operator, which is why every magnetic effect in that subject, including the interference shift around a solenoid whose field the electrons never enter, follows from a single substitution.
The other object built here is usually introduced as the total energy. It is the total energy under two hypotheses, both of which fail in ordinary situations, and a bead on a motor-driven hoop is enough to break them. The conserved quantity and the energy come apart there, and it is the first that the formulation runs on.
3 · Hamilton's equations, derived
We have a new function . We do not yet have equations of motion in terms of it. There are two ways to get them, and they are worth seeing both, because they emphasise different things.
3.1 · First route: take the differential
Chapter 0.6's total differential is the statement that for any smooth function of several variables, the first-order change is the sum of the partial derivatives times the changes in their arguments. Our aim is to find out which variables genuinely depends on, so apply that rule to the right-hand side of (1.3.14), treating , and as the independent variables there:
Now look at the two terms containing . Their coefficients are and , and by the definition (1.3.11) those are equal and opposite. They cancel identically.
That cancellation is the entire trick, and it deserves a sentence of its own: the Legendre transform is engineered so that the variable you are trying to get rid of drops out of the differential. It is the same cancellation as (1.3.5), and the same one as in the thermodynamics callout. Not an analogy. One theorem, appearing for the third time.
With those two terms gone, what survives is
which contains only , and . That confirms really is a function of and not secretly of the velocities. And once we know that, we may write its differential a second way, straight from the definition of a partial derivative:
Two expressions for the same differential, in terms of the same independent increments. Their coefficients must therefore match one by one. Comparing , then , then gives three identities:
Only the middle one still mentions , and we want equations in alone, so kill it with the Euler–Lagrange equation, which says . And there they are:
Hamilton's equations. Count what happened. Our second-order equations became first-order ones. Nothing was gained or lost in information, because a second-order equation always splits into two first-order ones, and the standard way to do that is to call a new variable. What is special here is which new variable was chosen. Choosing rather than itself is what produces the near-symmetry of (1.3.24), and that symmetry is worth everything that follows.
One immediate dividend, free of charge. Suppose has no explicit time dependence, and ask how changes along a trajectory. The chain rule plus (1.3.24) answers in one line:
is conserved, by cancellation between the two halves of (1.3.24). That is Chapter 1.2's Beltrami identity again, now in three lines instead of five. Notice which feature did the work. The minus sign that made the two equations not symmetric is precisely what made the cancellation happen. Remember that, because §5 and §6 are both that minus sign.
3.2 · Second route: vary the action in phase space
The first derivation used the Euler–Lagrange equation as an input. Here is a derivation that does not, and that treats and on an equal footing from the start.
Rearrange (1.3.14) as and substitute it into the action. Then declare and to be independent functions, with no relation between them assumed. The object to be made stationary is
This is a functional on a space of paths in the -dimensional space of , and we ask for its stationary points exactly as Chapter 1.2 did. So vary both: , . Note the asymmetry that is about to matter. In (1.3.26) the velocity appears and does not. Collecting the terms linear in ,
Only one term carries a derivative of a variation, and Chapter 1.2's move three handles it. We integrate by parts to get , so that every remaining term carries or as a plain factor. Collecting those coefficients separately,
Kill the boundary term by requiring , which fixes the positions at both ends, as before. Note carefully that no condition on is needed. Since never appears differentiated, it generates no boundary term and is free at the endpoints.
That is not a technicality, and it is worth saying what it buys. It says the natural boundary-value problem in phase space is "specify where you start and where you finish", not "specify the whole initial state". Chapter 5.6's path integral sums over paths with fixed endpoints, and it inherits exactly this asymmetry.
Now apply the fundamental lemma (Chapter 1.2 §3.4) twice, once with arbitrary and , and once the other way round. Each bracket must vanish separately, which gives
and that is (1.3.24) again. The two derivations agree, as they must, and each tells you something the other does not. Route one tells you the equations are the Legendre transform of Euler–Lagrange. Route two tells you they are themselves a variational principle, one in which position and momentum are independent variables of equal status.
3.3 · The minus sign, and what it is the seed of
Stare at (1.3.24) until the near-symmetry bothers you. The equations for and are the same equation with , except for one minus sign. That asymmetry is irreducible. You cannot rescale it away, because rescaling just moves the sign to the other equation.
Since we cannot remove it, the next best thing is to give it a home of its own where we can see exactly what it does. Start by stacking the coordinates into one list of numbers,
so that a single letter now names the whole state. The minus sign is a relation between the first half of that list and the second half, so it should live in a matrix. Define
with the three properties on the right all read straight off the block form. With that one matrix in hand, both halves of (1.3.24) collapse into a single line:
Check it: the top block of is and the bottom block is . ✓
Now read that line geometrically, because it explains a fact we have already proved algebraically. The gradient of points uphill, at right angles to the level sets. rotates that direction by a quarter turn. So the system moves along the level sets, which is why is conserved. The flow is everywhere perpendicular to the gradient, by construction.
is called the symplectic form, and the antisymmetry is the minus sign, isolated. Compare it with the object that runs Chapter 3.3, the metric , which is symmetric and measures lengths. A symplectic form measures oriented areas instead, and it is the reason phase space has a geometry at all.
Every structural fact in the rest of this chapter is a consequence of the antisymmetry of : Liouville's theorem, the Poisson bracket, and the of quantum mechanics. So is the fact that phase space is always even-dimensional. An antisymmetric matrix in odd dimensions has , forcing , so no invertible exists there.
Trading one second-order equation for two first-order ones is available to anybody and costs nothing, since calling the velocity a new variable does it. All of the content here is in which new variable was chosen. Taking the conjugate partner rather than the velocity produces two equations that are the same equation with the roles of coordinate and momentum exchanged, apart from a single minus sign sitting in one of them.
That sign cannot be removed. Reversing the sign of the momentum only moves it to the other equation, and it is what makes the two halves of the calculation cancel when you ask whether the energy changes, so that the energy does not change. Written compactly the sign becomes an antisymmetric object standing between the two halves of the space, whose effect is to turn a gradient through a quarter turn, so the motion runs along the level surfaces of the energy rather than across them.
Everything structural in the rest of the chapter is that antisymmetry and nothing else. A symmetric object of the same general kind measures lengths and will run the chapters on gravity; this one measures oriented areas instead, which is a different geometry and the reason the space must have an even number of dimensions, an antisymmetric table in an odd number of them being necessarily degenerate.
4 · Phase space
The space whose points are the numbers is phase space. Chapter 0.8 introduced it for the oscillator and promised it would come back. This is where it becomes the setting rather than a picture.
4.1 · One point fixes everything, and trajectories cannot cross
Hamilton's equations (1.3.32) are first order. That is the structural difference from the Lagrangian formulation, and it changes what a "state" is. Chapter 0.8's Picard–Lindelöf theorem says that a first-order system with Lipschitz has exactly one solution through each initial point. Applied here, that gives us the following.
A single point of phase space determines the entire past and future of the system. Not a point plus a velocity. A point.
Configuration space does not have this property. A pendulum passing through might be moving left or right, fast or slow, so "" is not a state. Phase space is precisely the space of things that are states, which is why Chapter 0.8 said the list the existence theorem asks for is the list phase space is made of.
There is an immediate geometric consequence, and it is what makes phase portraits legible.
Claim. If has no explicit time dependence, two distinct trajectories can never cross.
Proof. Suppose the trajectory passes through the point at time and the trajectory passes through the same at time . Because does not depend on , the right-hand side of (1.3.32) depends only on the point, so shifting a solution in time gives another solution. Define , which also solves the equations and satisfies . Now and are two solutions with the same value at , so by uniqueness they are the same solution. Hence is retimed, and it is the same curve.
So through every point of phase space there passes exactly one curve, and phase space is filled by these curves like the streamlines of a steady fluid, with no branching, no merging and no intersections. The only points where curves appear to meet are fixed points, where and the "trajectory" is a single stationary point. We will meet one shortly.
4.2 · The harmonic oscillator, again
The quickest way to get a feel for a flow is to draw one we already know. Take , which is , since both hypotheses of §2.3 hold. Differentiating it in and in gives Hamilton's equations for the oscillator:
The first says , recovering the mechanical momentum. Substituting it into the second gives , the oscillator equation. Nothing new so far. But now read the pair as a flow instead of as a route back to a second-order equation.
The level sets of are the ellipses of Chapter 0.8 §4.4, and by §3.3 the motion runs along them. To see what the motion actually is, we would like the ellipses to be circles, so rescale the two axes by equal and opposite factors:
which has the virtue of preserving areas, since . That is a property §8 will name. In these variables , and rewriting (1.3.33) in them gives
That is a rigid rotation of the whole plane at angular velocity , clockwise. Every point, whatever its energy, goes round at the same rate, which is the phase-space statement of the fact that a harmonic oscillator's period is independent of amplitude. Chapter 0.1 observed that is a rotation by 90°. Here the rotation is generated by , which squares to for the same reason. The oscillator flow is multiplication by , drawn.
4.3 · The pendulum, and the separatrix
Now a system where the flow is not uniform. Take the pendulum of Chapter 1.2 §6.2, with and . Both hypotheses of §2.3 hold, so its Hamiltonian is the total energy:
where the constant has been chosen so that at the bottom. Hamilton's equations are and , reproducing . Write .
Since is conserved, trajectories are its level sets, and there are three kinds.
Libration. For the bob cannot reach the top. Setting in (1.3.36) gives a turning point at , which has a solution only when . The level set is a closed curve encircling the origin, and the motion is back-and-forth swinging. Near the origin, expanding recovers the oscillator of §4.2 and the level curves are Chapter 0.8's ellipses.
Rotation. For there is no turning point. Here never vanishes, so the bob goes over the top and keeps going. The level set is an open curve running across the whole range of , and since and are the same physical configuration, phase space here is really a cylinder and the curve is a loop around it.
The separatrix. Between those two behaviours sits one energy at which the bob arrives at the inverted position with exactly zero momentum. Putting , into (1.3.36) gives its value:
We would like the curve itself, not only its energy, so solve (1.3.36) for at this energy, using to make the square root come out cleanly:
These are the two arcs Chapter 0.8 §8.2 found. They meet at , , and that meeting point is a fixed point, since and . So §4.1's no-crossing theorem is not violated. To see what the flow does near it, linearise by writing , so that :
whose solutions are , one growing and one decaying. This is a hyperbolic fixed point, or saddle. The level curves near it are hyperbolas, and the separatrix is the pair of straight lines they asymptote to. The decaying solution is the trajectory that creeps toward the inverted position and takes infinite time to arrive. The growing one is the same trajectory run backwards.
Compare the origin, where the same calculation gives and the level curves are closed. That is an elliptic fixed point, or centre. Two fixed points, two signs, two completely different local pictures, and the entire qualitative structure of the pendulum is the statement that they are joined by the separatrix.
4.4 · What the enclosed area means
For a closed orbit, meaning a libration, the natural number to attach to it is the area it encloses:
where the line integral around the closed curve equals the enclosed area by Green's theorem (Chapter 0.7 §5.1). That is what a line integral of around a loop is. Two facts make this the right quantity to care about.
First: its derivative is the period. To see it, ask how the area responds when the energy changes. Solve for on the upper branch and differentiate (1.3.40) under the integral sign. Since at fixed is the reciprocal of at fixed , and , the integrand turns into :
The area is the antiderivative of the period. For the oscillator, Chapter 0.8 computed , so ✓, and since does not depend on there, the area is exactly linear in the energy. For the pendulum it is not. Numerical integration of (1.3.40) at , gives at and at , against periods of and from Chapter 0.8's elliptic-integral formula. Agreement to ten digits, which is the sort of thing worth checking once.
Second: it is what quantum mechanics rations. ⚑ Quoted forward to Chapters 4.8 and 4.10. The Bohr–Sommerfeld condition states that in the semiclassical regime the allowed orbits are those with
with Planck's constant. For the oscillator this reads , that is , which Chapter 4.8 will derive exactly, with ladder operators and no semiclassical approximation, and get precisely this answer. That is the one case in which the condition is exact rather than semiclassical, and it is the only part of this flag 4.8 discharges. The general statement, and the measurement of how far it is from the truth when the potential is not a parabola, is Chapter 4.10 §6.
So the classical phase-space ellipse is not deleted by quantum mechanics. It is quantised. The continuum of orbits becomes a discrete stack of them, one per unit of of enclosed area, and the ground state is the one with half a unit. Phase-space area is measured in units of Planck's constant. That sentence is the reason has the units it does, namely , the units of , and it is the single most useful thing to remember about phase space.
Grind box — the area–period relation done carefully, and the adiabatic invariant
The derivation, without hand-waving. Take a one-dimensional system with and a closed orbit at energy . On the upper half of the orbit write . The enclosed area is , traversed so the integral is positive. Differentiating with respect to moves the curve. The endpoints, which are the turning points, move too, but there, so the boundary contribution vanishes and we may differentiate under the integral:
Now, differentiating the identity with respect to at fixed gives , so . Hence the integrand is , and integrating once round the orbit gives one period.
The action variable. Define . Then , the angular frequency. Problem 3 constructs a change of variables in which is literally the new momentum and the new coordinate is an angle advancing uniformly at rate . These are action–angle variables, and they turn any one-dimensional bound motion into the trivial system , .
⚑ One quoted fact, because it explains why (1.3.42) was ever a sensible guess. If a parameter of the system (a pendulum's length, say) is changed slowly compared with the period, then and both change but does not, to all orders in the slowness. is an adiabatic invariant. Ehrenfest's argument was that only an adiabatic invariant can sensibly be quantised, since if you could change by gently stretching a wire you could move a system off its allowed levels without a transition, and that is how Bohr and Sommerfeld chose which quantity to set equal to . The modern statement is Chapter 4.17's adiabatic theorem: a quantum system stays in the -th eigenstate under slow change, which is the same sentence with replaced by .
Numbers. For the pendulum at , , quadrature of (1.3.40) against the elliptic-integral period gives
| (numerical) | relative difference | ||
|---|---|---|---|
The period diverges as , since the separatrix takes infinite time, and (1.3.41) says the area therefore has infinite slope there while remaining finite. Both statements are visible in the interactive below. Near the separatrix, neighbouring orbits have wildly different periods, so a blob of initial conditions is sheared apart. That shearing is the whole show.
A single point of this space fixes the entire past and future of the system, which is the whole reason for building it. A configuration by itself is not a state, since a pendulum passing through the bottom of its swing might be going either way at any speed. A configuration together with its conjugate partner is a state — this is the space of states the existence theorem always asked for.
Two things follow. Exactly one solution passes through each point, so trajectories never cross and the space fills with curves like the streamlines of a steady fluid, the only apparent meetings being where nothing moves at all. And the energy is unchanging along each curve, so a pendulum's whole qualitative behaviour reduces to three kinds of curve: closed loops for swinging, open ones for going over the top, and the single curve dividing them, along which the bob takes forever to arrive upside down.
The number worth attaching to a closed loop is the area it encloses, whose rate of change with energy is the period. A promise made when the oscillator was first drawn is collected here: quantum mechanics does not delete these loops, it rations them, permitting only those enclosing a whole number of units of Planck's constant, with half a unit for the lowest. That is why the constant has the units it does.
5 · Liouville's theorem
This is the centrepiece, and it costs almost nothing, because Chapter 0.7 already did the work.
Forget for a moment that (1.3.24) describes mechanics and read it as what it literally is: a formula assigning a velocity vector to every point of a -dimensional space. That is a vector field in the sense of Chapter 0.7 §1, and writing out its components gives
Chapter 0.7 spent a section establishing what the divergence of such a field means, and the answer was not "something to do with flux" but something much more useful. The divergence is the trace of the Jacobian, hence the fractional rate at which a blob carried by the flow changes its volume:
So the question "does a Hamiltonian flow squeeze or spread the blobs it carries" is now a question about one derivative. Compute the divergence of (1.3.43), which is the sum of over all coordinates:
by the equality of mixed partial derivatives, which is Clairaut's theorem, Chapter 0.6 §6.1. Every term cancels against its partner, and the cancellation happens because of the minus sign of §3.3 and nothing else. Now substitute that zero into (1.3.44) and read off what it says about volume:
Liouville's theorem: Hamiltonian flow preserves phase-space volume exactly. Not approximately, not on average, and not for special systems. For every Hamiltonian, including time-dependent ones, since nothing in the calculation used , including chaotic ones, and including the entire universe if you are willing to write down its Hamiltonian.
Take any region of phase space, that is, a set of possible initial conditions. Let every point of it evolve for a time , and call the resulting set . Then .
The shape of is entirely unconstrained. It may be stretched, sheared, folded, wound into a spiral of a million turns, or drawn out into a filament thinner than any resolution you can afford. The theorem says nothing whatever about shape. It says the measure is exactly invariant. Arbitrary distortion together with exactly conserved measure is the combination that makes the theorem worth having.
Three payoffs, in increasing order of importance.
Statistical mechanics has a measure. To do statistical mechanics you must say what "equally likely" means for a continuous system, and that requires a measure on phase space that the dynamics does not distort. Liouville supplies exactly one, namely . The microcanonical ensemble, uniform on a surface of constant energy, is well defined because the flow is incompressible. Had the flow contracted volumes, an initially uniform ensemble would pile up and "uniform" would not be a stable notion. Every ensemble average you have ever seen rests on (1.3.46).
There is no classical attractor. A damped oscillator spirals into the origin, so a blob of initial conditions shrinks to a point and volume is destroyed. Liouville therefore says damping is never fundamental. It is what a Hamiltonian system looks like when you have thrown away the degrees of freedom that the energy leaked into: the air, the wire, the thermal bath. The full system, including the bath, conserves volume. This is the reason a fundamental theory can never contain a friction term, and why every dissipative equation in physics is an effective description.
Information is not destroyed, only stirred. This is the one to keep. Suppose you know the state to within a small blob of volume , which is your measurement precision. After evolution the blob has some horrible filamentary shape, but its volume is still . In principle, running the equations backwards recovers the initial blob exactly, so Hamiltonian evolution is reversible and loses nothing. Two forward pointers, both of which are this sentence in different clothes.
- Chapter 4.6: quantum time evolution is generated by , which is a unitary operator. It preserves inner products, hence probabilities, hence distinguishability, so two distinct initial states stay distinct forever. That is the quantum statement of exactly this theorem, with "volume in phase space" replaced by "overlap of state vectors" and the symplectic form replaced by the factor of .
- Chapters 3.8 and 7.9: the black-hole information problem is what happens when the two statements appear to disagree. Hawking's calculation says a black hole evaporates into radiation that depends only on its mass, charge and spin, so the volume of states that could have formed it collapses to a point and information is destroyed. Something in that argument must be wrong, because unitarity is not a convenience you can trade away, and the search for what is wrong produced holography. The problem is a problem precisely because (1.3.46) and its quantum successor are load-bearing.
Liouville's theorem seems to forbid the second law of thermodynamics. Entropy is supposed to increase. Entropy is a measure of the phase-space volume compatible with what you know. But the volume is exactly constant. Something has to give.
Look again at the figure. After a few periods the blob is a filament winding through a large region. Its area is what it was. But ask a different question: how much of the plane is within of some point of the blob, for a small but nonzero ? That number grows enormously, because a long thin filament has a huge neighbourhood. Any description of the system with finite resolution sees the coarse-grained blob rather than the blob, and any thermodynamic description is of that kind, since you cannot record coordinates. The coarse-grained volume is not conserved. It increases, monotonically in practice, and its logarithm is the entropy.
So the arrow of time is not in the dynamics. The dynamics is exactly reversible and exactly volume-preserving, forwards and backwards, and the figure runs equally happily either way. The arrow is in the coarse-graining. It is in the fact that stirring makes a distribution look uniform at any finite resolution, and that we are obliged to describe systems at finite resolution. Gibbs's own image was a drop of ink stirred into water. The ink occupies the same volume it always did, but you will never unstir it, because unstirring requires knowing where every filament went.
Two consequences worth having. First, "entropy increases" is a statement about our description rather than a term in any equation of motion. There is no anywhere in (1.3.24). Second, ⚑ quoted: Poincaré's recurrence theorem says that a Hamiltonian system in a bounded phase-space region must return arbitrarily close to its initial state, and the proof is Liouville. If it never returned, the images of a small blob would be disjoint forever and their total volume would exceed the finite volume available. The recurrence times are absurd, being exponentially large in for a mole of gas, which is why this does not overturn thermodynamics. But it is not a loophole. It is a theorem, and it exists because of the three lines above.
The centrepiece of the chapter costs three lines and is hard to credit at first hearing. Take any region of this space, regard it as a collection of possible starting conditions, let every point evolve for as long as you like, and the volume occupied afterwards is exactly what it was. Not nearly, not on average, and not for well-behaved systems only.
The shape is unconstrained while the measure is exact, and that combination is what makes the theorem worth having. A blob released near a pendulum's dividing curve is drawn into a filament wound many times round, thinner than any resolution you can afford, and its area does not move. Damping can therefore never be fundamental, since a system spiralling into rest destroys volume; friction is a larger volume-preserving system with some coordinates thrown away. And information is stirred rather than destroyed, because running the equations backwards recovers the original blob exactly.
That last sentence travels furthest. Its quantum successor says evolution preserves overlaps, so two states that begin distinguishable stay distinguishable forever, and an evaporating black hole is a crisis rather than a curiosity because the usual calculation appears to contradict it. The arrow of time, meanwhile, is not in the dynamics, which runs equally happily either way, but in the fact that a filament with an enormous neighbourhood looks like a filled region at finite resolution.
6 · Poisson brackets
We now extract from Hamilton's equations the structure that will survive quantisation. Everything in this section is elementary. Its importance is entirely in what it becomes.
6.1 · The equation of motion for any observable whatsoever
An observable is any function on phase space: the energy, a component of angular momentum, the distance between two particles, the electric dipole moment, anything you could in principle measure. Ask how it changes along a trajectory. By the multivariable chain rule (Chapter 0.6),
which is true of any function on any space and says nothing about mechanics yet. The mechanics enters when we say what and actually are, so substitute Hamilton's equations (1.3.24) for both of them:
The bracketed combination has appeared without being asked for, so let's give it a name. For any two functions on phase space define the Poisson bracket
where the second form uses the symplectic matrix of (1.3.31). Check the block structure and you will find it reproduces the sum exactly. With that abbreviation, (1.3.48) shortens to
This is not a formula about a particular system. It is the equation of motion for every observable of every Hamiltonian system, and it contains Hamilton's equations as the special cases and (verify: and ✓).
It also contains the master conservation law, immediately. Set in (1.3.50) and the left-hand side vanishes exactly when the bracket does:
Conservation has become an algebraic condition. You no longer solve the equations of motion to find out whether something is conserved. You compute one bracket. Taking gives by antisymmetry alone, recovering (1.3.25).
6.2 · The fundamental brackets
The first brackets to compute are the ones between the coordinates themselves, since everything else is built from those. For , note that , , and , so only the first term of (1.3.49) survives:
The last two vanish because in each case one of the two factors in every term is a derivative of a with respect to a , or the other way round. These three lines are the canonical commutation relations of classical mechanics. Every coordinate commutes with every other coordinate. Every momentum commutes with every other momentum. And a coordinate fails to commute with its own conjugate momentum, by exactly .
6.3 · The algebra
The bracket has four properties, and they are the properties rather than incidental features.
Bilinearity. for constants , and likewise in the second slot. Immediate from (1.3.49), since differentiation is linear.
Antisymmetry. , and hence . Immediate from . Swapping and in transposes a number, which does nothing, and transposing the middle factor flips the sign.
Leibniz. . Immediate from the product rule, since every term of (1.3.49) differentiates exactly once. This says acts as a derivation, meaning it behaves like a directional derivative, and §7 will explain that by exhibiting the direction.
The Jacobi identity. The fourth is the only one that is not a line of inspection, and it says that the three ways of nesting a bracket inside a bracket sum to nothing:
This one is not obvious and is proved in the grind box. It is the compatibility condition that makes everything hang together. Bilinear, antisymmetric, and Jacobi is the definition of a Lie algebra, the structure Chapter 6.1 builds a whole part of the book on.
One concrete consequence is available right now: if and are both conserved, so is . To see it, put in (1.3.53) and use . The first two terms die, leaving . This is Poisson's theorem, and it is a machine for manufacturing conservation laws from ones you already have. It occasionally produces something new and usually produces something you knew, and either way it does not cost you an integration.
Grind box — proving the Jacobi identity, and why no second derivatives can survive
Write everything in the compact notation of (1.3.49). Let , , be the phase-space coordinates, , and the constant antisymmetric matrix of (1.3.31), so that
with repeated indices summed. The two facts we will use are that is constant, so it passes through derivatives, and antisymmetric.
The counting argument first, because it tells you why the identity is true. Each term of (1.3.53) is a bracket of a bracket, so it contains exactly two derivatives distributed over three functions. Every term therefore contains a second derivative of exactly one of and first derivatives of the other two. The claim is that the second-derivative terms cancel among themselves, separately for , for and for , and that once they are gone nothing is left, since every term has a second derivative somewhere. So it suffices to check one of the three cancellations, the other two following by relabelling.
The cancellation. Expand the first and third terms by the product rule:
Collect the two terms carrying second derivatives of : they are
(The middle term of (1.3.53), , contains second derivatives of and of only, so it contributes nothing here.) Now relabel the dummy indices in by the substitution , , , :
using once and the symmetry (Clairaut) once. So . By the cyclic symmetry of (1.3.53), the second-derivative-of- terms and the second-derivative-of- terms cancel by the identical relabelling. Nothing remains.
Note what the proof used: antisymmetry of , constancy of , and equality of mixed partials. Nothing else, and in particular nothing about mechanics. The identity is a fact about any constant antisymmetric bilinear pairing of gradients.
Verification, since a proof by index gymnastics deserves a check. Taking the deliberately unstructured functions , , on a three-degree-of-freedom phase space, symbolic differentiation gives a Jacobi residual of exactly , and the Leibniz residual is exactly as well.
6.4 · The payoff, and the reason Part I exists
Here is what the last three pages were for.
Chapter 4.2 §8 will take the classical structure you now own, replace observables by self-adjoint operators on a Hilbert space, and make the single substitution
Nothing else changes. Every equation of this section survives verbatim.
Watch what that does to the results we have. The fundamental brackets (1.3.52) become
which is the foundational relation of quantum mechanics, from which the uncertainty principle follows in three lines (Chapter 4.9) and for which is the standard realisation (Chapter 4.6). Do the same to the equation of motion (1.3.50) and it becomes
the Heisenberg equation of motion, which is the whole of quantum dynamics in the Heisenberg picture. And the conservation criterion (1.3.51) becomes " is conserved if and only if it commutes with the Hamiltonian", the statement that gets used, without comment, in every quantum mechanics course from the second week onward.
The bracket properties go across too, and they are exactly the properties a commutator has. Bilinearity: immediate. Antisymmetry: by inspection. Leibniz: , which you can verify in one line by inserting and cancelling . Jacobi: holds for commutators identically, by expanding all twelve products and watching them cancel in pairs. The classical Poisson algebra and the quantum commutator algebra are the same Lie algebra. That is the sense in which the substitution "changes nothing else".
So the answer to "why did we spend a chapter on a change of variables" is that the change of variables is the thing that quantises. You cannot quantise the Lagrangian formulation directly, because there is no bracket there, only a variational principle. That is why Feynman had to invent an entirely different route, the path integral of Chapter 5.6, to quantise from an action. Canonical quantisation, by contrast, is a one-line substitution into a structure you now have.
"Replace brackets by commutators" is a correspondence rather than a theorem, and it is not quite consistent. ⚑ Quoted: the Groenewold–van Hove theorem proves that there is no map from classical observables to operators that sends to for all while also sending and . The obstruction shows up as soon as you go beyond quadratic. The bracket and the corresponding commutator disagree once you decide how to order 's and 's. So the substitution is exact for observables at most quadratic in and , which covers the harmonic oscillator, free particles, angular momentum and therefore most of Chapter 4, and beyond that it needs a convention, an ordering prescription.
This is not a scandal. It is a statement that classical mechanics does not uniquely determine its quantum parent, which we should have expected, since the arrow of explanation runs the other way. Quantum mechanics is the theory. Classical mechanics is its shadow, and shadows do not determine the objects that cast them. Keep the correspondence as what it is: an extremely good guess, exact where it matters most, and the reason the structure of this chapter is worth owning.
Out of the equations of motion falls a way of combining two quantities on this space to produce a third, and its importance lies almost entirely in what it later becomes. The immediate use is already substantial. The rate of change of any measurable quantity whatever is its combination with the energy, so asking whether something is conserved stops being a matter of solving the motion and becomes a matter of computing one expression and seeing whether it is zero.
Among the coordinates themselves the answers are as simple as they could be. Two coordinates give nothing, two momenta give nothing, and a coordinate paired with its own conjugate momentum gives exactly one. Those three lines and the four properties the operation possesses are the entire structure, and the fourth property is what makes the whole an algebra of the kind the last parts of this book are built from.
Here is what the chapter was for. Replace this operation by the difference between doing two things in one order and in the other, divided by the imaginary unit and Planck's constant, alter nothing else, and classical mechanics becomes quantum mechanics. The relation between a coordinate and its momentum becomes the founding relation of the subject, and the conservation criterion becomes the familiar remark that a quantity is conserved when it commutes with the energy.
7 · Symmetries generate motion
This is the deepest section in the chapter, and it is short. It answers a question you may not have thought to ask: why should conserved quantities and symmetries have anything to do with one another?
Chapter 1.4 will prove Noether's theorem, which says that every continuous symmetry of the action implies a conserved quantity. That is a one-way street, and it leaves the conserved quantity looking like a by-product, something the symmetry emits. The Hamiltonian picture reveals that the relationship is far tighter than that.
7.1 · Every observable generates a flow
The idea is to notice that Hamilton's equations never used any special property of . So take any function on phase space and put it where used to be. That is, define an infinitesimal change of every observable by
Applied to the coordinates themselves this reads and , which is the same structure as (1.3.24) with playing the role of a small time and the role of the Hamiltonian. So defines a flow on phase space, and (1.3.56) is how any observable changes along it. We say generates the transformation.
Before doing examples, we should check that these flows respect the structure of §3.3, and they do. Compute the fundamental bracket of the shifted variables, to first order in , using bilinearity and then Clairaut:
So the transformation is automatically canonical, meaning it preserves the fundamental brackets. Mixed partials again, for the third time in this chapter, after (1.3.45) and the Jacobi identity. Whatever you pick, the flow it generates is a symmetry of the symplectic structure. Now the three cases that matter.
7.2 · The three canonical examples
The Hamiltonian generates time translation. Take and . Then (1.3.56) is for any without explicit time dependence, by (1.3.50). So the flow generated by is the dynamics itself. Evolving the system forward by and applying the transformation generated by are the same operation. The Hamiltonian is not merely the energy. It is the generator of translations in time.
Momentum generates spatial translation. Take , the momentum conjugate to , and work out its bracket with an arbitrary observable. Every term of (1.3.49) except one vanishes, because depends on nothing else:
And is precisely the first-order change in when you move the system a distance in the -direction, which is Chapter 0.1's linear approximation. So generates translations along . In particular and , so the momentum shifts nothing and the position shifts by , which is exactly what a translation does.
Angular momentum generates rotation. Take . This one has four relevant coordinates rather than one, so compute the brackets with all four directly from (1.3.49):
Feeding those into (1.3.56) gives , , and identically , . Written out as a matrix acting on the position, the first pair is
which is a rotation by the angle about the -axis, to first order, and the momenta rotate the same way, as vectors must. So generates rotations about . Nothing was assumed about the system. This is a statement about the function and the bracket, valid for any Hamiltonian whatsoever.
7.3 · The two-way street
Now put the pieces together, and note that everything needed is antisymmetry.
Let have no explicit time dependence. Then
are the same quantity up to sign. Therefore
Read it left to right and it is the converse of Noether's theorem. Read it right to left and it is Noether's theorem. And read as a single line it says something stronger than either. The conserved quantity is not merely associated with the symmetry, and it is not a by-product of it. The conserved quantity is the generator of the symmetry. They are one object. Momentum is spatial translation. Angular momentum is rotation. Energy is time translation. The bracket is simultaneously "the rate of change of " and "the change in under ", and there is nothing else to say.
This is the single most portable idea in the book, and here is where it goes.
- Chapter 1.4 proves Noether's theorem for a general Lagrangian, including the field-theory version where the conserved quantity becomes a conserved current. This section is the Hamiltonian shadow of it, and it is the direction Noether does not give you.
- Chapter 4.2 repeats every word with operators. There, an observable generates the one-parameter family of unitary transformations , and differentiating that at gives , which is (1.3.56) under §6.4's substitution. "Observables generate unitary transformations" is this section, quantised. It is why , the generator of translations, and why appears in the exponent of the time-evolution operator.
- Chapters 6.1 and 6.2 take the set of generators seriously as an object in its own right. The brackets of generators with each other close into an algebra. Problem 2 computes and finds arriving classically, and a Lie algebra with its representations is what a "symmetry group" concretely is. Chapter 6.3 then makes the parameter depend on position, discovers that this forces a new field into existence, and that field is the photon. The entire gauge principle is this section with .
Any function on this space, not only the energy, can be put into the equations of motion in the energy's place, and doing so makes it push everything around in a definite way. The energy pushes the system forward in time, which is what the equations of motion said all along, so it is not merely the energy: it generates a translation of the clock.
Do the same with the momentum along a direction and everything slides a little that way. Do it with the angular momentum about an axis and everything turns about that axis, positions and momenta alike, as vectors should. Nothing was assumed about the system anywhere in this, and it collects a promise made when exponentials of self-partnered maps first appeared: a generator and the family it builds are two aspects of one object.
Read the combination of a quantity with the energy in both directions. Left to right it is the rate at which that quantity changes, so its vanishing means the quantity is conserved. Right to left it is the change in the energy under the flow that quantity generates, so its vanishing means the energy is untouched by it. They are one expression up to a sign — so a conserved quantity is not accompanied by a symmetry and not emitted by one, but is the symmetry, seen as something that moves things.
8 · Canonical transformations and Hamilton–Jacobi
This section is a signpost rather than a development. Two ideas, stated precisely enough to be usable, with the machinery left for the chapters that need it.
8.1 · Which changes of variable are allowed
Chapter 1.2 §7 proved that the Euler–Lagrange equations keep their form under any change of coordinates. Hamilton's equations are more demanding, because they involve the momenta, and the momenta are not free to be relabelled independently. They were defined by (1.3.11).
So we need a test for which relabellings are legitimate, and §6.2 has already supplied the thing to test. Call a transformation canonical if it preserves the fundamental brackets,
with the brackets still computed with respect to the old variables. There is a second way to say the same thing, in the notation of §3.3. If is the Jacobian matrix of the transformation, then (1.3.61) is precisely the statement
which is the definition of a symplectic matrix. Such transformations preserve Hamilton's equations, the new equations being , for a suitable . They preserve every Poisson bracket whatsoever, since (1.3.49) is built from . And taking determinants of (1.3.62) gives . ⚑ The sharper statement that always, never , is quoted here. It follows from a Pfaffian argument.
So canonical transformations preserve phase-space volume, and Liouville's theorem is the special case in which the transformation is "evolve for a time ". The flow map is canonical, which Problem 3 asks you to check.
Generating functions, in one paragraph. Knowing which transformations are allowed is not the same as knowing how to build one, and there is a practical recipe. Notice that the phase-space action (1.3.26) may differ between old and new variables by a total time derivative without changing the equations of motion (Chapter 1.2, Problem 4). So demand
and let . Expanding by the chain rule and matching the coefficients of the independent quantities , and gives
Any you write down generates a canonical transformation. Three cousins , and follow by Legendre-transforming in one or both arguments, which is §1 again in yet another costume. The grind box works the most useful one.
8.2 · Hamilton–Jacobi
Push the idea to its limit. Suppose you could find a canonical transformation making the new Hamiltonian identically zero. Then and , so all new variables are constants and the system is completely solved.
The price is one equation for the generating function that does it. From (1.3.64), with the generating function called instead of , means with substituted into :
This is the Hamilton–Jacobi equation: one first-order partial differential equation for one function , entirely equivalent to the ordinary differential equations we started with. It is rarely the easiest way to solve a mechanics problem. It is on this page for two reasons.
First, is the action. Not a function that resembles the action. The action. To see it, compute its total time derivative along a trajectory, using and (1.3.65):
so along the path. The generating function that trivialises the dynamics is Chapter 1.2's action regarded as a function of the endpoint. That is why the letter is the same, and it closes a loop opened two chapters ago.
Second, it is the classical limit of the Schrödinger equation. ⚑ Quoted, with the derivation deferred to Chapter 4.10. The way to see it is to feed the wave equation a wave whose phase is , so substitute the ansatz
into and separate the real and imaginary parts. The real part, at leading order in , is
which is exactly (1.3.65) for with , and the terms that were dropped are down by one power of . (The imaginary part gives , which is Chapter 0.7's continuity equation for probability, with velocity .)
So the phase of a quantum wavefunction is the classical action divided by , which is also what Chapter 1.2 §5 quoted about the path integral weight , and what Chapter 5.6 will prove. Classical mechanics is the geometry of surfaces of constant phase, in exactly the way that ray optics is the geometry of surfaces of constant phase of a light wave. Hamilton knew this in 1834, which is ninety-two years before anyone wrote down a wave equation for matter.
Grind box — the generating function, and the oscillator solved by Hamilton–Jacobi
The transform between generating functions. is awkward because the identity transformation is not expressible in it. Legendre-transform the argument: set , treating as the conjugate slope. Then , giving
Now gives and , the identity, as required. This is the version used in practice, and the one Hamilton–Jacobi uses.
The oscillator, solved from scratch. Take . Since has no explicit , look for a separated solution , so that and (1.3.65) becomes
So , and . Take the constant to be the new momentum . Then the new coordinate is
and is constant because . Setting and solving for :
The oscillator solution, with the amplitude expressed through the energy exactly as Chapter 0.8 had it. Note what the method did. It turned solving a differential equation into evaluating an integral, and the integral was the one whose derivative is the period, (1.3.41) again, since .
And the action variable. Instead of , take the new momentum to be . Then and Hamilton's equations give and : an angle advancing uniformly and a constant. Every one-dimensional bound system can be brought to this form, which is the content of "action–angle variables" and the starting point for perturbation theory in celestial mechanics. It is also, ⚑ quoted forward, the starting point for the KAM theorem, which says what survives when you switch a small perturbation on and is the reason the solar system's stability is a hard question rather than an obvious one.
Not every relabelling is permitted here, and the restriction is real. The equations survived any smooth change of coordinates because the momenta were free to follow; once the momenta are independent variables, only changes preserving the pairing between a coordinate and its partner leave the equations alone. Those changes preserve areas and volumes, and evolving the system forward is itself one of them.
Push the freedom to its limit and ask for a relabelling making the new energy identically zero. Every new variable is then constant and the system solved, the price being one partial differential equation for the function generating the change. That function does not merely resemble the number attached to each history. It is that number, regarded as depending on where the path ends, which closes a loop opened two chapters ago.
One remark makes this part look different in retrospect. Feed a wave whose phase is that number divided by Planck's constant into the equation of quantum mechanics, and the leading term is that same partial differential equation, the discarded terms smaller by one factor of the constant. Nearly everything is an approximation, and this is one: an expansion in a named small quantity, with a known first discarded term. Classical mechanics is the geometry of surfaces of constant phase, as ray optics is, and Hamilton had that structure ninety-two years before anyone wrote a wave equation for matter.
9 · Worked examples
Take the Lagrangian (1.3.12) for a particle of mass and charge in given potentials and :
(a) It is the right Lagrangian. Before transforming it, check it. The canonical momentum is (1.3.13), , so the Euler–Lagrange equation reads
We want an equation for alone, and the left-hand side is carrying an extra term. The total time derivative of along the trajectory is (chain rule, Chapter 0.6). Move it to the right:
The first bracket is the electric field. The second is the antisymmetric part of the Jacobian of , which Chapter 0.7 §4.3 identified with the curl. Writing , the identity holds component by component, so the whole right-hand side collapses to
the Lorentz force. ✓ (Verified symbolically for all three components.) The Lagrangian is correct, and note the structure: the term is linear in velocity, so it produces a force linear in velocity, which is what a magnetic force is.
(b) The Hamiltonian. By §2.3's grind box, . But must be expressed in the momenta, and from (1.3.13) we have . Substituting that in,
Look at what happened. The vector potential vanished from when written in velocities and reappeared inside a square when written in momenta. The magnetic field does no work, which is the first statement, and yet it is not absent from the dynamics, which is the second. Both are true, and the Hamiltonian formalism is what makes them compatible.
(c) It reproduces the same physics. Hamilton's equations give , which is again, and , which after substituting the first equation reduces to the Lorentz force. (Verified symbolically. The residual is exactly zero once is used.)
(d) Why this box exists. The boxed Hamiltonian is the object Chapter 6.3 quantises. The rule there is called minimal coupling, and it will be stated like this: to couple a charged particle to electromagnetism, replace
everywhere in the free Hamiltonian, and nothing else. Chapter 6.3 derives that rule from the requirement that the theory be invariant under a phase rotation whose parameter varies from point to point, and then repeats the derivation with the phase replaced by an matrix to obtain the strong force. All of it is the substitution above. You have now derived that substitution, from a Legendre transform, in classical mechanics, with no quantum mechanics and no gauge theory in sight. When Chapter 6.3 tells you that the covariant derivative is forced, you will recognise it as this worked example wearing indices.
Take (1.3.36) with and , so . We work out the complete structure of the phase portrait, which is the thing the interactive draws underneath the blob.
(a) Fixed points. and require and or (modulo ). Two fixed points per period, and no others.
(b) Their character. To classify a fixed point we linearise the flow at it. Writing the state as and , the Jacobian of is
Its trace is zero everywhere, which is Liouville's theorem (1.3.45) seen locally, since the trace of the Jacobian is the divergence. Its determinant is , so the eigenvalues satisfy :
| Fixed point | Eigenvalues | Type | |
|---|---|---|---|
| (hanging) | centre — closed orbits | ||
| (inverted) | saddle — hyperbolic |
Because the trace vanishes, the eigenvalues always come in pairs. So a Hamiltonian system in one degree of freedom can only have centres and saddles, never spirals or nodes. Spiralling in would contract volume, and Liouville forbids it. Chapter 0.8's classification of damped oscillators had four cases, and here two of them are illegal.
(c) The separatrix energy, exactly. The saddle sits at , so its energy is . Restoring units, , which is (1.3.37), and which is exactly the potential energy of a bob lifted from the bottom to the top, namely . The separatrix is the level set through the saddle, given by (1.3.38) as , which is the orange curve in the figure. Its maximum height is at , so the pendulum needs exactly the angular velocity at the bottom to just reach the top.
(d) Time on the separatrix. On it, , which separates and integrates:
which diverges logarithmically as . Inverting, , the famous homoclinic orbit. It leaves the inverted position at and returns to it at , taking infinite time at both ends and a finite time in between. It is a single trajectory whose past and future limit points are the same point, which is why it can appear to "cross itself" at the saddle without violating §4.1.
(e) Why this makes the interactive violent. Just inside the separatrix, the period is in the elliptic-integral notation of Chapter 0.8, and diverges logarithmically as its argument approaches . At : at , at , at , and at . A blob of radius centred at spans momenta from to , hence energies from to , hence periods from to , a spread of 45%. After ten circuits the fast edge has lapped the slow edge several times, and the blob is a filament. Its area, by (1.3.46), has not moved.
10 · Your turn
Problem 1 — Legendre-transform a Lagrangian, and check the involution
A particle of mass moves in a plane under a central potential . In polar coordinates, (Chapter 1.2 §7.1). (a) Find both canonical momenta and construct . (b) Write down Hamilton's equations and identify the conserved quantities without solving anything. (c) Verify explicitly that Legendre-transforming back with respect to the momenta returns .
Solution
(a) The momenta are
Note that is the angular momentum, and that it is not . It carries an extra factor of , which is §2.1's warning made concrete. Inverting, and . Both hypotheses of §2.3 hold, since is a homogeneous quadratic and is velocity-independent, so . Doing it by the definition anyway,
(b) Differentiating that in each of its four slots gives Hamilton's equations:
Two conservation laws can be read straight off. The coordinate does not appear in , so is constant, which is angular momentum. And has no explicit , so is constant, which is energy. That is two of the four constants of the motion, obtained by inspection rather than integration, which is the practical advantage of the Hamiltonian form.
Substituting the constant into the equations leaves a one-degree-of-freedom problem with effective potential . That is the centrifugal barrier, appearing here as an honest term in rather than a fictitious force. With this is the Kepler problem, and the phase portrait in has exactly the structure of §4.3: a centre at the circular orbit, closed curves around it for bound elliptical orbits, and open curves for hyperbolic ones.
(c) The involution. Transform back with respect to . The conjugate slopes are what we differentiate for, so take them first:
reproducing (1.3.6), in that the transform of has the velocities as its slopes, exactly mirroring . With the slopes in hand the transform is one substitution:
Which is (1.3.8), verified on a real example. The formula and the formula are the same formula, and the apparent asymmetry between them is entirely in which variables you regard as independent.
Problem 2 — the angular momentum algebra, arriving three parts early
With , so that and cyclically, show that
where is for cyclic index orders, for anticyclic, and if any two indices agree. Then show , and say what both results become in Chapter 4.11.
Solution
The computation. Take , . The rest follow by cycling , which leaves the definitions invariant. The bracket needs gradients, so write down the four of them for and :
Now assemble (1.3.49) term by term over the three coordinates:
Only the column survived, which is worth noticing. The bracket of two components of angular momentum "about" and is built entirely from the third direction. Cycling gives and , and antisymmetry supplies the rest, so in every case, including . (Verified symbolically.)
The Casimir. For the second part we need no new brackets, only the ones just computed. Use bilinearity and Leibniz on :
The magnitude of the angular momentum commutes with every component, while no two components commute with each other.
What this is. The three functions close under the bracket, meaning the bracket of any two is a linear combination of the three. A set of generators with that property is a Lie algebra, and this particular one, three generators with structure constants , is , the same algebra as . You have just derived it inside classical mechanics, with no rotation matrices and no quantum mechanics.
By §6.4's substitution it becomes and , the two relations from which Chapters 4.11 and 4.12 derive, with no further input, that angular momentum is quantised, that its magnitude takes the values , that can be a half-integer, and hence that spin exists. Note what §7.2 already told you about these generators: they generate rotations. So "angular momentum is quantised" and "the rotation group has discrete representations" are the same fact, which is what Chapter 6.2 is about. Everything in that sentence was visible here, classically, as one bracket.
Problem 3 — canonical transformations have unit Jacobian
(a) For one degree of freedom, show that a transformation satisfies if and only if its Jacobian determinant is , and explain how this makes Liouville's theorem a corollary. (b) Verify it on a concrete case: show that
is canonical, and find what it does to the oscillator Hamiltonian.
Solution
(a) With one degree of freedom the Jacobian matrix of the transformation is
Now write out the bracket the definition (1.3.49) asks for, and compare the two expressions:
They are literally the same expression. So , with no computation at all. In degrees of freedom the statement is (1.3.62), . Taking determinants and using gives , and the sharper is the fact quoted in §8.1.
Liouville as a corollary. Fix a time and consider the map that sends each initial condition to its state at time . This map is canonical. The brackets computed with respect to are preserved by the flow, because , which is zero by the Jacobi identity (1.3.53) applied to , and together with constant. Hence , and by Chapter 0.6's change-of-variables formula the map preserves areas. That is Liouville's theorem again, derived algebraically rather than through the divergence. The two proofs are §5 and this one, and it is worth having both, because §5's generalises to any divergence-free flow while this one explains why the phase-space structure, and not merely the volume, is what is being conserved.
(b) Compute the four partials, writing , and , , so , , and :
using and . Assembling them into the determinant,
(This is the Jacobian of the inverse map. Since , the forward Jacobian is too.) The payoff. Substitute into the oscillator Hamiltonian:
The Hamiltonian has become a single term with no in it. Hamilton's equations are therefore and , whose solution is and constant . Substituting back gives . These are the action–angle variables of §4.4's grind box, with the action and the angle around the ellipse. The whole solution of the oscillator has been reduced to choosing good coordinates, and the reason we were allowed to choose them is .
Problem 4 — the area of an oscillator orbit, and what quantum mechanics does to it
For , compute the phase-space area enclosed by the orbit of energy directly from the integral , and show it equals . Confirm it against (1.3.41). Then apply the Bohr–Sommerfeld condition and say precisely what it predicts.
Solution
The integral. Solving for gives , with turning points at , . Going round the loop, the upper branch contributes and the lower contributes the same again with both signs reversed, so
using . The remaining integral is the area of a half-disc of radius , namely (substitute if you want it by hand: ). Putting that value in,
Same answer as the ellipse-area route of Chapter 0.8 §4.4, which is times the product of the semi-axes and , as it must be. Numerically, at , , : the semi-axis product route gives and gives .
The check. , which is the period . ✓ (1.3.41) holds, and the fact that is exactly linear in is the statement that the oscillator's period does not depend on its amplitude, the property that makes it the one system everything is expanded around.
Bohr–Sommerfeld. Impose (1.3.42), , on the area we have just computed:
Three things to notice, all of which Chapter 4.8 will confirm by an exact operator calculation that uses none of this reasoning.
The levels are equally spaced, by . That follows from (1.3.41), since equal increments of area correspond to equal increments of energy exactly when the period is energy-independent. A pendulum at large amplitude, whose period grows with energy, has levels that converge. And indeed real molecular vibrations, which are anharmonic, have converging levels that eventually run into the dissociation limit. You can read that off the phase portrait without solving anything.
The ground state is not zero. , the zero-point energy, enclosing half a unit of . A classical oscillator can sit at the origin of phase space, a single point of zero area. A quantum one cannot, because the smallest allowed area is . Chapter 4.9 says the same thing as . The uncertainty principle is the statement that a state occupies a minimum area of phase space, and that minimum is Planck's constant. Every "why can't the electron just fall into the nucleus" question is answered by this picture.
The classical limit is visible. For large the spacing is negligible compared with , the rungs of the ladder crowd together, and the discrete set of ellipses is indistinguishable from the continuum. Numerically, an oscillator with carrying sits at . This is why nobody noticed.
You traded second-order equations for first-order ones, by a Legendre transform you now understand geometrically. It is the dictionary between describing a convex curve by its points and describing it by its tangent lines, the same operation that turns internal energy into free energy. Out of it came three things. First, the canonical momentum, which is not in general: in a magnetic field it is , and it is that object rather than the mechanical momentum that quantum mechanics turns into an operator. Second, the Hamiltonian, which equals only under two hypotheses you can state. Third, Hamilton's equations, derived twice.
You then met phase space, in which one point fixes the whole future, trajectories cannot cross, and the pendulum's entire qualitative behaviour is a centre, a saddle, and the separatrix joining them. You proved Liouville's theorem in three lines from Chapter 0.7's identity of divergence with the trace of the Jacobian. Hamiltonian flow is exactly incompressible, so information is stirred and never destroyed. You saw it in the figure, where a blob became unrecognisable while its area sat there to within a few parts in ten thousand of integrator error. And you saw that the arrow of time is therefore in the coarse-graining rather than in the dynamics.
And you built the Poisson bracket: bilinear, antisymmetric, Leibniz, Jacobi. With it came for every observable there is, the fundamental relation , and the two-way street of §7, in which a conserved quantity does not merely accompany a symmetry but is the generator of it.
Where this gets spent. The bracket goes to Chapter 4.9, which replaces it by and changes nothing else, so that becomes and (1.3.50) becomes the Heisenberg equation. It goes also to Chapter 5.3, which does the same substitution for a field and gets particles. Phase space and Liouville go to Chapter 4.6, where unitarity is the same statement in a Hilbert space, and to statistical mechanics, which cannot define an ensemble without them. The generators of §7 go to Chapter 1.4 (Noether, from the other side), Chapter 4.2 (observables generate unitaries), and Chapter 6.1, where the brackets of generators with each other become a Lie algebra and the whole of Part VI. The canonical momentum goes to Chapter 6.3, where is derived from a symmetry principle and named minimal coupling. And Hamilton–Jacobi goes to Chapter 4.10 as the classical limit of the Schrödinger equation, the last stop before the wavefunction.
Chapter 1.4 now closes Part I by proving Noether's theorem in the Lagrangian language, where it applies to fields and to relativity, and where it will be the tool that constructs the conserved currents of every theory in this book.