Part IV · Quantum Mechanics — Chapter 4.9
Commutators, Uncertainty, and Symmetry
The uncertainty relation was proved in Part 0, with no physics anywhere in it, and it has been waiting three parts for one substitution. This chapter makes the substitution in a line. Then it spends the rest of its length on the fact that the same commutator which bounds a product of spreads is also what moves every observable in time and what generates every continuous symmetry. Very little here is new. Almost all of it is one object, met from three sides.
Six results built elsewhere meet in this chapter, and all six were built with this meeting in mind. Chapter 0.9 §6.4 proved that a function and its Fourier transform cannot both be narrow, and left one substitution for quantum mechanics to add and nothing else. Chapter 0.5 §1.4 proved the Cauchy–Schwarz inequality and its insight box wrote down, three parts early, the two vectors this chapter would feed into it. Chapter 4.2 §5.3 wrote the variance of an observable as the squared length of a vector, which is the form Cauchy–Schwarz can act on. Chapter 1.3 §6.1 derived the classical equation of motion for an arbitrary observable, and §6.4 announced the substitution that would turn it into the quantum one. Chapter 1.4 §7 showed that a conserved charge generates the symmetry it came from. Chapter 4.6 §2 produced the operator that moves a state forward in time.
Nothing on that list needs to be built again. What this chapter does is spend it, and the shape of the spending is worth having before it starts, because it is the same object doing three jobs. The commutator bounds how sharp two quantities can be at once. The commutator with the Hamiltonian moves any observable in time. And a commutator with a generator produces the symmetry that generator belongs to. Those are not three facts that happen to involve the same bracket. They are one fact seen from three sides, and by the end of §6 you should be able to say which side you are looking at from the shape of the equation alone.
Here is the route. Section 1 is the shortest in the chapter, and one multiplication pays four of the promises Chapter 0.9 made, with the remaining two settled in §2 and §2.6. Its result is one you have met: Chapter 4.6 §10.2 wrote it down for one Gaussian packet and Chapter 4.8 §6.1 sat the oscillator ground state exactly on it. What §1 adds is the reason the multiplication is legitimate for every state at once. Section 2 does the general case in three steps, using Chapter 0.5's inequality on Chapter 4.2's two vectors, and then spends a careful page on what the result does not say, which is where most of the damage in popular accounts is done. Section 3 answers the question Chapter 4.2 §4.3 explicitly left for this chapter: how you know that a set of commuting observables is complete. Section 4 moves the time dependence off the state and onto the operator, shows that this is a change of basis with a moving domain, and differentiates to get the Heisenberg equation, which is set beside Chapter 1.3's classical equation term by term. Section 5 takes expectation values of that equation and gets Ehrenfest's relations, which Chapter 4.6 §8.7 stated and deferred, together with the caveat that decides how much they are worth. Section 6 runs the symmetry chain in full for translation, rotation and time, and hands Chapter 4.11 the one commutator its entire chapter is built on. It then adds a symmetry that leaves a label rather than a generator, and closes by naming the one thing the commutator has not told you: what becomes of any of it when the scales grow large, which is Chapter 4.10's subject and not this chapter's. Sections 7 and 8 are worked examples and problems.
Conventions. always means the standard deviation of the outcomes of measuring on a large number of systems prepared identically in the state , and never anything about a single measurement. Hats stay on operators and come off their values, with and the one pair that needs care, sorted out where they meet in §2.3. Time evolution is with the minus sign, as Chapter 4.6 §2.4 fixed it. One subject is quoted rather than derived, and it carries the chapter's single mark, in §2.6: measurement disturbance, whose two named theorems are about apparatus rather than about states, and which this book states with their hypotheses and does not prove. The one mark covers both of them, and there are no others. Three marks standing in earlier chapters are leaned on and cited rather than raised again, and the closing brick says which.
Tools you'll need — Chapter 0.5 above all: §1.4 for Cauchy–Schwarz and its equality condition, §7.1 for the exponential of times a Hermitian operator, and §8 for commuting operators and simultaneous diagonalisation. Chapter 0.9 §6.4 for the bandwidth theorem and §6.5 for the Gaussian that saturates it. Chapter 4.2 §4.3 for compatible observables and quantum numbers, §5.3 for the variance written as a length, §7.2 and §7.3 for the argument from unitarity to a Hermitian generator, §7.5 for the first-order action of a generator, §8 for the canonical commutator, and §9 for the tensor product. Chapter 4.6 §2 for and Stone's theorem in the form used here, §5 for momentum in the position representation, §8.7 for the Ehrenfest relations as stated there, and §9.2 for what a stationary state is. Chapter 1.3 §2.1 for what a conjugate momentum is, §6.1 for the classical equation of motion, and §7 for the flow an observable generates. Chapter 1.4 §3 for the three classical symmetries and §7.3 for the algebra of the rotation charges. Chapter 0.4 §4 for change of basis and the similarity transformation. Chapter 4.4 §3 for domains and §5.4 for momentum on a ring. Chapter 4.5 §6 for the projection-valued measure and §9 for Stone. Chapter 4.3 §8.1 for the density of continuous functions in .
1 · One substitution:
You have seen the answer before you have seen this section. Chapter 4.6 §10.2 multiplied one Gaussian packet's two widths together and got , and Chapter 4.8 §6.1 sat the oscillator ground state on the same number. Both were particular states, taken one shape at a time. What is still owed is the bound for every state, and it costs one multiplication, because Chapter 0.9 did the hard part with no physics in it at all.
1.1 · What was proved in Part 0, and what was missing from it
Chapter 0.9 §6.4 proved the bandwidth theorem: for any normalised with finite widths, the spread of across positions and the spread of its Fourier transform across wavenumbers satisfy . The proof used the definition of the transform, Plancherel, the derivative theorem, Cauchy–Schwarz and one integration by parts, and Chapter 0.9 listed the ingredients in a bulleted list precisely because the shortness of the list was the point. No particle appears in it, no measurement, no observer, and no . A radio engineer shaping a pulse obeys it, and so does a seismic trace.
What was missing was a reason to care, in physics. The theorem constrains a wave, and a wave has a wavenumber. Nothing in Part 0 said that a particle has one.
1.2 · The substitution, and the line it takes
Chapter 4.6 supplied the missing sentence, and supplied it twice over. Its §5.5 showed that the map between the position and momentum representations of a state is Chapter 0.9's Fourier transform with , and that the map is unitary, so a state normalised in one description is normalised in the other. The density of the measured momentum is therefore the density of the wavenumber re-expressed under . Its §10.6 marked de Broglie's relation as the one experimental input of that chapter, and that mark stands there rather than being raised again here. Those are the two facts §10.2 used on its Gaussian, and neither of them mentions the Gaussian. Since is linear, and a linear change of variable multiplies a standard deviation by its slope, we have for any state with both widths finite. So multiply Chapter 0.9's inequality through by :
That is the whole of it. The inequality was never quantum. What is quantum, and what is strange, is the claim that a particle's momentum is the wavenumber of a wave. Grant that one claim and the Heisenberg uncertainty principle for position and momentum stops being an additional mystery and becomes arithmetic you did in Part 0.
Now read the display as a prohibition with a scale on it, because the scale is what makes it a statement about the world. Confine an electron to a region the width of an atom, , and (4.9.1) forces . That is a spread in speed of , and an energy of order . The electronvolt is the scale of atomic binding energies, and this is where that scale comes from: confinement to the size of an atom, priced by (4.9.1). Chapter 4.7 §3.5 read a confined particle's ground-state energy in exactly that way and got the same order back.
It is worth saying what has not been proved by that line, because §2 exists to prove it. Equation (4.9.1) is a statement about one pair of quantities, obtained through a Fourier transform, and the Fourier transform was available only because position and momentum are related in that particular way. Two arbitrary observables are not each other's transform, and for them the argument above says nothing at all.
1.3 · Every conjugate pair, and where the dimensions came from
Before leaving, one more paragraph, because it removes a puzzle that would otherwise sit in your way for the rest of Part IV. Chapter 1.3 §2.1 was careful to say that the canonical momentum is not , and that for the pendulum with coordinate it comes out as , an angular momentum measured in rather than in . It then said what really is: whatever pairs with so that has the dimensions of action.
Chapter 4.1 §5.7 found that the new constant of quantum mechanics has the dimensions of action, and Chapter 4.6 §2.2 spent that fact to fix the in the exponent. Put the two together and the bound is dimensionally consistent for every conjugate pair and not only for a Cartesian one, because both sides carry an action either way. An angle paired with an angular momentum works as well as a length paired with a linear momentum.
Whether the inequality holds for those other pairs is a separate question, and a more delicate one than it looks. The angle is the case where the general theorem's fine print does real work, and Worked example 2 is where that is shown rather than described.
A statement that has been called the deepest in twentieth-century physics has just been obtained by multiplying an inequality by a constant. That is not a trick, and it is worth being clear about why it is honest.
The toolkit chapter on Fourier analysis proved that a signal cannot be both brief and pure in pitch. Squeeze it in time and its spectrum widens by the reciprocal factor, always, with the product of the two spreads never falling below one half. That result is about waves. It was known to people designing radios before anyone applied it to matter, and nothing in its proof mentions particles or measurement or observation.
Quantum mechanics contributes exactly one sentence to the story: a particle's momentum is the wavenumber of a wave, multiplied by a constant with the units of action. That sentence is a physical claim with an experiment behind it, and the earlier chapter marked it as such. Accept it, multiply, and the famous inequality drops out. The mystery, if you want one, is entirely in the sentence. It is not in the inequality.
One small piece of tidying comes free, and it is worth being exact about how much of it is free. The constant has the units of action, and the classical chapters had already established that a coordinate and the momentum belonging to it always multiply together to give an action, whatever the coordinate happens to be. So the bound written down for an angle and its angular momentum at least makes dimensional sense, in the same form, with the same constant and no conversion factor anywhere. That much is not luck. It is what "conjugate" was defined to mean.
Whether the inequality is then true for such a pair is a separate question, and the dimensions do not settle it. For the angle the answer is no, until the statement is repaired. The next section says which step breaks, and Worked example 2 exhibits the failure and mends it.
2 · The general relation, and what it is not
Here is where this section is going. We want the same bound for an arbitrary pair of observables, with the right-hand side telling us which pair we are dealing with. The whole derivation is Chapter 0.5's Cauchy–Schwarz applied to two vectors that Chapter 4.2 has already built, and it takes three steps. Then we spend rather longer on what the result means, because the inequality is short and the number of things it is routinely taken to say is not.
2.1 · The two vectors, already in hand
Chapter 4.2 §5.3 derived the variance of an observable from the Born rule and then wrote it in a second form:
It said in place why the second form was the one this chapter would need. The spread of an observable is the length of a vector, and Cauchy–Schwarz is a statement about lengths. So name the two vectors, which is the only preparation the argument requires:
By (4.9.2) these have lengths and . Chapter 0.5's insight box wrote down this pair three parts ago and said that nothing would be added to it here except the meaning of the symbols.
2.2 · Cauchy–Schwarz, once
The inequality of Chapter 0.5 §1.4 says for any two vectors of an inner-product space. Apply it to (4.9.3) and take the square root of both sides, which is legitimate because both sides are non-negative:
The uncertainty relation is now proved. Everything remaining is the identification of the right-hand side, which is where the commutator will appear, and it appears because we are about to ask which part of survives when the two observables are swapped.
2.3 · Splitting the overlap into two real numbers
We want expressed in the observables rather than in the vectors, so move the first operator across the inner product using its Hermiticity. Writing and , both real by Chapter 4.2 §4.1, and abbreviating and ,
Two symbols differing by one hat are now in play, and they are not each other's value, so fix them apart before they appear in the same display. is the operator , and its expectation is zero by construction. is a length, the norm of , and it is the standard deviation the Conventions paragraph fixed. The hat here separates an operator from a number built out of it, not from its own expectation.
That product of two Hermitian operators is not itself Hermitian, and the standard repair is to split it into a part that is and a part that is not. Any product splits that way, by adding and subtracting half of the reversed product:
where is the anticommutator, and the constants and have dropped out of the commutator because a number commutes with everything. Now read off the character of each piece. The anticommutator of two Hermitian operators is Hermitian, since , so its expectation is real. The commutator is anti-Hermitian, since , so its expectation is purely imaginary. The two terms of (4.9.6) are therefore the real and imaginary parts of one complex number, and a complex number's modulus squared is the sum of the squares of its parts:
Both terms are non-negative, so dropping the first can only weaken the statement. Drop it, feed what is left into (4.9.4), and take the square root:
Three steps, and the third of them was arithmetic on a complex number. Set and and use Chapter 4.2 §8's : the right-hand side becomes in every state whatsoever, and (4.9.1) comes back. The two derivations share no step. Section 1 went through the Fourier transform and de Broglie's relation, this one through Cauchy–Schwarz and one postulated commutator, and they agree. That agreement is worth more than either route alone.
2.4 · The sharper form, which cost nothing
Look again at what was thrown away. Keeping the anticommutator term in (4.9.7) gives a stronger inequality for the same work:
The discarded quantity is the covariance of the two observables, in exactly the sense used of any two random quantities, symmetrised because the operators do not commute. So (4.9.8) is what you get by pretending the two are uncorrelated when they need not be. For position and momentum in a state with no correlation between them the two forms agree, and Problem 1 exhibits a perfectly ordinary Gaussian in which they do not. The geometry of that difference is drawn in the figure at the end of this section, with two caveats and a pair of numerical checks standing between here and there.
2.5 · The hypothesis, which is not decoration
Every line above assumed it was legal to move an operator across an inner product. Chapter 4.4 spent a whole chapter on when that is legal, and the answer was not "always". Written out, the derivation needs to lie in the domain of and to lie in the domain of , so that (4.9.5) is an identity rather than a hope. Call it the domain hypothesis. For a bounded operator it costs nothing, since Chapter 4.4 §2.1 showed bounded and continuous to be the same condition and a continuous operator on a dense domain extends by continuity to the whole space. For unbounded ones it is a genuine condition.
Something the domain hypothesis quietly requires is easier to miss, because it looks like a formality and is not. Both sides of (4.9.8) assume that and exist, and a normalised state is under no obligation to supply them. Chapter 4.3 §5.5 gave the structural reason, that does not sit inside , and said the physical reading belonged in this chapter. Here it is. Normalisation is a statement about and guarantees nothing about , so a perfectly legitimate state can have no mean position and no at all, as that chapter's Worked example 3 exhibits. For such a state the relation is true and empty, in exactly the way Chapter 0.9 §6.5's grind box said the bandwidth theorem is true and empty when a width is infinite. Nothing has gone wrong. The inequality compares two numbers, and it says nothing in a state where one of them is not a number.
The domain hypothesis itself fails in the most quoted example after position and momentum. Take an angle and the angular momentum conjugate to it. The naive reading of §1.3 would give . That statement is false, and Worked example 2 exhibits a state in which the left-hand side is exactly zero, identifies the line of the derivation that breaks, and computes the boundary term the breakage leaves behind. The repair is to replace the angle by a bounded periodic function of it, after which the hypothesis holds and the theorem applies with no modification. Chapter 4.4 §5.4 built the operator involved, so the machinery for the repair is already on the shelf.
2.6 · What the relation does not say
The relation is about preparation, not about disturbance. Read (4.9.8) back through its own definitions. Each is a standard deviation of measured values over an ensemble of systems prepared identically in , with one measurement performed on each. No system in that ensemble is measured twice. Nothing in the derivation refers to an apparatus, to an order of measurements, or to any effect one measurement has on another, because no such notion appears anywhere in Chapters 0.5 or 4.2. The content is that the state cannot drive the product of the two spreads below what the commutator allows in that same state. It is a constraint on what can be prepared.
Chapter 0.9's warning box made the same point about the wave version and promised that this chapter would keep the two ideas apart. Here is the separation, stated as sharply as it can be. Measurement disturbance is real, it is a different quantity, and it has theorems of its own. Those theorems are not proved in this book, and this is the one place in this chapter where something is quoted.
Ozawa's relation (2003). Model a measurement as a unitary interaction between the system and a probe, followed by a sharp reading of a meter observable on the probe. Define the error as the root-mean-square difference, in the given input state, between what the meter reports and what would have given, and the disturbance as the root-mean-square change the interaction inflicts on . With those definitions the naive product is false, and there are explicit measurement models that violate it. What holds instead carries two extra terms, , in which and are the very spreads of (4.9.8). The hypotheses are the indirect-measurement model described above, the root-mean-square definitions of and , and enough regularity for the relevant expectations to exist. It is a state-dependent statement.
Read what the two extra terms buy, because that is where the naive product went wrong. They let the state's own spreads carry the bound. When and are wide, the cross terms and can reach between them, and is then free to be as small as the apparatus can make it. So an accurate measurement of need not inflict a large disturbance on , which is precisely what the naive product forbade and what the explicit models exhibit.
The Busch–Lahti–Werner relation (2013). Define error and disturbance differently, as worst-case figures of merit obtained by calibrating the apparatus against states in which the reference observable is arbitrarily sharp, rather than as averages in one input state. With those definitions, and for the canonical pair position and momentum, the product of the calibration error and the calibration disturbance is bounded below by after all. The hypotheses are the calibration definition of error and the pair being canonically conjugate.
The two are not in conflict, and the reason is worth carrying. They bound different quantities. "How accurate is this apparatus" admits more than one honest definition, and which inequality you get depends on which one you chose. The diagnostic, when you next meet a claim in the wild, is to ask which state the error was computed in. An error quoted for the one input state at hand is Ozawa's quantity, and there the naive product is false. An error quoted as a worst case over calibration states in which the reference observable is arbitrarily sharp is Busch–Lahti–Werner's, and there the naive product holds for position and momentum after all. Neither is (4.9.8), and neither is derivable from it, which is exactly why Chapter 0.9 said there were separate theorems and why naming none of them here would have been an under-delivery. What this book proves is the preparation statement. Both relations above are covered by this box's mark, and it is the only one in the chapter.
2.7 · The relation, measured
Two checks, neither of which uses the derivation above, and the first of which is the one to remember. A third, on angular momentum, waits until §6.3, where the commutator it needs has been built.
The oscillator. Chapter 4.8 obtains and for the energy eigenstates by algebra, giving an uncertainty product of . Computing the same two integrals numerically for the Hermite functions at forty-digit working precision returns for the ground state and for in units where , printed to twelve figures, and departing from and by less than , which is the quadrature's error at that working precision rather than a property of the states. Read those two numbers side by side. The ground state sits exactly on the floor set by (4.9.1), and the state sits seven times above it. The relation is a floor and not a prediction. It says how small the product cannot be, and it says nothing whatever about how large it is in any particular state. Chapter 4.7 §3.5 made the same point from the other end before this theorem existed, reading the infinite well's ground-state energy as an uncertainty and getting against a floor of . An instance computed before the theorem is worth more than the theorem alone, and both of those instances were.
Random pairs. Two hundred thousand random Hermitian pairs and random states in dimensions two to five, all built independently of the derivation, satisfy both (4.9.8) and (4.9.9) in every trial. The smallest slack in the sharper form (4.9.9) is , which is roundoff and means the sharper form is attained; the smallest slack in (4.9.8) over the same trials is . The identity (4.9.7), on which the whole split rests, holds over twenty thousand random cases with a worst residual of .
You report standard deviations for a living, and is one of yours. Prepare systems the same way, measure once on each, and compute the sample standard deviation of the numbers. That is , in the same arithmetic you would apply to tumour volumes or trough concentrations, and Chapter 4.2 §5.3 derived it from the Born rule rather than defining it. Nobody has ever measured a standard deviation on one patient, and nobody has ever measured on one system. This is why the inequality cannot be about what happens to an individual when you poke it.
Chapter 4.5 §6.5's familiar-ground box named the place where the parallel stops and left it for this chapter, so here it is. When you see two correlated readings vary across a cohort, you assume without thinking that there is a joint distribution underneath: each patient has some true pair of values, you are seeing a marginal of it, and the variation reflects covariates you did not measure. Every method you have for such data is built on that assumption, and it is nearly always right.
The word that will not carry across is covariate. Finding one does not narrow a marginal. It splits the cohort into strata and narrows the distribution inside each, and the marginal you started with is exactly where it was, because nobody has been intervened upon. On this side there is nothing to stratify and there is something to do. A preparation is an act, so the question is not what you might discover about the systems. It is whether some other preparation makes both spreads small at once.
For position and momentum, none does. The commutator is in every state whatsoever, so the right-hand side of (4.9.8) is however you prepare, and the product has a floor no state gets under. There is no experiment to design and no covariate to go looking for. That much is proved here.
Say it for this pair and not in general, because the floor is not always a number. For a general pair the right-hand side of (4.9.8) is evaluated in the state, and §6.3 exhibits two components of angular momentum in a state where it comes out zero while neither spread is small. There the theorem has fallen silent rather than been met, and what still forbids both from being sharp is the non-zero commutator of §3.1 rather than anything the inequality has supplied.
The stronger statement, that there is no joint distribution underneath at all, is not proved by this inequality and should not be claimed from it. Chapter 4.5 §6.5's familiar-ground box raises it and leaves it open, Chapter 4.11 shows that the failure to be simultaneously sharp is structural rather than a matter of ignorance, and Chapter 4.20 settles it with an experiment. What you may take from this chapter is the operational half, and it is the half that changes how you read the inequality: the spread is a property of the preparation, not a measure of what you have not yet found out.
The general statement takes three steps and it is worth watching how little each one costs. The spread of a measurement, defined honestly as a standard deviation over repeated identical preparations, was shown two chapters ago to be the length of a particular vector built from the state and the quantity being measured. So take two quantities, build the two vectors, and apply the inequality proved in the toolkit chapter, which says that the overlap of two vectors is never bigger than the product of their lengths. That single application is the uncertainty principle.
What remains is bookkeeping on the overlap. It is a complex number, and its two parts have different meanings: the real part measures how the two quantities co-vary, and the imaginary part measures how badly the two operations fail to commute. Since a complex number is at least as big as either part alone, throwing away the real part costs nothing and leaves the familiar form, with the failure to commute sitting on the right-hand side. Keep the real part instead and you get a slightly stronger statement for the same work.
Two warnings, both earned. The first is that this is a statement about preparation. Every symbol in it refers to a spread across many systems prepared identically and each measured once. It says nothing about an apparatus knocking a particle about, and the theorems that do say something about that are separate results with different hypotheses, named above and not proved here.
The second is that the bound is a floor rather than a forecast. The lowest oscillator state sits exactly on it. The fourth one sits seven times above it, and is no less legitimate for that. And when the right-hand side happens to vanish in some particular state, the theorem has not announced that both quantities are sharp. It has fallen silent.
a natural place to stop · the bound is proved and spent; what follows is the commutator when it vanishes, and then when it moves things
3 · Compatible observables, and a complete set
Section 2 asked what happens when a commutator is large. This section asks what happens when it is zero, and then answers a question Chapter 4.2 §4.3 raised, named this section for, and deliberately left open. The answer is a criterion, and the criterion is what Chapters 4.11 and 4.13 will apply every time they label a state.
3.1 · The two halves, now both proved
Chapter 0.5 §8 proved that two Hermitian operators admit a common orthonormal eigenbasis exactly when they commute, and Chapter 4.2 §4.3 renamed both halves. If the two are compatible: there are states in which both quantities have definite values, and enough of them to span the space. If there is no common eigenbasis, hence no state at all in which both are sharp.
That second half is the qualitative content of the uncertainty principle: a non-zero commutator rules out every state in which both are sharp, without saying by how much. Equation (4.9.8) is now the quantitative version, and the two fit together with one caution worth fixing now. A non-zero commutator forbids a common eigenbasis outright. The inequality, by contrast, is evaluated in a particular state, and can vanish in a state even when the operator does not. In that state the bound has fallen silent while the prohibition still stands, and §6.3 puts a number on exactly that case.
3.2 · The question Chapter 4.2 left here
Chapter 4.2 §4.3 gave the definition this part of the book runs on. A set of mutually commuting observables is a complete set of commuting observables when their common eigenspaces are all one-dimensional, so that the list of eigenvalues fixes the state up to phase. That list is what the word quantum numbers means, and writing a hydrogen state as is naming the eigenvalues of three commuting operators and nothing more.
What it left open was how one knows a set is complete. It is a fair question and it is not answered by the definition, because the definition quantifies over eigenspaces you would have to find first. What you want is a test you can run on the operators.
3.3 · Maximality is the same condition
Here is the test, and its usefulness is that it never mentions eigenspaces. One phrase in it does all the work, so fix its meaning before the statement arrives. A function of a commuting set means what Chapter 0.5 §7 meant for a single operator, extended from one operator to a family: choose a number for each tuple of joint eigenvalues, and let the operator multiply by that number on the corresponding common eigenspace, so that with the projection onto it. Nothing else counts as a function of the set.
Theorem. A set of mutually commuting observables on a finite-dimensional space has all its common eigenspaces one-dimensional if and only if every observable that commutes with all of them is a function of them.
A set with that property is called maximal: you cannot add a genuinely new commuting observable, because anything you might add is already there. Both directions are short, and both use the same fact about a common eigenspace.
One-dimensional common eigenspaces imply maximality. Let commute with every , and let be a common eigenvector with eigenvalues . Then , so sits in the same common eigenspace. That eigenspace is one-dimensional and contains , so is a multiple of . Every common eigenvector is therefore an eigenvector of , and since those eigenvectors form a basis by Chapter 0.5 §8, we may define on the tuples and write , which is a function of the in exactly the sense just fixed.
A common eigenspace of dimension two or more destroys maximality. Let be such an eigenspace and pick a unit vector inside it. Take . It commutes with each , because and, using Hermiticity, as well. But any function of the acts on all of as a single multiple of the identity, since carries one tuple of eigenvalues, whereas annihilates the part of orthogonal to , which is not empty. So commutes with everything in the set and is not a function of the set.
3.4 · How you show it in practice: exhibit the count
What §3.3 buys is not the working test. It makes the definition safe, by showing that the property you cannot check without finding eigenspaces is the same property as one you can state about the operators alone, so nothing turns on which of the two anyone happens to mean. The working test comes from the definition and the basis, and needs neither direction of the theorem. Since the joint eigenvectors form a basis, a set is complete exactly when the number of distinct eigenvalue tuples equals the dimension of the space. So you do not test maximality directly. You list the tuples and you count.
- Angular momentum. On a single multiplet of dimension , the set produces the tuples for , which is of them. The count matches and the set is complete. Chapter 4.11 constructs the multiplet and then makes exactly this remark.
- Hydrogen. The set produces, at fixed principal quantum number , the tuples with and , so the count is . That sum telescopes to , since consecutive squares differ by consecutive odd numbers. Chapter 4.13 computes the dimension of the -th energy level and finds , at which point the count matches and the set is complete; Chapter 4.14 explains why the degeneracy is that number and not something else. What belongs here is the criterion those chapters apply.
- A set that fails. Problem 4 builds two qubits and shows that the total of the two "which-state" observables is not by itself complete. It then exhibits an operator commuting with that total which is not a function of it, and completes the set in two different ways whose members do not commute with each other. That last point is the entire structure of Chapter 4.12's coupled and uncoupled bases, met early and with no angular momentum in it.
3.5 · What the criterion becomes in infinite dimensions
Name the difficulty rather than stepping over it. The proof in §3.3 used a basis of eigenvectors, and Chapter 4.5 §2 showed that an observable in infinite dimensions may have no eigenvectors at all. The replacement is the projection-valued measure of Chapter 4.5 §6. Commuting observables have commuting spectral measures, and the joint measure assigns a projection to each region of the joint spectrum. The set is complete when that joint measure has multiplicity one, meaning that no region can be split further by any projection commuting with all of them, which is the one-dimensional condition with "dimension" replaced by "cannot be cut".
This book needs the finite-dimensional statement and needs it repeatedly, because every use in Chapters 4.11 to 4.16 is a count inside a finite-dimensional eigenspace of . The infinite-dimensional version is stated so the definition is not silently changed later, and the momentum components of a free particle in three dimensions are the standard example: one joint generalised eigenstate per momentum vector, nothing left to cut, and Chapter 4.5 §7 is where such objects were given their meaning.
One separate caution is worth parking here before §4 begins, since it belongs to the criterion rather than to infinite dimensions and Chapter 4.2 §4.3 stated it with a figure to go with it. The theorem promises that a common eigenbasis exists. It does not promise that a numerical eigensolver hands you that one, because inside a degenerate eigenspace the solver returned some basis rather than the basis the second observable prefers. Chapter 4.13's degeneracies are exactly where that bites.
What you now hold is a test you can run on paper: list the tuples, count them, compare with the dimension. Chapters 4.11 and 4.13 run it on angular momentum and on hydrogen, and Problem 4 runs it on two qubits, on a set that passes and a set that fails. The chapter now changes subject from a single instant to time.
Two quantities that commute can be sharp together, and the states in which both are sharp are numerous enough to describe everything. That much was proved in the toolkit chapter. The question left open was practical: given a handful of such quantities, how do you know you have enough of them to tell every state apart?
The definition says you have enough when knowing all their values pins the state down. The useful reformulation says you have enough when the list cannot be extended: any further quantity that commutes with all of yours is already some combination of them, so adding it would tell you nothing new. The two conditions are the same condition, and the proof of that is half a page.
In practice nobody checks either one directly. You write down the possible combinations of values, you count them, and you compare that count with the number of independent states available. If the two numbers agree, the labels are enough. That is the whole method, and it is what lets an atomic state be written as three numbers in a bracket, those three labels being the whole of what distinguishes it from every other state.
4 · The Heisenberg picture, and the Heisenberg equation
Everything so far has been about a single instant. This section puts the time back in, and the first thing to settle is where the time is allowed to sit. The answer is that there is a choice, that the choice is a change of basis rather than a change of physics, and that making it produces an equation of motion for observables which is the exact quantum copy of the one Chapter 1.3 derived classically. That equation is what §5 and §6 both run on.
4.1 · Where the time can be put
Every prediction the theory makes is an expectation value. Chapter 4.2 §5.3 derived from the Born rule, and Chapter 4.2 §5 showed that a probability is such a number with a projection, so there is nothing else to compute. Take one at time and substitute the evolution that Chapter 4.6 §2 produced:
Now look at the middle of that expression and notice that the brackets can be grouped two ways. Group the 's with the state on either side and you have the picture used since Chapter 4.6: the state moves and the operator sits still. Group them with the operator instead and you have a different bookkeeping with the same value, in which the state sits still and the operator moves. Define
and (4.9.10) reads . This is the Heisenberg picture, and the one before it is the Schrödinger picture. Since every prediction of the theory is one of these numbers and the number has not changed, the two pictures agree on everything measurable by construction, and not as a theorem that could have come out otherwise.
4.2 · It is a change of basis, and the domain the basis does not carry
Look at the form of (4.9.11) rather than at what it does. Chapter 0.4 §4.2 derived the transformation of a matrix under a change of basis and got , calling two matrices related that way similar and summarising the content as: similar matrices are the same map, seen twice. Here , and because is unitary. So (4.9.11) is a similarity transformation, and as algebra it is nothing else. The Heisenberg picture is the Schrödinger picture written in a basis that rotates along with the state.
Chapter 0.4's warning box listed the consequences of confusing a map with its matrix, and one of them was "why the Schrödinger and Heisenberg pictures look like different physics instead of different bases". Collect it here, and pay for it, because in infinite dimensions the algebra leaves something out and Chapters 4.4 and 4.5 were spent on exactly that thing. An operator is a formula together with a domain, and Chapter 4.4 §3.1 fixed that two operators are the same operator only when the formulae agree and the domains do. Conjugation carries the domain along with the formula, so , which for an unbounded is a different subspace at every time. That is what makes a genuinely different operator at each instant rather than one operator wearing a rotating coat, and it is why the differentiation of §4.3 is a strong derivative taken on vectors that stay inside the moving domain rather than a derivative in the operator norm. In finite dimensions the slogan costs a sentence about similar matrices. Here it costs the domain as well, and the domain is what the algebra is silent about.
The claim has consequences you can check rather than admire, and the one that matters is the list of readings. Chapter 4.5 §2.1 took invertibility rather than an eigenvector as the primitive, exactly so that this kind of question could be asked about and , which have no determinant, no trace and in general no eigenvectors. Use it here. Since , the operator is a bounded everywhere-defined inverse of the left-hand side exactly when is one for . The two resolvent sets are therefore the same set, so the spectrum of equals the spectrum of at every time. The list of values a measurement of can return does not move. If it did, the two pictures would disagree about what an apparatus can read, and no amount of algebra would repair that.
4.3 · Differentiating, and the equation that falls out
We have as an explicit formula in , and what we want is its rate of change, so differentiate (4.9.11) with the product rule. Chapter 4.6 §3.1 supplies the derivative of the evolution operator, , and taking the adjoint of that gives since is self-adjoint. Allowing its own explicit time dependence as well, the three terms are
The first two terms combine, and the operations have to go in one particular order. Insert between the operator and the Hamiltonian in each term first, which converts each factor separately into its Heisenberg form and leaves . Only now is there a commutator to read off, and it is . Two sign changes then bring that to the form written below and they cancel: reversing the bracket into contributes one, and contributes the other. For a Hamiltonian with no explicit time dependence commutes with , so and no subscript is needed on it. What is left is the Heisenberg equation:
The two terms answer to different things and it is worth keeping them apart from the start. The second is present only if the observable was defined with a time in it, as is, and it has nothing to do with the dynamics. The first is the dynamics, entirely, and it is a commutator.
4.4 · Beside the classical equation, term by term
Now set that against what Chapter 1.3 §6.1 derived, before any quantum mechanics existed in this book, for the rate of change of an arbitrary function on phase space:
Chapter 1.3 called this the equation of motion for every observable of every Hamiltonian system, and noted that it contains Hamilton's equations as the two special cases and . Compare the two equations piece by piece:
- The left-hand sides are the same, with a function replaced by an operator.
- The explicit-time terms are the same, with the same meaning and the same irrelevance to the dynamics.
- The Hamiltonian appears in the same slot in both.
- The only difference anywhere is which bracket. The Poisson bracket has become .
That substitution is not being introduced here. Chapter 1.3 §6.4 wrote it down and Chapter 1.3's closing brick sent the bracket to this chapter to be replaced by , while Chapter 4.2 §7.5 derived the same replacement from the first-order expansion of a unitary. What is new is the equation itself, obtained from and the product rule rather than by analogy. Chapter 1.3 also warned that the substitution cannot be extended to every pair of classical observables at once, and that warning is still standing; Chapter 4.10 §8 is where it is proved.
4.5 · Conservation, in one line, and stronger than the classical version
Set in (4.9.13) and the left-hand side vanishes exactly when the commutator does:
Chapter 1.3 obtained the same statement with the same shape, and there too conservation stopped being something you established by solving the equations of motion and became something you established by computing one bracket. The quantum version says slightly more than the classical one, and the extra is worth having. What is constant is the operator , not merely its expectation. To see how much that buys, run it through the projections rather than the moments, since constant moments do not by themselves pin a distribution down for an unbounded observable. Because makes 's spectral projections of Chapter 4.5 §6 commute with , the probability of a reading landing in any region is the same at every time. The entire probability distribution of the measured values of is frozen, not merely its mean. Chapter 4.6's Worked example 1 found precisely this for the momentum of a free particle and obtained it there by a special argument. Here it is the general case.
4.6 · Two commutators the rest of the chapter needs
Before using (4.9.13) on anything, we need to be able to compute the commutator of a Hamiltonian with something. One rule does almost all of the work. Expanding both sides and cancelling the two middle terms gives, for any three operators,
which is the product rule with the ordering respected. Since the commutator with a fixed operator therefore behaves like a derivative, everything below is one application of it. With from Chapter 4.2 §8, applying (4.9.16) to gives
The other one needs the position representation, and Chapter 4.6 §5 established that there and is multiplication by . Acting on any differentiable in the domain, the product rule gives while , and the second terms cancel:
Both are verified symbolically. Now feed the Hamiltonian of Chapter 4.6 §4 into (4.9.13) twice, once with and once with . Since commutes with and commutes with , only one term survives in each case, and the factors of cancel:
Those are Hamilton's equations of Chapter 1.3 §3 with hats on everything, and they are exact operator identities with no approximation anywhere in them. Section 5 takes their expectation values and Problem 2 solves them in closed form for the oscillator.
4.7 · Which picture is used where
Neither picture is more correct, so the choice is made on convenience, and the convenience runs in opposite directions in the two halves of this book. Everything in Part IV is easier in the Schrödinger picture, because a bound-state problem is an eigenvalue problem for a fixed operator and Chapters 4.7 and 4.8 want the operator to hold still while they diagonalise it.
Part V reverses that. A field is an operator attached to each point of spacetime, and its time dependence has to sit in the operator rather than in the state. Otherwise the time coordinate is treated differently from the three spatial ones, and the relativity of Part II is invisible in the notation. Chapters 5.2 and 5.3 are written in the Heisenberg picture throughout for that reason, and (4.9.13) is the equation they solve. Building it here costs one section; discovering it there would cost a chapter.
Everything the theory predicts is an average of the form "state, operator, state". There are two places the time can be kept in such an expression, and both give the same number. Keep it in the state and the state evolves while the measured quantities stand still, which is the picture used so far. Move it onto the quantity instead and the state is frozen while the operators evolve.
This is not two theories. The manoeuvre is exactly the change of basis met in the linear algebra chapter, where one map acquires a different array of numbers in a different basis while remaining one map. Failing to see that is why the two pictures look like rival physics. The check is that the list of possible measured values, which is a property of the map and not of the basis, does not budge.
Differentiating the moved operator gives its equation of motion, and the answer is a bracket with the energy. Set it beside the classical equation of motion for an arbitrary quantity, derived in Part I from Hamilton's equations, and the two are identical except that one uses the classical bracket and the other uses the commutator divided by . Nothing else differs. That single replacement is what the phrase "canonical quantisation" refers to.
The immediate reward is that conservation becomes a computation rather than an investigation. A quantity is conserved when it commutes with the energy, and the statement is stronger than its classical counterpart: not merely the average but the whole distribution of measured values stays put.
a natural place to stop · the equation of motion is built; what follows is what it says about averages, and then about symmetries
5 · Ehrenfest, and the potentials for which it is exact
Chapter 4.6 §8.7 wrote down two relations and left the second unproved, for want of machinery it did not have. The machinery is (4.9.13), and with it both relations take two lines. The section is not really about deriving them. It is about the caveat that comes with them, which decides how much they are worth and which is the reason the harmonic oscillator occupies the position in physics that it does.
5.1 · Both relations, from one equation
Take the expectation of (4.9.13) in the Heisenberg state, which is fixed, so the derivative passes straight through the bracket. Translated back into the Schrödinger picture, where it is more familiar, the statement is that for any observable
Everything needed is already in (4.9.19). Neither nor carries an explicit time, so the last term is absent. Taking expectations of those two operator identities gives Ehrenfest's relations, written here in three dimensions because the promotion is free and componentwise: §4.6's two commutators used one axis at a time, keeps different axes from mixing, and becomes one component at a time.
Those are the two relations Chapter 4.6 §8.7 stated and deferred, and the route to them here is uniform: one equation of motion, two choices of observable. Chapter 4.6's Worked example 1 obtained the first of them for a free particle by a special argument in three lines, and its agreement with the general case is a check on both.
5.2 · The warning Chapter 1.1 issued, and why it stands
Chapter 1.1 §4.4 argued that force is not a fundamental concept in quantum mechanics: there is no force operator, nobody lists its eigenvalues, and it appears in no commutation relation. It then named (4.9.21) as the closest thing to a force that quantum mechanics contains, and it added a warning at the same time, which is worth quoting because softening it would spoil the point. Read carefully, it said, this "is not a fundamental law but a derived statement about expectation values, in which the potential is the primitive object and the force is what you get by differentiating it".
Everything in this section confirms that reading. The potential entered through , which is where Chapter 4.6 §4 put it. The gradient appeared in (4.9.18) as the by-product of a commutator, and the object it acts on is an operator, not a trajectory. Nothing anywhere obeys Newton's second law. Two averages obey an equation that looks like it.
5.3 · The step that is not there
Equation (4.9.21) is often read as saying that the centre of a wave packet obeys Newton's second law. It does not say that, and the difference is one symbol. Compare:
(a) , the average of the force over the packet. This is a theorem and it holds for every state and every potential.
(b) , the force evaluated at the average position. This is Newton's second law for the centre of the packet, and it is a different statement.
Getting from (a) to (b) requires , which is the claim that the average of a function is the function of the average. That is false for any function with curvature, for exactly the reason a curved dose–response relation has a mean response that is not the response at the mean dose. Spread the doses symmetrically about their mean and it fails as badly as ever, because what breaks the substitution is the bend in the curve and not any lopsidedness in the doses. The distinction is not a technicality here. It is the entire difference between quantum motion and classical motion at the level of averages, and §5.5 measures it.
5.4 · At most quadratic, and nothing else
To see how big the gap is, expand in a Taylor series about the mean position, which Chapter 0.3 licenses for a smooth potential, and take the expectation term by term. The first-order term vanishes because by construction, so the first surviving correction is the second-order one:
Read the correction term. It is controlled by the third derivative of the potential and by the square of the packet width, which by (4.9.1) cannot be sent to zero at finite momentum spread. Two consequences follow, and they are the point of the section.
The gap vanishes identically when , which is to say when is a polynomial of degree at most two. Then is an affine function, the average of an affine function is the affine function of the average, and every term after the first in (4.9.22) is absent rather than small. Statement (b) becomes exactly true, for every state, however wide, however skewed, and however far from classical it looks. That is the quadratics and nothing else: the free particle, a uniform field, the harmonic oscillator, and the upside-down oscillator of an unstable equilibrium, where the coefficient is negative and the classical reading is exact all the same.
The gap does not vanish otherwise. If is non-zero somewhere then a packet sitting there has a discrepancy of order , and since cannot be sent to zero without sending to infinity, there is no state that escapes it. Ehrenfest's theorem survives untouched, because it was statement (a). What fails is the classical reading.
One more thing needs saying, because the exactness in the quadratic case is weaker than it sounds. An oscillator energy eigenstate has at every time, and it satisfies the classical equation with both sides zero forever. Nothing about it resembles a swinging pendulum. Exact-in-the-mean is a statement about one number extracted from the state, and a state can satisfy it while looking nothing like a trajectory. The state that does look like a trajectory is Chapter 4.8 §7's coherent state, which keeps its width while its centre traces the classical orbit, and Chapter 4.10 takes up the question of when a general system admits anything of the kind.
5.5 · The two cases, measured against each other
The way to settle this is to integrate the Schrödinger equation twice with everything held fixed except the potential, and watch the residuals. The integrator is the split-operator scheme of Chapter 4.6 §10.7, in which every factor has modulus one, so the norm is conserved by construction and the readouts test the physics rather than the arithmetic. Units are , the grid is points on , the initial state is a Gaussian of width released from rest at , and the run is to with . The two potentials are and , chosen so that the first has and the second has . The time derivatives in the table are fourth-order centred differences of the recorded series, which is the one detail the first two rows depend on and the one a reader reproducing the table would otherwise have to guess.
| largest residual over the run | ||
|---|---|---|
| , bit for bit | ||
Read the table by rows and then by columns. The first two rows are Ehrenfest's relations, and they hold in both potentials with residuals of order or smaller against quantities of order one. Those residuals belong to the integrator rather than to the theorem, and here is how you know: halving divides them by four, three times running, which is a second-order error in the step size doing what second-order errors do. That test fixes the order and not the source, and the source here is the splitting's own second-order error seen through the difference stencil. The theorem is exact and the arithmetic is not.
The third row is the substitution the warning box refused to make. In the quadratic potential and are the same floating-point number at every one of the six thousand steps, because is the identity there and averaging commutes with doing nothing. In the quartic potential the two differ by as much as against forces of order one, with a mean gap of over the run. The fourth row is the consequence: the centre of the quartic packet is not obeying Newton's second law, it is missing it by an amount comparable to the force itself, and that number does not move when is halved because it is physics rather than error. Meanwhile the quadratic run's centre follows to throughout, although the packet is not a coherent state and its width is breathing the whole time.
Feeding position and momentum into the equation of motion produces two statements that look exactly like the classical laws with averages written over everything. The rate of change of the average position is the average momentum over the mass, and the rate of change of the average momentum is minus the average force. This is the result an earlier chapter stated without proof and attached a warning to, and the warning is the more valuable half.
The warning is this. "The average of the force" is not "the force at the average position". A packet has width, it samples the force over a range, and unless the force varies in a straight line across that range the two quantities differ. The size of the difference is set by the curvature of the force and by the square of the packet's width, and the width cannot be reduced to nothing without the momentum spread exploding.
So the classical-looking reading is exact only when the force is a straight line in position, which means the potential is at most quadratic. The cases are exactly the quadratics: no force at all, a uniform force, a spring, and the inverted spring of a ball balanced on a hilltop. Every other potential in physics fails the test, and that is a large part of why the spring turns up everywhere in this subject as the case that can be trusted.
Two runs of the same integration, differing only in the potential, make the split visible. Both satisfy the theorem to within a millionth. Only the quadratic one has the centre of its packet obeying Newton's second law; the quartic one misses by roughly as much as the force itself. And exactness in the mean is a modest property in any case: a stationary state of the spring sits with its average position at zero forever, satisfying the classical equation with both sides vanishing while behaving nothing like a swinging weight.
a natural place to stop · the commutator has bounded and it has moved things; what remains is what it generates
6 · Symmetries, generators, and conserved quantities
Chapter 1.4 ended Part I with a claim it called the real destination: a conserved quantity and the symmetry it comes from are one object under two names. Chapter 4.2 §7.5 restated that claim with operators in place of functions and expanded a unitary to first order to show the commutator appearing where the Poisson bracket had been. What it could not do was supply the generator's existence, because in infinite dimensions differentiating a one-parameter family of unitaries is not a formal manoeuvre, and it said so at the time. That link now exists, so the correspondence Chapter 4.2 could only state can be assembled here, and then run on the three cases Chapter 1.4 §3 ran classically.
6.1 · The chain, and which links were missing
A symmetry is a transformation of states that leaves every probability alone. Follow that through the machinery already built, one link at a time. Every link below was built somewhere else, so this subsection is where they are put in order rather than where anything is computed. The computing is in §6.2 and §6.3, on the three cases. The second link is the one Chapter 4.2 could not supply, and it is the reason this section can exist at all.
- Symmetry gives unitary. A map preserving all probabilities preserves all norms, and Chapter 4.2 §7.2 showed that a linear norm-preserving map preserves every inner product and is unitary. Linearity is assumed rather than derived at this link, exactly as Chapter 4.2 §7.1 assumed it, and every family in this section is visibly linear, so the assumption is discharged case by case instead of in general.
- A continuous family gives a self-adjoint generator. If the transformations form a strongly continuous one-parameter group, Stone's theorem produces a unique self-adjoint with . That is the theorem Chapter 4.5 §9 proved forward and quoted backward, and Chapter 4.6 §2 used it on time. This is the link Chapter 4.2 could not supply, and the reason it could not is that Chapter 0.5 §7.1's finite-dimensional argument assumes a bounded operator with an eigenbasis, which the generators here are not and do not have.
- The generator is an observable. Self-adjoint is exactly the condition Chapter 4.2's second postulate requires, in the corrected form Chapter 4.4 §4.2 gave it.
- The symmetry acting on an observable is a commutator. Chapter 4.2 §7.5 expanded to first order and obtained .
- Symmetry of the dynamics means conservation. The family is a symmetry of the system when it leaves alone, which by the previous line means . By (4.9.15) that is precisely the statement that is conserved.
Read the last two links in the other order and the converse is the same equation. If is conserved then it commutes with , so the family it generates leaves alone and is a symmetry. A conserved quantity is the generator of its symmetry, which is what Chapter 1.4 §7 proved classically, repeated here with the generator's existence supplied rather than assumed. The three cases below are Chapter 1.4 §3's three, and the third of them is already done.
6.2 · Translation, and momentum
Define the family that slides a wavefunction along, writing it in the notation of §6.1, so that . It is unitary because , Lebesgue measure being translation-invariant by Chapter 4.3 §2.1, and it is onto because undoes it. The group law holds because sliding twice is sliding once by the sum. Strong continuity takes one line: for a continuous of compact support, by uniform continuity, and Chapter 4.3 §8.1 proved such functions dense, so the general case follows from .
Stone therefore hands us a generator, and to find it we differentiate the family at the origin, which is what Stone's formula for the domain in Chapter 4.6 §2.1 instructs. Since gives , and differentiating at gives , the two must agree:
Momentum is the generator of translations, with the operator Chapter 4.6 §5 identified arriving here for a second and independent reason. Chapter 1.4 §7.1 computed and read the same sentence off it classically.
Now close the loop with a computation rather than a slogan. By (4.9.18), , which is the zero operator exactly when is constant. So momentum is conserved precisely when the potential is unchanged by translation, and the two statements are one commutator apart. Chapter 1.4 §3.2 reached the same conclusion from the action, and the two derivations have no step in common.
6.3 · Rotation, and the commutator Chapter 4.11 is built on
Do the same in three dimensions with rotations about the -axis. The family is , unitary because a rotation preserves volume, and differentiating at gives, with ,
Angular momentum is the generator of rotations, matching Chapter 1.4 §7.1's classical computation with , which also rotated the momentum as a bonus. The operator is the classical expression with hats, and the ordering ambiguity that usually attends such a substitution is absent here because and commute.
What that costs is one line and what it buys is the whole of Chapter 4.11, so compute the thing it buys. Rotations about different axes do not commute, and the algebraic residue of that geometric fact is the commutator of the generators. Take and and expand with (4.9.16). Of the four terms, two vanish outright because they contain no conjugate pair. The survivors are the first and the last, and . Both enter the sum with a plus sign, the first because it carried no minus sign in the expansion at all and the last because the two minus signs it carried multiply together. Adding them,
with the other two relations following by the cyclic relabelling . Verified in the Weyl algebra: normal-ordering the products using and nothing else returns zero for all three differences.
That verification deserves a sentence of its own, because of what it does not use. Chapter 4.2 §8 postulated the operator commutator for the single pair and was explicit that the general correspondence between Poisson brackets and commutators is not derivable and is in fact false. Equation (4.9.25) is what the correspondence predicts for a pair it was never postulated for, since Chapter 1.4 §7.3 computed classically, and applying to that turns it into (4.9.25) exactly. The prediction is correct, and it was obtained here from the one postulated pair by algebra alone. Chapter 4.10 §8 explains why such agreement cannot be made universal.
Two consequences, and then the hand-off. First, feeding (4.9.25) into (4.9.8) gives , so no state has all three components sharp unless all three vanish, which is what Chapter 1.4's grind box anticipated when it noted that would survive quantisation.
Second, the check §2.7 held back. It was the third of that section's three, and the commutator it needed now exists. On the three-dimensional representation Chapter 4.11 builds from (4.9.25), the state with the largest gives against a bound , so the relation is saturated to fifteen figures. The middle state of the same set gives against a bound of exactly zero, since there. That second case is the one to keep. A bound that vanishes has not announced that both quantities are sharp. It has fallen silent, while (4.9.25) still forbids a common eigenvector, which is §3.1's distinction made numerical.
And now the hand-off: equation (4.9.25) is the entire input to Chapter 4.11. That chapter takes it as the definition of what an angular momentum is, forgets where it came from, and extracts the whole spectrum from it by the ladder move Chapter 4.8 ran on the oscillator. Nothing about rotations in space is used again.
6.4 · Time, which was done first
The third case needs no work, because Chapter 4.6 did it before the pattern was visible. Stone applied to time evolution produced as the generator, so the Hamiltonian generates time translation, and makes energy conserved by (4.9.15) with no calculation at all. Chapter 1.4 §3.1 obtained conservation of energy from invariance of the action under a shift of the time origin, and Chapter 1.4 §7.1's third example observed that time evolution is the transformation generated by the energy, so it is not a different kind of process from a symmetry. All three of Chapter 1.4's classics have now been repeated with operators, and the table below is that chapter's §3 in quantum form.
| symmetry | unitary family | generator | conserved when |
|---|---|---|---|
| translation in space | is translation-invariant | ||
| rotation | depends on alone | ||
| translation in time | always, for a closed system |
6.5 · A symmetry with no generator of its own
The chain of §6.1 needs a continuous family, and not every symmetry sits inside one that tells you anything. Parity, built in Chapter 4.7 §2 and written , is the map . It is unitary, it is self-adjoint, and it commutes with whenever is even. What it is not is a member of a family that could produce something new. Geometrically the reflection has determinant , so no continuous family of spatial transformations reaches it from the identity. Algebraically a family can be manufactured, since the projections onto the even and the odd functions are functions of and exponentiating the odd one runs from to . But Stone applied to that family returns a generator which is itself a function of , so it is no observable that was not already parity, and nothing has been generated. What survives is the conclusion without the machinery: parity commutes with , so by (4.9.15) it is conserved, and since leaves it only two eigenvalues, what is conserved is a label taking two values rather than a quantity taking a continuum of them.
Chapter 1.4 §5 recorded the classical version of the same limitation and called it an honest exception, and this is what it looks like on this side. Noether's theorem needs a continuous symmetry that reaches the identity through the transformations themselves, and so does anything Stone can usefully hand back. A discrete symmetry produces a conserved quantum number with no independent generator behind it. Chapters 4.17 and 4.18 use exactly that, since a selection rule is a statement about a discrete label rather than about a continuous charge.
6.6 · What the commutator has not told us
Take stock, because the chapter has said three things about one object and it is worth being explicit that a fourth is missing.
The commutator of two observables bounds how sharp they can be at once ((4.9.8)), and it vanishes exactly when they can be sharp together (§3). The commutator with the Hamiltonian moves every observable in time ((4.9.13)), and it vanishes exactly when the observable is conserved ((4.9.15)). The commutator with any self-adjoint operator generates the symmetry that operator belongs to (§6.1). Three jobs, one object, and the classical shadow of all three is the Poisson bracket of Chapter 1.3.
What none of it says is what happens when is small compared with the action in play. Every equation in this chapter carries in a place where setting it to zero either destroys the statement or makes it vacuous: (4.9.8) becomes the empty assertion that a product of spreads is non-negative, and (4.9.13) becomes . Something more careful is needed, and it is a genuine limit rather than a substitution.
That is Chapter 4.10, and it comes with a warning worth meeting before the proof. The correspondence between classical and quantum observables looks, from (4.9.13) and Chapter 4.2 §7.5, like a dictionary that should extend to everything. It does not. Chapter 4.10 §8 proves that no consistent dictionary exists taking every classical observable to an operator while turning every Poisson bracket into the corresponding commutator. Chapter 4.2 kept its distance from that general claim twice, in §7.5 when it separated the derived algebraic statement from the undeliverable general one and in §8 when it postulated the correspondence for a single pair rather than for all of them. Now that you have watched the correspondence succeed for angular momentum in (4.9.25), the claim that it must eventually fail is worth arriving with rather than being told afterwards.
A symmetry is anything you can do to a system without changing any prediction. Chase that definition through the machinery and a chain of four links appears, each of them already built. Preserving predictions means preserving lengths, and provided the transformation is linear, which every family the section used visibly is, that forces it to be a rotation of the space of states. Linearity is an assumption at this link rather than a result, and the main text says so where it is used. A continuous family of such rotations has, by the theorem imported earlier in this part, a unique observable behind it that generates it. The effect of the family on any other quantity is a commutator with that generator. And leaving the energy alone, which is what makes the family a symmetry of the dynamics rather than a mere relabelling, is the same equation as the generator being conserved.
So a conserved quantity and a symmetry are one thing, exactly as the Part I chapters concluded for classical mechanics, and the argument is shorter here because the theorem doing the heavy lifting was proved elsewhere.
Three examples, and they are the same three the classical chapter used. Sliding everything along is generated by momentum, and momentum is conserved when the potential is flat. Turning everything about an axis is generated by angular momentum, and the commutator of two such generators is another one, which is the algebraic fingerprint of the fact that turns about different axes do not commute. Waiting is generated by the energy, which is therefore conserved with no computation needed. Discrete symmetries, like reflection, leave no generator behind, because any continuous family you could run through such a symmetry has to be built out of the symmetry itself and hands back nothing that was not already there. What they leave is a two-valued label rather than a continuous charge.
One thing has been missing throughout. Everything here describes how quantum quantities behave among themselves. None of it says how the classical world reappears when the scales get large, and the reason no equation here answers that is that setting the constant to zero in any of them produces nonsense. The next chapter takes the limit properly, and it also proves that the translation between classical and quantum quantities, which has worked every time it has been tried here, cannot be made to work for everything at once.
7 · Worked examples
The bound is attained by some states and not by others. (a) Find the condition on for equality in (4.9.8) for a general pair. (b) Solve that condition for and . (c) Identify what you get. (d) Say what changes if you ask instead for equality in the sharper form (4.9.9).
(a) Two inequalities were used and both must be tight. Chapter 0.5 §1.4 recorded the equality condition for Cauchy–Schwarz: the two vectors are parallel, so for some complex . The second inequality was the discarding of the anticommutator term in (4.9.7), so that term must vanish. Compute it with :
So equality holds exactly when with purely imaginary. Both conditions are conditions on the state, and neither mentions the observables beyond (4.9.3).
(b) Write with real, put and , and work in the position representation where . The condition becomes a first-order linear equation,
which separates in the manner of Chapter 0.8 §2 and integrates to
(c) That is a Gaussian of width , centred at , multiplied by a plane wave carrying mean momentum . Its momentum spread is , and the product is for every . This is the third time the Gaussian has been produced as the unique minimiser, by a third route. Chapter 0.9 §6.5 ran the bandwidth theorem's equality conditions backwards, Chapter 4.6 §10 found it as the packet whose spreading formula is exact, and here it comes out of Cauchy–Schwarz. Chapter 4.8 §7's coherent states are exactly this family with , which is why they are the states that behave most like a classical oscillation.
(d) Only the second condition is dropped, so equality in (4.9.9) requires with complex and otherwise unrestricted. Repeating the integration with gives the same Gaussian with a complex width, that is, with a quadratic phase across it. Problem 1 computes such a state's three numbers and finds a product exceeding while the sharper form stays tight, which is the right-hand pair of points in the figure of §2.7.
On the circle, with and on the periodic domain, which is the member of the circle of self-adjoint domains Chapter 4.4 §5.4 found, let be multiplication by . (a) Evaluate both sides of (4.9.8) in the state . (b) Locate the step of §2 that fails. (c) Compute the defect exactly. (d) Repair the statement.
(a) is an eigenstate of with eigenvalue , so exactly. Its density is uniform on , so and , giving . The left-hand side is therefore zero. Computed formally, on smooth functions, so the right-hand side would be . Zero is not at least . One of the hypotheses has to have failed, and the useful question is which.
(b) It is (4.9.5), the step that moved across the inner product. That move is the definition of self-adjointness applied to the vector , and it is legitimate only when that vector lies in the operator's domain. Here is not periodic: it starts at and ends at . So it is not in the domain of , which single-valuedness on the circle fixes as the periodic functions, the member of Chapter 4.4 §5.4's circle of domains, and the whole of §2 is unavailable.
(c) Chapter 4.4 §5.1's boundary form says what is lost. For smooth on , integrating by parts once gives . Put and , so , whose bracket is . Hence
and the two differ by exactly . That is the whole of the formal commutator. The would-be right-hand side of the relation is not a property of the state at all. It is a boundary term, and it is there because the operator was applied outside its domain.
(d) Replace the angle by a bounded periodic function of it, which keeps the domain intact because multiplying a periodic function by a periodic function leaves it periodic. With the same one-line calculation as (4.9.18) gives , so §2 applies unchanged and delivers
Test it on the offending state: for a uniform density, so the bound is zero and violates nothing. The theorem was never wrong. The statement it was applied to was not one of its instances, and this is what Chapter 4.4 was for.
The relation is quoted as often as the position–momentum one. (a) Say why it cannot be an instance of (4.9.8). (b) Derive, from this chapter and nothing else, a true statement of that shape. (c) Interpret the quantity that plays the part of . (d) Put a number on it.
(a) Every symbol in (4.9.8) refers to two observables, and time is not one. It is the parameter labelling the family , not an operator on the space, and Chapter 4.6 §1 treated it that way throughout. There is no whose commutator with one could evaluate, so (4.9.8) has nothing to say and quoting it here is a category error rather than an approximation.
(b) What is available is (4.9.20), which is the only place in the theory where a time derivative and a commutator meet. Take any observable with no explicit time dependence and apply (4.9.8) to the pair :
the last step being (4.9.20) read from right to left. Now divide through by the rate, which is legitimate in any state where the rate is non-zero, and define
(c) Read the definition of rather than the inequality. It is the time the mean of needs in order to shift by one standard deviation of , which is the shortest time in which a measurement of could tell that anything had happened. So the honest statement is not about an uncertainty in time. It is a speed limit: a state with a small energy spread cannot change quickly, in any observable at all, since was arbitrary. A stationary state has and changes in nothing, which is the extreme case and agrees with Chapter 4.6 §9.2. Checked numerically on two hundred thousand random Hamiltonians, observables and states: no violation, with the smallest slack , and a two-level system prepared as saturates it exactly at .
(d) A state whose energy is uncertain by eV cannot change appreciably in less than , using . Read the same relation with the lifetime identified as the of some observable that registers the decay, and it says an excited state living for a time cannot have a sharper energy than about , which is why spectral lines have widths. That identification is an extra step rather than a re-reading, and it is part of why the statement is an order of magnitude and not a bound. Chapter 4.17 computes those widths from the transition rate rather than bounding them, and the two answers agree in order of magnitude, which is the check worth making on any argument of this shape.
8 · Your turn
Problem 1 — the sharper form, and a Gaussian that is not on the floor
Let with and , a Gaussian with a quadratic phase across it. (a) Compute , and the covariance . (b) Show that (4.9.9) is an equality while (4.9.8) is not, and find the factor by which the product exceeds . (c) Locate this state in the figure of §2.7 and say what the horizontal displacement of its point means physically. (d) Chapter 4.6 §10 evolved a free Gaussian and found it spreading. Without redoing that calculation, say what an initially real Gaussian has acquired by the time it has spread, and which of the three numbers in (a) detects it.
Solution
(a) The density is , so and . Since , . For the covariance, integrate by parts once: , so the covariance is .
(b) , while the right-hand side of (4.9.9) is , which is the same number. The sharper form is exactly tight. Meanwhile , which exceeds for every . Equality in (4.9.8) needs a real width, which is Worked example 1(b).
(c) The point sits on the unit circle, at height and horizontal coordinate . The two drawn cases are and . The horizontal displacement is the correlation between position and momentum: in this state a particle found on the right is more likely to be moving right, which is what a quadratic phase encodes and what the modulus of the wavefunction cannot show.
(d) It has acquired exactly this quadratic phase, with growing from zero. That is what free spreading is: the fast components outrun the slow ones, so position and momentum become correlated. Neither nor separately reveals the correlation, since the second is constant for a free particle by Chapter 4.6's Worked example 1(b); the covariance detects it, and it is the reason a free packet's product grows without the state ever ceasing to saturate (4.9.9).
Problem 2 — the Heisenberg equation solved exactly, for the one potential where that is possible
Take . (a) Write out (4.9.19) for this potential and solve the pair of operator equations. (b) Verify that at all times, and say why it had to come out that way. (c) Take expectations and check §5.4's claim that the classical equation is exact here for every state. (d) Compute for an arbitrary initial state and identify the states whose width does not move.
Solution
(a) With the equations are and , a linear system with constant coefficients in which the coefficients are numbers, so Chapter 0.8 §4's solution applies with operators in place of the initial values:
Differentiating returns the equations, and at both reduce to the Schrödinger-picture operators. This is the only potential for which the operator equations close, because it is the only one where is linear in .
(b) Expanding, . It had to: by (4.9.11) both operators are conjugated by the same unitary, and , so every algebraic relation among operators survives the change of picture unchanged. That is the same invariance argument as §4.2's about the spectrum.
(c) Taking expectations gives , which satisfies identically. No property of the state was used anywhere, which is §5.4's claim: for a quadratic potential the mean obeys the classical equation exactly, whatever the state.
(d) Squaring and subtracting the square of the mean,
The two time-dependent pieces combine into a constant exactly when the covariance vanishes and , since then the and terms add to and the cross term is absent. Every energy eigenstate meets that condition, which it had to, since Chapter 4.6 §9.2 makes every observable's distribution constant in a stationary state, as part (a) of Problem 3 records. Among the minimum-uncertainty states of Worked example 1 the condition picks out alone, which is the ground state and the coherent states built on it, Chapter 4.8 §7's family and the reason Chapter 4.6 §10.8 found the oscillator packet's width unchanging. Any state failing the condition breathes, and the equation above shows it does so at rather than .
Problem 3 — the virial theorem, which is one commutator
(a) Show that in a stationary state the expectation of any observable without explicit time dependence is constant. (b) Take and compute for . (c) Deduce the quantum virial theorem and check it against the oscillator. (d) Apply it to in three dimensions and say what it gives for a hydrogen state.
Solution
(a) Chapter 4.6 §9.2 showed the two phase factors cancel in for , so the expectation carries no . Equivalently, and more usefully here, (4.9.20) gives , and in an eigenstate of the Hamiltonian may be moved onto either side as the number , so the two terms cancel and the derivative is zero.
(b) Since is a constant, and have the same commutator with anything, so compute with . Two applications of (4.9.16) and one of (4.9.17) and (4.9.18) give
(c) Put (b) into (a). The left-hand side is zero in a stationary state, so
For we have , so . Chapter 4.5's Problem 2(c) computed both halves independently for the Hermite functions and got exactly that, which is a check with no algebra in common.
(d) The three-dimensional version replaces by , the derivation running component by component. For , Euler's relation for a function homogeneous of degree gives , so and therefore . Every bound hydrogen state has kinetic energy equal to minus its total energy and potential energy twice its total energy, before any wavefunction has been written down. Chapter 4.13 computes those wavefunctions and Chapter 4.16 uses this relation to evaluate the fine-structure corrections without doing new integrals. Chapter 1.4's Problem 4 obtained the classical version as a statement about time averages; here it holds state by state, with no averaging over an orbit, because a stationary state is already what a time average is trying to be.
Problem 4 — completeness, by counting, on two qubits
Take two of Chapter 4.2 §10.4's qubits, so the space has dimension four, with and the which-state observables of the two factors, eigenvalues . (a) Show is a complete set by exhibiting the count. (b) Show that the total is not complete on its own, and exhibit an observable commuting with it that is not a function of it. (c) Let be the swap, . Show is an observable, that , and that is complete. (d) Show that the two complete sets cannot be used together, and name the chapter where this situation recurs.
Solution
(a) The four product states are joint eigenvectors with tuples , , , . Four distinct tuples in a four-dimensional space, so by §3.4 the set is complete and the two eigenvalues are a full set of quantum numbers.
(b) has eigenvalues , and the eigenvalue carries the two-dimensional space spanned by and . So the tuple count is three against a dimension of four and the set fails. For an explicit witness take with . It commutes with because is a -eigenvector, and it is not a function of , because any function of acts as a single number on the whole zero-eigenvalue space while annihilates . That is §3.3's second argument, made concrete.
(c) is unitary and , so and it is self-adjoint with eigenvalues . It commutes with because treats the two factors alike. The joint eigenvectors and tuples are , , , and . Four distinct tuples again, so the set is complete: the swap supplies exactly the label was missing, and it does so inside the degenerate eigenspace, which is Chapter 0.5 §8.2's Step 4 in action.
(d) Compute on : the first order gives and the second gives , so the commutator is . The two sets are individually complete and mutually incompatible, so a state labelled by one has no definite labels in the other. This is exactly the relationship between the uncoupled and coupled bases of Chapter 4.12, which adds two angular momenta and finds the same two ways of labelling the same four states, with the swap replaced by the total angular momentum. Chapter 4.18 then explains why the antisymmetric combination is singled out by nature rather than by choice.
The uncertainty relation cost one line and then three. Chapter 0.9 proved the bandwidth theorem with no physics in it and said that quantum mechanics would add a single substitution. This chapter made it: turns into by multiplication, and Chapter 4.6 §10.6's mark on de Broglie's relation is the only physical input anywhere in it. The general case took three steps and no new ideas. Chapter 4.2 §5.3's variance is the length of a vector, and Chapter 0.5 §1.4's Cauchy–Schwarz bounds the overlap of two vectors by the product of their lengths. Splitting that overlap into its real and imaginary parts then puts the anticommutator in one and the commutator in the other. Keeping both gives the sharper form; discarding the first gives . Chapter 0.5 said in advance that nothing would be added here except the meaning of the symbols, and that is what happened. The hypothesis that makes the derivation legal is a domain condition, and Worked example 2 exhibits it failing for an angle, computes the boundary term that the failure leaves behind, and repairs the statement.
What the relation is about, said once and kept. Each is a spread of outcomes over systems prepared identically and measured once each. Chapter 4.5 §6.5 raises the deeper question and leaves it open, and §2's familiar-ground box says exactly how far this chapter can carry it. What is proved here is that for position and momentum, where the commutator is in every state and the floor is however you prepare, no preparation narrows both spreads at once, so there is no unmeasured covariate to go looking for. For a general pair the floor is a property of the state and can be zero, at which point the theorem falls silent rather than permitting anything. The stronger claim, that no joint distribution exists underneath the two marginals at all, is Chapter 4.20's and is not derivable from the inequality. Measurement disturbance is a separate phenomenon with theorems of its own, and those theorems are the single thing this chapter quotes rather than derives, named in §2.6 with their hypotheses. Numbers rather than assurances. The ground state sits on the floor at and the state sits seven times above it at . Two hundred thousand random pairs violate neither form. And for a spin-one multiplet the bound is saturated in the top state, while in the middle one it falls silent and the product is .
Completeness is a count. Chapter 4.2 §4.3 defined a complete set of commuting observables and asked, in this section's name, how one knows a set is complete. The answer proved here is that one-dimensional common eigenspaces and maximality are the same condition, so a set is complete exactly when every observable commuting with all of it is already a function of it. The way anyone establishes that in practice is to list the eigenvalue tuples and check the count against the dimension. Chapters 4.11 and 4.13 apply it to and to , and Problem 4 runs the whole argument on two qubits, including a set that fails, the operator that witnesses the failure, and two complete sets that cannot be used together.
The same bracket moves things. Grouping the evolution operators with the observable instead of the state is Chapter 0.4 §4.2's similarity transformation, so the Heisenberg picture is a change of basis and not a rival theory, which is the sentence Chapter 0.4's warning box asked this chapter for. What the algebra does not carry is the domain, which moves as , so is a different operator at each time in Chapter 4.4 §3.1's sense and §4.3's derivative is a strong one on vectors that stay inside it. The list of readings is unmoved even so, by Chapter 4.5 §2.1's invertibility definition of the spectrum rather than by any determinant or trace. Differentiating gives , which is Chapter 1.3 §6.1's classical equation with the Poisson bracket replaced by and nothing else altered, and conservation becomes with the whole distribution frozen rather than only its mean. Applied to and it gives Hamilton's equations as exact operator identities, and their expectations are Ehrenfest's relations, which Chapter 4.6 §8.7 stated and deferred. Chapter 1.1 §4.4 called these a derived statement about averages rather than a fundamental law, and §5 collected that warning rather than softening it. Here it is: is not unless vanishes identically, which happens for the quadratic potentials and for nothing else, the free particle, the uniform field and the oscillator upright or inverted. Two split-operator runs differing only in the potential measure the split. Both satisfy Ehrenfest to within a millionth, with the residuals falling by four each time is halved. The quadratic run has equal to zero in the last bit at every step and its centre on to ; the quartic run misses by up to against forces of order one, and that number does not move when the step size does.
And the same bracket generates. A linear symmetry forces unitarity by Chapter 4.2 §7.2, and linearity is assumed at that link rather than derived, discharged case by case because every family used here is visibly linear. Then a strongly continuous family has a self-adjoint generator by Stone, the generator acts on observables through a commutator by Chapter 4.2 §7.5, and leaving alone is the same equation as being conserved. Chapter 4.2 stated that correspondence and this chapter assembled it, the missing link having been the generator's existence in infinite dimensions. Translation gives , rotation gives , time gives , and those are Chapter 1.4 §3's three classics repeated with operators. The rotation case delivers from the single postulated commutator by algebra alone, agreeing with what Chapter 1.4 §7.3's Poisson brackets predict under a substitution that was never postulated for that pair. A discrete symmetry leaves no generator, because any family running through it must be built out of the symmetry itself, and what it leaves instead is a two-valued label.
Three marks made elsewhere are leaned on here and cited rather than raised again, and this is the list. De Broglie's relation, Chapter 4.6 §10.6, is the one physical input to §1. The converse half of Stone's theorem, Chapter 4.5 §9.3, is what §6 needs to know that a continuous symmetry has a generator at all. And the identification , Chapter 4.6 §4.2, is what makes §4.6's two operator equations equations about a particle rather than about an unnamed self-adjoint operator. The canonical commutator of Chapter 4.2 §8 is used constantly and is a postulate rather than a quotation, which is a different thing and is why it carries no mark.
Where this gets spent. Chapter 4.11 takes (4.9.25) as its entire input, forgets that it came from rotations, and extracts the angular momentum spectrum from it by the ladder move Chapter 4.8 used on the oscillator. Chapter 4.12 meets Problem 4's two incompatible complete sets again, as the coupled and uncoupled bases. Chapter 4.13 labels hydrogen with and uses §3's count to know that three labels are enough, and Chapter 4.16 uses Problem 3's virial relation to avoid new integrals. Chapter 4.17 turns Worked example 3's speed limit into line widths computed rather than bounded. Chapters 5.2 and 5.3 write everything in the Heisenberg picture, because a field has to carry its time dependence in the operator. The one thing this chapter has not touched is what becomes of all of it when is small against the action in play, since setting to zero in any equation here gives either nonsense or nothing. That is Chapter 4.10, and its §8 proves that the correspondence which worked for angular momentum in (4.9.25) cannot be made to work for every observable at once. You should now meet that theorem already believing it is needed.