Part IV · Quantum Mechanics — Chapter 4.9

Commutators, Uncertainty, and Symmetry

The uncertainty relation was proved in Part 0, with no physics anywhere in it, and it has been waiting three parts for one substitution. This chapter makes the substitution in a line. Then it spends the rest of its length on the fact that the same commutator which bounds a product of spreads is also what moves every observable in time and what generates every continuous symmetry. Very little here is new. Almost all of it is one object, met from three sides.

Where we are

Six results built elsewhere meet in this chapter, and all six were built with this meeting in mind. Chapter 0.9 §6.4 proved that a function and its Fourier transform cannot both be narrow, and left one substitution for quantum mechanics to add and nothing else. Chapter 0.5 §1.4 proved the Cauchy–Schwarz inequality and its insight box wrote down, three parts early, the two vectors this chapter would feed into it. Chapter 4.2 §5.3 wrote the variance of an observable as the squared length of a vector, which is the form Cauchy–Schwarz can act on. Chapter 1.3 §6.1 derived the classical equation of motion for an arbitrary observable, and §6.4 announced the substitution that would turn it into the quantum one. Chapter 1.4 §7 showed that a conserved charge generates the symmetry it came from. Chapter 4.6 §2 produced the operator U^(t)\hat U(t) that moves a state forward in time.

Nothing on that list needs to be built again. What this chapter does is spend it, and the shape of the spending is worth having before it starts, because it is the same object doing three jobs. The commutator bounds how sharp two quantities can be at once. The commutator with the Hamiltonian moves any observable in time. And a commutator with a generator produces the symmetry that generator belongs to. Those are not three facts that happen to involve the same bracket. They are one fact seen from three sides, and by the end of §6 you should be able to say which side you are looking at from the shape of the equation alone.

Here is the route. Section 1 is the shortest in the chapter, and one multiplication pays four of the promises Chapter 0.9 made, with the remaining two settled in §2 and §2.6. Its result is one you have met: Chapter 4.6 §10.2 wrote it down for one Gaussian packet and Chapter 4.8 §6.1 sat the oscillator ground state exactly on it. What §1 adds is the reason the multiplication is legitimate for every state at once. Section 2 does the general case in three steps, using Chapter 0.5's inequality on Chapter 4.2's two vectors, and then spends a careful page on what the result does not say, which is where most of the damage in popular accounts is done. Section 3 answers the question Chapter 4.2 §4.3 explicitly left for this chapter: how you know that a set of commuting observables is complete. Section 4 moves the time dependence off the state and onto the operator, shows that this is a change of basis with a moving domain, and differentiates to get the Heisenberg equation, which is set beside Chapter 1.3's classical equation term by term. Section 5 takes expectation values of that equation and gets Ehrenfest's relations, which Chapter 4.6 §8.7 stated and deferred, together with the caveat that decides how much they are worth. Section 6 runs the symmetry chain in full for translation, rotation and time, and hands Chapter 4.11 the one commutator its entire chapter is built on. It then adds a symmetry that leaves a label rather than a generator, and closes by naming the one thing the commutator has not told you: what becomes of any of it when the scales grow large, which is Chapter 4.10's subject and not this chapter's. Sections 7 and 8 are worked examples and problems.

Conventions. ΔA\Delta A always means the standard deviation of the outcomes of measuring A^\hat A on a large number of systems prepared identically in the state ψ\psi, and never anything about a single measurement. Hats stay on operators and come off their values, with ΔA\Delta A and ΔA^\Delta\hat A the one pair that needs care, sorted out where they meet in §2.3. Time evolution is U^(t)=eiH^t/\hat U(t)=\ee^{-\ii\hat Ht/\hbar} with the minus sign, as Chapter 4.6 §2.4 fixed it. One subject is quoted rather than derived, and it carries the chapter's single mark, in §2.6: measurement disturbance, whose two named theorems are about apparatus rather than about states, and which this book states with their hypotheses and does not prove. The one mark covers both of them, and there are no others. Three marks standing in earlier chapters are leaned on and cited rather than raised again, and the closing brick says which.

Tools you'll need  — Chapter 0.5 above all: §1.4 for Cauchy–Schwarz and its equality condition, §7.1 for the exponential of i\ii times a Hermitian operator, and §8 for commuting operators and simultaneous diagonalisation. Chapter 0.9 §6.4 for the bandwidth theorem and §6.5 for the Gaussian that saturates it. Chapter 4.2 §4.3 for compatible observables and quantum numbers, §5.3 for the variance written as a length, §7.2 and §7.3 for the argument from unitarity to a Hermitian generator, §7.5 for the first-order action of a generator, §8 for the canonical commutator, and §9 for the tensor product. Chapter 4.6 §2 for U^(t)\hat U(t) and Stone's theorem in the form used here, §5 for momentum in the position representation, §8.7 for the Ehrenfest relations as stated there, and §9.2 for what a stationary state is. Chapter 1.3 §2.1 for what a conjugate momentum is, §6.1 for the classical equation of motion, and §7 for the flow an observable generates. Chapter 1.4 §3 for the three classical symmetries and §7.3 for the algebra of the rotation charges. Chapter 0.4 §4 for change of basis and the similarity transformation. Chapter 4.4 §3 for domains and §5.4 for momentum on a ring. Chapter 4.5 §6 for the projection-valued measure and §9 for Stone. Chapter 4.3 §8.1 for the density of continuous functions in L2L^{2}.

1 · One substitution: p=kp=\hbar k

You have seen the answer before you have seen this section. Chapter 4.6 §10.2 multiplied one Gaussian packet's two widths together and got /2\hbar/2, and Chapter 4.8 §6.1 sat the oscillator ground state on the same number. Both were particular states, taken one shape at a time. What is still owed is the bound for every state, and it costs one multiplication, because Chapter 0.9 did the hard part with no physics in it at all.

1.1 · What was proved in Part 0, and what was missing from it

Chapter 0.9 §6.4 proved the bandwidth theorem: for any normalised ff with finite widths, the spread of ff across positions and the spread of its Fourier transform across wavenumbers satisfy ΔxΔk12\Delta x\,\Delta k\ge\tfrac12. The proof used the definition of the transform, Plancherel, the derivative theorem, Cauchy–Schwarz and one integration by parts, and Chapter 0.9 listed the ingredients in a bulleted list precisely because the shortness of the list was the point. No particle appears in it, no measurement, no observer, and no \hbar. A radio engineer shaping a pulse obeys it, and so does a seismic trace.

What was missing was a reason to care, in physics. The theorem constrains a wave, and a wave has a wavenumber. Nothing in Part 0 said that a particle has one.

1.2 · The substitution, and the line it takes

Chapter 4.6 supplied the missing sentence, and supplied it twice over. Its §5.5 showed that the map between the position and momentum representations of a state is Chapter 0.9's Fourier transform with k=p/k=p/\hbar, and that the map is unitary, so a state normalised in one description is normalised in the other. The density of the measured momentum is therefore the density of the wavenumber re-expressed under p=kp=\hbar k. Its §10.6 marked de Broglie's relation as the one experimental input of that chapter, and that mark stands there rather than being raised again here. Those are the two facts §10.2 used on its Gaussian, and neither of them mentions the Gaussian. Since kkk\mapsto\hbar k is linear, and a linear change of variable multiplies a standard deviation by its slope, we have Δp=Δk\Delta p=\hbar\,\Delta k for any state with both widths finite. So multiply Chapter 0.9's inequality through by \hbar:

ΔxΔp  =  ΔxΔk    2. \Delta x\,\Delta p \;=\; \hbar\,\Delta x\,\Delta k \;\ge\; \frac{\hbar}{2}. (4.9.1)

That is the whole of it. The inequality was never quantum. What is quantum, and what is strange, is the claim that a particle's momentum is the wavenumber of a wave. Grant that one claim and the Heisenberg uncertainty principle for position and momentum stops being an additional mystery and becomes arithmetic you did in Part 0.

Now read the display as a prohibition with a scale on it, because the scale is what makes it a statement about the world. Confine an electron to a region the width of an atom, Δx=1010 m\Delta x=10^{-10}\ \mathrm{m}, and (4.9.1) forces Δp5.27×1025 kgms1\Delta p\ge5.27\times10^{-25}\ \mathrm{kg\,m\,s^{-1}}. That is a spread in speed of 5.8×105 ms15.8\times10^{5}\ \mathrm{m\,s^{-1}}, and an energy of order (Δp)2/2me=0.95 eV(\Delta p)^{2}/2m_{\mathrm e}=0.95\ \mathrm{eV}. The electronvolt is the scale of atomic binding energies, and this is where that scale comes from: confinement to the size of an atom, priced by (4.9.1). Chapter 4.7 §3.5 read a confined particle's ground-state energy in exactly that way and got the same order back.

It is worth saying what has not been proved by that line, because §2 exists to prove it. Equation (4.9.1) is a statement about one pair of quantities, obtained through a Fourier transform, and the Fourier transform was available only because position and momentum are related in that particular way. Two arbitrary observables are not each other's transform, and for them the argument above says nothing at all.

1.3 · Every conjugate pair, and where the dimensions came from

Before leaving, one more paragraph, because it removes a puzzle that would otherwise sit in your way for the rest of Part IV. Chapter 1.3 §2.1 was careful to say that the canonical momentum is not mvm\vv v, and that for the pendulum with coordinate θ\theta it comes out as pθ=m2θ˙p_{\theta}=m\ell^{2}\dot\theta, an angular momentum measured in Js\mathrm{J\,s} rather than in kgms1\mathrm{kg\,m\,s^{-1}}. It then said what pip_{i} really is: whatever pairs with qiq^{i} so that pidqip_{i}\,\dd q^{i} has the dimensions of action.

Chapter 4.1 §5.7 found that the new constant of quantum mechanics has the dimensions of action, and Chapter 4.6 §2.2 spent that fact to fix the \hbar in the exponent. Put the two together and the bound ΔqΔp/2\Delta q\,\Delta p\ge\hbar/2 is dimensionally consistent for every conjugate pair and not only for a Cartesian one, because both sides carry an action either way. An angle paired with an angular momentum works as well as a length paired with a linear momentum.

Whether the inequality holds for those other pairs is a separate question, and a more delicate one than it looks. The angle is the case where the general theorem's fine print does real work, and Worked example 2 is where that is shown rather than described.

In plain terms 4.9.1

A statement that has been called the deepest in twentieth-century physics has just been obtained by multiplying an inequality by a constant. That is not a trick, and it is worth being clear about why it is honest.

The toolkit chapter on Fourier analysis proved that a signal cannot be both brief and pure in pitch. Squeeze it in time and its spectrum widens by the reciprocal factor, always, with the product of the two spreads never falling below one half. That result is about waves. It was known to people designing radios before anyone applied it to matter, and nothing in its proof mentions particles or measurement or observation.

Quantum mechanics contributes exactly one sentence to the story: a particle's momentum is the wavenumber of a wave, multiplied by a constant with the units of action. That sentence is a physical claim with an experiment behind it, and the earlier chapter marked it as such. Accept it, multiply, and the famous inequality drops out. The mystery, if you want one, is entirely in the sentence. It is not in the inequality.

One small piece of tidying comes free, and it is worth being exact about how much of it is free. The constant has the units of action, and the classical chapters had already established that a coordinate and the momentum belonging to it always multiply together to give an action, whatever the coordinate happens to be. So the bound written down for an angle and its angular momentum at least makes dimensional sense, in the same form, with the same constant and no conversion factor anywhere. That much is not luck. It is what "conjugate" was defined to mean.

Whether the inequality is then true for such a pair is a separate question, and the dimensions do not settle it. For the angle the answer is no, until the statement is repaired. The next section says which step breaks, and Worked example 2 exhibits the failure and mends it.

2 · The general relation, and what it is not

Here is where this section is going. We want the same bound for an arbitrary pair of observables, with the right-hand side telling us which pair we are dealing with. The whole derivation is Chapter 0.5's Cauchy–Schwarz applied to two vectors that Chapter 4.2 has already built, and it takes three steps. Then we spend rather longer on what the result means, because the inequality is short and the number of things it is routinely taken to say is not.

2.1 · The two vectors, already in hand

Chapter 4.2 §5.3 derived the variance of an observable from the Born rule and then wrote it in a second form:

(ΔA)2  =  A^2A^2  =  (A^A^)ψ2. (\Delta A)^{2} \;=\; \avg{\hat A^{2}} - \avg{\hat A}^{2} \;=\; \norm{\big(\hat A-\avg{\hat A}\big)\ket\psi}^{2}. (4.9.2)

It said in place why the second form was the one this chapter would need. The spread of an observable is the length of a vector, and Cauchy–Schwarz is a statement about lengths. So name the two vectors, which is the only preparation the argument requires:

f  =  (A^A^)ψ,g  =  (B^B^)ψ. \ket f \;=\; \big(\hat A-\avg{\hat A}\big)\ket\psi, \qquad\qquad \ket g \;=\; \big(\hat B-\avg{\hat B}\big)\ket\psi. (4.9.3)

By (4.9.2) these have lengths f=ΔA\norm f=\Delta A and g=ΔB\norm g=\Delta B. Chapter 0.5's insight box wrote down this pair three parts ago and said that nothing would be added to it here except the meaning of the symbols.

2.2 · Cauchy–Schwarz, once

The inequality of Chapter 0.5 §1.4 says f,g2f,fg,g\abs{\avg{f,g}}^{2}\le\avg{f,f}\avg{g,g} for any two vectors of an inner-product space. Apply it to (4.9.3) and take the square root of both sides, which is legitimate because both sides are non-negative:

ΔAΔB    f,g. \Delta A\,\Delta B \;\ge\; \abs{\avg{f,g}}. (4.9.4)

The uncertainty relation is now proved. Everything remaining is the identification of the right-hand side, which is where the commutator will appear, and it appears because we are about to ask which part of f,g\avg{f,g} survives when the two observables are swapped.

2.3 · Splitting the overlap into two real numbers

We want f,g\avg{f,g} expressed in the observables rather than in the vectors, so move the first operator across the inner product using its Hermiticity. Writing a=A^a=\avg{\hat A} and b=B^b=\avg{\hat B}, both real by Chapter 4.2 §4.1, and abbreviating ΔA^=A^a\Delta\hat A=\hat A-a and ΔB^=B^b\Delta\hat B=\hat B-b,

f,g  =  ψΔA^ΔB^ψ. \avg{f,g} \;=\; \bra\psi\,\Delta\hat A\,\Delta\hat B\,\ket\psi. (4.9.5)

Two symbols differing by one hat are now in play, and they are not each other's value, so fix them apart before they appear in the same display. ΔA^\Delta\hat A is the operator A^a\hat A-a, and its expectation is zero by construction. ΔA\Delta A is a length, the norm of ΔA^ψ\Delta\hat A\ket\psi, and it is the standard deviation the Conventions paragraph fixed. The hat here separates an operator from a number built out of it, not from its own expectation.

That product of two Hermitian operators is not itself Hermitian, and the standard repair is to split it into a part that is and a part that is not. Any product splits that way, by adding and subtracting half of the reversed product:

ΔA^ΔB^  =  12{ΔA^,ΔB^}  +  12[A^,B^], \Delta\hat A\,\Delta\hat B \;=\; \half\big\{\Delta\hat A,\Delta\hat B\big\} \;+\; \half\big[\hat A,\hat B\big], (4.9.6)

where {X^,Y^}=X^Y^+Y^X^\{\hat X,\hat Y\}=\hat X\hat Y+\hat Y\hat X is the anticommutator, and the constants aa and bb have dropped out of the commutator because a number commutes with everything. Now read off the character of each piece. The anticommutator of two Hermitian operators is Hermitian, since (X^Y^+Y^X^)=Y^X^+X^Y^(\hat X\hat Y+\hat Y\hat X)^{\dagger}=\hat Y\hat X+\hat X\hat Y, so its expectation is real. The commutator is anti-Hermitian, since [X^,Y^]=Y^X^X^Y^=[X^,Y^][\hat X,\hat Y]^{\dagger}=\hat Y\hat X-\hat X\hat Y =-[\hat X,\hat Y], so its expectation is purely imaginary. The two terms of (4.9.6) are therefore the real and imaginary parts of one complex number, and a complex number's modulus squared is the sum of the squares of its parts:

f,g2  =  14{ΔA^,ΔB^}2  +  14[A^,B^]2. \abs{\avg{f,g}}^{2} \;=\; \tfrac14\avg{\big\{\Delta\hat A,\Delta\hat B\big\}}^{2} \;+\; \tfrac14\abs{\avg{\big[\hat A,\hat B\big]}}^{2}. (4.9.7)

Both terms are non-negative, so dropping the first can only weaken the statement. Drop it, feed what is left into (4.9.4), and take the square root:

ΔAΔB  spread of the outcomes      12[A^,B^]  how badly they fail to commute   \ann{\Delta A\,\Delta B}{spread of the outcomes} \;\ge\; \ann{\half\abs{\avg{\big[\hat A,\hat B\big]}}}{how badly they fail to commute} (4.9.8)

Three steps, and the third of them was arithmetic on a complex number. Set A^=x^\hat A=\hat x and B^=p^\hat B=\hat p and use Chapter 4.2 §8's [x^,p^]=i[\hat x,\hat p]=\ii\hbar: the right-hand side becomes 12i=/2\tfrac12\abs{\ii\hbar}=\hbar/2 in every state whatsoever, and (4.9.1) comes back. The two derivations share no step. Section 1 went through the Fourier transform and de Broglie's relation, this one through Cauchy–Schwarz and one postulated commutator, and they agree. That agreement is worth more than either route alone.

2.4 · The sharper form, which cost nothing

Look again at what was thrown away. Keeping the anticommutator term in (4.9.7) gives a stronger inequality for the same work:

(ΔA)2(ΔB)2    14{ΔA^,ΔB^}2  +  14[A^,B^]2. (\Delta A)^{2}(\Delta B)^{2} \;\ge\; \tfrac14\avg{\big\{\Delta\hat A,\Delta\hat B\big\}}^{2} \;+\; \tfrac14\abs{\avg{\big[\hat A,\hat B\big]}}^{2}. (4.9.9)

The discarded quantity 12{ΔA^,ΔB^}\tfrac12\avg{\{\Delta\hat A,\Delta\hat B\}} is the covariance of the two observables, in exactly the sense used of any two random quantities, symmetrised because the operators do not commute. So (4.9.8) is what you get by pretending the two are uncorrelated when they need not be. For position and momentum in a state with no correlation between them the two forms agree, and Problem 1 exhibits a perfectly ordinary Gaussian in which they do not. The geometry of that difference is drawn in the figure at the end of this section, with two caveats and a pair of numerical checks standing between here and there.

2.5 · The hypothesis, which is not decoration

Every line above assumed it was legal to move an operator across an inner product. Chapter 4.4 spent a whole chapter on when that is legal, and the answer was not "always". Written out, the derivation needs ΔB^ψ\Delta\hat B\ket\psi to lie in the domain of A^\hat A and ΔA^ψ\Delta\hat A\ket\psi to lie in the domain of B^\hat B, so that (4.9.5) is an identity rather than a hope. Call it the domain hypothesis. For a bounded operator it costs nothing, since Chapter 4.4 §2.1 showed bounded and continuous to be the same condition and a continuous operator on a dense domain extends by continuity to the whole space. For unbounded ones it is a genuine condition.

Something the domain hypothesis quietly requires is easier to miss, because it looks like a formality and is not. Both sides of (4.9.8) assume that ΔA\Delta A and ΔB\Delta B exist, and a normalised state is under no obligation to supply them. Chapter 4.3 §5.5 gave the structural reason, that L2(R)L^{2}(\R) does not sit inside L1(R)L^{1}(\R), and said the physical reading belonged in this chapter. Here it is. Normalisation is a statement about ψ2\int\abs\psi^{2} and guarantees nothing about xψ2\int x\abs\psi^{2}, so a perfectly legitimate state can have no mean position and no Δx\Delta x at all, as that chapter's Worked example 3 exhibits. For such a state the relation is true and empty, in exactly the way Chapter 0.9 §6.5's grind box said the bandwidth theorem is true and empty when a width is infinite. Nothing has gone wrong. The inequality compares two numbers, and it says nothing in a state where one of them is not a number.

The domain hypothesis itself fails in the most quoted example after position and momentum. Take an angle and the angular momentum conjugate to it. The naive reading of §1.3 would give ΔφΔLz/2\Delta\varphi\,\Delta L_{z}\ge\hbar/2. That statement is false, and Worked example 2 exhibits a state in which the left-hand side is exactly zero, identifies the line of the derivation that breaks, and computes the boundary term the breakage leaves behind. The repair is to replace the angle by a bounded periodic function of it, after which the hypothesis holds and the theorem applies with no modification. Chapter 4.4 §5.4 built the operator involved, so the machinery for the repair is already on the shelf.

2.6 · What the relation does not say

⚑ Quoted, not derived — what the relation does not say, and two theorems about the question it gets confused with

The relation is about preparation, not about disturbance. Read (4.9.8) back through its own definitions. Each Δ\Delta is a standard deviation of measured values over an ensemble of systems prepared identically in ψ\psi, with one measurement performed on each. No system in that ensemble is measured twice. Nothing in the derivation refers to an apparatus, to an order of measurements, or to any effect one measurement has on another, because no such notion appears anywhere in Chapters 0.5 or 4.2. The content is that the state ψ\psi cannot drive the product of the two spreads below what the commutator allows in that same state. It is a constraint on what can be prepared.

Chapter 0.9's warning box made the same point about the wave version and promised that this chapter would keep the two ideas apart. Here is the separation, stated as sharply as it can be. Measurement disturbance is real, it is a different quantity, and it has theorems of its own. Those theorems are not proved in this book, and this is the one place in this chapter where something is quoted.

Ozawa's relation (2003). Model a measurement as a unitary interaction between the system and a probe, followed by a sharp reading of a meter observable on the probe. Define the error ε(A)\varepsilon(A) as the root-mean-square difference, in the given input state, between what the meter reports and what A^\hat A would have given, and the disturbance η(B)\eta(B) as the root-mean-square change the interaction inflicts on B^\hat B. With those definitions the naive product ε(A)η(B)12[A^,B^]\varepsilon(A)\eta(B)\ge\tfrac12\abs{\avg{[\hat A,\hat B]}} is false, and there are explicit measurement models that violate it. What holds instead carries two extra terms, ε(A)η(B)+ε(A)ΔB+ΔAη(B)12[A^,B^]\varepsilon(A)\eta(B)+\varepsilon(A)\,\Delta B+\Delta A\,\eta(B)\ge\tfrac12\abs{\avg{[\hat A,\hat B]}}, in which ΔA\Delta A and ΔB\Delta B are the very spreads of (4.9.8). The hypotheses are the indirect-measurement model described above, the root-mean-square definitions of ε\varepsilon and η\eta, and enough regularity for the relevant expectations to exist. It is a state-dependent statement.

Read what the two extra terms buy, because that is where the naive product went wrong. They let the state's own spreads carry the bound. When ΔA\Delta A and ΔB\Delta B are wide, the cross terms ε(A)ΔB\varepsilon(A)\Delta B and ΔAη(B)\Delta A\,\eta(B) can reach 12[A^,B^]\tfrac12\abs{\avg{[\hat A,\hat B]}} between them, and ε(A)η(B)\varepsilon(A)\eta(B) is then free to be as small as the apparatus can make it. So an accurate measurement of AA need not inflict a large disturbance on BB, which is precisely what the naive product forbade and what the explicit models exhibit.

The Busch–Lahti–Werner relation (2013). Define error and disturbance differently, as worst-case figures of merit obtained by calibrating the apparatus against states in which the reference observable is arbitrarily sharp, rather than as averages in one input state. With those definitions, and for the canonical pair position and momentum, the product of the calibration error and the calibration disturbance is bounded below by /2\hbar/2 after all. The hypotheses are the calibration definition of error and the pair being canonically conjugate.

The two are not in conflict, and the reason is worth carrying. They bound different quantities. "How accurate is this apparatus" admits more than one honest definition, and which inequality you get depends on which one you chose. The diagnostic, when you next meet a claim in the wild, is to ask which state the error was computed in. An error quoted for the one input state at hand is Ozawa's quantity, and there the naive product is false. An error quoted as a worst case over calibration states in which the reference observable is arbitrarily sharp is Busch–Lahti–Werner's, and there the naive product holds for position and momentum after all. Neither is (4.9.8), and neither is derivable from it, which is exactly why Chapter 0.9 said there were separate theorems and why naming none of them here would have been an under-delivery. What this book proves is the preparation statement. Both relations above are covered by this box's mark, and it is the only one in the chapter.

2.7 · The relation, measured

Two checks, neither of which uses the derivation above, and the first of which is the one to remember. A third, on angular momentum, waits until §6.3, where the commutator it needs has been built.

The oscillator. Chapter 4.8 obtains x^2\avg{\hat x^{2}} and p^2\avg{\hat p^{2}} for the energy eigenstates by algebra, giving an uncertainty product of (n+12)(n+\half)\hbar. Computing the same two integrals numerically for the Hermite functions at forty-digit working precision returns ΔxΔp=0.500000000000\Delta x\,\Delta p=0.500000000000 for the ground state and 3.5000000000003.500000000000 for n=3n=3 in units where =m=ω=1\hbar=m=\omega=1, printed to twelve figures, and departing from 12\tfrac12\hbar and 72\tfrac72\hbar by less than 104010^{-40}, which is the quadrature's error at that working precision rather than a property of the states. Read those two numbers side by side. The ground state sits exactly on the floor set by (4.9.1), and the n=3n=3 state sits seven times above it. The relation is a floor and not a prediction. It says how small the product cannot be, and it says nothing whatever about how large it is in any particular state. Chapter 4.7 §3.5 made the same point from the other end before this theorem existed, reading the infinite well's ground-state energy as an uncertainty and getting ΔxΔp=0.567862\Delta x\,\Delta p=0.567862\,\hbar against a floor of /2\hbar/2. An instance computed before the theorem is worth more than the theorem alone, and both of those instances were.

Random pairs. Two hundred thousand random Hermitian pairs and random states in dimensions two to five, all built independently of the derivation, satisfy both (4.9.8) and (4.9.9) in every trial. The smallest slack in the sharper form (4.9.9) is 3.0×1014-3.0\times10^{-14}, which is roundoff and means the sharper form is attained; the smallest slack in (4.9.8) over the same trials is 7.8×10107.8\times10^{-10}. The identity (4.9.7), on which the whole split rests, holds over twenty thousand random cases with a worst residual of 2.5×10142.5\times10^{-14}.

The uncertainty relation as a fact about lengths. Left: the two vectors of (4.9.3) for the first excited state of the oscillator, in units where =m=ω=1\hbar=m=\omega=1. Their lengths are Δx\Delta x and Δp\Delta p, both 3/2=1.2247\sqrt{3/2}=1.2247. The dashed line drops g\ket g onto the line through f\ket f, and Chapter 0.5's grind box read Cauchy–Schwarz as exactly this: a projection is never longer than the vector being projected. The projection here has length f,g/Δx=0.4082\abs{\avg{f,g}}/\Delta x=0.4082, which is 1/31/3 of g=Δp\norm g=\Delta p, so f,g=0.500\abs{\avg{f,g}}=0.500 cannot exceed ΔxΔp=1.500\Delta x\,\Delta p=1.500. The drawing is the real plane spanned by the projection and the leftover, which is all the geometry the inequality uses. Right: the same statement seen in the complex plane. Every state puts the overlap f,g\avg{f,g} somewhere in the disc of radius ΔAΔB\Delta A\,\Delta B, and the picture is drawn with that radius scaled to one. Both coordinates carry that scaling: the horizontal one is the covariance 12{ΔA^,ΔB^}\tfrac12\avg{\{\Delta\hat A,\Delta\hat B\}} divided by ΔAΔB\Delta A\,\Delta B, and the vertical one is 12[A^,B^]\tfrac12\abs{\avg{[\hat A,\hat B]}} divided by the same product, by (4.9.6). So (4.9.9) says the point lies in the disc and (4.9.8) says its height is at most one. Marked: the Gaussian ground state at the top of the circle, saturating both; two chirped Gaussians of Problem 1, on the circle but not at the top, saturating the sharper form only; and the n=1n=1 and n=3n=3 oscillator states well inside, saturating neither. The gap between a point and the top of the circle is precisely the covariance that (4.9.8) throws away.
Familiar ground — a spread over identical preparations, and the one thing it is not

You report standard deviations for a living, and ΔA\Delta A is one of yours. Prepare NN systems the same way, measure A^\hat A once on each, and compute the sample standard deviation of the NN numbers. That is ΔA\Delta A, in the same arithmetic you would apply to NN tumour volumes or NN trough concentrations, and Chapter 4.2 §5.3 derived it from the Born rule rather than defining it. Nobody has ever measured a standard deviation on one patient, and nobody has ever measured ΔA\Delta A on one system. This is why the inequality cannot be about what happens to an individual when you poke it.

Chapter 4.5 §6.5's familiar-ground box named the place where the parallel stops and left it for this chapter, so here it is. When you see two correlated readings vary across a cohort, you assume without thinking that there is a joint distribution underneath: each patient has some true pair of values, you are seeing a marginal of it, and the variation reflects covariates you did not measure. Every method you have for such data is built on that assumption, and it is nearly always right.

The word that will not carry across is covariate. Finding one does not narrow a marginal. It splits the cohort into strata and narrows the distribution inside each, and the marginal you started with is exactly where it was, because nobody has been intervened upon. On this side there is nothing to stratify and there is something to do. A preparation is an act, so the question is not what you might discover about the systems. It is whether some other preparation makes both spreads small at once.

For position and momentum, none does. The commutator is i\ii\hbar in every state whatsoever, so the right-hand side of (4.9.8) is /2\hbar/2 however you prepare, and the product has a floor no state gets under. There is no experiment to design and no covariate to go looking for. That much is proved here.

Say it for this pair and not in general, because the floor is not always a number. For a general pair the right-hand side of (4.9.8) is evaluated in the state, and §6.3 exhibits two components of angular momentum in a state where it comes out zero while neither spread is small. There the theorem has fallen silent rather than been met, and what still forbids both from being sharp is the non-zero commutator of §3.1 rather than anything the inequality has supplied.

The stronger statement, that there is no joint distribution underneath at all, is not proved by this inequality and should not be claimed from it. Chapter 4.5 §6.5's familiar-ground box raises it and leaves it open, Chapter 4.11 shows that the failure to be simultaneously sharp is structural rather than a matter of ignorance, and Chapter 4.20 settles it with an experiment. What you may take from this chapter is the operational half, and it is the half that changes how you read the inequality: the spread is a property of the preparation, not a measure of what you have not yet found out.

In plain terms 4.9.2

The general statement takes three steps and it is worth watching how little each one costs. The spread of a measurement, defined honestly as a standard deviation over repeated identical preparations, was shown two chapters ago to be the length of a particular vector built from the state and the quantity being measured. So take two quantities, build the two vectors, and apply the inequality proved in the toolkit chapter, which says that the overlap of two vectors is never bigger than the product of their lengths. That single application is the uncertainty principle.

What remains is bookkeeping on the overlap. It is a complex number, and its two parts have different meanings: the real part measures how the two quantities co-vary, and the imaginary part measures how badly the two operations fail to commute. Since a complex number is at least as big as either part alone, throwing away the real part costs nothing and leaves the familiar form, with the failure to commute sitting on the right-hand side. Keep the real part instead and you get a slightly stronger statement for the same work.

Two warnings, both earned. The first is that this is a statement about preparation. Every symbol in it refers to a spread across many systems prepared identically and each measured once. It says nothing about an apparatus knocking a particle about, and the theorems that do say something about that are separate results with different hypotheses, named above and not proved here.

The second is that the bound is a floor rather than a forecast. The lowest oscillator state sits exactly on it. The fourth one sits seven times above it, and is no less legitimate for that. And when the right-hand side happens to vanish in some particular state, the theorem has not announced that both quantities are sharp. It has fallen silent.

a natural place to stop  ·  the bound is proved and spent; what follows is the commutator when it vanishes, and then when it moves things

3 · Compatible observables, and a complete set

Section 2 asked what happens when a commutator is large. This section asks what happens when it is zero, and then answers a question Chapter 4.2 §4.3 raised, named this section for, and deliberately left open. The answer is a criterion, and the criterion is what Chapters 4.11 and 4.13 will apply every time they label a state.

3.1 · The two halves, now both proved

Chapter 0.5 §8 proved that two Hermitian operators admit a common orthonormal eigenbasis exactly when they commute, and Chapter 4.2 §4.3 renamed both halves. If [A^,B^]=0[\hat A,\hat B]=0 the two are compatible: there are states in which both quantities have definite values, and enough of them to span the space. If [A^,B^]0[\hat A,\hat B]\neq0 there is no common eigenbasis, hence no state at all in which both are sharp.

That second half is the qualitative content of the uncertainty principle: a non-zero commutator rules out every state in which both are sharp, without saying by how much. Equation (4.9.8) is now the quantitative version, and the two fit together with one caution worth fixing now. A non-zero commutator forbids a common eigenbasis outright. The inequality, by contrast, is evaluated in a particular state, and [A^,B^]\avg{[\hat A,\hat B]} can vanish in a state even when the operator [A^,B^][\hat A,\hat B] does not. In that state the bound has fallen silent while the prohibition still stands, and §6.3 puts a number on exactly that case.

3.2 · The question Chapter 4.2 left here

Chapter 4.2 §4.3 gave the definition this part of the book runs on. A set {A^1,,A^m}\{\hat A_{1},\dots,\hat A_{m}\} of mutually commuting observables is a complete set of commuting observables when their common eigenspaces are all one-dimensional, so that the list of eigenvalues (a1,,am)(a_{1},\dots,a_{m}) fixes the state up to phase. That list is what the word quantum numbers means, and writing a hydrogen state as n,,m\ket{n,\ell,m_{\ell}} is naming the eigenvalues of three commuting operators and nothing more.

What it left open was how one knows a set is complete. It is a fair question and it is not answered by the definition, because the definition quantifies over eigenspaces you would have to find first. What you want is a test you can run on the operators.

3.3 · Maximality is the same condition

Here is the test, and its usefulness is that it never mentions eigenspaces. One phrase in it does all the work, so fix its meaning before the statement arrives. A function of a commuting set means what Chapter 0.5 §7 meant for a single operator, extended from one operator to a family: choose a number c(a1,,am)c(a_{1},\dots,a_{m}) for each tuple of joint eigenvalues, and let the operator multiply by that number on the corresponding common eigenspace, so that C^=c(a1,,am)P^(a)\hat C=\sum c(a_{1},\dots,a_{m})\hat P_{(a)} with P^(a)\hat P_{(a)} the projection onto it. Nothing else counts as a function of the set.

Theorem. A set of mutually commuting observables on a finite-dimensional space has all its common eigenspaces one-dimensional if and only if every observable that commutes with all of them is a function of them.

A set with that property is called maximal: you cannot add a genuinely new commuting observable, because anything you might add is already there. Both directions are short, and both use the same fact about a common eigenspace.

One-dimensional common eigenspaces imply maximality. Let C^\hat C commute with every A^i\hat A_{i}, and let v\ket v be a common eigenvector with eigenvalues (a1,,am)(a_{1},\dots,a_{m}). Then A^iC^v=C^A^iv=aiC^v\hat A_{i}\hat C\ket v=\hat C\hat A_{i}\ket v=a_{i}\hat C\ket v, so C^v\hat C\ket v sits in the same common eigenspace. That eigenspace is one-dimensional and contains v\ket v, so C^v\hat C\ket v is a multiple of v\ket v. Every common eigenvector is therefore an eigenvector of C^\hat C, and since those eigenvectors form a basis by Chapter 0.5 §8, we may define cc on the tuples and write C^=c(a1,,am)P^(a)\hat C=\sum c(a_{1},\dots,a_{m})\hat P_{(a)}, which is a function of the A^i\hat A_{i} in exactly the sense just fixed.

A common eigenspace of dimension two or more destroys maximality. Let WW be such an eigenspace and pick a unit vector u\ket u inside it. Take C^=uu\hat C=\ket u\bra u. It commutes with each A^i\hat A_{i}, because A^iuu=aiuu\hat A_{i}\ket u\bra u=a_{i}\ket u\bra u and, using Hermiticity, uuA^i=aiuu\ket u\bra u\hat A_{i}=a_{i}\ket u\bra u as well. But any function of the A^i\hat A_{i} acts on all of WW as a single multiple of the identity, since WW carries one tuple of eigenvalues, whereas C^\hat C annihilates the part of WW orthogonal to u\ket u, which is not empty. So C^\hat C commutes with everything in the set and is not a function of the set. \blacksquare

3.4 · How you show it in practice: exhibit the count

What §3.3 buys is not the working test. It makes the definition safe, by showing that the property you cannot check without finding eigenspaces is the same property as one you can state about the operators alone, so nothing turns on which of the two anyone happens to mean. The working test comes from the definition and the basis, and needs neither direction of the theorem. Since the joint eigenvectors form a basis, a set is complete exactly when the number of distinct eigenvalue tuples equals the dimension of the space. So you do not test maximality directly. You list the tuples and you count.

  • Angular momentum. On a single multiplet of dimension 2j+12j+1, the set {J^2,J^z}\{\hat J^{2},\hat J_{z}\} produces the tuples (j(j+1)2,m)(j(j+1)\hbar^{2},m\hbar) for m=j,,+jm=-j,\dots,+j, which is 2j+12j+1 of them. The count matches and the set is complete. Chapter 4.11 constructs the multiplet and then makes exactly this remark.
  • Hydrogen. The set {H^,L^2,L^z}\{\hat H,\hat L^{2},\hat L_{z}\} produces, at fixed principal quantum number nn, the tuples with =0,,n1\ell=0,\dots,n-1 and m=,,+m_{\ell}=-\ell,\dots,+\ell, so the count is =0n1(2+1)\sum_{\ell=0}^{n-1}(2\ell+1). That sum telescopes to n2n^{2}, since consecutive squares differ by consecutive odd numbers. Chapter 4.13 computes the dimension of the nn-th energy level and finds n2n^{2}, at which point the count matches and the set is complete; Chapter 4.14 explains why the degeneracy is that number and not something else. What belongs here is the criterion those chapters apply.
  • A set that fails. Problem 4 builds two qubits and shows that the total of the two "which-state" observables is not by itself complete. It then exhibits an operator commuting with that total which is not a function of it, and completes the set in two different ways whose members do not commute with each other. That last point is the entire structure of Chapter 4.12's coupled and uncoupled bases, met early and with no angular momentum in it.

3.5 · What the criterion becomes in infinite dimensions

Name the difficulty rather than stepping over it. The proof in §3.3 used a basis of eigenvectors, and Chapter 4.5 §2 showed that an observable in infinite dimensions may have no eigenvectors at all. The replacement is the projection-valued measure of Chapter 4.5 §6. Commuting observables have commuting spectral measures, and the joint measure assigns a projection to each region of the joint spectrum. The set is complete when that joint measure has multiplicity one, meaning that no region can be split further by any projection commuting with all of them, which is the one-dimensional condition with "dimension" replaced by "cannot be cut".

This book needs the finite-dimensional statement and needs it repeatedly, because every use in Chapters 4.11 to 4.16 is a count inside a finite-dimensional eigenspace of H^\hat H. The infinite-dimensional version is stated so the definition is not silently changed later, and the momentum components of a free particle in three dimensions are the standard example: one joint generalised eigenstate per momentum vector, nothing left to cut, and Chapter 4.5 §7 is where such objects were given their meaning.

One separate caution is worth parking here before §4 begins, since it belongs to the criterion rather than to infinite dimensions and Chapter 4.2 §4.3 stated it with a figure to go with it. The theorem promises that a common eigenbasis exists. It does not promise that a numerical eigensolver hands you that one, because inside a degenerate eigenspace the solver returned some basis rather than the basis the second observable prefers. Chapter 4.13's degeneracies are exactly where that bites.

What you now hold is a test you can run on paper: list the tuples, count them, compare with the dimension. Chapters 4.11 and 4.13 run it on angular momentum and on hydrogen, and Problem 4 runs it on two qubits, on a set that passes and a set that fails. The chapter now changes subject from a single instant to time.

In plain terms 4.9.3

Two quantities that commute can be sharp together, and the states in which both are sharp are numerous enough to describe everything. That much was proved in the toolkit chapter. The question left open was practical: given a handful of such quantities, how do you know you have enough of them to tell every state apart?

The definition says you have enough when knowing all their values pins the state down. The useful reformulation says you have enough when the list cannot be extended: any further quantity that commutes with all of yours is already some combination of them, so adding it would tell you nothing new. The two conditions are the same condition, and the proof of that is half a page.

In practice nobody checks either one directly. You write down the possible combinations of values, you count them, and you compare that count with the number of independent states available. If the two numbers agree, the labels are enough. That is the whole method, and it is what lets an atomic state be written as three numbers in a bracket, those three labels being the whole of what distinguishes it from every other state.

4 · The Heisenberg picture, and the Heisenberg equation

Everything so far has been about a single instant. This section puts the time back in, and the first thing to settle is where the time is allowed to sit. The answer is that there is a choice, that the choice is a change of basis rather than a change of physics, and that making it produces an equation of motion for observables which is the exact quantum copy of the one Chapter 1.3 derived classically. That equation is what §5 and §6 both run on.

4.1 · Where the time can be put

Every prediction the theory makes is an expectation value. Chapter 4.2 §5.3 derived A^=ψA^ψ\avg{\hat A}=\bra\psi\hat A\ket\psi from the Born rule, and Chapter 4.2 §5 showed that a probability is such a number with A^\hat A a projection, so there is nothing else to compute. Take one at time tt and substitute the evolution ψ(t)=U^(t)ψ(0)\ket{\psi(t)}=\hat U(t)\ket{\psi(0)} that Chapter 4.6 §2 produced:

A^(t)  =  ψ(0)U^(t)A^U^(t)ψ(0). \avg{\hat A}(t) \;=\; \bra{\psi(0)}\,\hat U^{\dagger}(t)\,\hat A\,\hat U(t)\,\ket{\psi(0)}. (4.9.10)

Now look at the middle of that expression and notice that the brackets can be grouped two ways. Group the U^\hat U's with the state on either side and you have the picture used since Chapter 4.6: the state moves and the operator sits still. Group them with the operator instead and you have a different bookkeeping with the same value, in which the state sits still and the operator moves. Define

A^H(t)    U^(t)A^U^(t),ψH    ψ(0), \hat A_{H}(t) \;\equiv\; \hat U^{\dagger}(t)\,\hat A\,\hat U(t), \qquad\qquad \ket{\psi_{H}} \;\equiv\; \ket{\psi(0)}, (4.9.11)

and (4.9.10) reads A^(t)=ψHA^H(t)ψH\avg{\hat A}(t)=\bra{\psi_{H}}\hat A_{H}(t)\ket{\psi_{H}}. This is the Heisenberg picture, and the one before it is the Schrödinger picture. Since every prediction of the theory is one of these numbers and the number has not changed, the two pictures agree on everything measurable by construction, and not as a theorem that could have come out otherwise.

4.2 · It is a change of basis, and the domain the basis does not carry

Look at the form of (4.9.11) rather than at what it does. Chapter 0.4 §4.2 derived the transformation of a matrix under a change of basis and got A=P1APA'=P^{-1}AP, calling two matrices related that way similar and summarising the content as: similar matrices are the same map, seen twice. Here P=U^(t)P=\hat U(t), and U^1=U^\hat U^{-1}=\hat U^{\dagger} because U^\hat U is unitary. So (4.9.11) is a similarity transformation, and as algebra it is nothing else. The Heisenberg picture is the Schrödinger picture written in a basis that rotates along with the state.

Chapter 0.4's warning box listed the consequences of confusing a map with its matrix, and one of them was "why the Schrödinger and Heisenberg pictures look like different physics instead of different bases". Collect it here, and pay for it, because in infinite dimensions the algebra leaves something out and Chapters 4.4 and 4.5 were spent on exactly that thing. An operator is a formula together with a domain, and Chapter 4.4 §3.1 fixed that two operators are the same operator only when the formulae agree and the domains do. Conjugation carries the domain along with the formula, so dom(A^H(t))=U^(t)dom(A^)\operatorname{dom}\big(\hat A_{H}(t)\big)=\hat U^{\dagger}(t)\operatorname{dom}(\hat A), which for an unbounded A^\hat A is a different subspace at every time. That is what makes A^H(t)\hat A_{H}(t) a genuinely different operator at each instant rather than one operator wearing a rotating coat, and it is why the differentiation of §4.3 is a strong derivative taken on vectors that stay inside the moving domain rather than a derivative in the operator norm. In finite dimensions the slogan costs a sentence about similar matrices. Here it costs the domain as well, and the domain is what the algebra is silent about.

The claim has consequences you can check rather than admire, and the one that matters is the list of readings. Chapter 4.5 §2.1 took invertibility rather than an eigenvector as the primitive, exactly so that this kind of question could be asked about x^\hat x and H^\hat H, which have no determinant, no trace and in general no eigenvectors. Use it here. Since A^Hλ=U^(A^λ)U^\hat A_{H}-\lambda=\hat U^{\dagger}(\hat A-\lambda)\hat U, the operator U^(A^λ)1U^\hat U^{\dagger}(\hat A-\lambda)^{-1}\hat U is a bounded everywhere-defined inverse of the left-hand side exactly when (A^λ)1(\hat A-\lambda)^{-1} is one for A^λ\hat A-\lambda. The two resolvent sets are therefore the same set, so the spectrum of A^H(t)\hat A_{H}(t) equals the spectrum of A^\hat A at every time. The list of values a measurement of AA can return does not move. If it did, the two pictures would disagree about what an apparatus can read, and no amount of algebra would repair that.

4.3 · Differentiating, and the equation that falls out

We have A^H(t)\hat A_{H}(t) as an explicit formula in tt, and what we want is its rate of change, so differentiate (4.9.11) with the product rule. Chapter 4.6 §3.1 supplies the derivative of the evolution operator, dU^dt=iH^U^\dv{\hat U}{t}=-\tfrac{\ii}{\hbar}\hat H\hat U, and taking the adjoint of that gives dU^dt=+iU^H^\dv{\hat U^{\dagger}}{t}=+\tfrac{\ii}{\hbar}\hat U^{\dagger}\hat H since H^\hat H is self-adjoint. Allowing A^\hat A its own explicit time dependence as well, the three terms are

dA^Hdt  =  iU^H^A^U^    iU^A^H^U^  +  U^(A^t)U^. \dv{\hat A_{H}}{t} \;=\; \frac{\ii}{\hbar}\hat U^{\dagger}\hat H\hat A\hat U \;-\; \frac{\ii}{\hbar}\hat U^{\dagger}\hat A\hat H\hat U \;+\; \hat U^{\dagger}\Big(\pdv{\hat A}{t}\Big)\hat U. (4.9.12)

The first two terms combine, and the operations have to go in one particular order. Insert U^U^=I^\hat U\hat U^{\dagger}=\hat I between the operator and the Hamiltonian in each term first, which converts each factor separately into its Heisenberg form and leaves i(H^HA^HA^HH^H)\tfrac{\ii}{\hbar}\big(\hat H_{H}\hat A_{H}-\hat A_{H}\hat H_{H}\big). Only now is there a commutator to read off, and it is i[H^H,A^H]\tfrac{\ii}{\hbar}\big[\hat H_{H},\hat A_{H}\big]. Two sign changes then bring that to the form written below and they cancel: reversing the bracket into [A^H,H^H]\big[\hat A_{H},\hat H_{H}\big] contributes one, and i=1i\tfrac{\ii}{\hbar}=-\tfrac{1}{\ii\hbar} contributes the other. For a Hamiltonian with no explicit time dependence H^\hat H commutes with U^\hat U, so H^H=H^\hat H_{H}=\hat H and no subscript is needed on it. What is left is the Heisenberg equation:

dA^Hdt  =  1i[A^H,H^]  motion the dynamics supplies    +  (A^t)H  motion you put in by hand   \dv{\hat A_{H}}{t} \;=\; \ann{\frac{1}{\ii\hbar}\big[\hat A_{H},\hat H\big]}{motion the dynamics supplies} \;+\; \ann{\Big(\pdv{\hat A}{t}\Big)_{H}}{motion you put in by hand} (4.9.13)

The two terms answer to different things and it is worth keeping them apart from the start. The second is present only if the observable was defined with a time in it, as x^cosωt\hat x\cos\omega t is, and it has nothing to do with the dynamics. The first is the dynamics, entirely, and it is a commutator.

4.4 · Beside the classical equation, term by term

Now set that against what Chapter 1.3 §6.1 derived, before any quantum mechanics existed in this book, for the rate of change of an arbitrary function on phase space:

dfdt  =  {f,H}  +  ft. \dv{f}{t} \;=\; \{f,H\} \;+\; \pdv{f}{t}. (4.9.14)

Chapter 1.3 called this the equation of motion for every observable of every Hamiltonian system, and noted that it contains Hamilton's equations as the two special cases f=qif=q^{i} and f=pif=p_{i}. Compare the two equations piece by piece:

  • The left-hand sides are the same, with a function replaced by an operator.
  • The explicit-time terms are the same, with the same meaning and the same irrelevance to the dynamics.
  • The Hamiltonian appears in the same slot in both.
  • The only difference anywhere is which bracket. The Poisson bracket {  ,  }\{\;,\;\} has become 1i[  ,  ]\tfrac{1}{\ii\hbar}[\;,\;].

That substitution is not being introduced here. Chapter 1.3 §6.4 wrote it down and Chapter 1.3's closing brick sent the bracket to this chapter to be replaced by 1i[  ,  ]\tfrac{1}{\ii\hbar}[\;,\;], while Chapter 4.2 §7.5 derived the same replacement from the first-order expansion of a unitary. What is new is the equation itself, obtained from U^(t)\hat U(t) and the product rule rather than by analogy. Chapter 1.3 also warned that the substitution cannot be extended to every pair of classical observables at once, and that warning is still standing; Chapter 4.10 §8 is where it is proved.

4.5 · Conservation, in one line, and stronger than the classical version

Set A^t=0\pdv{\hat A}{t}=0 in (4.9.13) and the left-hand side vanishes exactly when the commutator does:

A^ is conserved[A^,H^]=0. \hat A \text{ is conserved} \qquad\Longleftrightarrow\qquad \big[\hat A,\hat H\big]=0. (4.9.15)

Chapter 1.3 obtained the same statement with the same shape, and there too conservation stopped being something you established by solving the equations of motion and became something you established by computing one bracket. The quantum version says slightly more than the classical one, and the extra is worth having. What is constant is the operator A^H\hat A_{H}, not merely its expectation. To see how much that buys, run it through the projections rather than the moments, since constant moments do not by themselves pin a distribution down for an unbounded observable. Because [A^,H^]=0[\hat A,\hat H]=0 makes A^\hat A's spectral projections of Chapter 4.5 §6 commute with U^(t)\hat U(t), the probability P(E)ψ(t)2\norm{P(E)\ket{\psi(t)}}^{2} of a reading landing in any region EE is the same at every time. The entire probability distribution of the measured values of AA is frozen, not merely its mean. Chapter 4.6's Worked example 1 found precisely this for the momentum of a free particle and obtained it there by a special argument. Here it is the general case.

4.6 · Two commutators the rest of the chapter needs

Before using (4.9.13) on anything, we need to be able to compute the commutator of a Hamiltonian with something. One rule does almost all of the work. Expanding both sides and cancelling the two middle terms gives, for any three operators,

[A^,B^C^]  =  [A^,B^]C^  +  B^[A^,C^], \big[\hat A,\hat B\hat C\big] \;=\; \big[\hat A,\hat B\big]\hat C \;+\; \hat B\big[\hat A,\hat C\big], (4.9.16)

which is the product rule with the ordering respected. Since the commutator with a fixed operator therefore behaves like a derivative, everything below is one application of it. With [x^,p^]=i[\hat x,\hat p]=\ii\hbar from Chapter 4.2 §8, applying (4.9.16) to p^p^\hat p\hat p gives

[x^,p^2]  =  [x^,p^]p^+p^[x^,p^]  =  2ip^. \big[\hat x,\hat p^{2}\big] \;=\; \big[\hat x,\hat p\big]\hat p + \hat p\big[\hat x,\hat p\big] \;=\; 2\ii\hbar\,\hat p. (4.9.17)

The other one needs the position representation, and Chapter 4.6 §5 established that there p^=ix\hat p=-\ii\hbar\,\partial_{x} and V(x^)V(\hat x) is multiplication by V(x)V(x). Acting on any differentiable ψ\psi in the domain, the product rule gives p^(Vψ)=iVψiVψ\hat p(V\psi)=-\ii\hbar V'\psi-\ii\hbar V\psi' while Vp^ψ=iVψV\hat p\psi=-\ii\hbar V\psi', and the second terms cancel:

[p^,V(x^)]  =  iV(x^). \big[\hat p,V(\hat x)\big] \;=\; -\ii\hbar\,V'(\hat x). (4.9.18)

Both are verified symbolically. Now feed the Hamiltonian H^=p^2/2m+V(x^)\hat H=\hat p^{2}/2m+V(\hat x) of Chapter 4.6 §4 into (4.9.13) twice, once with A^=x^\hat A=\hat x and once with A^=p^\hat A=\hat p. Since x^\hat x commutes with V(x^)V(\hat x) and p^\hat p commutes with p^2\hat p^{2}, only one term survives in each case, and the factors of i\ii\hbar cancel:

dx^Hdt  =  p^Hm,dp^Hdt  =  V(x^H). \dv{\hat x_{H}}{t} \;=\; \frac{\hat p_{H}}{m}, \qquad\qquad \dv{\hat p_{H}}{t} \;=\; -\,V'(\hat x_{H}). (4.9.19)

Those are Hamilton's equations of Chapter 1.3 §3 with hats on everything, and they are exact operator identities with no approximation anywhere in them. Section 5 takes their expectation values and Problem 2 solves them in closed form for the oscillator.

4.7 · Which picture is used where

Neither picture is more correct, so the choice is made on convenience, and the convenience runs in opposite directions in the two halves of this book. Everything in Part IV is easier in the Schrödinger picture, because a bound-state problem is an eigenvalue problem for a fixed operator and Chapters 4.7 and 4.8 want the operator to hold still while they diagonalise it.

Part V reverses that. A field is an operator attached to each point of spacetime, and its time dependence has to sit in the operator rather than in the state. Otherwise the time coordinate is treated differently from the three spatial ones, and the relativity of Part II is invisible in the notation. Chapters 5.2 and 5.3 are written in the Heisenberg picture throughout for that reason, and (4.9.13) is the equation they solve. Building it here costs one section; discovering it there would cost a chapter.

In plain terms 4.9.4

Everything the theory predicts is an average of the form "state, operator, state". There are two places the time can be kept in such an expression, and both give the same number. Keep it in the state and the state evolves while the measured quantities stand still, which is the picture used so far. Move it onto the quantity instead and the state is frozen while the operators evolve.

This is not two theories. The manoeuvre is exactly the change of basis met in the linear algebra chapter, where one map acquires a different array of numbers in a different basis while remaining one map. Failing to see that is why the two pictures look like rival physics. The check is that the list of possible measured values, which is a property of the map and not of the basis, does not budge.

Differentiating the moved operator gives its equation of motion, and the answer is a bracket with the energy. Set it beside the classical equation of motion for an arbitrary quantity, derived in Part I from Hamilton's equations, and the two are identical except that one uses the classical bracket and the other uses the commutator divided by i\ii\hbar. Nothing else differs. That single replacement is what the phrase "canonical quantisation" refers to.

The immediate reward is that conservation becomes a computation rather than an investigation. A quantity is conserved when it commutes with the energy, and the statement is stronger than its classical counterpart: not merely the average but the whole distribution of measured values stays put.

a natural place to stop  ·  the equation of motion is built; what follows is what it says about averages, and then about symmetries

5 · Ehrenfest, and the potentials for which it is exact

Chapter 4.6 §8.7 wrote down two relations and left the second unproved, for want of machinery it did not have. The machinery is (4.9.13), and with it both relations take two lines. The section is not really about deriving them. It is about the caveat that comes with them, which decides how much they are worth and which is the reason the harmonic oscillator occupies the position in physics that it does.

5.1 · Both relations, from one equation

Take the expectation of (4.9.13) in the Heisenberg state, which is fixed, so the derivative passes straight through the bracket. Translated back into the Schrödinger picture, where it is more familiar, the statement is that for any observable

ddtA^  =  1i[A^,H^]  +  A^t. \dv{}{t}\avg{\hat A} \;=\; \frac{1}{\ii\hbar}\avg{\big[\hat A,\hat H\big]} \;+\; \avg{\pdv{\hat A}{t}}. (4.9.20)

Everything needed is already in (4.9.19). Neither x^\hat x nor p^\hat p carries an explicit time, so the last term is absent. Taking expectations of those two operator identities gives Ehrenfest's relations, written here in three dimensions because the promotion is free and componentwise: §4.6's two commutators used one axis at a time, [x^i,p^j]=iδij[\hat x_{i},\hat p_{j}]=\ii\hbar\delta_{ij} keeps different axes from mixing, and VV' becomes V\nabla V one component at a time.

ddtx^  =  p^m,ddtp^  =  V. \dv{}{t}\avg{\hat{\vv x}} \;=\; \frac{\avg{\hat{\vv p}}}{m}, \qquad\qquad \dv{}{t}\avg{\hat{\vv p}} \;=\; -\avg{\nabla V}. (4.9.21)

Those are the two relations Chapter 4.6 §8.7 stated and deferred, and the route to them here is uniform: one equation of motion, two choices of observable. Chapter 4.6's Worked example 1 obtained the first of them for a free particle by a special argument in three lines, and its agreement with the general case is a check on both.

5.2 · The warning Chapter 1.1 issued, and why it stands

Chapter 1.1 §4.4 argued that force is not a fundamental concept in quantum mechanics: there is no force operator, nobody lists its eigenvalues, and it appears in no commutation relation. It then named (4.9.21) as the closest thing to a force that quantum mechanics contains, and it added a warning at the same time, which is worth quoting because softening it would spoil the point. Read carefully, it said, this "is not a fundamental law but a derived statement about expectation values, in which the potential VV is the primitive object and the force is what you get by differentiating it".

Everything in this section confirms that reading. The potential entered through H^\hat H, which is where Chapter 4.6 §4 put it. The gradient appeared in (4.9.18) as the by-product of a commutator, and the object it acts on is an operator, not a trajectory. Nothing anywhere obeys Newton's second law. Two averages obey an equation that looks like it.

5.3 · The step that is not there

⚠ The substitution everyone makes without noticing

Equation (4.9.21) is often read as saying that the centre of a wave packet obeys Newton's second law. It does not say that, and the difference is one symbol. Compare:

(a) ddtp^=V(x^)\dv{}{t}\avg{\hat p}=-\avg{V'(\hat x)}, the average of the force over the packet. This is a theorem and it holds for every state and every potential.

(b) ddtp^=V(x^)\dv{}{t}\avg{\hat p}=-V'(\avg{\hat x}), the force evaluated at the average position. This is Newton's second law for the centre of the packet, and it is a different statement.

Getting from (a) to (b) requires V(x^)=V(x^)\avg{V'(\hat x)}=V'(\avg{\hat x}), which is the claim that the average of a function is the function of the average. That is false for any function with curvature, for exactly the reason a curved dose–response relation has a mean response that is not the response at the mean dose. Spread the doses symmetrically about their mean and it fails as badly as ever, because what breaks the substitution is the bend in the curve and not any lopsidedness in the doses. The distinction is not a technicality here. It is the entire difference between quantum motion and classical motion at the level of averages, and §5.5 measures it.

5.4 · At most quadratic, and nothing else

To see how big the gap is, expand VV' in a Taylor series about the mean position, which Chapter 0.3 licenses for a smooth potential, and take the expectation term by term. The first-order term vanishes because x^x^=0\avg{\hat x-\avg{\hat x}}=0 by construction, so the first surviving correction is the second-order one:

V(x^)  =  V(x^)  +  12V(x^)(Δx)2  +   \avg{V'(\hat x)} \;=\; V'\big(\avg{\hat x}\big) \;+\; \half V'''\big(\avg{\hat x}\big)\,(\Delta x)^{2} \;+\; \cdots (4.9.22)

Read the correction term. It is controlled by the third derivative of the potential and by the square of the packet width, which by (4.9.1) cannot be sent to zero at finite momentum spread. Two consequences follow, and they are the point of the section.

The gap vanishes identically when V0V'''\equiv0, which is to say when VV is a polynomial of degree at most two. Then VV' is an affine function, the average of an affine function is the affine function of the average, and every term after the first in (4.9.22) is absent rather than small. Statement (b) becomes exactly true, for every state, however wide, however skewed, and however far from classical it looks. That is the quadratics and nothing else: the free particle, a uniform field, the harmonic oscillator, and the upside-down oscillator of an unstable equilibrium, where the coefficient is negative and the classical reading is exact all the same.

The gap does not vanish otherwise. If VV''' is non-zero somewhere then a packet sitting there has a discrepancy of order 12V(Δx)2\half V'''(\Delta x)^{2}, and since Δx\Delta x cannot be sent to zero without sending Δp\Delta p to infinity, there is no state that escapes it. Ehrenfest's theorem survives untouched, because it was statement (a). What fails is the classical reading.

One more thing needs saying, because the exactness in the quadratic case is weaker than it sounds. An oscillator energy eigenstate has x^=0\avg{\hat x}=0 at every time, and it satisfies the classical equation with both sides zero forever. Nothing about it resembles a swinging pendulum. Exact-in-the-mean is a statement about one number extracted from the state, and a state can satisfy it while looking nothing like a trajectory. The state that does look like a trajectory is Chapter 4.8 §7's coherent state, which keeps its width while its centre traces the classical orbit, and Chapter 4.10 takes up the question of when a general system admits anything of the kind.

5.5 · The two cases, measured against each other

The way to settle this is to integrate the Schrödinger equation twice with everything held fixed except the potential, and watch the residuals. The integrator is the split-operator scheme of Chapter 4.6 §10.7, in which every factor has modulus one, so the norm is conserved by construction and the readouts test the physics rather than the arithmetic. Units are =m=1\hbar=m=1, the grid is 40964096 points on [20,20][-20,20], the initial state is a Gaussian of width 0.60.6 released from rest at x=1.3x=1.3, and the run is to t=6t=6 with dt=103\dd t=10^{-3}. The two potentials are V=x2/2V=x^{2}/2 and V=x4/4V=x^{4}/4, chosen so that the first has V=0V'''=0 and the second has V=6xV'''=6x. The time derivatives in the table are fourth-order centred differences of the recorded series, which is the one detail the first two rows depend on and the one a reader reproducing the table would otherwise have to guess.

largest residual over the runV=x2/2V=x^{2}/2V=x4/4V=x^{4}/4
ddtx^p^\abs{\dv{}{t}\avg{\hat x}-\avg{\hat p}}2.2×1072.2\times10^{-7}1.0×1061.0\times10^{-6}
ddtp^+V(x^)\abs{\dv{}{t}\avg{\hat p}+\avg{V'(\hat x)}}1.1×1071.1\times10^{-7}3.1×1063.1\times10^{-6}
V(x^)V(x^)\abs{\avg{V'(\hat x)}-V'(\avg{\hat x})}00, bit for bit1.6657721.665772
ddtp^+V(x^)\abs{\dv{}{t}\avg{\hat p}+V'(\avg{\hat x})}1.1×1071.1\times10^{-7}1.6657711.665771

Read the table by rows and then by columns. The first two rows are Ehrenfest's relations, and they hold in both potentials with residuals of order 10610^{-6} or smaller against quantities of order one. Those residuals belong to the integrator rather than to the theorem, and here is how you know: halving dt\dd t divides them by four, three times running, which is a second-order error in the step size doing what second-order errors do. That test fixes the order and not the source, and the source here is the splitting's own second-order error seen through the difference stencil. The theorem is exact and the arithmetic is not.

The third row is the substitution the warning box refused to make. In the quadratic potential V(x^)\avg{V'(\hat x)} and V(x^)V'(\avg{\hat x}) are the same floating-point number at every one of the six thousand steps, because VV' is the identity there and averaging commutes with doing nothing. In the quartic potential the two differ by as much as 1.671.67 against forces of order one, with a mean gap of 0.620.62 over the run. The fourth row is the consequence: the centre of the quartic packet is not obeying Newton's second law, it is missing it by an amount comparable to the force itself, and that number does not move when dt\dd t is halved because it is physics rather than error. Meanwhile the quadratic run's centre follows x^(t)=1.3cost\avg{\hat x}(t)=1.3\cos t to 2.6×1072.6\times10^{-7} throughout, although the packet is not a coherent state and its width is breathing the whole time.

In plain terms 4.9.5

Feeding position and momentum into the equation of motion produces two statements that look exactly like the classical laws with averages written over everything. The rate of change of the average position is the average momentum over the mass, and the rate of change of the average momentum is minus the average force. This is the result an earlier chapter stated without proof and attached a warning to, and the warning is the more valuable half.

The warning is this. "The average of the force" is not "the force at the average position". A packet has width, it samples the force over a range, and unless the force varies in a straight line across that range the two quantities differ. The size of the difference is set by the curvature of the force and by the square of the packet's width, and the width cannot be reduced to nothing without the momentum spread exploding.

So the classical-looking reading is exact only when the force is a straight line in position, which means the potential is at most quadratic. The cases are exactly the quadratics: no force at all, a uniform force, a spring, and the inverted spring of a ball balanced on a hilltop. Every other potential in physics fails the test, and that is a large part of why the spring turns up everywhere in this subject as the case that can be trusted.

Two runs of the same integration, differing only in the potential, make the split visible. Both satisfy the theorem to within a millionth. Only the quadratic one has the centre of its packet obeying Newton's second law; the quartic one misses by roughly as much as the force itself. And exactness in the mean is a modest property in any case: a stationary state of the spring sits with its average position at zero forever, satisfying the classical equation with both sides vanishing while behaving nothing like a swinging weight.

a natural place to stop  ·  the commutator has bounded and it has moved things; what remains is what it generates

6 · Symmetries, generators, and conserved quantities

Chapter 1.4 ended Part I with a claim it called the real destination: a conserved quantity and the symmetry it comes from are one object under two names. Chapter 4.2 §7.5 restated that claim with operators in place of functions and expanded a unitary to first order to show the commutator appearing where the Poisson bracket had been. What it could not do was supply the generator's existence, because in infinite dimensions differentiating a one-parameter family of unitaries is not a formal manoeuvre, and it said so at the time. That link now exists, so the correspondence Chapter 4.2 could only state can be assembled here, and then run on the three cases Chapter 1.4 §3 ran classically.

6.1 · The chain, and which links were missing

A symmetry is a transformation of states that leaves every probability alone. Follow that through the machinery already built, one link at a time. Every link below was built somewhere else, so this subsection is where they are put in order rather than where anything is computed. The computing is in §6.2 and §6.3, on the three cases. The second link is the one Chapter 4.2 could not supply, and it is the reason this section can exist at all.

  • Symmetry gives unitary. A map preserving all probabilities preserves all norms, and Chapter 4.2 §7.2 showed that a linear norm-preserving map preserves every inner product and is unitary. Linearity is assumed rather than derived at this link, exactly as Chapter 4.2 §7.1 assumed it, and every family in this section is visibly linear, so the assumption is discharged case by case instead of in general.
  • A continuous family gives a self-adjoint generator. If the transformations form a strongly continuous one-parameter group, Stone's theorem produces a unique self-adjoint G^\hat G with U^ϵ=eiϵG^/\hat U_{\epsilon}=\ee^{-\ii\epsilon\hat G/\hbar}. That is the theorem Chapter 4.5 §9 proved forward and quoted backward, and Chapter 4.6 §2 used it on time. This is the link Chapter 4.2 could not supply, and the reason it could not is that Chapter 0.5 §7.1's finite-dimensional argument assumes a bounded operator with an eigenbasis, which the generators here are not and do not have.
  • The generator is an observable. Self-adjoint is exactly the condition Chapter 4.2's second postulate requires, in the corrected form Chapter 4.4 §4.2 gave it.
  • The symmetry acting on an observable is a commutator. Chapter 4.2 §7.5 expanded U^ϵf^U^ϵ\hat U^{\dagger}_{\epsilon}\hat f\hat U_{\epsilon} to first order and obtained δf^=ϵ1i[f^,G^]\delta\hat f=\epsilon\cdot\tfrac{1}{\ii\hbar}[\hat f,\hat G].
  • Symmetry of the dynamics means conservation. The family is a symmetry of the system when it leaves H^\hat H alone, which by the previous line means [H^,G^]=0[\hat H,\hat G]=0. By (4.9.15) that is precisely the statement that G^\hat G is conserved.

Read the last two links in the other order and the converse is the same equation. If G^\hat G is conserved then it commutes with H^\hat H, so the family it generates leaves H^\hat H alone and is a symmetry. A conserved quantity is the generator of its symmetry, which is what Chapter 1.4 §7 proved classically, repeated here with the generator's existence supplied rather than assumed. The three cases below are Chapter 1.4 §3's three, and the third of them is already done.

6.2 · Translation, and momentum

Define the family that slides a wavefunction along, writing it U^a\hat U_{a} in the notation of §6.1, so that (U^aψ)(x)=ψ(xa)(\hat U_{a}\psi)(x)=\psi(x-a). It is unitary because ψ(xa)2dx=ψ2dx\int\abs{\psi(x-a)}^{2}\dd x=\int\abs{\psi}^{2}\dd x, Lebesgue measure being translation-invariant by Chapter 4.3 §2.1, and it is onto because U^a\hat U_{-a} undoes it. The group law holds because sliding twice is sliding once by the sum. Strong continuity takes one line: for a continuous ψ\psi of compact support, U^aψψ0\norm{\hat U_{a}\psi-\psi}\to0 by uniform continuity, and Chapter 4.3 §8.1 proved such functions dense, so the general case follows from U^a=1\norm{\hat U_{a}}=1.

Stone therefore hands us a generator, and to find it we differentiate the family at the origin, which is what Stone's formula for the domain in Chapter 4.6 §2.1 instructs. Since U^a=eiaG^/\hat U_{a}=\ee^{-\ii a\hat G/\hbar} gives dU^ada0=iG^\dv{\hat U_{a}}{a}\big|_{0}=-\tfrac{\ii}{\hbar}\hat G, and differentiating ψ(xa)\psi(x-a) at a=0a=0 gives ψ(x)-\psi'(x), the two must agree:

iG^ψ  =  ψG^  =  ix  =  p^. -\frac{\ii}{\hbar}\hat G\psi \;=\; -\psi' \qquad\Longrightarrow\qquad \hat G \;=\; -\ii\hbar\,\pdv{}{x} \;=\; \hat p. (4.9.23)

Momentum is the generator of translations, with the operator Chapter 4.6 §5 identified arriving here for a second and independent reason. Chapter 1.4 §7.1 computed {x,px}=1\{x,p_{x}\}=1 and read the same sentence off it classically.

Now close the loop with a computation rather than a slogan. By (4.9.18), [p^,H^]=[p^,V(x^)]=iV(x^)[\hat p,\hat H]=[\hat p,V(\hat x)]=-\ii\hbar V'(\hat x), which is the zero operator exactly when VV is constant. So momentum is conserved precisely when the potential is unchanged by translation, and the two statements are one commutator apart. Chapter 1.4 §3.2 reached the same conclusion from the action, and the two derivations have no step in common.

6.3 · Rotation, and the commutator Chapter 4.11 is built on

Do the same in three dimensions with rotations about the zz-axis. The family is (R^(θ)ψ)(r)=ψ(R1(θ)r)(\hat R(\theta)\psi)(\vv r)=\psi(R^{-1}(\theta)\vv r), unitary because a rotation preserves volume, and differentiating at θ=0\theta=0 gives, with R1(θ)r=(xcosθ+ysinθ,xsinθ+ycosθ,z)R^{-1}(\theta)\vv r=(x\cos\theta+y\sin\theta,\, -x\sin\theta+y\cos\theta,\,z),

ddθ0ψ(R1r)  =  yψxxψy  =  iL^zψ,L^z  =  x^p^yy^p^x. \dv{}{\theta}\Big|_{0}\psi\big(R^{-1}\vv r\big) \;=\; y\,\pdv{\psi}{x} - x\,\pdv{\psi}{y} \;=\; -\frac{\ii}{\hbar}\,\hat L_{z}\psi, \qquad \hat L_{z} \;=\; \hat x\hat p_{y}-\hat y\hat p_{x}. (4.9.24)

Angular momentum is the generator of rotations, matching Chapter 1.4 §7.1's classical computation with Lz=xpyypxL_{z}=xp_{y}-yp_{x}, which also rotated the momentum as a bonus. The operator is the classical expression with hats, and the ordering ambiguity that usually attends such a substitution is absent here because x^\hat x and p^y\hat p_{y} commute.

What that costs is one line and what it buys is the whole of Chapter 4.11, so compute the thing it buys. Rotations about different axes do not commute, and the algebraic residue of that geometric fact is the commutator of the generators. Take L^x=y^p^zz^p^y\hat L_{x}=\hat y\hat p_{z}-\hat z\hat p_{y} and L^y=z^p^xx^p^z\hat L_{y}=\hat z\hat p_{x}-\hat x\hat p_{z} and expand with (4.9.16). Of the four terms, two vanish outright because they contain no conjugate pair. The survivors are the first and the last, [y^p^z,z^p^x]=y^[p^z,z^]p^x=iy^p^x[\hat y\hat p_{z},\hat z\hat p_{x}]=\hat y[\hat p_{z},\hat z]\hat p_{x}=-\ii\hbar\,\hat y\hat p_{x} and [z^p^y,x^p^z]=x^[z^,p^z]p^y=+ix^p^y[\hat z\hat p_{y},\hat x\hat p_{z}]=\hat x[\hat z,\hat p_{z}]\hat p_{y} =+\ii\hbar\,\hat x\hat p_{y}. Both enter the sum with a plus sign, the first because it carried no minus sign in the expansion at all and the last because the two minus signs it carried multiply together. Adding them,

[L^x,L^y]  =  i(x^p^yy^p^x)  =  iL^z, \big[\hat L_{x},\hat L_{y}\big] \;=\; \ii\hbar\big(\hat x\hat p_{y}-\hat y\hat p_{x}\big) \;=\; \ii\hbar\,\hat L_{z}, (4.9.25)

with the other two relations following by the cyclic relabelling xyzxx\to y\to z\to x. Verified in the Weyl algebra: normal-ordering the products using [x^i,p^j]=iδij[\hat x_{i},\hat p_{j}]=\ii\hbar\delta_{ij} and nothing else returns zero for all three differences.

That verification deserves a sentence of its own, because of what it does not use. Chapter 4.2 §8 postulated the operator commutator for the single pair (q^,p^)(\hat q,\hat p) and was explicit that the general correspondence between Poisson brackets and commutators is not derivable and is in fact false. Equation (4.9.25) is what the correspondence predicts for a pair it was never postulated for, since Chapter 1.4 §7.3 computed {Li,Lj}=εijkLk\{L_{i},L_{j}\}=\varepsilon_{ijk}L_{k} classically, and applying {  ,  }1i[  ,  ]\{\;,\;\}\mapsto\tfrac{1}{\ii\hbar}[\;,\;] to that turns it into (4.9.25) exactly. The prediction is correct, and it was obtained here from the one postulated pair by algebra alone. Chapter 4.10 §8 explains why such agreement cannot be made universal.

Two consequences, and then the hand-off. First, feeding (4.9.25) into (4.9.8) gives ΔLxΔLy2L^z\Delta L_{x}\,\Delta L_{y}\ge\tfrac{\hbar}{2}\abs{\avg{\hat L_{z}}}, so no state has all three components sharp unless all three vanish, which is what Chapter 1.4's grind box anticipated when it noted that {Lx,Ly}0\{L_{x},L_{y}\}\neq0 would survive quantisation.

Second, the check §2.7 held back. It was the third of that section's three, and the commutator it needed now exists. On the three-dimensional representation Chapter 4.11 builds from (4.9.25), the state with the largest L^z\hat L_{z} gives ΔLxΔLy=0.5000000000000002\Delta L_{x}\,\Delta L_{y}=0.500000000000000\,\hbar^{2} against a bound 2L^z=0.5000000000000002\tfrac{\hbar}{2}\abs{\avg{\hat L_{z}}}=0.500000000000000\,\hbar^{2}, so the relation is saturated to fifteen figures. The middle state of the same set gives ΔLxΔLy=1.02\Delta L_{x}\,\Delta L_{y}=1.0\,\hbar^{2} against a bound of exactly zero, since L^z=0\avg{\hat L_{z}}=0 there. That second case is the one to keep. A bound that vanishes has not announced that both quantities are sharp. It has fallen silent, while (4.9.25) still forbids a common eigenvector, which is §3.1's distinction made numerical.

And now the hand-off: equation (4.9.25) is the entire input to Chapter 4.11. That chapter takes it as the definition of what an angular momentum is, forgets where it came from, and extracts the whole spectrum from it by the ladder move Chapter 4.8 ran on the oscillator. Nothing about rotations in space is used again.

6.4 · Time, which was done first

The third case needs no work, because Chapter 4.6 did it before the pattern was visible. Stone applied to time evolution produced H^\hat H as the generator, so the Hamiltonian generates time translation, and [H^,H^]=0[\hat H,\hat H]=0 makes energy conserved by (4.9.15) with no calculation at all. Chapter 1.4 §3.1 obtained conservation of energy from invariance of the action under a shift of the time origin, and Chapter 1.4 §7.1's third example observed that time evolution is the transformation generated by the energy, so it is not a different kind of process from a symmetry. All three of Chapter 1.4's classics have now been repeated with operators, and the table below is that chapter's §3 in quantum form.

symmetryunitary familygeneratorconserved when
translation in spaceψ(x)ψ(xa)\psi(x)\mapsto\psi(x-a)p^\hat pVV is translation-invariant
rotationψ(r)ψ(R1r)\psi(\vv r)\mapsto\psi(R^{-1}\vv r)L^z\hat L_{z}VV depends on rr alone
translation in timeψU^(t)ψ\ket\psi\mapsto\hat U(t)\ket\psiH^\hat Halways, for a closed system

6.5 · A symmetry with no generator of its own

The chain of §6.1 needs a continuous family, and not every symmetry sits inside one that tells you anything. Parity, built in Chapter 4.7 §2 and written Π^\hat\Pi, is the map ψ(x)ψ(x)\psi(x)\mapsto\psi(-x). It is unitary, it is self-adjoint, and it commutes with H^\hat H whenever VV is even. What it is not is a member of a family that could produce something new. Geometrically the reflection rr\vv r\mapsto-\vv r has determinant 1-1, so no continuous family of spatial transformations reaches it from the identity. Algebraically a family can be manufactured, since the projections onto the even and the odd functions are functions of Π^\hat\Pi and exponentiating the odd one runs from I^\hat I to Π^\hat\Pi. But Stone applied to that family returns a generator which is itself a function of Π^\hat\Pi, so it is no observable that was not already parity, and nothing has been generated. What survives is the conclusion without the machinery: parity commutes with H^\hat H, so by (4.9.15) it is conserved, and since Π^2=I^\hat\Pi^{2}=\hat I leaves it only two eigenvalues, what is conserved is a label taking two values rather than a quantity taking a continuum of them.

Chapter 1.4 §5 recorded the classical version of the same limitation and called it an honest exception, and this is what it looks like on this side. Noether's theorem needs a continuous symmetry that reaches the identity through the transformations themselves, and so does anything Stone can usefully hand back. A discrete symmetry produces a conserved quantum number with no independent generator behind it. Chapters 4.17 and 4.18 use exactly that, since a selection rule is a statement about a discrete label rather than about a continuous charge.

6.6 · What the commutator has not told us

Take stock, because the chapter has said three things about one object and it is worth being explicit that a fourth is missing.

The commutator of two observables bounds how sharp they can be at once ((4.9.8)), and it vanishes exactly when they can be sharp together (§3). The commutator with the Hamiltonian moves every observable in time ((4.9.13)), and it vanishes exactly when the observable is conserved ((4.9.15)). The commutator with any self-adjoint operator generates the symmetry that operator belongs to (§6.1). Three jobs, one object, and the classical shadow of all three is the Poisson bracket of Chapter 1.3.

What none of it says is what happens when \hbar is small compared with the action in play. Every equation in this chapter carries \hbar in a place where setting it to zero either destroys the statement or makes it vacuous: (4.9.8) becomes the empty assertion that a product of spreads is non-negative, and (4.9.13) becomes 0/00/0. Something more careful is needed, and it is a genuine limit rather than a substitution.

That is Chapter 4.10, and it comes with a warning worth meeting before the proof. The correspondence between classical and quantum observables looks, from (4.9.13) and Chapter 4.2 §7.5, like a dictionary that should extend to everything. It does not. Chapter 4.10 §8 proves that no consistent dictionary exists taking every classical observable to an operator while turning every Poisson bracket into the corresponding commutator. Chapter 4.2 kept its distance from that general claim twice, in §7.5 when it separated the derived algebraic statement from the undeliverable general one and in §8 when it postulated the correspondence for a single pair rather than for all of them. Now that you have watched the correspondence succeed for angular momentum in (4.9.25), the claim that it must eventually fail is worth arriving with rather than being told afterwards.

In plain terms 4.9.6

A symmetry is anything you can do to a system without changing any prediction. Chase that definition through the machinery and a chain of four links appears, each of them already built. Preserving predictions means preserving lengths, and provided the transformation is linear, which every family the section used visibly is, that forces it to be a rotation of the space of states. Linearity is an assumption at this link rather than a result, and the main text says so where it is used. A continuous family of such rotations has, by the theorem imported earlier in this part, a unique observable behind it that generates it. The effect of the family on any other quantity is a commutator with that generator. And leaving the energy alone, which is what makes the family a symmetry of the dynamics rather than a mere relabelling, is the same equation as the generator being conserved.

So a conserved quantity and a symmetry are one thing, exactly as the Part I chapters concluded for classical mechanics, and the argument is shorter here because the theorem doing the heavy lifting was proved elsewhere.

Three examples, and they are the same three the classical chapter used. Sliding everything along is generated by momentum, and momentum is conserved when the potential is flat. Turning everything about an axis is generated by angular momentum, and the commutator of two such generators is another one, which is the algebraic fingerprint of the fact that turns about different axes do not commute. Waiting is generated by the energy, which is therefore conserved with no computation needed. Discrete symmetries, like reflection, leave no generator behind, because any continuous family you could run through such a symmetry has to be built out of the symmetry itself and hands back nothing that was not already there. What they leave is a two-valued label rather than a continuous charge.

One thing has been missing throughout. Everything here describes how quantum quantities behave among themselves. None of it says how the classical world reappears when the scales get large, and the reason no equation here answers that is that setting the constant to zero in any of them produces nonsense. The next chapter takes the limit properly, and it also proves that the translation between classical and quantum quantities, which has worked every time it has been tried here, cannot be made to work for everything at once.

7 · Worked examples

Worked example 1 — which states sit exactly on the floor, and the Gaussian arriving for the third time

The bound ΔxΔp/2\Delta x\,\Delta p\ge\hbar/2 is attained by some states and not by others. (a) Find the condition on ψ\ket\psi for equality in (4.9.8) for a general pair. (b) Solve that condition for x^\hat x and p^\hat p. (c) Identify what you get. (d) Say what changes if you ask instead for equality in the sharper form (4.9.9).

(a) Two inequalities were used and both must be tight. Chapter 0.5 §1.4 recorded the equality condition for Cauchy–Schwarz: the two vectors are parallel, so g=λf\ket g=\lambda\ket f for some complex λ\lambda. The second inequality was the discarding of the anticommutator term in (4.9.7), so that term must vanish. Compute it with g=λf\ket g=\lambda\ket f:

{ΔA^,ΔB^}  =  f,g+g,f  =  (λ+λˉ)f2  =  2Re(λ)(ΔA)2. \avg{\big\{\Delta\hat A,\Delta\hat B\big\}} \;=\; \avg{f,g}+\avg{g,f} \;=\; \big(\lambda+\bar\lambda\big)\norm f^{2} \;=\; 2\,\mathrm{Re}(\lambda)\,(\Delta A)^{2}.

So equality holds exactly when g=λf\ket g=\lambda\ket f with λ\lambda purely imaginary. Both conditions are conditions on the state, and neither mentions the observables beyond (4.9.3).

(b) Write λ=iμ\lambda=\ii\mu with μ\mu real, put A^=x^\hat A=\hat x and B^=p^\hat B=\hat p, and work in the position representation where p^=ix\hat p=-\ii\hbar\,\partial_{x}. The condition (p^p^)ψ=iμ(x^x^)ψ(\hat p-\avg{\hat p})\psi=\ii\mu(\hat x-\avg{\hat x})\psi becomes a first-order linear equation,

ψψ  =  ip^    μ(xx^), \frac{\psi'}{\psi} \;=\; \frac{\ii\avg{\hat p}}{\hbar} \;-\; \frac{\mu}{\hbar}\big(x-\avg{\hat x}\big),

which separates in the manner of Chapter 0.8 §2 and integrates to

ψ(x)  =  Cexp ⁣(ip^x)exp ⁣(μ(xx^)22),μ>0  for normalisability. \psi(x) \;=\; C\,\exp\!\Big(\frac{\ii\avg{\hat p}x}{\hbar}\Big)\, \exp\!\Big(-\frac{\mu\,(x-\avg{\hat x})^{2}}{2\hbar}\Big), \qquad \mu \gt 0 \ \text{ for normalisability}.

(c) That is a Gaussian of width Δx=/2μ\Delta x=\sqrt{\hbar/2\mu}, centred at x^\avg{\hat x}, multiplied by a plane wave carrying mean momentum p^\avg{\hat p}. Its momentum spread is Δp=μ/2\Delta p=\sqrt{\mu\hbar/2}, and the product is /2\hbar/2 for every μ\mu. This is the third time the Gaussian has been produced as the unique minimiser, by a third route. Chapter 0.9 §6.5 ran the bandwidth theorem's equality conditions backwards, Chapter 4.6 §10 found it as the packet whose spreading formula is exact, and here it comes out of Cauchy–Schwarz. Chapter 4.8 §7's coherent states are exactly this family with μ=mω\mu=m\omega, which is why they are the states that behave most like a classical oscillation.

(d) Only the second condition is dropped, so equality in (4.9.9) requires g=λf\ket g=\lambda\ket f with λ\lambda complex and otherwise unrestricted. Repeating the integration with λ=ν+iμ\lambda=\nu+\ii\mu gives the same Gaussian with a complex width, that is, with a quadratic phase across it. Problem 1 computes such a state's three numbers and finds a product exceeding /2\hbar/2 while the sharper form stays tight, which is the right-hand pair of points in the figure of §2.7.

Worked example 2 — angle and angular momentum, where the hypothesis fails and what is left behind

On the circle, with L2[0,2π)L^{2}[0,2\pi) and L^z=id/dφ\hat L_{z}=-\ii\hbar\,\dd/\dd\varphi on the periodic domain, which is the θ=0\theta=0 member of the circle of self-adjoint domains Chapter 4.4 §5.4 found, let φ^\hat\varphi be multiplication by φ\varphi. (a) Evaluate both sides of (4.9.8) in the state ψm=eimφ/2π\psi_{m}=\ee^{\ii m\varphi}/\sqrt{2\pi}. (b) Locate the step of §2 that fails. (c) Compute the defect exactly. (d) Repair the statement.

(a) ψm\psi_{m} is an eigenstate of L^z\hat L_{z} with eigenvalue mm\hbar, so ΔLz=0\Delta L_{z}=0 exactly. Its density is uniform on [0,2π)[0,2\pi), so φ^=π\avg{\hat\varphi}=\pi and φ^2=4π2/3\avg{\hat\varphi^{2}}=4\pi^{2}/3, giving Δφ=π/3=1.8138\Delta\varphi=\pi/\sqrt3=1.8138. The left-hand side is therefore zero. Computed formally, [φ^,L^z]=i[\hat\varphi,\hat L_{z}]=\ii\hbar on smooth functions, so the right-hand side would be /2\hbar/2. Zero is not at least /2\hbar/2. One of the hypotheses has to have failed, and the useful question is which.

(b) It is (4.9.5), the step that moved A^\hat A across the inner product. That move is the definition of self-adjointness applied to the vector ΔB^ψ\Delta\hat B\ket\psi, and it is legitimate only when that vector lies in the operator's domain. Here φ^ψm=φeimφ/2π\hat\varphi\psi_{m}=\varphi\,\ee^{\ii m\varphi}/\sqrt{2\pi} is not periodic: it starts at 00 and ends at 2π/2π2\pi/\sqrt{2\pi}. So it is not in the domain of L^z\hat L_{z}, which single-valuedness on the circle fixes as the periodic functions, the θ=0\theta=0 member of Chapter 4.4 §5.4's circle of domains, and the whole of §2 is unavailable.

(c) Chapter 4.4 §5.1's boundary form says what is lost. For smooth u,vu,v on [0,2π][0,2\pi], integrating by parts once gives L^zu,vu,L^zv=i[uˉv]02π\avg{\hat L_{z}u,v}-\avg{u,\hat L_{z}v}=\ii\hbar\big[\bar u v\big]_{0}^{2\pi}. Put u=ψmu=\psi_{m} and v=φ^ψmv=\hat\varphi\psi_{m}, so uˉv=φ/2π\bar uv=\varphi/2\pi, whose bracket is 11. Hence

L^zψm,φ^ψm  =  πm,ψm,L^zφ^ψm  =  πmi, \avg{\hat L_{z}\psi_{m},\,\hat\varphi\psi_{m}} \;=\; \pi m\hbar, \qquad \avg{\psi_{m},\,\hat L_{z}\hat\varphi\psi_{m}} \;=\; \pi m\hbar - \ii\hbar,

and the two differ by exactly i\ii\hbar. That i\ii\hbar is the whole of the formal commutator. The would-be right-hand side of the relation is not a property of the state at all. It is a boundary term, and it is there because the operator was applied outside its domain.

(d) Replace the angle by a bounded periodic function of it, which keeps the domain intact because multiplying a periodic function by a periodic function leaves it periodic. With A^=sinφ^\hat A=\sin\hat\varphi the same one-line calculation as (4.9.18) gives [sinφ^,L^z]=icosφ^[\sin\hat\varphi,\hat L_{z}]=\ii\hbar\cos\hat\varphi, so §2 applies unchanged and delivers

Δ(sinφ)  ΔLz    2cosφ^. \Delta(\sin\varphi)\;\Delta L_{z} \;\ge\; \frac{\hbar}{2}\,\abs{\avg{\cos\hat\varphi}}.

Test it on the offending state: cosφ^=0\avg{\cos\hat\varphi}=0 for a uniform density, so the bound is zero and ΔLz=0\Delta L_{z}=0 violates nothing. The theorem was never wrong. The statement it was applied to was not one of its instances, and this is what Chapter 4.4 was for.

Worked example 3 — energy and time, which is not an instance of the theorem

The relation ΔEΔt/2\Delta E\,\Delta t\ge\hbar/2 is quoted as often as the position–momentum one. (a) Say why it cannot be an instance of (4.9.8). (b) Derive, from this chapter and nothing else, a true statement of that shape. (c) Interpret the quantity that plays the part of Δt\Delta t. (d) Put a number on it.

(a) Every symbol in (4.9.8) refers to two observables, and time is not one. It is the parameter labelling the family U^(t)\hat U(t), not an operator on the space, and Chapter 4.6 §1 treated it that way throughout. There is no t^\hat t whose commutator with H^\hat H one could evaluate, so (4.9.8) has nothing to say and quoting it here is a category error rather than an approximation.

(b) What is available is (4.9.20), which is the only place in the theory where a time derivative and a commutator meet. Take any observable A^\hat A with no explicit time dependence and apply (4.9.8) to the pair (A^,H^)(\hat A,\hat H):

ΔAΔE    12[A^,H^]  =  2ddtA^, \Delta A\,\Delta E \;\ge\; \half\abs{\avg{\big[\hat A,\hat H\big]}} \;=\; \frac{\hbar}{2}\,\abs{\dv{}{t}\avg{\hat A}},

the last step being (4.9.20) read from right to left. Now divide through by the rate, which is legitimate in any state where the rate is non-zero, and define

τA    ΔAddtA^τAΔE    2. \tau_{A} \;\equiv\; \frac{\Delta A}{\abs{\dv{}{t}\avg{\hat A}}} \qquad\Longrightarrow\qquad \tau_{A}\,\Delta E \;\ge\; \frac{\hbar}{2}.

(c) Read the definition of τA\tau_{A} rather than the inequality. It is the time the mean of AA needs in order to shift by one standard deviation of AA, which is the shortest time in which a measurement of AA could tell that anything had happened. So the honest statement is not about an uncertainty in time. It is a speed limit: a state with a small energy spread cannot change quickly, in any observable at all, since A^\hat A was arbitrary. A stationary state has ΔE=0\Delta E=0 and changes in nothing, which is the extreme case and agrees with Chapter 4.6 §9.2. Checked numerically on two hundred thousand random Hamiltonians, observables and states: no violation, with the smallest slack 2.7×10102.7\times10^{-10}, and a two-level system prepared as (0+i1)/2(\ket0+\ii\ket1)/\sqrt2 saturates it exactly at τAΔE=/2\tau_{A}\Delta E=\hbar/2.

(d) A state whose energy is uncertain by 11 eV cannot change appreciably in less than /(2×1 eV)=3.29×1016s\hbar/(2\times1\ \mathrm{eV})=3.29\times10^{-16}\,\mathrm{s}, using =6.5821×1016 eVs\hbar=6.5821\times10^{-16}\ \mathrm{eV\,s}. Read the same relation with the lifetime identified as the τA\tau_{A} of some observable that registers the decay, and it says an excited state living for a time τ\tau cannot have a sharper energy than about /2τ\hbar/2\tau, which is why spectral lines have widths. That identification is an extra step rather than a re-reading, and it is part of why the statement is an order of magnitude and not a bound. Chapter 4.17 computes those widths from the transition rate rather than bounding them, and the two answers agree in order of magnitude, which is the check worth making on any argument of this shape.

8 · Your turn

Problem 1 — the sharper form, and a Gaussian that is not on the floor

Let ψ(x)=Cexp(ax2)\psi(x)=C\exp(-ax^{2}) with a=αiβa=\alpha-\ii\beta and α>0\alpha\gt0, a Gaussian with a quadratic phase across it. (a) Compute Δx\Delta x, Δp\Delta p and the covariance 12{Δx^,Δp^}\tfrac12\avg{\{\Delta\hat x,\Delta\hat p\}}. (b) Show that (4.9.9) is an equality while (4.9.8) is not, and find the factor by which the product exceeds /2\hbar/2. (c) Locate this state in the figure of §2.7 and say what the horizontal displacement of its point means physically. (d) Chapter 4.6 §10 evolved a free Gaussian and found it spreading. Without redoing that calculation, say what an initially real Gaussian has acquired by the time it has spread, and which of the three numbers in (a) detects it.

Solution

(a) The density is ψ2e2αx2\abs\psi^{2}\propto\ee^{-2\alpha x^{2}}, so x^=0\avg{\hat x}=0 and (Δx)2=1/4α(\Delta x)^{2}=1/4\alpha. Since ψ=2axψ\psi'=-2ax\psi, (Δp)2=2ψ2=42a2x^2=2(α2+β2)/α(\Delta p)^{2}=\hbar^{2}\int\abs{\psi'}^{2}=4\hbar^{2}\abs a^{2}\avg{\hat x^{2}} =\hbar^{2}(\alpha^{2}+\beta^{2})/\alpha. For the covariance, integrate by parts once: x^p^+p^x^=i(2xψψ+ψ2)=i(14ax^2)=β/α\avg{\hat x\hat p+\hat p\hat x}=-\ii\hbar\int\big(2x\psi^{*}\psi'+\abs\psi^{2}\big) =-\ii\hbar\big(1-4a\avg{\hat x^{2}}\big)=\hbar\beta/\alpha, so the covariance is β/2α\hbar\beta/2\alpha.

(b) (ΔxΔp)2=2(α2+β2)/4α2(\Delta x\Delta p)^{2}=\hbar^{2}(\alpha^{2}+\beta^{2})/4\alpha^{2}, while the right-hand side of (4.9.9) is (β/2α)2+(/2)2(\hbar\beta/2\alpha)^{2}+(\hbar/2)^{2}, which is the same number. The sharper form is exactly tight. Meanwhile ΔxΔp=(/2)1+(β/α)2\Delta x\,\Delta p=(\hbar/2)\sqrt{1+(\beta/\alpha)^{2}}, which exceeds /2\hbar/2 for every β0\beta\neq0. Equality in (4.9.8) needs a real width, which is Worked example 1(b).

(c) The point sits on the unit circle, at height 1/1+(β/α)21/\sqrt{1+(\beta/\alpha)^{2}} and horizontal coordinate (β/α)/1+(β/α)2(\beta/\alpha)/\sqrt{1+(\beta/\alpha)^{2}}. The two drawn cases are β/α=1\beta/\alpha=1 and 22. The horizontal displacement is the correlation between position and momentum: in this state a particle found on the right is more likely to be moving right, which is what a quadratic phase encodes and what the modulus of the wavefunction cannot show.

(d) It has acquired exactly this quadratic phase, with β\beta growing from zero. That is what free spreading is: the fast components outrun the slow ones, so position and momentum become correlated. Neither Δx\Delta x nor Δp\Delta p separately reveals the correlation, since the second is constant for a free particle by Chapter 4.6's Worked example 1(b); the covariance detects it, and it is the reason a free packet's product grows without the state ever ceasing to saturate (4.9.9).

Problem 2 — the Heisenberg equation solved exactly, for the one potential where that is possible

Take H^=p^2/2m+12mω2x^2\hat H=\hat p^{2}/2m+\half m\omega^{2}\hat x^{2}. (a) Write out (4.9.19) for this potential and solve the pair of operator equations. (b) Verify that [x^H(t),p^H(t)]=i[\hat x_{H}(t),\hat p_{H}(t)]=\ii\hbar at all times, and say why it had to come out that way. (c) Take expectations and check §5.4's claim that the classical equation is exact here for every state. (d) Compute (Δx(t))2(\Delta x(t))^{2} for an arbitrary initial state and identify the states whose width does not move.

Solution

(a) With V=mω2x^V'=m\omega^{2}\hat x the equations are x^˙H=p^H/m\dot{\hat x}_{H}=\hat p_{H}/m and p^˙H=mω2x^H\dot{\hat p}_{H}=-m\omega^{2}\hat x_{H}, a linear system with constant coefficients in which the coefficients are numbers, so Chapter 0.8 §4's solution applies with operators in place of the initial values:

x^H(t)=x^cosωt+p^mωsinωt,p^H(t)=p^cosωtmωx^sinωt. \hat x_{H}(t) = \hat x\cos\omega t + \frac{\hat p}{m\omega}\sin\omega t, \qquad \hat p_{H}(t) = \hat p\cos\omega t - m\omega\hat x\sin\omega t.

Differentiating returns the equations, and at t=0t=0 both reduce to the Schrödinger-picture operators. This is the only potential for which the operator equations close, because it is the only one where VV' is linear in x^\hat x.

(b) Expanding, [x^H,p^H]=cos2ωt[x^,p^]sin2ωt[p^,x^]=i(cos2+sin2)=i[\hat x_{H},\hat p_{H}]=\cos^{2}\omega t\,[\hat x,\hat p] -\sin^{2}\omega t\,[\hat p,\hat x]=\ii\hbar(\cos^{2}+\sin^{2})=\ii\hbar. It had to: by (4.9.11) both operators are conjugated by the same unitary, and U^A^U^U^B^U^=U^A^B^U^\hat U^{\dagger}\hat A\hat U\,\hat U^{\dagger}\hat B\hat U=\hat U^{\dagger}\hat A\hat B\hat U, so every algebraic relation among operators survives the change of picture unchanged. That is the same invariance argument as §4.2's about the spectrum.

(c) Taking expectations gives x^(t)=x^cosωt+(p^/mω)sinωt\avg{\hat x}(t)=\avg{\hat x}\cos\omega t +(\avg{\hat p}/m\omega)\sin\omega t, which satisfies x^¨=ω2x^\ddot{\avg{\hat x}}=-\omega^{2}\avg{\hat x} identically. No property of the state was used anywhere, which is §5.4's claim: for a quadratic potential the mean obeys the classical equation exactly, whatever the state.

(d) Squaring and subtracting the square of the mean,

(Δx(t))2=(Δx)2cos2ωt+(Δp)2m2ω2sin2ωt+sin2ωt2mω{Δx^,Δp^}. (\Delta x(t))^{2} = (\Delta x)^{2}\cos^{2}\omega t + \frac{(\Delta p)^{2}}{m^{2}\omega^{2}}\sin^{2}\omega t + \frac{\sin2\omega t}{2m\omega}\avg{\{\Delta\hat x,\Delta\hat p\}}.

The two time-dependent pieces combine into a constant exactly when the covariance vanishes and (Δp)2=m2ω2(Δx)2(\Delta p)^{2}=m^{2}\omega^{2}(\Delta x)^{2}, since then the cos2\cos^{2} and sin2\sin^{2} terms add to (Δx)2(\Delta x)^{2} and the cross term is absent. Every energy eigenstate meets that condition, which it had to, since Chapter 4.6 §9.2 makes every observable's distribution constant in a stationary state, as part (a) of Problem 3 records. Among the minimum-uncertainty states of Worked example 1 the condition picks out Δx=/2mω\Delta x=\sqrt{\hbar/2m\omega} alone, which is the ground state and the coherent states built on it, Chapter 4.8 §7's family and the reason Chapter 4.6 §10.8 found the oscillator packet's width unchanging. Any state failing the condition breathes, and the equation above shows it does so at 2ω2\omega rather than ω\omega.

Problem 3 — the virial theorem, which is one commutator

(a) Show that in a stationary state the expectation of any observable without explicit time dependence is constant. (b) Take G^=12(x^p^+p^x^)\hat G=\half(\hat x\hat p+\hat p\hat x) and compute [G^,H^][\hat G,\hat H] for H^=p^2/2m+V(x^)\hat H=\hat p^{2}/2m+V(\hat x). (c) Deduce the quantum virial theorem and check it against the oscillator. (d) Apply it to V=k/rV=-k/r in three dimensions and say what it gives for a hydrogen state.

Solution

(a) Chapter 4.6 §9.2 showed the two phase factors cancel in ψA^ψ\bra\psi\hat A\ket\psi for ψ=ueiEt/\psi=u\,\ee^{-\ii Et/\hbar}, so the expectation carries no tt. Equivalently, and more usefully here, (4.9.20) gives ddtA^=1i[A^,H^]\dv{}{t}\avg{\hat A}=\tfrac{1}{\ii\hbar}\avg{[\hat A,\hat H]}, and in an eigenstate of H^\hat H the Hamiltonian may be moved onto either side as the number EE, so the two terms cancel and the derivative is zero.

(b) Since x^p^p^x^=i\hat x\hat p-\hat p\hat x=\ii\hbar is a constant, G^\hat G and x^p^\hat x\hat p have the same commutator with anything, so compute with x^p^\hat x\hat p. Two applications of (4.9.16) and one of (4.9.17) and (4.9.18) give

[x^p^,H^]  =  [x^,p^2]p^2m+x^[p^,V]  =  i(p^2mx^V(x^)). \big[\hat x\hat p,\hat H\big] \;=\; \frac{\big[\hat x,\hat p^{2}\big]\hat p}{2m} + \hat x\big[\hat p,V\big] \;=\; \ii\hbar\left(\frac{\hat p^{2}}{m} - \hat x\,V'(\hat x)\right).

(c) Put (b) into (a). The left-hand side is zero in a stationary state, so

2T^  =  x^V(x^),T^=p^22m. 2\avg{\hat T} \;=\; \avg{\hat x\,V'(\hat x)}, \qquad \hat T=\frac{\hat p^{2}}{2m}.

For V=12mω2x2V=\half m\omega^{2}x^{2} we have xV=2VxV'=2V, so T^=V^=E/2\avg{\hat T}=\avg{\hat V}=E/2. Chapter 4.5's Problem 2(c) computed both halves independently for the Hermite functions and got exactly that, which is a check with no algebra in common.

(d) The three-dimensional version replaces x^V\hat x V' by r^V\hat{\vv r}\cdot\nabla V, the derivation running component by component. For V=k/rV=-k/r, Euler's relation for a function homogeneous of degree 1-1 gives rV=V\vv r\cdot\nabla V=-V, so 2T^=V^2\avg{\hat T}=-\avg{\hat V} and therefore E=T^+V^=T^E=\avg{\hat T}+\avg{\hat V}=-\avg{\hat T}. Every bound hydrogen state has kinetic energy equal to minus its total energy and potential energy twice its total energy, before any wavefunction has been written down. Chapter 4.13 computes those wavefunctions and Chapter 4.16 uses this relation to evaluate the fine-structure corrections without doing new integrals. Chapter 1.4's Problem 4 obtained the classical version as a statement about time averages; here it holds state by state, with no averaging over an orbit, because a stationary state is already what a time average is trying to be.

Problem 4 — completeness, by counting, on two qubits

Take two of Chapter 4.2 §10.4's qubits, so the space has dimension four, with Z^1\hat Z_{1} and Z^2\hat Z_{2} the which-state observables of the two factors, eigenvalues ±1\pm1. (a) Show {Z^1,Z^2}\{\hat Z_{1},\hat Z_{2}\} is a complete set by exhibiting the count. (b) Show that the total T^=Z^1+Z^2\hat T=\hat Z_{1}+\hat Z_{2} is not complete on its own, and exhibit an observable commuting with it that is not a function of it. (c) Let S^\hat S be the swap, S^ab=ba\hat S\ket{ab}=\ket{ba}. Show S^\hat S is an observable, that [T^,S^]=0[\hat T,\hat S]=0, and that {T^,S^}\{\hat T,\hat S\} is complete. (d) Show that the two complete sets cannot be used together, and name the chapter where this situation recurs.

Solution

(a) The four product states 00,01,10,11\ket{00},\ket{01},\ket{10},\ket{11} are joint eigenvectors with tuples (+1,+1)(+1,+1), (+1,1)(+1,-1), (1,+1)(-1,+1), (1,1)(-1,-1). Four distinct tuples in a four-dimensional space, so by §3.4 the set is complete and the two eigenvalues are a full set of quantum numbers.

(b) T^\hat T has eigenvalues 2,0,0,22,0,0,-2, and the eigenvalue 00 carries the two-dimensional space spanned by 01\ket{01} and 10\ket{10}. So the tuple count is three against a dimension of four and the set fails. For an explicit witness take C^=χχ\hat C=\ket{\chi}\bra{\chi} with χ=(0110)/2\ket\chi=(\ket{01}-\ket{10})/\sqrt2. It commutes with T^\hat T because χ\ket\chi is a T^\hat T-eigenvector, and it is not a function of T^\hat T, because any function of T^\hat T acts as a single number on the whole zero-eigenvalue space while C^\hat C annihilates (01+10)/2(\ket{01}+\ket{10})/\sqrt2. That is §3.3's second argument, made concrete.

(c) S^\hat S is unitary and S^2=I^\hat S^{2}=\hat I, so S^=S^1=S^\hat S^{\dagger}=\hat S^{-1} =\hat S and it is self-adjoint with eigenvalues ±1\pm1. It commutes with T^\hat T because T^\hat T treats the two factors alike. The joint eigenvectors and tuples are 00(2,+1)\ket{00}\to(2,+1), 11(2,+1)\ket{11}\to(-2,+1), (01+10)/2(0,+1)(\ket{01}+\ket{10})/\sqrt2\to(0,+1), and χ(0,1)\ket\chi\to(0,-1). Four distinct tuples again, so the set is complete: the swap supplies exactly the label T^\hat T was missing, and it does so inside the degenerate eigenspace, which is Chapter 0.5 §8.2's Step 4 in action.

(d) Compute [Z^1,S^][\hat Z_{1},\hat S] on 01\ket{01}: the first order gives Z^110=10\hat Z_{1}\ket{10}=-\ket{10} and the second gives S^01=+10\hat S\ket{01}=+\ket{10}, so the commutator is 2100-2\ket{10}\neq0. The two sets are individually complete and mutually incompatible, so a state labelled by one has no definite labels in the other. This is exactly the relationship between the uncoupled and coupled bases of Chapter 4.12, which adds two angular momenta and finds the same two ways of labelling the same four states, with the swap replaced by the total angular momentum. Chapter 4.18 then explains why the antisymmetric combination χ\ket\chi is singled out by nature rather than by choice.

The brick you just laid — one bracket, doing three jobs

The uncertainty relation cost one line and then three. Chapter 0.9 proved the bandwidth theorem with no physics in it and said that quantum mechanics would add a single substitution. This chapter made it: p=kp=\hbar k turns ΔxΔk12\Delta x\,\Delta k\ge\tfrac12 into ΔxΔp/2\Delta x\,\Delta p\ge\hbar/2 by multiplication, and Chapter 4.6 §10.6's mark on de Broglie's relation is the only physical input anywhere in it. The general case took three steps and no new ideas. Chapter 4.2 §5.3's variance is the length of a vector, and Chapter 0.5 §1.4's Cauchy–Schwarz bounds the overlap of two vectors by the product of their lengths. Splitting that overlap into its real and imaginary parts then puts the anticommutator in one and the commutator in the other. Keeping both gives the sharper form; discarding the first gives ΔAΔB12[A^,B^]\Delta A\,\Delta B\ge\tfrac12\abs{\avg{[\hat A,\hat B]}}. Chapter 0.5 said in advance that nothing would be added here except the meaning of the symbols, and that is what happened. The hypothesis that makes the derivation legal is a domain condition, and Worked example 2 exhibits it failing for an angle, computes the boundary term that the failure leaves behind, and repairs the statement.

What the relation is about, said once and kept. Each Δ\Delta is a spread of outcomes over systems prepared identically and measured once each. Chapter 4.5 §6.5 raises the deeper question and leaves it open, and §2's familiar-ground box says exactly how far this chapter can carry it. What is proved here is that for position and momentum, where the commutator is i\ii\hbar in every state and the floor is /2\hbar/2 however you prepare, no preparation narrows both spreads at once, so there is no unmeasured covariate to go looking for. For a general pair the floor is a property of the state and can be zero, at which point the theorem falls silent rather than permitting anything. The stronger claim, that no joint distribution exists underneath the two marginals at all, is Chapter 4.20's and is not derivable from the inequality. Measurement disturbance is a separate phenomenon with theorems of its own, and those theorems are the single thing this chapter quotes rather than derives, named in §2.6 with their hypotheses. Numbers rather than assurances. The ground state sits on the floor at ΔxΔp=0.500000000000\Delta x\,\Delta p=0.500000000000 and the n=3n=3 state sits seven times above it at 3.5000000000003.500000000000. Two hundred thousand random pairs violate neither form. And for a spin-one multiplet the bound is saturated in the top state, while in the middle one it falls silent and the product is 1.021.0\,\hbar^{2}.

Completeness is a count. Chapter 4.2 §4.3 defined a complete set of commuting observables and asked, in this section's name, how one knows a set is complete. The answer proved here is that one-dimensional common eigenspaces and maximality are the same condition, so a set is complete exactly when every observable commuting with all of it is already a function of it. The way anyone establishes that in practice is to list the eigenvalue tuples and check the count against the dimension. Chapters 4.11 and 4.13 apply it to {J^2,J^z}\{\hat J^{2},\hat J_{z}\} and to {H^,L^2,L^z}\{\hat H,\hat L^{2},\hat L_{z}\}, and Problem 4 runs the whole argument on two qubits, including a set that fails, the operator that witnesses the failure, and two complete sets that cannot be used together.

The same bracket moves things. Grouping the evolution operators with the observable instead of the state is Chapter 0.4 §4.2's similarity transformation, so the Heisenberg picture is a change of basis and not a rival theory, which is the sentence Chapter 0.4's warning box asked this chapter for. What the algebra does not carry is the domain, which moves as dom(A^H(t))=U^(t)dom(A^)\operatorname{dom}(\hat A_{H}(t))=\hat U^{\dagger}(t)\operatorname{dom}(\hat A), so A^H(t)\hat A_{H}(t) is a different operator at each time in Chapter 4.4 §3.1's sense and §4.3's derivative is a strong one on vectors that stay inside it. The list of readings is unmoved even so, by Chapter 4.5 §2.1's invertibility definition of the spectrum rather than by any determinant or trace. Differentiating gives dA^dt=1i[A^,H^]+A^t\dv{\hat A}{t}=\tfrac{1}{\ii\hbar}[\hat A,\hat H]+\pdv{\hat A}{t}, which is Chapter 1.3 §6.1's classical equation with the Poisson bracket replaced by 1i[  ,  ]\tfrac{1}{\ii\hbar}[\;,\;] and nothing else altered, and conservation becomes [A^,H^]=0[\hat A,\hat H]=0 with the whole distribution frozen rather than only its mean. Applied to x^\hat x and p^\hat p it gives Hamilton's equations as exact operator identities, and their expectations are Ehrenfest's relations, which Chapter 4.6 §8.7 stated and deferred. Chapter 1.1 §4.4 called these a derived statement about averages rather than a fundamental law, and §5 collected that warning rather than softening it. Here it is: V(x^)\avg{V'(\hat x)} is not V(x^)V'(\avg{\hat x}) unless VV''' vanishes identically, which happens for the quadratic potentials and for nothing else, the free particle, the uniform field and the oscillator upright or inverted. Two split-operator runs differing only in the potential measure the split. Both satisfy Ehrenfest to within a millionth, with the residuals falling by four each time dt\dd t is halved. The quadratic run has VV(x^)\avg{V'}-V'(\avg{\hat x}) equal to zero in the last bit at every step and its centre on 1.3cost1.3\cos t to 2.6×1072.6\times10^{-7}; the quartic run misses by up to 1.6657721.665772 against forces of order one, and that number does not move when the step size does.

And the same bracket generates. A linear symmetry forces unitarity by Chapter 4.2 §7.2, and linearity is assumed at that link rather than derived, discharged case by case because every family used here is visibly linear. Then a strongly continuous family has a self-adjoint generator by Stone, the generator acts on observables through a commutator by Chapter 4.2 §7.5, and leaving H^\hat H alone is the same equation as being conserved. Chapter 4.2 stated that correspondence and this chapter assembled it, the missing link having been the generator's existence in infinite dimensions. Translation gives p^\hat p, rotation gives L^z\hat L_{z}, time gives H^\hat H, and those are Chapter 1.4 §3's three classics repeated with operators. The rotation case delivers [L^x,L^y]=iL^z[\hat L_{x},\hat L_{y}]=\ii\hbar\hat L_{z} from the single postulated commutator by algebra alone, agreeing with what Chapter 1.4 §7.3's Poisson brackets predict under a substitution that was never postulated for that pair. A discrete symmetry leaves no generator, because any family running through it must be built out of the symmetry itself, and what it leaves instead is a two-valued label.

Three marks made elsewhere are leaned on here and cited rather than raised again, and this is the list. De Broglie's relation, Chapter 4.6 §10.6, is the one physical input to §1. The converse half of Stone's theorem, Chapter 4.5 §9.3, is what §6 needs to know that a continuous symmetry has a generator at all. And the identification H^=p^2/2m+V(x^)\hat H=\hat p^{2}/2m+V(\hat x), Chapter 4.6 §4.2, is what makes §4.6's two operator equations equations about a particle rather than about an unnamed self-adjoint operator. The canonical commutator of Chapter 4.2 §8 is used constantly and is a postulate rather than a quotation, which is a different thing and is why it carries no mark.

Where this gets spent. Chapter 4.11 takes (4.9.25) as its entire input, forgets that it came from rotations, and extracts the angular momentum spectrum from it by the ladder move Chapter 4.8 used on the oscillator. Chapter 4.12 meets Problem 4's two incompatible complete sets again, as the coupled and uncoupled bases. Chapter 4.13 labels hydrogen with {H^,L^2,L^z}\{\hat H,\hat L^{2},\hat L_{z}\} and uses §3's count to know that three labels are enough, and Chapter 4.16 uses Problem 3's virial relation to avoid new integrals. Chapter 4.17 turns Worked example 3's speed limit into line widths computed rather than bounded. Chapters 5.2 and 5.3 write everything in the Heisenberg picture, because a field has to carry its time dependence in the operator. The one thing this chapter has not touched is what becomes of all of it when \hbar is small against the action in play, since setting \hbar to zero in any equation here gives either nonsense or nothing. That is Chapter 4.10, and its §8 proves that the correspondence which worked for angular momentum in (4.9.25) cannot be made to work for every observable at once. You should now meet that theorem already believing it is needed.