Part IV · Quantum Mechanics — Chapter 4.5

The Spectral Theorem in Infinite Dimensions

An observable can have no eigenvectors at all. So "diagonalise it" has to be replaced by something that still says what the possible values are, and that replacement is what the rest of Part IV runs on.

Where we are

Chapter 4.4 said what an operator is. It is a formula together with a domain, the domain is forced rather than chosen, and the word that matters is self-adjoint rather than Hermitian, because the two come apart the moment the domain is not everything. That chapter's §1.3 divided Chapter 0.5's outstanding bill in two and named this chapter as the second payment: Chapter 4.4 is about what an operator may act on, and this one is about the values it may return.

Name the difficulty before the route, because it is not a technicality and it does not go away. In Chapter 0.5 an observable was a Hermitian matrix, and the spectral theorem produced an orthonormal basis of its eigenvectors. Position on L2(R)L^{2}(\R) has no eigenvectors. Not few, not hard to find, not eigenvectors of a delicate kind. It has none, and the same is true of momentum. So a theorem that hands you a basis of eigenvectors cannot be what survives, and the replacement has to say what the readings of an instrument are when there is nothing in the space to attach them to.

Here is the route. Section 1 states the difficulty exactly and says what will replace the finite-dimensional theorem. Section 2 widens the word spectrum so that an operator with no eigenvectors still has one, and proves two things about it that the rest of the chapter needs. Section 3 states the theorem that replaces A=UDUA=UDU^{\dagger}, and it is the one substantial result this book quotes rather than proves. Sections 4 and 5 then check it by hand on the three operators the rest of Part IV is built out of, which is what makes the mark honest; the third check needs the Hermite functions to be a genuine basis, and §5 proves that rather than quoting it. Section 6 turns the theorem into the projection-valued measure and the integral λdP(λ)\int\lambda\,\dd P(\lambda) that replaces Chapter 0.5's sum. Section 7 gives x\ket x and p\ket p a precise meaning, which Chapter 4.3 §8.5 explicitly declined to do and sent here. Section 8 is a checklist of what is now safe to write down, and Section 9 proves the half of Stone's theorem that makes "time evolution is unitary" and "the Hamiltonian is self-adjoint" the same statement.

Conventions. The inner product is linear in its second slot, as Chapter 0.5 §1.1 chose it. Every integral is Chapter 4.3's. Operators carry hats and eigenvalues do not, and I^\hat I is the identity. Three results are quoted rather than proved and each is marked where it arrives: the spectral theorem in multiplication form in §3.2, the rigged Hilbert space in §7.5, and the converse half of Stone's theorem in §9.3. There are no others. Three marks standing elsewhere are leaned on and cited rather than re-raised: Chapter 4.3 §2.3's construction of Lebesgue measure, Chapter 4.4 §2.3's closed graph theorem, and Chapter 4.4 §6.2's classification of extensions. The closing brick says what each mark bought.

Tools you'll need  — Chapter 0.5 above all: §6 for the finite-dimensional spectral theorem in both its forms, §6.4 for the projection form and the two promises made there, §7 for functions of an operator, and its closing paragraph, which this chapter finishes paying. Chapter 4.4 §3 for domains and the adjoint's domain, §4 for self-adjointness, §5.4 for the circle of momentum operators on an interval, and §6 for the extension count. Chapter 4.3 §4.3 for dominated convergence, §5.3 for the fact that a vector of L2L^{2} has no value at a point, §6 for completeness, §7.3 for the four equivalent statements about a basis, and §8.5 for the warning this chapter answers. Chapter 0.9 §2.1 for the box limit and the manufacturing of the mode spacing, §2.3 for Plancherel and the unitarity of the Fourier transform, §3.1 for the derivative theorem, and §5 for the delta defined by what it does. Chapter 4.2 §4 for the second postulate, §5 for the Born rule, and §7 for time evolution and its generator.

1 · The half of the bill that is still unpaid

This section does three things and then gets out of the way. It reads back what Chapter 0.5 left outstanding and says which part of it belongs here. It states the difficulty in its sharpest form, which is an operator with no eigenvectors at all rather than an operator whose eigenvectors are hard to find. And it says, before any machinery arrives, what the three replacements are going to be, so that you can watch each one land. The part to hold on to is the second: the failure is not that a proof is missing, it is that the conclusion is false as stated.

1.1 · Which instalment this is

Chapter 0.5 ended by naming four steps in which its proofs had leaned on the dimension being finite, and it said the bill would come due in Chapters 4.4 and 4.5. Chapter 4.4 §1.2 took the four in order and settled two of them. The interchange of sums was paid outright, using a corollary of monotone convergence that Chapter 4.3 had already proved. The failure of injective implies surjective was worked in full, first as the shift operator and then, in that chapter's Worked example 3, as the reason a particle held on a half-line has no momentum observable.

Two remain, and Chapter 4.4 §1.3 assigned both of them here. They are the induction, whose very first step assumes an eigenvector exists, and rank–nullity, whose replacement is the observation that an image can be dense without being closed. Those look like two separate repairs and they are one, as §2 will show: coming arbitrarily close to every vector and reaching every vector are different achievements, and the gap between them is exactly where an operator keeps values that have no eigenvectors attached to them.

So this chapter is the second and last instalment of the bill Chapter 0.5 named. Counting instead from Chapter 4.3, which built the space rather than the operators on it, this is the third chapter of the run, and Chapter 4.4's closing brick counts it that way. Either count gives the same shape, and it is the shape Chapter 4.2's closing brick announced in advance: Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it.

1.2 · The difficulty, in its sharpest form

Chapter 0.5 §6.3 proved its theorem by an induction whose first step was to produce one eigenvector, using the fact that a polynomial over C\C has a root. Here is what happens to that step. Take x^\hat x, multiplication by xx, on L2(R)L^{2}(\R), and ask for a vector ψ\psi and a number λ\lambda with x^ψ=λψ\hat x\psi=\lambda\psi. Written out, that is the requirement

(xλ)ψ(x)  =  0for almost every x. (x-\lambda)\,\psi(x) \;=\; 0 \qquad\text{for almost every } x. (4.5.1)

At every point where xλx\ne\lambda the first factor is not zero, so ψ(x)\psi(x) must be. The set where x=λx=\lambda is a single point, which has measure zero by Chapter 4.3 §2.4, so ψ\psi vanishes almost everywhere, and by the quotient of Chapter 4.3 §5.3 that means ψ\psi is the zero vector. There is no eigenvector for any λ\lambda whatsoever, real or complex. This is the step where the finite-dimensional argument stops working, and it is worth slowing down over, because the usual reassurances do not apply: nothing here is a matter of the eigenvector being awkward or the proof being incomplete.

Momentum fails in the mirror-image way, and Chapter 0.9 §2.1 had already flagged it. The equation iψ=pψ-\ii\hbar\psi'=p\psi does have solutions, namely ψ(x)=eipx/\psi(x)=\ee^{\ii px/\hbar}, and they are perfectly good functions. They are not vectors of L2(R)L^{2}(\R), because eipx/2=1\abs{\ee^{\ii px/\hbar}}^{2}=1 and the integral of 11 over the line is infinite. So for x^\hat x the eigenvalue equation has only the zero solution, and for p^\hat p it has solutions that live outside the space. Both operators are self-adjoint, by Chapter 4.4 §5.3 and its Worked example 1, and both are observables in the sense of Chapter 4.2's second postulate. Neither has a single eigenvector.

Say plainly what that costs. Chapter 0.5's conclusion was that a Hermitian operator has an orthonormal basis of eigenvectors, and Chapter 4.2 spent that conclusion everywhere: the possible outcomes of a measurement are the eigenvalues, the state expands in eigenstates, and the Born rule is a statement about projections onto eigenspaces. For position and momentum every one of those sentences is about an empty set. The theorem is not weakened here. It is inapplicable, which is a different and more serious thing, and Chapter 0.5's warning box used exactly that distinction about defective matrices.

1.3 · The three replacements, named in advance

What follows repairs this in three moves, and here they are as a list before any of them is built, because each is a weakening of something you already know and it helps to see which part is being given up.

  • The eigenvalues become the spectrum. Instead of asking which λ\lambda have eigenvectors, ask which λ\lambda make A^λ\hat A-\lambda fail to be invertible in a strong enough sense. In finite dimensions those two questions have the same answer, so nothing is being redefined. In infinite dimensions the second question has answers when the first has none, and the set of them is the spectrum.
  • A=UDUA=UDU^{\dagger} becomes unitary equivalence to a multiplication operator. The finite-dimensional statement is that some unitary change of basis turns AA into a list of numbers attached to independent directions. The infinite-dimensional statement is that some unitary map turns A^\hat A into multiplication by a real function on some space of square-integrable functions. A list of numbers is a function on a finite set, so this is the same sentence with a wider idea of what a list can be.
  • The sum kλkPk\sum_k\lambda_kP_k becomes an integral λdP(λ)\int\lambda\,\dd P(\lambda). Chapter 0.5 §6.4 wrote the theorem a third way, as a weighted sum of projections onto eigenspaces, and said that this basis-independent form is the one that survives here. It does survive, with the sum replaced by an integral against a family of projections indexed by the real line. That family is what the physics is waiting for. The Born rule is a statement about projections, so until a projection exists for a continuous observable there is no probability to compute, and a position measurement has no theory of its outcome at all.

One of those three is quoted rather than proved, and it is the middle one. That is the single substantial mathematical flag of Part IV, and §3.2 states it with its hypotheses attached. The answer to the obvious objection is the point of §§4 and 5: the theorem is then verified by hand on x^\hat x, on p^\hat p and on the oscillator Hamiltonian, which are the three operators everything later is built out of. You will finish this chapter holding a quoted theorem that you have checked yourself on the cases that matter, and §5.6 says exactly how far that checking reaches.

The three replacements are not where the chapter stops, and knowing now what they are for is what carries you through the checking. Once an observable has become multiplication by a real function, every function of that observable can be built the same way, and the one Part IV needs most is the exponential eiH^t/\ee^{-\ii\hat Ht/\hbar}. Section 9 shows that this operator is unitary, that evolving by ss and then by tt is evolving by s+ts+t, and that differentiating it at the origin hands H^\hat H back. That is Stone's theorem, and what it buys is that time evolution follows from self-adjointness instead of being postulated alongside it. Chapter 4.6 begins there.

In plain terms 4.5.1

The previous chapter established which mathematical objects are allowed to represent measurable quantities, and it did that by being careful about what each object is permitted to act on. What it did not supply is the list of numbers the instrument can actually read. In the finite world that list was the set of special multipliers belonging to the special directions, and finding it was the whole content of the central theorem of the toolkit.

Here is the trouble, and it is not a subtlety. For a particle on a line there are no special directions. Ask for a state in which the position has one definite value and the honest answer is that no such state exists, because a state concentrated at a single point is indistinguishable from nothing at all. Ask for a state of definite momentum and the answer is a perfectly good wave that spreads over the whole line with undiminished size, which means it is not a state either. So the recipe of finding the special directions and reading off their multipliers has nothing to work on, and yet a position measurement plainly returns a number.

The repair takes three steps, and each one loosens a requirement rather than replacing it. Ask which candidate values cannot be undone instead of which ones have states attached. Ask for a change of description that turns the quantity into ordinary multiplication instead of into a list. And replace the sum over separate values by an integral over a continuum of them. The one step that this book quotes rather than builds is the middle one, and the two sections after it check that quotation by hand on the three quantities everything later is built out of.

2 · The spectrum, when there are no eigenvectors

Section 1 left a hole where the eigenvalues used to be, and this section fills it. We take the finite-dimensional definition of an eigenvalue, notice that it can be phrased in a way that never mentions an eigenvector, adopt that phrasing as the definition, and check that in finite dimensions it gives back exactly the old answer. Then we run it on position and momentum and get, in both cases, the whole real line. The result to carry forward is the last one: for a self-adjoint operator the spectrum is real and it has only two parts rather than three, and both of those facts are proved here rather than assumed.

2.1 · The definition an eigenvalue already had

In Chapter 0.5 a number λ\lambda was an eigenvalue of AA when Av=λvAv=\lambda v for some v0v\ne0. That says AλA-\lambda kills a non-zero vector, which is the same as saying AλA-\lambda is not injective, which for a square matrix is the same as saying AλA-\lambda is not invertible. The chain of equivalences repays a pause, because every link in it is a finite-dimensional fact. Killing nothing and reaching everything are the same condition only when there is a dimension count to convert one into the other, and Chapter 4.4 §1.2 showed that count failing in the flesh with the shift operator.

So the finite-dimensional definition can be read two ways, and the two readings come apart here. One of them mentions an eigenvector and the other does not, and we want the one that does not, because §1.2 has just shown there may be no eigenvector to mention. Taking invertibility as the primitive costs nothing in finite dimensions and buys everything in infinite ones.

2.2 · The definition, widened

Let A^\hat A be an operator with a dense domain in a Hilbert space HH, in the sense Chapter 4.4 §3 fixed. Since A^\hat A acts only on dom(A^)\operatorname{dom}(\hat A), so does A^λ\hat A-\lambda, and asking whether it is invertible means asking whether it maps that domain onto the whole space in a way that can be undone continuously. Write that out as the condition on λ\lambda:

λρ(A^)A^λ: dom(A^)H  is one-to-one and onto,with (A^λ)1 bounded. \begin{gathered} \lambda\in\rho(\hat A) \quad\Longleftrightarrow\quad \hat A-\lambda:\ \operatorname{dom}(\hat A)\to H \ \text{ is one-to-one and onto,} \\[4pt] \text{with } (\hat A-\lambda)^{-1} \text{ bounded.} \end{gathered} (4.5.2)

The set ρ(A^)\rho(\hat A) of such λ\lambda is the resolvent set, and (A^λ)1(\hat A-\lambda)^{-1} is the resolvent. Everything this chapter calls a value of an observable is defined by the failure of that condition. So the spectrum is what is left over:

σ(A^)  =  Cρ(A^). \sigma(\hat A) \;=\; \C\setminus\rho(\hat A). (4.5.3)

Three demands are being made at once in (4.5.2) and each can fail on its own, which is why the spectrum splits into parts rather than being a single kind of thing. Notice also that boundedness of the inverse is a genuine third demand and not a convenience. Chapter 4.4 §2.1 measured what an unbounded operator does to a convergent sequence, and an unbounded inverse would mean that vectors A^\hat A maps to nearly nothing are not nearly nothing themselves, which is precisely the situation an approximate measurement is in.

2.3 · Nothing has been redefined, only widened

Before using the new definition, confirm that it agrees with the old one where the old one applies, since otherwise the chapter would be changing the subject rather than continuing it. Let AA be an operator on a space of finite dimension nn, so dom(A)\operatorname{dom}(A) is everything and every linear map is bounded by Chapter 0.6 §2. If AλA-\lambda is injective then rank–nullity makes it surjective, and the inverse is a linear map on a finite-dimensional space and therefore bounded, so λ\lambda is in the resolvent set. If AλA-\lambda is not injective then λ\lambda is an eigenvalue in the old sense. The two possibilities are exhaustive, so

σ(A)  =  {eigenvalues of A}when dimH<. \sigma(A) \;=\; \{\text{eigenvalues of } A\} \qquad\text{when } \dim H\lt\infty. (4.5.4)

Every step of that used finite dimension, which is the point: the new definition is the old one wherever the old one was available, and the two separate only where Chapter 0.5's proofs separate from their hypotheses. The word spectrum has been widened and nothing has been renamed.

2.4 · The three ways the condition can fail

Since (4.5.2) makes three demands, a λ\lambda in the spectrum fails at least one of them, and sorting by which one fails first gives the standard division. Read the three as a sequence of increasingly mild failures.

  • Point spectrum σp\sigma_p: A^λ\hat A-\lambda is not one-to-one. Something non-zero is sent to zero, so there is an eigenvector, and λ\lambda is an eigenvalue in Chapter 0.5's sense. This is the part that survives from finite dimensions unchanged.
  • Continuous spectrum σc\sigma_c: A^λ\hat A-\lambda is one-to-one, its image is dense, and its image is not the whole space. The inverse then exists on a dense set and is unbounded there. This is the case Chapter 4.4 §1.2 predicted when it said the replacement for rank–nullity is the distinction between a dense image and a closed one.
  • Residual spectrum σr\sigma_r: A^λ\hat A-\lambda is one-to-one and its image is not even dense. There is then a direction that no vector of the form (A^λ)ψ(\hat A-\lambda)\psi has any component along, which is a strong kind of failure.

One case is absent from that list, and here is why. A λ\lambda could in principle leave A^λ\hat A-\lambda one-to-one and onto while failing only the third demand, with an inverse that is defined everywhere and unbounded. For a closed operator that cannot happen. The inverse of a closed operator has a closed graph, and a closed map defined on the whole of a complete space is bounded, which is the closed graph theorem Chapter 4.4 §2.3 quotes. That mark is leaned on here rather than raised again, and §2.7 supplies the hypothesis by proving that a self-adjoint operator is closed.

With that case gone the three are mutually exclusive and together they are the whole spectrum. Section 2.7 will show that for a self-adjoint operator the third is always empty, so the physically relevant division is two-way: values that have states attached to them, and values that do not.

2.5 · Position: the spectrum is the whole real line

Now run the definition on the operator §1.2 showed has no eigenvectors. Take x^\hat x on L2(R)L^{2}(\R), with dom(x^)={ψ:xψL2}\operatorname{dom}(\hat x)=\{\psi:x\psi\in L^{2}\} from Chapter 4.4's Worked example 1. We want to know for which λ\lambda the three demands hold, so start with the candidate inverse, which is forced: undoing multiplication by xλx-\lambda can only be division by it.

[(x^λ)1ψ](x)  =  ψ(x)xλ. \big[(\hat x-\lambda)^{-1}\psi\big](x) \;=\; \frac{\psi(x)}{x-\lambda}. (4.5.5)

Take λ\lambda off the real axis first, since that is the easy case and it fixes the pattern. Writing λ=a+ib\lambda=a+\ii b with b0b\ne0, the modulus xλ\abs{x-\lambda} is at least b\abs b for every real xx, so dividing by xλx-\lambda multiplies the modulus of ψ\psi pointwise by at most 1/b1/\abs b. That makes (4.5.5) defined on every ψ\psi in the space and bounded, with (x^λ)11/b\norm{(\hat x-\lambda)^{-1}}\le1/\abs b, and multiplying back by xλx-\lambda returns ψ\psi. All three demands hold, so no non-real number is in the spectrum.

Now take λ\lambda real, and the useful thing to compute is not the inverse but how badly it fails to be bounded. Concentrate a unit vector on a shrinking interval to the right of λ\lambda, by setting ψn=n1[λ,λ+1/n]\psi_n=\sqrt n\,\mathbf 1_{[\lambda,\,\lambda+1/n]}, which has ψn=1\norm{\psi_n}=1 for every nn. Apply x^λ\hat x-\lambda to it and measure the result:

(x^λ)ψn2  =  nλλ+1/n(xλ)2dx  =  13n2,so(x^λ)ψn=1n3. \norm{(\hat x-\lambda)\psi_n}^{2} \;=\; n\int_{\lambda}^{\lambda+1/n}(x-\lambda)^{2}\,\dd x \;=\; \frac{1}{3n^{2}}, \qquad\text{so}\qquad \norm{(\hat x-\lambda)\psi_n}=\frac{1}{n\sqrt3}. (4.5.6)

Unit vectors are being sent to vectors of length 1/n31/n\sqrt3, which tends to zero. Any inverse would have to send those short vectors back to unit ones, so its norm is at least n3n\sqrt3 for every nn and it cannot be bounded. Every real λ\lambda is therefore in the spectrum, and by §1.2 none of them is an eigenvalue. Two further facts put them all in the continuous part, and the second is the one usually skipped. The image is dense: given any ϕ\phi, cut out a shrinking neighbourhood of λ\lambda by setting ϕn=ϕ1{xλ>1/n}\phi_n=\phi\,\mathbf 1_{\{\abs{x-\lambda}\gt1/n\}}. Then ϕn/(xλ)\phi_n/(x-\lambda) is in the space, because the factor is bounded by nn there, and ϕnϕ\phi_n\to\phi in norm by dominated convergence with ϕ2\abs\phi^{2} as the dominating function, which is Chapter 4.3 §4.3. And the image is not the whole space, which the continuous case demands as well. Take ϕ=1[λ,λ+1]\phi=\mathbf 1_{[\lambda,\lambda+1]}, a perfectly good vector. The only candidate for a ψ\psi with (x^λ)ψ=ϕ(\hat x-\lambda)\psi=\phi is ϕ/(xλ)\phi/(x-\lambda), and λλ+1(xλ)2dx\int_{\lambda}^{\lambda+1}(x-\lambda)^{-2}\dd x diverges, so that candidate is not in the space and ϕ\phi is not in the image. So

σ(x^)  =  σc(x^)  =  R,σp(x^)=. \sigma(\hat x) \;=\; \sigma_c(\hat x) \;=\; \R, \qquad \sigma_p(\hat x)=\varnothing. (4.5.7)

Read that against what an instrument does. The possible readings of a position measurement are all the real numbers, which is the physically right answer, and the theory delivers it without producing a single state of definite position. The sequence ψn\psi_n is what stands in for the missing eigenvector, and it is not a formal device: ψn\psi_n is a state whose position is known to within 1/n1/n, which is what a real detector prepares.

⚠ What "no eigenvectors" does not mean

It does not mean position cannot be measured, and it does not mean the theory is silent about where the particle is. Everything an experiment can report survives, and (4.5.6) is where to see why. There is no state with x^ψ=λψ\hat x\psi=\lambda\psi exactly, and there are states with (x^λ)ψ\norm{(\hat x-\lambda)\psi} as small as you please, at unit norm. The difference between those two sentences is the difference between an eigenvector and an approximate eigenvector, and no measurement of finite resolution can tell them apart.

The right way to say it is that the theory refuses to answer a question that no apparatus asks. A detector reports that the particle was found in a region, never that it was found at a point, and Chapter 4.3 §5.3 already showed that the second report would be meaningless in this space for reasons that have nothing to do with quantum mechanics. Section 6.5 turns that into the form of the Born rule for a continuous variable, where the probability attaches to an interval and not to a point.

What genuinely is lost is the expansion. In Chapter 0.5 you could write any vector as a sum over eigenvectors with coefficients, and for x^\hat x there is no such sum, because there is nothing to sum over. Section 7 supplies the replacement and says exactly how much of the old notation survives the substitution.

2.6 · Momentum, by the same calculation in the other description

Momentum needs no separate argument, because Chapter 0.9 §2.3 supplies a unitary map that turns it into a multiplication operator, and the spectrum cannot notice a unitary change of description. Say why not, since the step is used repeatedly from here on. If W^\hat W is unitary and B^=W^A^W^1\hat B=\hat W\hat A\hat W^{-1}, then B^λ=W^(A^λ)W^1\hat B-\lambda=\hat W(\hat A-\lambda)\hat W^{-1}, so one of the two is one-to-one, onto, or boundedly invertible exactly when the other is, since composing with a unitary map changes neither injectivity nor surjectivity nor any norm. Unitarily equivalent operators therefore have equal spectra, part by part.

Apply that with W^\hat W the Fourier transform. Chapter 0.9 §3.1's derivative theorem says ψ~=ikψ~\widetilde{\psi'}=\ii k\tilde\psi, so transforming p^=id/dx\hat p=-\ii\hbar\,\dd/\dd x turns differentiation into multiplication:

p^ψ~(k)  =  iψ~(k)  =  iikψ~(k)  =  kψ~(k). \widetilde{\hat p\psi}(k) \;=\; -\ii\hbar\,\widetilde{\psi'}(k) \;=\; -\ii\hbar\cdot\ii k\,\tilde\psi(k) \;=\; \hbar k\,\tilde\psi(k). (4.5.8)

In the transformed description p^\hat p is multiplication by the real function k\hbar k, which is x^\hat x with k\hbar k in place of xx. Section 2.5 therefore applies word for word, and the conclusion transfers back through the unitary:

σ(p^)  =  σc(p^)  =  R,σp(p^)=. \sigma(\hat p) \;=\; \sigma_c(\hat p) \;=\; \R, \qquad \sigma_p(\hat p)=\varnothing. (4.5.9)

That is the fact Chapter 0.9 §2.1 flagged and could not then justify, now derived. Momentum has every real number among its possible readings and not one momentum eigenstate in the space. It is also the first appearance of the pattern the whole chapter turns on, which is that a hard question about an operator became an easy question about a function once a unitary map had been found. Section 3 says that such a map always exists.

2.7 · Two facts about self-adjoint operators, proved

The definition in (4.5.2) allows complex λ\lambda, and an observable had better not have complex readings, so that has to be settled rather than assumed. It follows from self-adjointness together with completeness, and the argument also disposes of the residual spectrum. Start from the estimate that does the work. Let A^\hat A be self-adjoint and let λ=a+ib\lambda=a+\ii b with b0b\ne0. Write B^=A^a\hat B=\hat A-a, which is symmetric because A^\hat A is and aa is real. Now expand the norm for any ψ\psi in the domain:

(A^λ)ψ2  =  B^ψ2    ibB^ψ,ψ  +  ibψ,B^ψ  +  b2ψ2=  B^ψ2+b2ψ2, \begin{aligned} \norm{(\hat A-\lambda)\psi}^{2} \;&=\; \norm{\hat B\psi}^{2} \;-\;\ii b\avg{\hat B\psi,\psi} \;+\;\ii b\avg{\psi,\hat B\psi} \;+\; b^{2}\norm{\psi}^{2} \\[4pt] &=\; \norm{\hat B\psi}^{2}+b^{2}\norm{\psi}^{2}, \end{aligned} (4.5.10)

The two middle terms cancelled because symmetry makes B^ψ,ψ\avg{\hat B\psi,\psi} and ψ,B^ψ\avg{\psi,\hat B\psi} the same number. Dropping the first term on the right gives (A^λ)ψbψ\norm{(\hat A-\lambda)\psi}\ge\abs b\,\norm\psi, so A^λ\hat A-\lambda is one-to-one and its inverse is bounded by 1/b1/\abs b wherever it is defined. What remains is to show the inverse is defined everywhere, which is where completeness enters, and the grind box does it in two steps: the image is dense, and it is closed.

Grind box — the image of A^λ\hat A-\lambda is the whole space when λ\lambda is not real

First, the adjoint is closed. Suppose undom(A^)u_n\in\operatorname{dom}(\hat A^{\dagger}) with unuu_n\to u and A^unw\hat A^{\dagger}u_n\to w. For every vdom(A^)v\in\operatorname{dom}(\hat A) the defining relation of Chapter 4.4 §3.3 gives A^un,v=un,A^v\avg{\hat A^{\dagger}u_n,v}=\avg{u_n,\hat Av}. Both sides converge, because Cauchy–Schwarz makes the inner product continuous in each slot, and the limits are w,v=u,A^v\avg{w,v}=\avg{u,\hat Av}. That is the statement that udom(A^)u\in\operatorname{dom}(\hat A^{\dagger}) with A^u=w\hat A^{\dagger}u=w. So a convergent sequence in the domain whose images also converge has its limit in the domain, which is what closed means, and a self-adjoint operator is its own adjoint and therefore closed.

Second, the image is dense. Suppose it is not. Then there is u0u\ne0 orthogonal to it, so u,(A^λ)v=0\avg{u,(\hat A-\lambda)v}=0 for every vv in the domain, which reads u,A^v=λˉu,v\avg{u,\hat Av}=\avg{\bar\lambda u,v} and says udom(A^)u\in\operatorname{dom}(\hat A^{\dagger}) with A^u=λˉu\hat A^{\dagger}u=\bar\lambda u. Self-adjointness turns that into A^u=λˉu\hat Au=\bar\lambda u, so λˉ\bar\lambda is an eigenvalue. But Chapter 4.4 §4.3 showed the eigenvalues of a symmetric operator are real, and λˉ\bar\lambda is not. So no such uu exists.

Third, the image is closed. Let yn=(A^λ)ψny_n=(\hat A-\lambda)\psi_n converge to yy. The estimate above applied to ψnψm\psi_n-\psi_m gives ψnψmynym/b\norm{\psi_n-\psi_m}\le\norm{y_n-y_m}/\abs b, so ψn\psi_n is Cauchy and converges to some ψ\psi by completeness of the space, which is Chapter 4.3 §6.2. Then ψnψ\psi_n\to\psi and A^ψn=yn+λψny+λψ\hat A\psi_n=y_n+\lambda\psi_n\to y+\lambda\psi, so closedness puts ψ\psi in the domain with A^ψ=y+λψ\hat A\psi=y+\lambda\psi, which is (A^λ)ψ=y(\hat A-\lambda)\psi=y. So yy is in the image.

Dense and closed is everything, so A^λ\hat A-\lambda maps its domain onto HH, and with the bound already in hand all three demands of (4.5.2) hold.

That settles the first fact. Every later section leans on it, so it goes on its own line.

A^=A^σ(A^)R,(A^λ)11Imλ. \hat A=\hat A^{\dagger} \qquad\Longrightarrow\qquad \sigma(\hat A)\subseteq\R, \qquad \norm{(\hat A-\lambda)^{-1}}\le\frac{1}{\abs{\operatorname{Im}\lambda}}. (4.5.11)

The second fact is the middle step of the grind box read again, and it costs nothing extra. If λ\lambda is real and in the residual spectrum, then by definition the image of A^λ\hat A-\lambda is not dense, so the same argument produces u0u\ne0 with A^u=λu\hat Au=\lambda u. That makes λ\lambda an eigenvalue, so A^λ\hat A-\lambda is not one-to-one, and the residual case required it to be. The two conditions contradict each other, so a self-adjoint operator has no residual spectrum. Every value an observable can return is therefore either an eigenvalue with a genuine state behind it, or a continuous-spectrum value with approximate states of the kind (4.5.6) exhibited, and there is no third possibility.

In plain terms 4.5.2

The word that had to be widened is the one naming the list of possible readings. In the finite world a candidate reading earned its place by having a state attached to it, a state in which the quantity has that value and no other. The widened test asks something weaker and more robust: subtract the candidate reading from the quantity and ask whether what is left can be reliably undone. If it can, that reading is impossible. If it cannot, the reading is possible, and this version of the question makes sense whether or not any state is attached.

Running the test on position gives every real number, which is the answer anyone would want, and it gives it without a single state of definite position existing. What stands in the missing state's place is a sequence of states in which the position is pinned down more and more tightly, none of them perfect and each of them something a laboratory could prepare. The theory is refusing only the idealisation, not the measurement.

Two further things fall out and both are reassurances the theory owed. The readings of a genuine observable are real numbers rather than complex ones, which had been assumed since the quantity was first written down and is now proved. And the ways the test can fail reduce from three to two, so every possible reading either has a state behind it or has a sequence of states closing in on it. Nothing sits in a third category with no physical account at all.

3 · The spectral theorem, in multiplication form

We now have a set of possible readings for an operator that has no eigenvectors, and no way yet to compute with it. This section supplies the statement that does the computing. It is the replacement for A=UDUA=UDU^{\dagger}, it is quoted rather than proved, and the two sections after it check it by hand on the three operators the rest of Part IV rests on. Before the statement arrives, see what shape it has to have. The shape is forced by §2, and the theorem then looks inevitable rather than arbitrary.

3.1 · What a list of numbers has to become

Chapter 0.5's conclusion, stripped of its notation, was that a Hermitian operator is a list of numbers attached to independent directions, seen from a rotated angle. The rotation is UU, the list is the diagonal of DD, and the directions are the columns of UU. Ask what each of those three pieces has to become here, and take them in the order that forces the answer.

The directions are the part that cannot survive, because §1.2 showed there may be none. So the theorem cannot be stated in terms of them, and whatever replaces the list has to be indexed by something other than a set of eigenvectors. The list itself is the piece to work on. A list of nn numbers is a function on the set {1,2,,n}\{1,2,\dots,n\}, and a diagonal matrix acts on a vector by multiplying its jjth component by the jjth entry of that function. Written that way, the operation DD performs is not "attach a number to each direction" but multiply by a function, and multiplication by a function makes sense on any set at all, finite or not.

That is the whole of the generalisation, and §2.6 has already shown one instance of it working. The Fourier transform turned p^\hat p into multiplication by k\hbar k, and k\hbar k is a function on the real line rather than a list on a finite set. Nothing about that calculation needed an eigenvector, and the spectrum came out of it immediately. The theorem says this is always available.

3.2 · The theorem, quoted with its hypotheses

⚑ Quoted, not derived — the spectral theorem for unbounded self-adjoint operators, in multiplication form

We use, without proof, the following. Let A^\hat A be a self-adjoint operator, domains included, on a separable Hilbert space HH. Then there exist a measure space (X,μ)(X,\mu), a unitary map U^:HL2(X,μ)\hat U:H\to L^{2}(X,\mu), and a real-valued measurable function gg on XX, such that

(U^A^U^1Ψ)(ξ)  =  g(ξ)Ψ(ξ),dom(A^)={ψH: gU^ψL2(X,μ)}. \big(\hat U\hat A\hat U^{-1}\Psi\big)(\xi) \;=\; g(\xi)\,\Psi(\xi), \qquad \operatorname{dom}(\hat A)=\big\{\psi\in H:\ g\cdot\hat U\psi\in L^{2}(X,\mu)\big\}.

In words: every self-adjoint operator is unitarily equivalent to multiplication by a real function on some L2L^{2} space, and the domain is exactly the set of vectors for which that multiplication lands back in the space.

The hypotheses are the content. Self-adjoint cannot be weakened to symmetric, and Chapter 4.4 §5.4 built the counterexample. Momentum on [0,L][0,L] with the wavefunction pinned at both ends is symmetric and has no eigenvalues at all. It is not unitarily equivalent to multiplication by anything, because multiplication by a real function is self-adjoint and self-adjointness is preserved by unitary maps. Separable is free here, since Chapter 4.3 §7.4 proved every space this book uses has a countable orthonormal basis. What the statement does not claim is uniqueness of the display: the triple (X,μ,g)(X,\mu,g) is not determined by A^\hat A, and §4 will exhibit the same operator in two different ones.

What the mark covers, and what it costs. It covers two statements and nothing else. The first is the existence of (X,μ,U^,g)(X,\mu,\hat U,g). The second is that the family of projections §6.1 reads off that data comes out the same whichever display is used, so that P(E)P(E) depends on A^\hat A and on EE and on nothing else. The second clause is quoted here alongside the first because it is the same theorem's other half, and because §6 cannot write the projection-valued measure of A^\hat A, or hang the Born rule on it, until that half is available. The standard proof of the existence runs in three stages. The Cayley transform (A^i)(A^+i)1(\hat A-\ii)(\hat A+\ii)^{-1} converts the unbounded self-adjoint operator into a bounded unitary one, which is legitimate precisely because (4.5.11) put ±i\pm\ii in the resolvent set. A continuous functional calculus for that unitary operator is then built from polynomials by Stone–Weierstrass. And the Riesz representation theorem converts the resulting positive linear functionals into the measures that make up μ\mu. It is three chapters of analysis and this book does not spend them, and the uniqueness clause is proved in the same place. Reed and Simon's Methods of Modern Mathematical Physics, volume I, chapters VII and VIII, is where to read it, and Rudin's Functional Analysis gives the same route more briskly.

What the mark does not cover. Sections 4 and 5 do not use it. They construct U^\hat U, μ\mu and gg explicitly for the three operators everything later is built out of, so that the quotation is checked rather than trusted where it matters most. Section 5.6 says how far that checking reaches and where a further check is still owed.

3.3 · Read against A=UDUA=UDU^{\dagger}, term by term

Put the statement side by side with the one it replaces, because every piece corresponds and the correspondence is what makes it recognisable rather than new. This is the infinite-dimensional reading of Chapter 0.5's A=UDUA=UDU^{\dagger}, and the reading goes like this.

  • UU becomes U^\hat U. In Chapter 0.5 the columns of UU were the eigenvectors, so UU was built out of the directions. Here U^\hat U is a unitary map onto a space of functions, and it is not built out of anything: its existence is the theorem. Chapter 0.9 §2.3 already supplied one such map, the Fourier transform, and proved it unitary.
  • DD becomes multiplication by gg. A diagonal matrix multiplies the jjth component by the jjth entry, and multiplication by gg multiplies the value at ξ\xi by g(ξ)g(\xi). The index set has changed from {1,,n}\{1,\dots,n\} to XX and the operation has not changed at all.
  • The diagonal entries become the values of gg. Chapter 0.5's list λ1,,λn\lambda_1,\dots, \lambda_n was the spectrum. Here the spectrum is the set of values gg takes, in a sense §3.4 makes precise, and gg may take each of them on a set of positive measure rather than at a single index.
  • DD real becomes gg real. Chapter 0.5 got real eigenvalues from Hermiticity in three lines. Here gg real is part of the statement, and §2.7 proved the consequence independently, which is a check on the two agreeing.
  • The eigenvectors have no counterpart. This is the one entry with nothing on the right, and it is the honest reason a new theorem was needed rather than a new proof of the old one.

One more difference has to be named, because it is where the physics enters. In Chapter 0.5 the measure was invisible, since a finite list carries counting measure and nobody has to mention it. Here μ\mu is part of the data, and §§4 and 5 will produce three operators whose measures are Lebesgue measure on the line, Lebesgue measure on the line again, and counting measure on the non-negative integers. The difference between a continuous spectrum and a discrete one is a difference of measure, and it is the only thing separating position from the oscillator.

3.4 · The spectrum, read off gg

The theorem is only useful if the spectrum can be read out of gg, so establish that now, since every later section quotes it. Multiplication by gg fails to be boundedly invertible at λ\lambda exactly when 1/(gλ)1/(g-\lambda) fails to be bounded, and by the argument of §2.5 that happens exactly when gg comes arbitrarily close to λ\lambda on sets of positive measure. Give that set a name. The essential range of gg is

ess ran(g)  =  {λR : μ({ξ:g(ξ)λ<ϵ})>0  for every ϵ>0}, \operatorname{ess\,ran}(g) \;=\; \big\{\lambda\in\R\ :\ \mu\big(\{\xi:\abs{g(\xi)-\lambda}\lt\epsilon\}\big)\gt0 \ \text{ for every } \epsilon\gt0\big\}, (4.5.12)

and the spectrum of A^\hat A is exactly that set, since unitary equivalence preserves spectra by §2.6. The qualifier essential is doing real work rather than decorating: a value that gg takes only on a set of measure zero is invisible to the space, because Chapter 4.3 §5.3's quotient has already declared such a set to carry no vectors. The distinction is the same one that made a wavefunction have no value at a point, appearing again on the other side of the unitary map.

The point spectrum reads off equally directly. A vector Ψ\Psi satisfies gΨ=λΨg\Psi=\lambda\Psi when (gλ)Ψ=0(g-\lambda)\Psi=0 almost everywhere, which forces Ψ\Psi to vanish wherever gλg\ne\lambda. So λ\lambda is an eigenvalue exactly when gg takes the value λ\lambda on a set of positive measure, and the eigenspace is the set of functions supported there. For x^\hat x the function is g(x)=xg(x)=x, which is constant on no set of positive measure at all, and §1.2's calculation is that observation made twice. Worked example 2 builds an operator where gg is constant on a half-line, and it has both kinds of spectrum at once.

In plain terms 4.5.3

The theorem that replaces the central result of the toolkit says something that sounds almost disappointing until you see what it buys. Every legitimate observable, however complicated, is the operation of multiplying by an ordinary real-valued function, once you have found the right way of describing the states. That is all. The complexity of the operator has been moved entirely into the change of description, and the operator itself has become the simplest thing a mathematician can write down.

Compare that with the finite version and the family resemblance is exact. There the claim was that a quantity is a list of numbers attached to a set of mutually perpendicular directions, and a list of numbers is nothing but a function defined on a finite set of labels. Widening the labels from a finite set to a continuum is the entire generalisation. What genuinely disappears is the set of special directions, and it disappears because for a particle on a line there are none, so the new statement has been written so as never to mention them.

This is the one result in this part of the book that is used but not proved, and the mark on it is the honest admission that its proof would cost three chapters of pure analysis. The next two sections are what makes borrowing it defensible. Rather than trust the general claim, we produce the change of description and the multiplying function by hand for each of the three quantities that everything later is built out of, so the borrowed statement is tested where it carries the most weight.

a natural place to stop  ·  the theorem is on the table; what follows is three checks of it by hand

4 · Checked three times

A quoted theorem is worth what you can check of it, so this section and the next check it. The plan is simple and the order is chosen so that the cost rises: position needs no work at all, momentum needs one unitary map that Chapter 0.9 already built and proved unitary, and the oscillator needs a basis whose completeness is proved from scratch in §5. Two of the three land here, and the third is set up here and finished there. Two subsections interrupt that count and neither is a fourth verification: §4.2 exhibits one operator through two different triples, which §3.2 said to expect, and §4.4 puts the second verification against arithmetic. At the end of §5 the three are collected, and that collection is what turns the mark of §3.2 from a promise into an accounting.

4.1 · Verification 1: position, where there is nothing to do

The theorem asks for a measure space, a unitary map and a real function. For x^\hat x on L2(R,dx)L^{2}(\R,\dd x) every one of them is already sitting there, because the operator was defined as multiplication in the first place. Take X=RX=\R with Lebesgue measure, take U^\hat U to be the identity, and take the function to be g(x)=xg(x)=x:

(x^ψ)(x)  =  xψ(x),(X,μ)=(R,dx),U^=I^,g(x)=x. \big(\hat x\psi\big)(x) \;=\; x\,\psi(x), \qquad (X,\mu)=(\R,\dd x), \qquad \hat U=\hat I, \qquad g(x)=x. (4.5.13)

The domain condition of §3.2 asks that gU^ψ=xψg\cdot\hat U\psi=x\psi lie in the space. That is exactly dom(x^)\operatorname{dom}(\hat x) as Chapter 4.4's Worked example 1 computed it, so nothing has to be adjusted. The essential range of g(x)=xg(x)=x is the whole real line, since every interval around every real number has positive length, and that reproduces (4.5.7) from a one-line reading of gg instead of from the sequence argument of §2.5.

This verification is trivial, and its triviality is the content rather than an embarrassment. What §3.2 asserts is that every self-adjoint operator looks like this one. Position is the model that the general statement is copied from, and checking it amounts to noticing that the model is a case of the statement it models.

4.2 · The same operator, two different answers

Section 3.2 said the triple (X,μ,g)(X,\mu,g) is not unique, and position is where that is easiest to see, so take a moment over it now rather than meeting it later as a surprise. Map the line onto a bounded interval by x=tanθx=\tan\theta, and carry states across in the way that preserves the norm:

(V^ψ)(θ)  =  ψ(tanθ)secθ,θ(π2,π2). (\hat V\psi)(\theta) \;=\; \psi(\tan\theta)\,\sec\theta, \qquad \theta\in\left(-\tfrac\pi2,\tfrac\pi2\right). (4.5.14)

That map is unitary, by the substitution x=tanθx=\tan\theta with dx=sec2θdθ\dd x=\sec^{2}\theta\,\dd\theta, which turns V^ψ2dθ\int\abs{\hat V\psi}^{2}\dd\theta into ψ2dx\int\abs\psi^{2}\dd x. Conjugating x^\hat x by it gives multiplication by tanθ\tan\theta on L2((π/2,π/2),dθ)L^{2}\big((-\pi/2,\pi/2),\dd\theta\big), so the same operator now comes with a bounded interval, a different measure and a different function. The essential range of tanθ\tan\theta over that interval is still the whole real line, so the spectrum is unchanged, as §2.6 requires. What the theorem fixes is the spectrum, not the triple that displays it, and a reader who expects (X,μ,g)(X,\mu,g) to be canonical will be looking for something that is not there.

4.3 · Verification 2: momentum, on a unitary you already built

Momentum is the first case where the unitary map has to be found rather than noticed, and it has already been found. Chapter 0.9 §2.3 proved that the Fourier transform preserves the norm and read that back as the statement that it is unitary, and §2.6 above transformed p^\hat p with it in one line. Collecting that as the theorem's data:

U^=F,(X,μ)=(R,dk),g(k)=k,(Fp^F1ψ~)(k)=kψ~(k). \hat U=\mathcal F, \qquad (X,\mu)=(\R,\dd k), \qquad g(k)=\hbar k, \qquad \big(\mathcal F\hat p\,\mathcal F^{-1}\tilde\psi\big)(k)=\hbar k\,\tilde\psi(k). (4.5.15)

Check the domain as well, since §3.2 makes it part of the statement and Chapter 4.4 spent a whole chapter on the point that a domain is not an afterthought. The theorem's condition is that kψ~\hbar k\tilde\psi lie in L2(dk)L^{2}(\dd k), and Chapter 4.4 §5.3 defined dom(p^)\operatorname{dom}(\hat p) as the functions that are integrals of an L2L^{2} derivative. Those are the same set, and Plancherel is what says so: ψL2\psi'\in L^{2} transforms to ikψ~L2\ii k\tilde\psi\in L^{2}, and the norms agree. So the description of the domain that took most of a page in Chapter 4.4 becomes, in the transformed picture, the single requirement that ψ~\tilde\psi decay fast enough for kψ~k\tilde\psi to stay square integrable.

Notice two things about how cheap that was. The whole verification is one line of algebra applied to a theorem you proved in Chapter 0.9, and the measure space is again the real line with Lebesgue measure, which is why position and momentum have identical spectra despite being different operators. Chapter 4.3 §8.5 had already said that the two are one Hilbert space seen through a unitary map; here that statement acquires its consequence for the values.

4.4 · The second verification, checked against arithmetic

Run the identity in (4.5.15) numerically, because it is the one place in this chapter where a claim about an unbounded operator can be put against a number. Work in units where =1\hbar=1 and lengths are measured in units of the width of the test state, and take a grid of 40964096 points across a box of length 4040. The test state is a Gaussian of unit width centred at x=1x=1 and given a boost of k=3k=3, which is ψ(x)e(x1)2/2e3ix\psi(x)\propto\ee^{-(x-1)^{2}/2}\ee^{3\ii x}, normalised on the grid.

Now compute p^ψ\hat p\psi two ways that share nothing. The first is the theorem's way: transform, multiply by k\hbar k, transform back. The second is the definition's way: approximate idψ/dx-\ii\hbar\,\dd\psi/\dd x by the centred difference i[ψ(x+h)ψ(xh)]/2h-\ii\hbar\big[\psi(x+h)-\psi(x-h)\big]/2h. The two answers agree to a relative discrepancy of

p^ψtransformp^ψdifferencep^ψ  =  1.96×104(4096 points), \frac{\norm{\hat p\psi\big|_{\text{transform}} - \hat p\psi\big|_{\text{difference}}}}{\norm{\hat p\psi}} \;=\; 1.96\times10^{-4} \qquad (4096\ \text{points}), (4.5.16)

and the size of that number is less informative than the way it moves. Halving the grid spacing divides it by four every time, running 3.13×1033.13\times10^{-3}, 7.82×1047.82\times10^{-4}, 1.96×1041.96\times10^{-4}, 4.89×1054.89\times10^{-5} over 10241024, 20482048, 40964096 and 81928192 points. A centred difference has an error proportional to the square of the spacing, so that factor of four identifies the entire discrepancy as the finite difference's error, with nothing left over to attribute to the transform. The expectation values say the same thing more bluntly: multiplying by k\hbar k in the transformed picture returns p^=3.0000000000\avg{\hat p}=3.0000000000, which is the boost that was put in, while the centred difference returns 2.99949935052.9994993505.

4.5 · Verification 3: the shape, and the one thing missing

The third operator is the harmonic oscillator Hamiltonian, and it is the case the rest of Part IV uses most, so it gets done properly. Write it with the particle's mass mm and the angular frequency ω\omega:

H^osc  =  22md2dx2  +  12mω2x2. \hat H_{\text{osc}} \;=\; -\frac{\hbar^{2}}{2m}\,\dvn{2}{}{x} \;+\; \half m\omega^{2}x^{2}. (4.5.17)

This one is different in kind from the first two, and the difference is the whole reason the theorem has to allow an arbitrary measure. Position and momentum have no eigenvectors, so their measure spaces had to be continuous. The oscillator does have eigenvectors, and they are the Hermite functions, so its measure space is going to be a set of isolated points carrying counting measure. If the theorem could not accommodate that, it would not cover the operator whose spectrum every physics course computes first.

The unitary map is the one Chapter 0.5 would have used: send a state to its list of coefficients against the eigenfunctions. What makes that map unitary is Parseval, and what makes Parseval available is that the eigenfunctions form a genuine orthonormal basis rather than merely an orthonormal set. Chapter 4.3 §7.3 is explicit that those are different claims, and it proved they are equivalent to two others, one of which is checkable. So the verification reduces to a single statement about the Hermite functions, and that statement is what §5 proves.

In plain terms 4.5.4

Borrowing a theorem is only respectable if you check it, and this book checks the borrowed one by hand on three quantities. The first is position, where the check is a matter of noticing that the quantity was already defined as multiplication by a function, so the general claim is being tested against the example it was modelled on. That is not a weakness in the check. It is the reason the general claim is plausible in the first place.

The second is momentum, and there the change of description is the transform between position and wavelength that was built in Chapter 0.9 and shown there to preserve every length and angle. In the wavelength description momentum is multiplication by an ordinary function, one line of algebra confirms it, and a numerical experiment on a four-thousand-point grid confirms it again to four figures. The discrepancy that remains shrinks by a factor of four each time the grid is refined, which is the signature of the crude derivative being wrong rather than the transform.

The third is the energy of an oscillating particle, and it differs from the other two in a way worth pausing on. This quantity does have states of definite value, an entire family of them, so its labels form a discrete list rather than a continuum. The general statement was written to allow either, which is why it speaks of a measure rather than of an interval. All that is missing is the guarantee that the family is large enough to describe every state, and the next section proves that rather than assuming it.

5 · The Hermite functions are complete

This section builds the one ingredient §4.5 was missing, and it builds it rather than quoting it, because Chapter 4.3's closing brick promised in writing that this chapter would. The route is short to describe. Define the Hermite functions by a generating function, and get two recurrences and a differential equation out of it. Use those to show the functions are eigenfunctions of H^osc\hat H_{\text{osc}} and that they are orthonormal. Then prove the claim that costs something, which is that no non-zero vector is orthogonal to all of them. The proof of that last step uses dominated convergence and Plancherel and nothing else. No complex analysis is required and nothing here carries a mark.

5.1 · The functions, and the variable that removes the constants

Constants clutter this calculation and there is a length in the problem that removes all of them, so introduce it first. The oscillator has one length that can be built out of \hbar, mm and ω\omega, namely x0=/mωx_0=\sqrt{\hbar/m\omega}. Writing ξ=x/x0\xi=x/x_0 turns (4.5.17) into

H^osc  =  ω2(d2dξ2+ξ2),ξ=xx0,x0=mω. \hat H_{\text{osc}} \;=\; \frac{\hbar\omega}{2}\left(-\dvn{2}{}{\xi}+\xi^{2}\right), \qquad \xi=\frac{x}{x_0}, \qquad x_0=\sqrt{\frac{\hbar}{m\omega}}. (4.5.18)

One notational borrowing, since that letter is already in use. From §3.2 onward ξ\xi has been the generic point of the abstract measure space XX, and for the whole of §5 it is the oscillator's dimensionless coordinate instead. It goes back to the abstract meaning at §6.1.

Both terms carry the same factor because that is what x0x_0 was chosen to arrange, and every statement below is about the bracket. Now define the polynomials that will appear. Rather than write them down and check their properties one at a time, define all of them at once by a single function whose expansion produces them. That is the device Chapter 0.9 §7.2 used for probability distributions:

e2ξtt2  =  n0Hn(ξ)tnn!. \ee^{\,2\xi t-t^{2}} \;=\; \sum_{n\ge0} H_n(\xi)\,\frac{t^{n}}{n!}. (4.5.19)

The left-hand side is an entire function of tt for each fixed ξ\xi, so the expansion exists and defines each HnH_n as a polynomial in ξ\xi. The first few are H0=1H_0=1, H1=2ξH_1=2\xi and H2=4ξ22H_2=4\xi^{2}-2, read off by expanding. The Hermite functions are these polynomials damped by a Gaussian and normalised. The constant is chosen now and justified in §5.3:

hn(ξ)  =  12nn!π  Hn(ξ)eξ2/2. h_n(\xi) \;=\; \frac{1}{\sqrt{2^{n}\,n!\,\sqrt\pi}}\;H_n(\xi)\,\ee^{-\xi^{2}/2}. (4.5.20)

These are real, they decay faster than any exponential, and they are square integrable, so each one is a genuine vector of L2(R)L^{2}(\R) in a way the momentum eigenfunctions of §1.2 were not. That is the first thing to notice about the oscillator: unlike position and momentum, it has candidate eigenvectors that actually live in the space.

5.2 · They are the eigenfunctions, and the eigenvalues are what you expect

We want H^oschn=(n+12)ωhn\hat H_{\text{osc}}h_n=(n+\half)\hbar\omega\,h_n, and by (4.5.18) that is the single statement hn+ξ2hn=(2n+1)hn-h_n''+\xi^{2}h_n=(2n+1)h_n. Everything needed comes out of (4.5.19) by differentiating it once in each variable, and the grind box does that. What comes out is two recurrences and, from them, a differential equation for HnH_n:

Hn+1=2ξHn2nHn1,Hn=2nHn1,Hn2ξHn+2nHn=0. H_{n+1}=2\xi H_n-2nH_{n-1}, \qquad H_n'=2nH_{n-1}, \qquad H_n''-2\xi H_n'+2nH_n=0. (4.5.21)
Grind box — the two recurrences and the equation, from the generating function

Write G(ξ,t)=e2ξtt2G(\xi,t)=\ee^{2\xi t-t^{2}}. Differentiating in tt gives tG=(2ξ2t)G\partial_tG=(2\xi-2t)G, and expanding both sides using (4.5.19),

n1Hntn1(n1)!  =  2ξn0Hntnn!    2n0Hntn+1n!. \sum_{n\ge1}H_n\frac{t^{n-1}}{(n-1)!} \;=\; 2\xi\sum_{n\ge0}H_n\frac{t^{n}}{n!} \;-\; 2\sum_{n\ge0}H_n\frac{t^{n+1}}{n!}.

Comparing the coefficient of tnt^{n} on the two sides gives Hn+1/n!=2ξHn/n!2Hn1/(n1)!H_{n+1}/n!=2\xi H_n/n!-2H_{n-1}/(n-1)!, and multiplying through by n!n! is the first recurrence.

Differentiating in ξ\xi instead gives ξG=2tG\partial_\xi G=2tG, so

n0Hntnn!  =  2n0Hntn+1n!, \sum_{n\ge0}H_n'\frac{t^{n}}{n!} \;=\; 2\sum_{n\ge0}H_n\frac{t^{n+1}}{n!},

and the coefficient of tnt^{n} reads Hn/n!=2Hn1/(n1)!H_n'/n!=2H_{n-1}/(n-1)!, which is the second recurrence.

Now eliminate. Substituting 2nHn1=Hn2nH_{n-1}=H_n' into the first recurrence gives Hn+1=2ξHnHnH_{n+1}=2\xi H_n-H_n'. Differentiate that in ξ\xi and use the second recurrence at index n+1n+1, which says Hn+1=2(n+1)HnH_{n+1}'=2(n+1)H_n:

2(n+1)Hn  =  2Hn+2ξHnHn, 2(n+1)H_n \;=\; 2H_n+2\xi H_n'-H_n'',

and cancelling 2Hn2H_n from both sides leaves Hn2ξHn+2nHn=0H_n''-2\xi H_n'+2nH_n=0.

Finally carry that to un=Hneξ2/2u_n=H_n\ee^{-\xi^{2}/2}. Differentiating twice, un=(Hn2ξHnHn+ξ2Hn)eξ2/2u_n''=\big(H_n''-2\xi H_n'-H_n+\xi^{2}H_n\big)\ee^{-\xi^{2}/2}, and the equation replaces Hn2ξHnH_n''-2\xi H_n' by 2nHn-2nH_n, leaving un=(ξ22n1)unu_n''=(\xi^{2}-2n-1)u_n.

The last line of the grind box rearranges to un+ξ2un=(2n+1)un-u_n''+\xi^{2}u_n=(2n+1)u_n, and hnh_n is unu_n times a constant, so the same holds for it. Putting that into (4.5.18),

H^oschn  =  (n+12)ωhn,n=0,1,2, \hat H_{\text{osc}}\,h_n \;=\; \left(n+\half\right)\hbar\omega\,h_n, \qquad n=0,1,2,\dots (4.5.22)

So the oscillator has eigenvectors, they are in the space, and the eigenvalues are the equally spaced ladder every physics course quotes. The half in the ground-state energy has arrived from nothing more than the recurrences. No physical argument produced it, and Chapter 4.8 will produce it a second way and get the same number.

5.3 · They are orthonormal

Orthonormality also comes out of (4.5.19), by multiplying two copies of it together and integrating. Take the generating function at two parameters ss and tt, multiply, and include the Gaussian weight, completing the square in the exponent so that Chapter 0.2's Gaussian integral applies:

e2ξss2e2ξtt2eξ2dξ  =  es2t2+(s+t)2e(ξst)2dξ  =  πe2st. \int_{-\infty}^{\infty}\ee^{2\xi s-s^{2}}\,\ee^{2\xi t-t^{2}}\,\ee^{-\xi^{2}}\,\dd\xi \;=\; \ee^{-s^{2}-t^{2}+(s+t)^{2}}\int_{-\infty}^{\infty}\ee^{-(\xi-s-t)^{2}}\dd\xi \;=\; \sqrt\pi\,\ee^{2st}. (4.5.23)

Now expand the two ends in powers and compare. The left-hand side, with each exponential replaced by its Hermite series, is m,nsmtnm!n!HmHneξ2dξ\sum_{m,n}\frac{s^{m}t^{n}}{m!\,n!}\int H_mH_n\ee^{-\xi^{2}}\dd\xi. The right-hand side is πn(2st)n/n!\sqrt\pi\sum_n(2st)^{n}/n!, which carries only terms with m=nm=n. Two power series in two variables agree only coefficient by coefficient, so

Hm(ξ)Hn(ξ)eξ2dξ  =  δmn2nn!π,sohm,hn=δmn. \int_{-\infty}^{\infty}H_m(\xi)H_n(\xi)\,\ee^{-\xi^{2}}\,\dd\xi \;=\; \delta_{mn}\,2^{n}\,n!\,\sqrt\pi, \qquad\text{so}\qquad \avg{h_m,h_n}=\delta_{mn}. (4.5.24)

The normalising constant in (4.5.20) was chosen to be exactly the square root of that number, which is why the second equality is immediate. So the Hermite functions are an orthonormal set. Chapter 4.3 §7.3 is emphatic that this is a weaker claim than being a basis, and it named the four statements that would upgrade it. The next subsection proves one of them.

5.4 · No vector is orthogonal to all of them

Of Chapter 4.3 §7.3's four equivalent statements, the one that can be checked here is (d): the only vector orthogonal to every hnh_n is the zero vector. Suppose then that fL2(R)f\in L^{2}(\R) satisfies hn,f=0\avg{h_n,f}=0 for every nn, and follow the consequences.

First convert the hypothesis into a statement about powers. The recurrence in (4.5.21) shows HnH_n has degree exactly nn, since it builds Hn+1H_{n+1} by multiplying HnH_n by 2ξ2\xi and subtracting something of lower degree. So the span of H0,,HNH_0,\dots,H_N is every polynomial of degree at most NN, and each power ξj\xi^{j} is a finite combination of Hermite polynomials. The hnh_n are real, so hn,f=hnf\avg{h_n,f}=\int h_nf, and the hypothesis therefore says

f(ξ)ξjeξ2/2dξ  =  0for every j=0,1,2, \int_{-\infty}^{\infty} f(\xi)\,\xi^{j}\,\ee^{-\xi^{2}/2}\,\dd\xi \;=\; 0 \qquad\text{for every } j=0,1,2,\dots (4.5.25)

Write q=feξ2/2q=f\,\ee^{-\xi^{2}/2}, which is in L2L^{2} because qf\abs q\le\abs f and in L1L^{1} by Cauchy–Schwarz, since feξ2/2feξ2/2\int\abs f\ee^{-\xi^{2}/2}\le\norm f\,\norm{\ee^{-\xi^{2}/2}} and both factors are finite. Our goal is to show its Fourier transform vanishes identically, so write the transform out and replace eikξ\ee^{-\ii k\xi} by its power series, which converges for every ξ\xi:

2π  q~(k)  =  q(ξ)eikξdξ  =  j0(ik)jj!q(ξ)ξjdξ  =  0, \sqrt{2\pi}\;\tilde q(k) \;=\; \int q(\xi)\,\ee^{-\ii k\xi}\,\dd\xi \;=\; \sum_{j\ge0}\frac{(-\ii k)^{j}}{j!}\int q(\xi)\,\xi^{j}\,\dd\xi \;=\; 0, (4.5.26)

where every integral in the sum vanishes by (4.5.25). The only step that needs an argument is the interchange of the sum with the integral, and it is dominated convergence, which is exactly the theorem Chapter 4.3 §4.3 proved and exactly the reason this proof belongs after that chapter rather than before it. The partial sums of the exponential series satisfy jJ(ikξ)j/j!ekξ\abs{\sum_{j\le J}(-\ii k\xi)^{j}/j!}\le\ee^{\abs{k\xi}} for every JJ, so every partial sum of the integrand is dominated by

Φ(ξ)  =  f(ξ)eξ2/2ekξ,Φ    f(eξ2+2kξdξ)1/2  <  , \Phi(\xi) \;=\; \abs{f(\xi)}\,\ee^{-\xi^{2}/2}\,\ee^{\abs{k\xi}}, \qquad \int\Phi \;\le\; \norm f\left(\int\ee^{-\xi^{2}+2\abs{k\xi}}\dd\xi\right)^{1/2} \;\lt\;\infty, (4.5.27)

the bound again being Cauchy–Schwarz, and the remaining integral being finite because completing the square turns it into 2ek20e(ξk)2dξ2\ee^{k^{2}}\int_0^{\infty}\ee^{-(\xi-\abs k)^{2}}\dd\xi. So a single integrable function dominates the whole sequence of partial sums, for each fixed kk, and the interchange is licensed.

The rest is one line of Chapter 0.9. Plancherel says the transform preserves the norm, so q=q~=0\norm q=\norm{\tilde q}=0, which makes qq the zero vector, so q=0q=0 almost everywhere. Since eξ2/2\ee^{-\xi^{2}/2} is never zero, f=0f=0 almost everywhere, which is to say ff is the zero vector of L2L^{2}. Statement (d) holds, and Chapter 4.3 §7.3's equivalence upgrades it to statement (b): every ff in the space equals nhn,fhn\sum_n\avg{h_n,f}h_n, with the series converging in norm. The Hermite functions are an orthonormal basis of L2(R)L^{2}(\R). \blacksquare

Chapter 4.3's closing brick said this proof would be statement (d) of its §7.3 run on an integral that only its §4.3 licenses, and that is precisely what happened. Nothing else was used, and in particular nothing about analytic functions of a complex variable, which this book has not built and does not build until Chapter 5.4.

5.5 · Verification 3, completed

With a genuine basis in hand the third verification is immediate. Send each state to its list of coefficients against the Hermite functions. That is Chapter 0.5's move, made with an infinite list:

U^ψ  =  (c0,c1,c2,),cn=hn,ψ,n0cn2=ψ2. \hat U\psi \;=\; \big(c_0,c_1,c_2,\dots\big), \qquad c_n=\avg{h_n,\psi}, \qquad \sum_{n\ge0}\abs{c_n}^{2}=\norm\psi^{2}. (4.5.28)

That map is unitary onto 2\ell^{2}, and both halves of the claim are already proved. It preserves the norm by Parseval, which is statement (c) of Chapter 4.3 §7.3, and it hits every square-summable list, which is that chapter's §7.2. Now identify the measure space. A sequence space is a space of functions on the set {0,1,2,}\{0,1,2,\dots\}, and the sum in (4.5.28) is the integral of c2\abs{c}^{2} against the measure that gives each point weight one. So 2\ell^{2} is not merely like an L2L^{2} space, it is one:

X={0,1,2,},μ=counting measure,g(n)=(n+12)ω. X=\{0,1,2,\dots\}, \qquad \mu=\text{counting measure}, \qquad g(n)=\left(n+\half\right)\hbar\omega. (4.5.29)

Multiplication by gg is what H^osc\hat H_{\text{osc}} becomes, by (4.5.22) applied term by term, and the domain condition of §3.2 reads n(n+12)22ω2cn2<\sum_n(n+\half)^{2}\hbar^{2}\omega^{2}\abs{c_n}^{2}\lt\infty, with the constants written out even though they cannot affect whether the sum converges. Every point of XX has measure one, hence positive measure, so §3.4 makes every value of gg an eigenvalue and there is nothing else in the essential range. The spectrum is pure point, and it is the ladder.

One honest remark about what has been fixed and by what. The differential expression in (4.5.17) is a formula, and Chapter 4.4 §7 showed at length that a formula does not determine an operator. What determines this one is (4.5.29): the operator is multiplication by a real function in the U^\hat U picture, which makes it self-adjoint by construction, and its action agrees with the formula on every hnh_n and hence on every finite combination of them. That is the operator the rest of Part IV means by H^osc\hat H_{\text{osc}}, and saying so here saves an argument in Chapter 4.8.

5.6 · The three verifications, collected

Put the three side by side, because the comparison is the point of having done them.

  • x^\hat x. XX is the line, μ\mu is Lebesgue measure, U^\hat U is the identity, and g(x)=xg(x)=x. Spectrum R\R, purely continuous, no eigenvectors.
  • p^\hat p. XX is the line, μ\mu is Lebesgue measure, U^\hat U is the Fourier transform of Chapter 0.9, and g(k)=kg(k)=\hbar k. Spectrum R\R, purely continuous, no eigenvectors.
  • H^osc\hat H_{\text{osc}}. XX is the non-negative integers, μ\mu is counting measure, U^\hat U is expansion in Hermite functions, and g(n)=(n+12)ωg(n)=(n+\half)\hbar\omega. Spectrum a ladder, purely point, an orthonormal basis of eigenvectors.

Now say what that buys, because it is the reason the mark in §3.2 is defensible and it must not be passed over. You are holding a theorem this book quotes rather than proves, and you have personally checked it in every case this book applies it to. Then say exactly what that covers, since it is easy to overstate. The theorem is verified here for these three operators and for nothing else. Every later use in Part IV is one of three kinds. It is a function of one of the three, built by §6.6's functional calculus, which never leaves the multiplication picture. Or it is an operator whose own chapter exhibits a complete orthonormal family of eigenfunctions, supplying the triple directly in the way §5.5 supplied it for the oscillator. Or it is p^2/2m+V(x^)\hat p^{2}/2m+V(\hat x), whose self-adjointness is a separate question about VV that Chapter 4.6 §4.4 answers case by case.

That third kind is the one to be careful about, because there is a shortcut here and it does not work. Unitary equivalence to a multiplication operator is not inherited by sums. For unbounded self-adjoint A^\hat A and B^\hat B the sum on dom(A^)dom(B^)\operatorname{dom}(\hat A)\cap\operatorname{dom}(\hat B) need not be self-adjoint, need not be essentially self-adjoint, and need not even be densely defined. So "built out of operators that have been verified" certifies nothing at all. Adding V(x^)V(\hat x) to p^2/2m\hat p^{2}/2m produces a fourth operator whose admissibility has to be established rather than inherited, and §8.5 below is the standing warning against assuming otherwise. That is the standard the book set for itself when it decided what a mark means, and here it is met rather than approximated.

The three also make one structural point that a single verification could not. The operators differ in their measure and in nothing else that matters: continuous spectrum and discrete spectrum are Lebesgue measure and counting measure, seen through the same statement. Chapter 0.5 never had to mention a measure because a finite list carries only one, and that silence is exactly what had to be broken to get here.

One picture, three measures. Each panel plots the multiplying function gg of §3.2 over its own measure space, with the resulting spectrum marked as a rail on the left. Top: x^\hat x, where XX is the line with Lebesgue measure and g(x)=xg(x)=x. Middle: p^\hat p, where XX is the line with Lebesgue measure and g(k)=kg(k)=\hbar k, drawn in units =1\hbar=1. Bottom: H^osc\hat H_{\text{osc}}, where XX is the non-negative integers with counting measure and g(n)=(n+12)ωg(n)=(n+\half)\hbar\omega, drawn in units ω=1\hbar\omega=1. The first two graphs are the same straight line, and the third is that line lifted by half a rung, which is the ground-state energy. What differs between the three panels is the set the line is drawn over, and that difference is the whole difference between a continuous spectrum and a discrete one. The top two rails are solid because every neighbourhood of every real number has positive length. The bottom one is a row of dots, because each integer carries weight one and the gaps between them carry nothing.

5.7 · The numerical confirmation

Parseval in the Hermite basis is the statement to check, since the whole of §5.5 rests on it, and there is a test function whose coefficients come out in closed form. Take the ground state displaced sideways by aa, meaning ψa(ξ)=π1/4e(ξa)2/2\psi_a(\xi)=\pi^{-1/4}\ee^{-(\xi-a)^{2}/2}, and pair it against (4.5.19). Completing the square in the exponent gives G(ξ,t)eξ2/2ψadξ=π1/4ea2/4eat\int G(\xi,t)\ee^{-\xi^{2}/2}\psi_a\,\dd\xi=\pi^{1/4}\ee^{-a^{2}/4}\ee^{at}. Matching powers of tt against the Hermite series then produces every coefficient at once:

cn  =  hn,ψa  =  ea2/4(a/2)nn!,socn2  =  ea2/2(a2/2)nn!. c_n \;=\; \avg{h_n,\psi_a} \;=\; \ee^{-a^{2}/4}\,\frac{(a/\sqrt2)^{n}}{\sqrt{n!}}, \qquad\text{so}\qquad \abs{c_n}^{2} \;=\; \ee^{-a^{2}/2}\,\frac{(a^{2}/2)^{n}}{n!}. (4.5.30)

Those squared coefficients are the Poisson probabilities with mean a2/2a^{2}/2, and the mathematics is identical rather than analogous: the same expression, the same normalisation, the same parameter in the same place. So Parseval's claim that they sum to one is, in this instance, the statement that a Poisson distribution is a distribution. That is a strong check, because it is a closed form to compare arithmetic against rather than a number produced by the same arithmetic.

Run it at a=2a=2, so the mean is 22. Computing the coefficients numerically on a grid of 40964096 points across ξ12\abs\xi\le12 reproduces c0,,c5=0.3678794c_0,\dots,c_5=0.3678794, 0.52026010.5202601, 0.52026010.5202601, 0.42479060.4247906, 0.30037230.3003723, 0.18997210.1899721. Those match (4.5.30) to sixteen digits. The truncated sum is

n<20cn2  =  1.00000000,with an exact remainder n20cn2=6.44×1014. \sum_{n\lt20}\abs{c_n}^{2} \;=\; 1.00000000, \qquad\text{with an exact remainder } \sum_{n\ge20}\abs{c_n}^{2}=6.44\times10^{-14}. (4.5.31)

That remainder is the exact Poisson tail summed in arithmetic of higher precision, not the difference of the printed sum from one, which is contaminated by rounding in the twelfth figure. The orthonormality of (4.5.24) can be checked on the same grid by forming the whole Gram matrix hm,hn\avg{h_m,h_n} for m,n<40m,n\lt40, and its largest departure from the identity is 4.0×10154.0\times10^{-15}. Read that last figure for what it is. It is the floor of double-precision arithmetic rather than anything about the Hermite functions, and it moves by a factor of ten either way with the quadrature rule and the exact placement of the grid points. What is worth knowing is how the departure grows once the grid is too narrow, and that part is truncation and is implementation-independent: at half-width 1212 it is still at the rounding floor, at 1111 it is 6.34×10116.34\times10^{-11}, at 1010 it is 8.16×1068.16\times10^{-6} and at 99 it is 2.17×1022.17\times10^{-2}. The reason is visible in (4.5.22). The state h39h_{39} has energy 39.5ω39.5\hbar\omega, so its classical turning point sits at ξ=79=8.89\xi=\sqrt{79}=8.89, and a grid that stops at 99 is cutting the function off where it is still large.

In plain terms 4.5.5

An oscillating particle is the one quantity in this chapter that behaves the way the finite theory said everything would. It has states of definite energy, they are honest members of the space, and their energies form an evenly spaced ladder starting not at zero but half a rung up. All of that came out of a single compact function whose expansion generates the whole family, with no physical input beyond the shape of the energy.

The claim that costs something is not that these states exist but that there are enough of them. Being mutually perpendicular is easy; being numerous enough to describe every state is the hard half, and the previous chapter was explicit that these are different claims and that the second is what everyone actually uses. The proof runs by contradiction. Assume some state is perpendicular to every member of the family, deduce that it is perpendicular to every power of position damped by a Gaussian, expand a wave in powers, and conclude that the transform of that state vanishes at every wavelength. A state with no content at any wavelength has no size, so it was nothing to begin with.

The step that has to be justified is exchanging an infinite sum with an integral, and the licence is the convergence theorem the previous chapter proved for exactly this purpose. That is worth registering as the shape of the whole enterprise: a result quoted in Part 0 and used freely ever since is now earned, and what earns it is a theorem that was itself built rather than borrowed.

a natural place to stop  ·  the checking is finished; next comes the projection a continuous observable needs before the Born rule can give it a probability at all

6 · The projection-valued measure, and the integral that replaces the sum

Chapter 0.5 wrote its theorem three ways, and only two of them have been dealt with. The eigenvector form is gone, the matrix form A=UDUA=UDU^{\dagger} has become §3.2, and the third form, the weighted sum of projections, is what this section builds. It gets a section of its own because it is the form the physics is written in: the Born rule is a statement about projections, and until a projection exists for a continuous observable there is no probability to compute. By the end you will have A^=λdP(λ)\hat A=\int\lambda\,\dd P(\lambda), the Born rule for a continuous variable, and the recipe for building functions of an operator that §9 needs.

6.1 · Projections, read off the multiplication form

Multiplication operators come with an obvious supply of projections, because multiplying by the indicator function of a set is a projection: doing it twice is the same as doing it once, and it is real, so it is its own adjoint. Transport that supply back through U^\hat U. For each Borel subset EE of the real line, take the part of the measure space where gg lands inside EE. From here ξ\xi is the generic point of XX again, as it was in §3.2, and §5's oscillator coordinate is finished with. Define

P(E)  =  U^1M1g1(E)U^,g1(E)={ξX: g(ξ)E}. P(E) \;=\; \hat U^{-1}\,M_{\mathbf 1_{g^{-1}(E)}}\,\hat U, \qquad g^{-1}(E)=\{\xi\in X:\ g(\xi)\in E\}. (4.5.32)

Each P(E)P(E) is an orthogonal projection on HH, since conjugating by a unitary preserves both P2=PP^{2}=P and P=PP^{\dagger}=P. Each is defined on the whole space and bounded by one, because multiplying by an indicator never increases a modulus. The family EP(E)E\mapsto P(E) is the projection-valued measure of A^\hat A. It deserves the word measure because it satisfies the axioms Chapter 4.3 §2.1 laid down, with projections in place of numbers:

P()=0,P(R)=I^,P(E)P(F)=P(EF),P(jEj)ψ=jP(Ej)ψ, \begin{gathered} P(\varnothing)=0, \qquad P(\R)=\hat I, \qquad P(E)P(F)=P(E\cap F), \\[6pt] P\Big(\bigcup_j E_j\Big)\psi=\sum_j P(E_j)\psi, \end{gathered} (4.5.33)

the last for any countable family of disjoint sets, with the series converging in norm. All four are the corresponding statements about indicator functions, carried across the unitary. The first three are pointwise identities, and the fourth is dominated convergence one more time: the partial sums of j1Ej\sum_j\mathbf 1_{E_j} converge pointwise to 1Ej\mathbf 1_{\bigcup E_j} and are dominated by it, so the images converge in L2L^{2} by Chapter 4.3 §4.3.

The definite article in the projection-valued measure is doing physics, and it has to be paid for, which is why §3.2's mark has two clauses rather than one. The definition (4.5.32) runs through a particular (X,μ,U^,g)(X,\mu,\hat U,g), and §4.2 has already displayed x^\hat x through two of them. The mark's second clause is what says the resulting PP comes out the same either way. Without it, the probability §6.5 assigns to an interval would depend on whether the calculation was done on the line or on (π/2,π/2)(-\pi/2,\pi/2) through §4.2's map, and there is no reading of the Born rule under which that could be allowed. Uniqueness of the spectral measure is a real theorem rather than a formality, this book quotes it rather than proving it, and §3.2 is where the quotation is recorded.

Two special values are worth noticing at once. Taking EE to be a single point {λ}\{\lambda\} gives the projection onto the eigenspace at λ\lambda, by §3.4, so P({λ})=0P(\{\lambda\})=0 exactly when λ\lambda is not an eigenvalue. And taking EE to be the complement of the spectrum gives zero, since gg never lands outside its essential range on a set of positive measure. So the whole of PP lives on σ(A^)\sigma(\hat A), which is the sense in which the spectrum is the complete list of readings.

6.2 · The measure a state induces

A projection-valued measure is not yet a probability, and the step that makes one is to pair it with a state. Fix a unit vector ψ\psi and define, for each Borel set EE,

μψ(E)  =  ψ,P(E)ψ  =  P(E)ψ2  =  g1(E)Ψ2dμ,Ψ=U^ψ, \mu_\psi(E) \;=\; \avg{\psi,\,P(E)\psi} \;=\; \norm{P(E)\psi}^{2} \;=\; \int_{g^{-1}(E)}\abs{\Psi}^{2}\,\dd\mu, \qquad \Psi=\hat U\psi, (4.5.34)

where the middle equality is P(E)2=P(E)P(E)^{2}=P(E) together with P(E)=P(E)P(E)^{\dagger}=P(E), and the right one is (4.5.32) written out. Read the right-hand expression and μψ\mu_\psi is visibly an ordinary measure in Chapter 4.3's sense: it is non-negative, it is countably additive because the integral is, its total mass is Ψ2dμ=ψ2=1\int\abs\Psi^{2}\dd\mu=\norm\psi^{2}=1, and by the end of §6.1 it assigns zero to everything outside the spectrum. So each state turns the observable into a probability distribution on the real line, concentrated on the possible readings, and that distribution is an object Chapter 4.3 knows how to integrate against.

6.3 · The integral that replaces the sum

Now compute the expectation value of A^\hat A in the state ψ\psi and watch μψ\mu_\psi appear. Work in the multiplication picture, where the operator is nothing but a factor:

ψ,A^ψ  =  XΨgΨdμ  =  Xg(ξ)Ψ(ξ)2dμ(ξ)  =  Rλ  dμψ(λ). \avg{\psi,\hat A\psi} \;=\; \int_X \overline{\Psi}\,g\,\Psi\,\dd\mu \;=\; \int_X g(\xi)\,\abs{\Psi(\xi)}^{2}\,\dd\mu(\xi) \;=\; \int_{\R}\lambda\;\dd\mu_\psi(\lambda). (4.5.35)

The last equality is the only one doing work, and it is the standard argument for pushing a measure forward along a map, run entirely inside Chapter 4.3. For an indicator 1E\mathbf 1_E the two sides agree by the definition (4.5.34), so they agree for simple functions by linearity. They then agree for every non-negative measurable function by monotone convergence, which is Chapter 4.3 §4.1, and for a general one by splitting into positive and negative parts as Chapter 4.3 §3 does. Nothing beyond that chapter is used.

That identity is what licenses writing the operator itself as an integral. Define λdP(λ)\int\lambda\,\dd P(\lambda) to be the operator whose expectation value in every state is the right-hand side of (4.5.35), on the set of states for which the second moment is finite. That does pick out one operator rather than a family, and here is the reason. On a complex space the expectation values determine the operator, by the move Chapter 4.2 §7.2 made for a unitary: expand u+v,A^(u+v)\avg{u+v,\hat A(u+v)}, then run the same line with vv replaced by iv\ii v, and every u,A^v\avg{u,\hat Av} is recovered. Then A^\hat A is that operator, and the two labels below say what the integral runs over and what is being integrated against:

A^  =  σ(A^)  over every reading the instrument can give  λ  dP(λ)  the projection onto the states that give it   \hat A \;=\; \ann{\int_{\sigma(\hat A)}}{over every reading the instrument can give} \lambda \; \ann{\dd P(\lambda)}{the projection onto the states that give it} (4.5.36)

That is Chapter 0.5's third form, and it now exists for an operator with no eigenvectors at all. Nothing on the right-hand side mentions one. The sum over a finite list of eigenvalues has become an integral over the spectrum, and the projections onto eigenspaces have become the P(E)P(E) of §6.1, which exist whether or not any state sits behind a value. That is the promise Chapter 0.5 §6.4 made about this form, and §6.4 below collects it against the words that chapter used.

The domain is not an extra stipulation, and that has to be checked rather than asserted, because Chapter 4.4 spent a chapter establishing that domains are forced. The condition λ2dμψ<\int\lambda^{2}\dd\mu_\psi\lt\infty is, by the same pushforward argument, Xg2Ψ2dμ<\int_X g^{2}\abs\Psi^{2}\dd\mu\lt\infty, which says gΨL2g\Psi\in L^{2}, which is exactly the domain condition §3.2 attached to the multiplication form. So the integral form carries its own domain and it is the same one.

6.4 · Chapter 0.5's sum, recovered

Take a finite-dimensional space and watch (4.5.36) collapse. There σ(A)\sigma(A) is a finite set of eigenvalues by (4.5.4), and each P({λk})P(\{\lambda_k\}) is the projection PkP_k onto the eigenspace EλkE_{\lambda_k} of Chapter 0.5 §6.4. The measure μψ\mu_\psi is then a sum of atoms, μψ=kPkψ2δλk\mu_\psi=\sum_k\norm{P_k\psi}^{2}\delta_{\lambda_k}, so an integral against it is a finite sum, and (4.5.36) reads

A  =  kλkPk,PkPl=δklPk,kPk=I^, A \;=\; \sum_k \lambda_k P_k, \qquad P_kP_l=\delta_{kl}P_k, \qquad \sum_k P_k=\hat I, (4.5.37)

which is Chapter 0.5 §6.4 verbatim, with its two side conditions arriving as the third and second identities of (4.5.33). That chapter said of exactly this formula that "this is the basis-independent form, and it is the one that survives to infinite dimensions in Chapter 4.5". It survives. The sense in which it survives is that the sum over a finite list of atoms has become an integral over a measure that may have no atoms at all.

The second promise made in the same place is longer and it is worth reading back before answering it. Chapter 0.5's closing warning box said that in infinite dimensions "the spectrum becomes continuous, the sum kλkPk\sum_k\lambda_kP_k becomes an integral λdP(λ)\int\lambda\,\dd P(\lambda) over projection-valued measures, and completeness must be re-proved rather than assumed. Chapter 4.5 pays this bill in full." Take the three clauses in order. The spectrum became continuous in §2.5 and §2.6, for position and for momentum, and the widened definition that made the word available is §2.2. The sum became the integral in (4.5.36), over the projection-valued measure of §6.1. And completeness was re-proved rather than assumed, three separate times: completeness of the space in Chapter 4.3 §6.2, completeness of the Fourier basis in Chapter 4.3 §8.3, and completeness of the Hermite basis in §5.4 above. That is the bill, paid clause by clause.

6.5 · The Born rule for a continuous variable

Chapter 4.2's third postulate said that the probability of the outcome λk\lambda_k in the state ψ\psi is Pkψ2\norm{P_k\ket\psi}^{2}, and it was stated for an observable with eigenvalues and eigenspaces. Position has neither, so the postulate as written says nothing about the one measurement every introductory treatment performs first. The repair needs no new assumption, only the projection that §6.1 supplies: the probability that a measurement of A^\hat A returns a value in the set EE is P(E)ψ2\norm{P(E)\psi}^{2}, which is μψ(E)\mu_\psi(E). Setting E={λk}E=\{\lambda_k\} gives back the original statement word for word, so this is the same postulate with a wider notion of outcome and not a second one.

Now run it on position, where (4.5.13) makes U^\hat U the identity and P([a,b])P([a,b]) multiplication by 1[a,b]\mathbf 1_{[a,b]}:

Prob(axb)  =  P([a,b])ψ2  =  abψ(x)2dx. \text{Prob}\big(a\le x\le b\big) \;=\; \norm{P([a,b])\psi}^{2} \;=\; \int_a^{b}\abs{\psi(x)}^{2}\,\dd x. (4.5.38)

That closes a loop the reader will have felt open since Chapter 4.3. That chapter's §5.3 warned that a vector of L2L^{2} has no value at any point, so ψ(x0)2\abs{\psi(x_0)}^{2} is not a probability of anything, and it said the Born rule for a continuous variable would have to be stated over a region instead. Here is why it is stated that way, and it is not a convention adopted for safety. The projection P({x0})P(\{x_0\}) onto a single point is the zero operator, because a point has measure zero, so the probability of finding the particle exactly at x0x_0 is zero for every state and every point. The integral over an interval is the only quantity the theory produces, and ψ2\abs\psi^{2} is a density rather than a probability.

Familiar ground — a continuous outcome has no atoms, and where the parallel stops

You already work with a distribution whose probability at every single point is zero. A survival time is one. The probability that a patient's time to progression is exactly 14.00014.000\ldots months is zero, and it is zero for every value, which does not make the distribution empty or the measurement meaningless. Every quantity anyone reports is an integral of the density over a region: the probability of progressing within a year, the median, the hazard over an interval, the area under a curve between two times.

The mathematics here is identical, not analogous. The measure μψ\mu_\psi of (4.5.34) is a probability measure on the real line in exactly Chapter 4.3's sense, ψ2\abs\psi^{2} is its density with respect to Lebesgue measure, and (4.5.38) is the cumulative probability over an interval. A continuous spectrum is a distribution with no atoms, and a point spectrum is a distribution that is all atoms. An observable can have both at once, which is a mixed distribution, and Worked example 2 builds one.

Two places the parallel stops, and both matter later. First, μψ\mu_\psi depends on which observable is being measured, and the same state gives different distributions for position and for momentum. Whether there is a joint distribution underneath them from which both could be recovered is a real question, and it is not settled here. Chapter 4.9 turns the incompatibility into the uncertainty relation, which bounds the two spreads without deciding that question; Chapter 4.11 shows the failure to be simultaneously sharp is structural rather than a matter of ignorance; and Chapter 4.20 settles it with an experiment. Second, the object being integrated in (4.5.36) is a family of projections rather than a family of numbers, and projections for different observables do not commute. There is no survival analysis in which the events fail to commute.

6.6 · Functions of an operator

One more construction comes free, and §9 needs it. Chapter 0.5 §7 defined f(A)f(A) in finite dimensions by diagonalising and applying ff to each diagonal entry, which in the present language is applying ff to the multiplying function. So define

f(A^)  =  U^1MfgU^  =  σ(A^)f(λ)dP(λ), f(\hat A) \;=\; \hat U^{-1}\,M_{f\circ g}\,\hat U \;=\; \int_{\sigma(\hat A)} f(\lambda)\,\dd P(\lambda), (4.5.39)

for any measurable ff, with the second expression meaning what §6.3 made it mean. If ff is bounded on the spectrum then f(A^)f(\hat A) is defined on the whole space with f(A^)supλσ(A^)f(λ)\norm{f(\hat A)}\le\sup_{\lambda\in\sigma(\hat A)}\abs{f(\lambda)}, since multiplying by a bounded function does exactly that to a norm. Three properties follow by composing functions pointwise, and they are the ones every later use needs: (fh)(A^)=f(A^)h(A^)(fh)(\hat A)=f(\hat A)h(\hat A), the map ff(A^)f\mapsto f(\hat A) is linear, and f(A^)=f(A^)\overline{f}(\hat A)=f(\hat A)^{\dagger}. In particular ff real makes f(A^)f(\hat A) self-adjoint, and f=1\abs f=1 makes f(A^)f(\hat A) unitary, which is the observation §9 is built on.

In plain terms 4.5.6

Once a quantity has become multiplication by a function, there is a natural way to ask how much of a given state sits in any range of values. Restrict the function to that range by multiplying by something that equals one inside it and zero outside, and what you have is a device that keeps the part of the state belonging to those readings and discards the rest. That device is a projection in the old sense, and there is one for every range rather than one for every value.

Pairing that family with a particular state produces an ordinary probability distribution over the possible readings, and from there everything the finite theory did carries over with sums replaced by integrals. The average reading is the average of the values against that distribution. The list of values times their projections becomes an integral of the same shape. And the probability rule that was stated for individual outcomes becomes a rule about ranges of outcomes, which is what it always had to be for a continuous quantity, since no continuous quantity has a positive probability of taking any one exact value.

This is where the promise made in the toolkit, Chapter 0.5, is finally settled. That chapter said the projection form of its central theorem was the one that would survive, that the sum would become an integral, and that the claim about having enough directions would have to be re-earned rather than assumed. All three have now happened, and the third happened three separate times, once for the space itself and once for each of the two families of waves the book expands in.

7 · What x\ket x and p\ket p actually mean

Every physics text writes x\ket x and p\ket p, inserts xxdx=I^\int\ket x\bra x\dd x=\hat I between two operators, and writes xp=eipx//2π\avg{x|p}=\ee^{\ii px/\hbar}/\sqrt{2\pi\hbar}. Chapter 4.3 §8.5 explicitly declined to say what those symbols are and sent the question here. This section answers it. The answer is a procedure rather than a definition: put the system in a box, or on a lattice, do the computation where every object is a genuine vector, and then take a limit in which the spacing is manufactured into the sums. The procedure always works, it is what every one of those equations abbreviates, and by the end of the section you will also know exactly which of the comfortable statements about x\ket x and p\ket p are false.

7.1 · Neither object is a vector, and both failures are already proved

Start by fixing what is wrong, since both halves have been established elsewhere and it is worth having them together. A state of definite momentum would be eipx/\ee^{\ii px/\hbar}, and Chapter 0.9 §2.1 pointed out that eipx/2dx\int\abs{\ee^{\ii px/\hbar}}^{2}\dd x diverges, so it is not in L2(R)L^{2}(\R). A state of definite position would be concentrated at a point, and Chapter 4.3 §5.3 showed that such a thing is the zero vector, because a point is a null set. So one object is too big for the space and the other is too small, and §1.2 turned both failures into the statement that neither operator has an eigenvector.

What makes this urgent rather than pedantic is that the notation works. Calculations performed with x\ket x and p\ket p give right answers, reliably, and have done since the 1920s. Something is being abbreviated, and the job here is to say what.

7.2 · Momentum, in a box: box normalisation worked in full

The move is to change the problem to one where the object exists, and Chapter 4.4 §5.4 has already built the replacement. It goes by the name box normalisation, and this subsection works it in full so that you have a procedure rather than a slogan. One notice about the label first. Everything from here to §7.5 is written in the wavenumber kk rather than the momentum pp, because the box's modes are labelled by kk and the constants only get in the way. Section 7.6 puts the \hbar back and turns k\ket k into the p\ket p this section was opened with.

Put the particle on the interval [L/2,L/2][-L/2,L/2] with the periodic condition ψ(L/2)=ψ(L/2)\psi(L/2)=\psi(-L/2), which is that section's p^θ\hat p_\theta at θ=0\theta=0 with the interval recentred, and which that section proved self-adjoint. This is not the operator §3.2 used as its counterexample. There the wavefunction was pinned to zero at both ends, and the result was symmetric with no eigenvalues at all. Here the two ends are joined rather than pinned, which is the difference Chapter 4.4 §5.4 turned into a self-adjointness question and settled. So on this interval momentum does have eigenvectors, and they are the Fourier modes:

un(x)=eiknxL,kn=2πnL,p^un=knun,nZ. u_n(x)=\frac{\ee^{\ii k_nx}}{\sqrt L}, \qquad k_n=\frac{2\pi n}{L}, \qquad \hat p\,u_n=\hbar k_n\,u_n, \qquad n\in\mathbb{Z}. (4.5.40)

Everything about this situation is legitimate and nothing is being smuggled. The unu_n satisfy the periodic condition, they are orthonormal by direct integration, and Chapter 4.3 §8.3 proved they are a genuine orthonormal basis. That chapter proved it on [π,π][-\pi,\pi], and the interval here is a unitary rescaling of that one, which carries a basis to a basis. So every statement of Chapter 4.3 §7.3 applies to them. In particular any state expands as ψ=ncnun\psi=\sum_nc_nu_n with cn=un,ψc_n=\avg{u_n,\psi}, and ncn2=ψ2\sum_n\abs{c_n}^{2}=\norm\psi^{2}.

Now prepare for the limit, and the preparation is Chapter 0.9 §2.1's move made once more. The mode spacing is Δk=2π/L\Delta k=2\pi/L, and a sum over nn becomes an integral over kk only if every term carries a factor of Δk\Delta k. So rescale the basis vectors by whatever makes that happen, which means dividing by Δk\sqrt{\Delta k}:

kn    unΔk  =  L2π  un,sokmkn  =  δmnΔk. \ket{k_n} \;\equiv\; \frac{u_n}{\sqrt{\Delta k}} \;=\; \sqrt{\frac{L}{2\pi}}\;u_n, \qquad\text{so}\qquad \avg{k_m|k_n} \;=\; \frac{\delta_{mn}}{\Delta k}. (4.5.41)

These rescaled vectors are still perfectly good vectors of the box's Hilbert space. Their lengths grow as LL grows, which is the first sign of what goes wrong in the limit. Watch what the rescaling does to the three statements that matter. The coefficient becomes the Fourier transform of Chapter 0.9, exactly and with no approximation:

knψ  =  L2π  1LL/2L/2ψ(x)eiknxdx=  12πL/2L/2ψ(x)eiknxdx  L  ψ~(kn). \begin{aligned} \avg{k_n|\psi} \;&=\; \sqrt{\frac{L}{2\pi}}\;\frac{1}{\sqrt L}\int_{-L/2}^{L/2}\psi(x)\,\ee^{-\ii k_nx}\,\dd x \\[4pt] &=\; \frac{1}{\sqrt{2\pi}}\int_{-L/2}^{L/2}\psi(x)\,\ee^{-\ii k_nx}\,\dd x \;\xrightarrow[L\to\infty]{}\; \tilde\psi(k_n). \end{aligned} (4.5.42)

The resolution of the identity becomes an integral. Chapter 4.3 §7.3's expansion statement, written in the Dirac notation Chapter 0.5 §2.2 installed, is nunun=I^\sum_n\ket{u_n}\bra{u_n}=\hat I. Each unun\ket{u_n}\bra{u_n} is Δkknkn\Delta k\ket{k_n}\bra{k_n} by (4.5.41), so the sum is already a Riemann sum:

I^  =  nΔk  knkn  L  kkdk. \hat I \;=\; \sum_n \Delta k\;\ket{k_n}\bra{k_n} \;\xrightarrow[L\to\infty]{}\; \int_{-\infty}^{\infty}\ket k\bra k\,\dd k. (4.5.43)

And Parseval becomes Plancherel, since ncn2=nΔkknψ2\sum_n\abs{c_n}^{2}=\sum_n\Delta k\,\abs{\avg{k_n|\psi}}^{2} is a Riemann sum for ψ~2dk\int\abs{\tilde\psi}^{2}\dd k. Three familiar identities, and every one of them is exactly true at finite LL before any limit is taken.

Now be exact about the three arrows themselves, because §7.3 sets a standard just below and this subsection has to meet it too. What is proved above is everything at finite LL: the unu_n are an orthonormal basis, (4.5.42) is an identity, and the resolution of the identity holds in the box. The arrows are not proved here, and it is worth saying why not. Each of them compares objects living in different spaces. The space L2[L/2,L/2]L^{2}[-L/2,L/2] changes as LL grows, ψ\psi lives on the whole line, and the labels knk_n move as well. No topology has been named in which such a comparison is a limit, and none is named here.

What makes the arrows safe to write is that their destinations are already known from the other side. Chapter 0.9 §2.1 built ψ~\tilde\psi as exactly this Riemann-sum limit, and Chapter 0.9 §2.3 proved Plancherel outright, without a box anywhere in the argument. So the box does not earn those identities and this subsection does not claim it does. What the box supplies is the thing that was actually wanted, which is a finite, legitimate statement that each of them abbreviates. Section 7.3's matching claim is different in kind and is proved there in four lines, because its objects all sit in one fixed space and the limit is an ordinary one.

The orthogonality relation is the one that needs care, because δmn/Δk\delta_{mn}/\Delta k has no limit as a number. It has a limit as an instruction, which is what Chapter 0.9 §5.1 insisted a delta always was. Test it the only way it is ever used, against a coefficient function ff, remembering that a sum over modes carries Δk\Delta k:

nΔk  f(kn)kmkn  =  nΔk  f(kn)δmnΔk  =  f(km), \sum_n \Delta k\; f(k_n)\,\avg{k_m|k_n} \;=\; \sum_n \Delta k\; f(k_n)\,\frac{\delta_{mn}}{\Delta k} \;=\; f(k_m), (4.5.44)

and the left-hand side is the Riemann sum for f(k)kkdk\int f(k)\avg{k'|k}\,\dd k. So in the limit the symbol kk\avg{k'|k} is whatever reproduces f(k)f(k') when integrated against ff, and that is precisely Chapter 0.9 §5.1's definition of δ(kk)\delta(k-k'). The equation kk=δ(kk)\avg{k'|k}=\delta(k-k') is therefore true, and true as a statement about what happens under an integral sign rather than as a statement about the value of an inner product.

7.3 · Position, on a lattice

Position needs the same treatment with a different replacement, and the reason a box does nothing for it takes one line. Confining the line to an interval leaves position as multiplication by xx, so its spectrum shrinks to that interval and stays continuous, with no eigenvector anywhere in it. Problem 1(c) runs that calculation. Bounding the space achieves nothing, because what has to be made discrete is the space itself rather than its extent. So chop the line into cells Cj=[ja,(j+1)a)C_j=[ja,(j+1)a) of width aa and take the normalised indicators ej=a1/21Cje_j=a^{-1/2}\mathbf 1_{C_j}, which are orthonormal because the cells are disjoint. They do not span the space, since their closed span is the functions that are constant on every cell, so the missing step is that the projection onto that span tends to the identity as a0a\to0.

That is worth proving rather than waving through, and it takes four lines. For a continuous ψ\psi vanishing outside a bounded set, the cell-average approximation differs from ψ\psi by at most the oscillation of ψ\psi over one cell, which tends to zero uniformly, so it converges in norm. Chapter 4.3 §8.1 proved such functions are dense in L2L^{2}. And each projection has norm one, so for a general ψ\psi choose a continuous ϕ\phi within ϵ\epsilon, note that the projection moves ϕ\phi by less than ϵ\epsilon once aa is small, and the triangle inequality bounds the total by 3ϵ3\epsilon. So the lattice recovers everything in the limit.

Now rescale as before. The spacing is aa rather than Δk\Delta k, so divide by a\sqrt a:

xj    eja  =  1Cja,xixj=δija,Πa  =  jaxjxj  a0  I^  =  xxdx. \begin{gathered} \ket{x_j} \;\equiv\; \frac{e_j}{\sqrt a} \;=\; \frac{\mathbf 1_{C_j}}{a}, \qquad \avg{x_i|x_j}=\frac{\delta_{ij}}{a}, \\[6pt] \Pi_a \;=\; \sum_j a\,\ket{x_j}\bra{x_j} \;\xrightarrow[a\to0]{}\; \hat I \;=\; \int\ket x\bra x\,\dd x. \end{gathered} (4.5.45)

Here Πa\Pi_a is the projection of the previous paragraph, so the middle statement is exactly true at every aa and becomes the identity only in the limit. The rescaled object xj\ket{x_j} is the cell indicator divided by the cell width, which is a bump of unit area narrowing as a0a\to0. That is exactly Chapter 0.9 §5.2's picture of the delta as a limit of narrowing bumps, arriving here from a different direction, and xixj=δij/a\avg{x_i|x_j}=\delta_{ij}/a tends to δ(xy)\delta(x-y) by the same test as (4.5.44). What is being written x\ket x is a bump, not a point, and the width is going to zero afterwards rather than being zero to begin with.

The pairing with a state also behaves. For continuous ψ\psi the number xjψ=a1Cjψ\avg{x_j|\psi}=a^{-1}\int_{C_j}\psi is the average of ψ\psi over the cell, which tends to ψ(x)\psi(x) as the cell shrinks onto xx. So xψ=ψ(x)\avg{x|\psi}=\psi(x) is true in the sense that the cell average converges to the value, and no other sense is available, because Chapter 4.3 §5.3 established that a vector of L2L^{2} has no values at points to converge to. For a general state what actually gets used is the norm statement of the previous paragraph, and that holds without any continuity.

7.4 · The procedure, stated once

Both calculations are the same three steps, and stating them once is the point of having done both, because from here on any manipulation with these symbols can be checked against the list rather than trusted.

  • Replace the continuum by something with a genuine orthonormal basis, carrying a spacing parameter. A periodic box for momentum, which is where the name box normalisation comes from, and a lattice for position. Chapter 4.4 §5.4 supplies the self-adjoint operator in the first case and the indicator functions supply the basis in the second.
  • Do the computation there. Every step is Chapter 4.3 §7.3 applied to an honest orthonormal basis, every object is a vector of a Hilbert space, and nothing needs excusing.
  • Rescale so that each term carries the spacing, then let the spacing go to zero. Sums become integrals as Riemann sums do, a Kronecker delta divided by the spacing becomes a Dirac delta in Chapter 0.9 §5.1's sense, and the answer is the equation everyone writes.

Every equation containing x\ket x or k\ket k is an abbreviation for the output of that procedure. The abbreviation is safe because the procedure always terminates and always gives the same answer, and it is worth using because writing the box out each time would triple the length of every calculation in Part IV. What it is not is a statement about vectors of L2(R)L^{2}(\R), and §7.6 says which comfortable readings are therefore false.

7.5 · The rigged Hilbert space

There is a framework in which x\ket x and k\ket k are objects in their own right rather than abbreviations, and it is stated here, because it explains what kind of thing they are and because you will meet its name elsewhere. The book quotes it and does not build it.

⚑ Quoted, not derived — the rigged Hilbert space, and Gelfand–Maurin

We use, without proof, the following. Let HH be a Hilbert space and let ΦH\Phi\subset H be a dense subspace carrying a finer notion of convergence than HH does. For L2(R)L^{2}(\R) the standard choice of Φ\Phi is the Schwartz functions, the smooth functions all of whose derivatives decay faster than any power. Write Φ\Phi' for the continuous linear functionals on Φ\Phi. Then

Φ    H    Φ, \Phi \;\subset\; H \;\subset\; \Phi',

each inclusion being continuous and dense, and the triple is called a rigged Hilbert space. Let A^\hat A be self-adjoint with A^ΦΦ\hat A\Phi\subseteq\Phi. Then for μψ\mu_\psi-almost every λ\lambda in the spectrum there is a functional FλΦF_\lambda\in\Phi' satisfying Fλ(A^ϕ)=λFλ(ϕ)F_\lambda(\hat A\phi)=\lambda F_\lambda(\phi) for every ϕΦ\phi\in\Phi, and every state of HH can be expanded against that family.

The concrete content, which is the part worth carrying. The functionals are the objects the notation names. Here x\ket x is the map ϕϕ(x)\phi\mapsto\phi(x), which is continuous on Φ\Phi because Schwartz functions are continuous, and k\ket k is the map ϕ12πϕeikxdx\phi\mapsto\frac{1}{\sqrt{2\pi}}\int\phi\,\ee^{-\ii kx}\dd x. Both are perfectly well-defined linear functionals on the smaller space, and neither is a vector of HH, which is the same conclusion §7.1 reached by a different route. An eigenvector that is a functional rather than a vector is the honest description of what these symbols are.

The hypotheses. Φ\Phi is chosen rather than canonical, and different choices give different Φ\Phi'. The operator must map Φ\Phi into itself, which is a real restriction and fails for many operators. And the eigenfunctionals exist for almost every λ\lambda with respect to a spectral measure, not for every λ\lambda, so the theorem does not hand you one at each point of the spectrum. Gelfand and Vilenkin's Generalized Functions, volume 4, is the source.

What this mark does not cover. No calculation in this book uses it. Every manipulation with x\ket x or k\ket k from here to the end of Part IV is an abbreviation of §7.4's procedure, which is proved, and the framework above is quoted so that you know what the abbreviation is abbreviating rather than because anything leans on it.

7.6 · What has been closed, and what has not

Chapter 0.9 §3.2 finished its account of why Fourier analysis works with a confession, and the confession named this chapter: "Chapter 0.5's spectral theorem was proved in finite dimensions, and the argument above applies it to an operator on a function space, where the eigenfunctions are not even in the space. That gap is real. Chapter 4.5 closes it, and the price is a genuinely more careful theory." The gap is now closed, and it is worth saying exactly which gap that was and what the price turned out to be.

The gap was that the manipulations of Chapter 0.9 treated eikx\ee^{\ii kx} as an eigenbasis of the derivative when those functions are not in L2(R)L^{2}(\R). That chapter's §2.1 named the same thing from the other side, saying of a basis labelled by a real number rather than by an integer that "that one change is the difficulty Chapter 4.5 has to work to make legitimate." Section 7.2 supplies the finite statement standing behind each of them: the expansion, the resolution of the identity, the Plancherel identity and the continuum normalisation are each exactly true in a box, and the notation records what the box gives in the limit. The price is the procedure itself, which is that these are limits and not identities between vectors. The word closes applies to the gap Chapter 0.9 named and to nothing wider. The general theory of distributions, in which objects like δ\delta' and the Fourier transform of a tempered distribution are built rather than approached through limits, is Chapter 5.4's, and this section claims none of it.

Chapter 4.3 §8.5's warning also stands, unamended, and it is worth restating because the notation invites forgetting it. Three things remain false. The families {x}\{\ket x\} and {p}\{\ket p\} are not orthonormal bases in Chapter 4.3 §7.3's sense, because that chapter's §7.4 proved every orthonormal basis of this space is countable and the points of a line are not countable. The function eikx\ee^{\ii kx} is not in L2(R)L^{2}(\R) and no limit makes it so. And a state concentrated at a point is the zero vector, not a small vector. What is true is what §7.4 says: a procedure exists, it is finite at every stage, and the notation records its output.

One piece of bookkeeping remains, and Chapter 4.6 needs it. Everything above was in the wavenumber kk, and Part IV writes momentum p=kp=\hbar k with the constants visible. Since kk=δ(kk)\avg{k'|k}=\delta(k-k') and δ(kk)=δ(pp)\delta(k-k')=\hbar\,\delta(p-p'), the momentum kets have to be rescaled by \sqrt\hbar to keep the normalisation. That gives

p  =  kk=p/,pp=δ(pp),xp=eipx/2π. \ket p \;=\; \frac{\ket k}{\sqrt\hbar}\bigg|_{k=p/\hbar}, \qquad \avg{p'|p}=\delta(p-p'), \qquad \avg{x|p}=\frac{\ee^{\ii px/\hbar}}{\sqrt{2\pi\hbar}}. (4.5.46)

Worked example 3 derives the last of those from the box and checks that it reproduces the momentum-space wavefunction ψ^(p)\hat\psi(p) that Chapter 4.3 §8.5 wrote with the \sqrt\hbar in place. That agreement is not automatic and it is the reason the symmetric Fourier convention was chosen in Chapter 0.9 §2.2.

In plain terms 4.5.7

Two symbols appear on nearly every page of every quantum mechanics book, and neither of them names anything in the space of states. One would be a wave of definite wavelength stretching over the whole line with undiminished size, which has infinite total intensity. The other would be a state concentrated at a single point, which in this space is indistinguishable from nothing. Calculations using them nevertheless come out right, every time, and that is the fact needing an explanation.

The explanation is that both symbols abbreviate a procedure rather than denoting an object. Confine the particle to a large ring, or chop the line into small cells, and in either case the troublesome object becomes an ordinary state and the space acquires an ordinary perpendicular family to expand in. Do the calculation there, where nothing is idealised. Then arrange the bookkeeping so that each term carries the spacing between neighbouring labels, and let the spacing shrink. Sums turn into integrals, and the spike that everyone writes appears as the limit of a narrowing bump of fixed area, which is how it was introduced in the first place.

Two things are worth keeping from this. The procedure always works, so the abbreviation is safe, and that is the whole licence for the notation. And the comfortable reading of the notation, in which these families are perpendicular bases like any other, is false rather than imprecise. There are only ever countably many directions in this space, and the points of a line cannot be counted. The general theory in which such objects are built properly rather than reached by a limit belongs to a later part of the book, and nothing here needs it.

a natural place to stop  ·  the notation is legitimised; what follows is a checklist and one theorem

8 · What is now safe to do

The rest of Part IV performs four manipulations constantly, and until this chapter none of them was licensed for an operator without eigenvectors. This section is the licence, written as a checklist with the condition attached to each item, because a condition stated once at the point of use is worth more than a theorem stated once and forgotten. The fourth item is Chapter 4.4's rather than this chapter's, and it is on the list because it is the one that fails most often in practice. The third depends on §9, which has not been proved yet, and the entry says so.

8.1 · Insert a resolution of the identity

Writing I^\hat I as a sum or an integral of projections and slipping it between two operators is the commonest move in the subject. It is legitimate in exactly three situations and there is no fourth. If {en}\{e_n\} is a genuine orthonormal basis, then nenen=I^\sum_n\ket{e_n}\bra{e_n}=\hat I is statement (b) of Chapter 4.3 §7.3, and the word basis is carrying the whole weight: an orthonormal set that is merely orthonormal gives an inequality rather than an identity, by Bessel. If the observable is self-adjoint, then dP(λ)=P(R)=I^\int\dd P(\lambda)=P(\R)=\hat I is (4.5.33), whether or not there are any eigenvectors. And if the family is {x}\{\ket x\} or {k}\{\ket k\}, the identity is (4.5.43) or (4.5.45), meaning the limit of a box or lattice statement in §7.4's sense.

8.2 · Expand a state in eigenstates

This is legitimate exactly when the spectrum is pure point and the eigenvectors are complete, and that is a genuine restriction rather than a formality. It holds for H^osc\hat H_{\text{osc}}, by §5.5, and it will hold for angular momentum in Chapter 4.11, in each case because a complete orthonormal family of eigenvectors gets exhibited. For hydrogen in Chapter 4.13 it holds on the bound-state subspace and not on the whole space, for the reason given at the end of this paragraph. It fails for x^\hat x and for p^\hat p, where §1.2 showed there is nothing to expand in, and the replacement is §7.4's procedure rather than a different expansion. It also fails, more quietly, for a Hamiltonian with both a discrete and a continuous part, where an expansion in bound states alone misses the scattering states entirely. Chapter 4.7 meets that case at a barrier, and hydrogen is an instance of it, since its spectrum carries the whole of [0,)[0,\infty) above the bound states.

8.3 · Write a function of an operator, including the exponential

For any self-adjoint A^\hat A and any measurable ff, the operator f(A^)f(\hat A) of (4.5.39) exists, and if ff is bounded on the spectrum then f(A^)f(\hat A) is bounded and defined on the whole space. The case the book uses most is ft(λ)=eiλt/f_t(\lambda)=\ee^{-\ii\lambda t/\hbar}, which has modulus one on the real line and therefore gives a bounded operator for every real tt. That eiH^t/\ee^{-\ii\hat Ht/\hbar} is not merely bounded but unitary, and that it is what Chapter 4.2 §7 meant by time evolution, is the content of §9 below, so this entry is provisional until then and complete afterwards.

8.4 · Integrate by parts and drop the boundary term

This one is Chapter 4.4's and it is included because it is the item that fails silently. The boundary term of Chapter 0.2 §3.2 may be dropped when both functions lie in the domain of the operator, and the whole of Chapter 4.4 §5 is the demonstration that "both functions are nice" is not the same condition. On the line the term vanishes for every pair in the domain. On an interval it vanishes only for the pairs a chosen phase permits. On a half-line it cannot be made to vanish at all, and the operator is not an observable. The rule to carry is that the boundary term is where the domain lives, so dropping it is a claim about domains and must be checked as one.

8.5 · What is still not safe

Four things need naming, since a checklist that only permits is a checklist that gets misread.

  • Multiplying two unbounded operators without checking domains. The product A^B^\hat A\hat B is defined only where B^ψ\hat B\psi lands in dom(A^)\operatorname{dom}(\hat A), and Chapter 4.4's Worked example 2 showed that even squaring a self-adjoint operator changes the domain in a way that matters physically.
  • Adding two self-adjoint operators and expecting a third. On dom(A^)dom(B^)\operatorname{dom}(\hat A)\cap\operatorname{dom}(\hat B) the sum need not be self-adjoint, need not be essentially self-adjoint, and need not even be densely defined. Nothing in §3.2 is inherited by sums, which is why §5.6 refuses to certify a Hamiltonian by the operators it was assembled from and why Chapter 4.6 §4.4 has to check p^2/2m+V(x^)\hat p^{2}/2m+V(\hat x) case by case.
  • Assuming a symmetric operator has a spectral theorem. It does not, and §3.2 named the counterexample. Every result in this chapter has self-adjoint in its hypothesis, which is a condition about domains and nothing else.
  • Treating x\ket x as a vector. It has no norm, it cannot be normalised, and xx\avg{x|x} is not a number. Every appearance of it is §7.4's procedure in shorthand.
In plain terms 4.5.8

What this chapter has really produced is permission. Four moves are made constantly in the rest of the subject, and before now none of them was justified for a quantity with no states of definite value. Slipping a complete set of alternatives into the middle of an expression is allowed when the set is genuinely complete, or when it comes from the family of projections built earlier, or when it is shorthand for the box-and-limit routine. Expanding a state in states of definite value is allowed only when such states exist and are numerous enough, which rules out position and momentum altogether.

Writing down an exponential of a quantity is allowed for any legitimate observable, and the next section shows that the particular exponential describing the passage of time preserves total probability exactly. Throwing away the leftover from an integration by parts is allowed only when both functions belong to the set the operator is permitted to act on, which is the whole subject of the previous chapter and the item that goes wrong most often.

The list of what remains forbidden is short and worth keeping. Two quantities cannot be multiplied without first asking where the product is defined, and, less obviously, they cannot be added either: the sum of two legitimate observables is not automatically a legitimate observable, which is why the energy of a particle in a potential has to be examined rather than assumed. A quantity that merely gives real averages does not qualify for any of this, since every result here demands the stronger condition. And the two most-used symbols in the subject are abbreviations rather than states, so nothing that treats them as states is safe.

9 · Stone's theorem

Chapter 4.2 §7.3 derived, from unitarity and the group law alone, that time evolution has the form eiG^t/\ee^{-\ii\hat Gt/\hbar} for some Hermitian G^\hat G with the dimensions of energy. It did that by differentiating the group law, and it said in place that the rigorous infinite-dimensional version of that step is Stone's theorem, quoted here. This section supplies it. One direction is three lines from §6.6 and is proved. The other is the hard one and is quoted. Together they say that "time evolution is unitary" and "the Hamiltonian is self-adjoint" are the same statement, which is the sentence Chapter 0.5 §7 has been pointing at since Part 0.

9.1 · The forward direction, built

One letter is about to carry two jobs, so separate them before anything is computed with it. The U^\hat U of §3.2 is the spectral unitary. It is fixed once and for all by H^\hat H, it never carries an argument, and it is the map (4.5.39) conjugates by. The U^(t)\hat U(t) this section builds is the evolution family, a different operator for each tt. The argument is the only thing separating them, so read each U^\hat U below by whether it has one.

Let H^\hat H be self-adjoint, and define U^(t)=eiH^t/\hat U(t)=\ee^{-\ii\hat Ht/\hbar} by (4.5.39), which is legitimate because λeiλt/\lambda\mapsto\ee^{-\ii\lambda t/\hbar} is measurable and bounded. Then work on the far side of the spectral unitary, in §3.2's multiplication picture, where H^\hat H is multiplication by gg. A state ψ\psi appears there as Ψ=U^ψ\Psi=\hat U\psi, which is the notation §6.2 set up, and U^(t)\hat U(t) appears as multiplication by eig(ξ)t/\ee^{-\ii g(\xi)t/\hbar}. So every claim becomes a claim about a function of modulus one. Three things follow immediately.

It is unitary, because multiplying by a function of modulus one leaves Ψ2dμ\int\abs\Psi^{2}\dd\mu unchanged and is undone by multiplying by the conjugate. It satisfies the group law, because eigs/eigt/=eig(s+t)/\ee^{-\ii g s/\hbar}\ee^{-\ii g t/\hbar}=\ee^{-\ii g(s+t)/\hbar} pointwise, and U^(0)=I^\hat U(0)=\hat I. And it is strongly continuous, meaning U^(t)ψψ\hat U(t)\psi\to\psi for every state as t0t\to0, which is the one that needs an argument:

U^(t)ψψ2  =  Xeig(ξ)t/12Ψ(ξ)2dμ  t0  0, \norm{\hat U(t)\psi-\psi}^{2} \;=\; \int_X \abs{\ee^{-\ii g(\xi)t/\hbar}-1}^{2}\,\abs{\Psi(\xi)}^{2}\,\dd\mu \;\xrightarrow[t\to0]{}\;0, (4.5.47)

because the integrand tends to zero pointwise and is dominated by 4Ψ24\abs\Psi^{2}, which is integrable. That is dominated convergence, Chapter 4.3 §4.3, used for the fourth time in this chapter. Note what was not claimed: U^(t)I^\norm{\hat U(t)-\hat I} does not tend to zero unless H^\hat H is bounded, so the continuity is state by state and not uniform. Since every observable in this book except spin is unbounded, that distinction is the normal case rather than the exception.

9.2 · The generator, recovered by differentiating

The remaining half of the forward direction is that H^\hat H can be read back off the flow, which is what makes it the generator rather than merely an ingredient. Take ψdom(H^)\psi\in\operatorname{dom}(\hat H) and form the difference quotient in the multiplication picture, since that turns an operator limit into a limit of an integral:

U^(t)ψψt+iH^ψ2  =  Xeigt/1t+ig2Ψ2dμ  t0  0. \norm{\frac{\hat U(t)\psi-\psi}{t}+\frac{\ii}{\hbar}\hat H\psi}^{2} \;=\; \int_X \abs{\frac{\ee^{-\ii g t/\hbar}-1}{t}+\frac{\ii g}{\hbar}}^{2}\abs{\Psi}^{2}\,\dd\mu \;\xrightarrow[t\to0]{}\;0. (4.5.48)

The integrand tends to zero pointwise, because that is the derivative of eigt/\ee^{-\ii gt/\hbar} at t=0t=0 computed for each fixed ξ\xi. And it is dominated: the elementary bound eiθ1θ\abs{\ee^{\ii\theta}-1}\le\abs\theta makes the first term at most g/\abs g/\hbar in modulus, so the whole bracket is at most 2g/2\abs g/\hbar, and g2Ψ2dμ\int g^{2}\abs\Psi^{2}\dd\mu is finite precisely because ψ\psi is in the domain. Dominated convergence again, and the limit is zero. So

iddt(eiH^t/ψ)t=0  =  H^ψfor every ψdom(H^), \ii\hbar\,\dv{}{t}\Big(\ee^{-\ii\hat Ht/\hbar}\psi\Big)\bigg|_{t=0} \;=\; \hat H\psi \qquad \text{for every } \psi\in\operatorname{dom}(\hat H), (4.5.49)

with the derivative taken in the norm of the space rather than pointwise in xx. That is the forward direction complete: a self-adjoint operator generates a strongly continuous one-parameter unitary group, and it is recoverable from that group by differentiating at the origin. Every step used §6.6 and Chapter 4.3 §4.3 and nothing else.

9.3 · The converse

⚑ Quoted, not derived — the converse half of Stone's theorem

We use, without proof, the following. Let {U^(t)}tR\{\hat U(t)\}_{t\in\R} be a family of operators on a Hilbert space satisfying three conditions: each U^(t)\hat U(t) is unitary; the family obeys the group law U^(s)U^(t)=U^(s+t)\hat U(s)\hat U(t)=\hat U(s+t) with U^(0)=I^\hat U(0)=\hat I; and the family is strongly continuous, meaning U^(t)ψψ\hat U(t)\psi\to\psi as t0t\to0 for every state ψ\psi. Then there is a unique self-adjoint operator H^\hat H with

U^(t)=eiH^t/,dom(H^)={ψ: limt0U^(t)ψψt exists in norm}. \hat U(t)=\ee^{-\ii\hat Ht/\hbar}, \qquad \operatorname{dom}(\hat H)=\left\{\psi:\ \lim_{t\to0}\frac{\hat U(t)\psi-\psi}{t}\ \text{exists in norm}\right\}.

The hypotheses are the content, again. Strong continuity cannot be weakened away and cannot be strengthened to continuity in the operator norm without destroying the theorem's use: a norm-continuous group has a bounded generator, and every Hamiltonian in this book is unbounded, so the norm-continuous version would apply to nothing. Unitarity at each tt is what makes the generator self-adjoint rather than merely symmetric, which is the whole difference this part of the book has been built on. And the domain is not an input: the theorem produces it as the set where the difference quotient converges, which is why the generator's domain is forced in exactly the sense Chapter 4.4 §2.3 established.

What the mark covers. Only the existence and uniqueness of H^\hat H given the group. Section 9.1 proved the other direction outright, so what is quoted is one implication of an equivalence rather than the whole statement. The proof constructs H^\hat H on the difference-quotient domain, shows that domain is dense by averaging U^(t)ψ\hat U(t)\psi against smooth weights, and then identifies the resulting symmetric operator as self-adjoint using the deficiency-index criterion Chapter 4.4 §6.2 quoted. It is about four pages and it is in Reed and Simon, volume I, §VIII.4.

9.4 · What the two halves say together

Read the two directions as one statement and it is a dictionary entry rather than a theorem. A strongly continuous one-parameter unitary group and a self-adjoint operator are the same object described twice. Chapter 4.2 §7 assumed the first, from linearity and conservation of probability, and deduced the second by differentiating the group law formally. Stone is what makes that differentiation legitimate in a space where the generator is unbounded and defined on a proper subset, which is the only case that arises.

Three consequences follow, and Part IV spends all three.

  • Probability is conserved as a theorem. U^(t)ψ=ψ\norm{\hat U(t)\psi}=\norm\psi for every tt, because U^(t)\hat U(t) is unitary, so a state that starts normalised stays normalised for all time and nothing has to be imposed to keep it so.
  • Self-adjointness is a physical requirement, not a technical one. Chapter 4.4 §5.5 showed that momentum on a half-line is symmetric with no self-adjoint extension, and its Worked example 3 showed the corresponding flow runs forwards only. Stone explains why those two facts are the same fact: no self-adjoint generator means no two-sided unitary group, and a system whose evolution cannot be run backwards is not a system this framework describes.
  • The Hamiltonian is what a clock reads off. Given the flow, (4.5.49) extracts H^\hat H uniquely, so there is no freedom to choose a different generator for the same evolution. Chapter 4.2's fifth postulate is then the single remaining physical statement, that this uniquely determined operator is the energy.

A note on signs, since the exponential now carries one. This book writes U^(t)=eiH^t/\hat U(t)=\ee^{-\ii\hat Ht/\hbar}, with the minus sign, matching itψ=H^ψ\ii\hbar\,\partial_t\ket\psi=\hat H\ket\psi. Chapter 4.6 §2 states the convention and what it buys. Some engineering literature uses the opposite time convention throughout, and a formula copied across from it will differ by the sign of every i\ii.

In plain terms 4.5.9

Time evolution has to preserve total probability, so whatever moves a state forward by a given interval must be a rotation of the space of states. Chapter 4.2 showed that every such rotation is an exponential of something, and that the something is a legitimate observable with the dimensions of energy. That argument was made by differentiating a product rule, and in a space where the observable is unbounded and defined only on part of the space, differentiating a product rule is not a step one may take without care.

This section supplies the care, and it splits into a half that is proved and a half that is quoted. The proved half says that starting from a legitimate energy observable, the exponential really is a rotation, really does compose correctly, really does move continuously, and really does hand the energy back when differentiated. All four are one-line statements once the observable has been turned into multiplication by a function, because an exponential of a real function has size one everywhere.

The quoted half runs the other way and says that any family of rotations that composes correctly and moves continuously arises from exactly one such observable. Put the two together and the statement is a translation rather than a discovery: a flow that conserves probability and a legitimate energy are the same thing named twice. What is left over as genuine physics is the single claim that this particular operator is the energy a calorimeter measures, and that claim was labelled a postulate when it was made and remains one.

10 · Worked examples

Worked example 1 — the free Hamiltonian, whose spectrum is a half-line and whose degeneracy has no eigenvectors to be degenerate between

Let H^0=p^2/2m\hat H_0=\hat p^{2}/2m on L2(R)L^{2}(\R). (a) Give the theorem's data (X,μ,U^,g)(X,\mu,\hat U,g). (b) Find the spectrum. (c) Show there are no eigenvectors. (d) Every text says each energy above zero is doubly degenerate. Say what that means here, and compute the density of μψ\mu_\psi in the energy variable.

(a) Take U^=F\hat U=\mathcal F as in (4.5.15), so p^\hat p becomes multiplication by k\hbar k on L2(R,dk)L^{2}(\R,\dd k). Then §6.6's rule with f(λ)=λ2/2mf(\lambda)=\lambda^{2}/2m makes H^0\hat H_0 multiplication by f(k)f(\hbar k), so

X=R,μ=dk,U^=F,g(k)=2k22m. X=\R, \qquad \mu=\dd k, \qquad \hat U=\mathcal F, \qquad g(k)=\frac{\hbar^{2}k^{2}}{2m}.

The domain is the set of ψ\psi with k2ψ~L2k^{2}\tilde\psi\in L^{2}, which is where the formula lands back in the space and is not something anyone chose.

(b) By §3.4 the spectrum is the essential range of gg. The function 2k2/2m\hbar^{2}k^{2}/2m takes every non-negative value, and every neighbourhood of every such value is hit on a set of positive length, so σ(H^0)=[0,)\sigma(\hat H_0)=[0,\infty). Negative energies are excluded, which is the right answer and is the first place in this book where a spectrum has a floor.

(c) Section 3.4 says EE is an eigenvalue exactly when gg takes the value EE on a set of positive measure. The equation 2k2/2m=E\hbar^{2}k^{2}/2m=E has at most the two solutions k=±2mE/k=\pm\sqrt{2mE}/\hbar, a set of two points, which has measure zero. So there are no eigenvalues and no eigenvectors, and a free particle has no state of definite energy any more than it has one of definite momentum.

(d) The degeneracy is a statement about the preimage, and that is the form of it that survives with no eigenvectors present. For a narrow energy window E=[E0,E0+ϵ]E=[E_0,E_0+\epsilon] the set g1(E)g^{-1}(E) has two components, one on each side of k=0k=0, so the projection P(E)P(E) of (4.5.32) keeps two separate bands of wavenumbers rather than one. Those are the right-moving and left-moving states, and "doubly degenerate" means the window has two components rather than that two eigenvectors share a value.

The density falls out of the same picture. Writing μψ(E)=g1(E)ψ~2dk\mu_\psi(E)=\int_{g^{-1}(E)}\abs{\tilde\psi}^{2}\dd k and changing variable by E=2k2/2mE=\hbar^{2}k^{2}/2m, so that dE=2kdk/m\dd E=\hbar^{2}k\,\dd k/m on each branch separately,

dμψ  =  ρ(E)dE,ρ(E)=m2k(ψ~(k)2+ψ~(k)2)k=2mE/. \dd\mu_\psi \;=\; \rho(E)\,\dd E, \qquad \rho(E)=\frac{m}{\hbar^{2}k}\Big(\abs{\tilde\psi(k)}^{2}+\abs{\tilde\psi(-k)}^{2}\Big)\bigg|_{k=\sqrt{2mE}/\hbar}.

The two terms are the two components, which is the degeneracy appearing as a sum. The factor m/2km/\hbar^{2}k is the density of states, and it diverges as E0E\to0, which is why the bottom of a continuous spectrum needs care in Chapter 4.17's transition rates.

Worked example 2 — one operator with both kinds of spectrum, and the atom you can see

On L2(R)L^{2}(\R) let A^\hat A be multiplication by g(x)=xg(x)=x for x<0x\lt0 and g(x)=0g(x)=0 for x0x\ge0. (a) Say why it is self-adjoint. (b) Find the spectrum and its parts. (c) Identify the eigenspace and the projection onto it. (d) Describe μψ\mu_\psi and say what kind of distribution it is.

(a) It is multiplication by a real measurable function, which is the model case of §3.2 with U^\hat U the identity, and gg is bounded above and below on any bounded set. Multiplication by a real function is symmetric because the function passes through the inner product unchanged, and the adjoint's domain is the same set by the argument Chapter 4.4's Worked example 1 ran for x^\hat x, so it is self-adjoint.

(b) The essential range is where the answer is. For λ<0\lambda\lt0, every neighbourhood of λ\lambda is the image of an interval of negative xx of positive length, so λ\lambda is in the spectrum. For λ=0\lambda=0, the set where g<ϵ\abs g\lt\epsilon contains the whole half-line [0,)[0,\infty), which has infinite measure. For λ>0\lambda\gt0, the set where gλ<ϵ\abs{g-\lambda}\lt\epsilon is empty once ϵ<λ\epsilon\lt\lambda, since gg never takes a positive value. So σ(A^)=(,0]\sigma(\hat A)=(-\infty,0].

(c) By §3.4 a value is an eigenvalue exactly when gg takes it on a set of positive measure. The value 00 is taken on all of [0,)[0,\infty), and no negative value is taken on more than a single point. So σp(A^)={0}\sigma_p(\hat A)=\{0\} and σc(A^)=(,0)\sigma_c(\hat A)=(-\infty,0), and the eigenspace at zero is the set of states supported on the right half-line. The projection onto it is multiplication by 1[0,)\mathbf 1_{[0,\infty)}, which is P({0})P(\{0\}) from (4.5.32), since g1({0})=[0,)g^{-1}(\{0\})=[0,\infty).

(d) Split any state as ψ=ψ+ψ+\psi=\psi_-+\psi_+ by the sign of xx. Then μψ\mu_\psi assigns the mass 0ψ2dx\int_0^{\infty}\abs\psi^{2}\dd x to the single point 00, and on the negative axis it has the density ψ(λ)2\abs{\psi(\lambda)}^{2} in the variable λ\lambda, since gg is the identity there. So μψ\mu_\psi is a mixed distribution: an atom at zero carrying the probability that the particle is on the right, plus a continuous part carrying the rest. That is the general shape, and the familiar-ground box of §6.5 is where the same object appears in survival data, where an atom is the probability of an event at a single exact time and a continuous part is everything else. An operator with a discrete spectrum is all atoms, one with a continuous spectrum has none, and this one has both.

Worked example 3 — xp\avg{x|p} from the box, with the constants made to match Chapter 4.3

(a) Derive xp\avg{x|p} from §7.2's box. (b) Show that pψ\avg{p|\psi} is the momentum-space wavefunction ψ^(p)\hat\psi(p) of Chapter 4.3 §8.5. (c) Insert ppdp\int\ket p\bra p\,\dd p into ψ,p^ψ\avg{\psi,\hat p\psi} and check the answer against §6.3. (d) Say why xx=δ(xx)\avg{x|x'}=\delta(x-x') and pp=δ(pp)\avg{p|p'}=\delta(p-p') cannot both be read as inner products of vectors.

(a) In the box, kn=L/2πun\ket{k_n}=\sqrt{L/2\pi}\,u_n by (4.5.41), and unu_n is continuous, so §7.3's cell average converges to its value at xx. That gives xkn=L/2πL1/2eiknx=eiknx/2π\avg{x|k_n}=\sqrt{L/2\pi}\cdot L^{-1/2}\ee^{\ii k_nx}=\ee^{\ii k_nx}/\sqrt{2\pi}, which does not depend on LL at all, so the limit is immediate. Rescaling to momentum by (4.5.46) divides by \sqrt\hbar:

xk=eikx2π,xp=eipx/2π. \avg{x|k}=\frac{\ee^{\ii kx}}{\sqrt{2\pi}}, \qquad \avg{x|p}=\frac{\ee^{\ii px/\hbar}}{\sqrt{2\pi\hbar}}.

Chapter 0.9 §5.3 wrote the first of these in one line as a summary of its §2, and the derivation above is what that line abbreviated.

(b) Insert the resolution of the identity in xx, which §8.1 permits as an abbreviation of §7.3, and take the conjugate of part (a):

pψ  =  pxxψdx  =  12πψ(x)eipx/dx  =  ψ^(p), \avg{p|\psi} \;=\; \int\avg{p|x}\avg{x|\psi}\,\dd x \;=\; \frac{1}{\sqrt{2\pi\hbar}}\int_{-\infty}^{\infty}\psi(x)\,\ee^{-\ii px/\hbar}\,\dd x \;=\; \hat\psi(p),

which is Chapter 4.3 §8.5's formula with the same 2π\sqrt{2\pi\hbar} in the same place. The agreement is not automatic. It happens because Chapter 0.9 §2.2 chose the symmetric convention, in which the transform is unitary, and any other convention would leave a factor here that the normalisation pp=δ(pp)\avg{p'|p}=\delta(p-p') would then have to absorb.

(c) Two insertions and the multiplication form:

ψ,p^ψ  =  ψpppψdp  =  pψ^(p)2dp. \avg{\psi,\hat p\psi} \;=\; \int \avg{\psi|p}\,p\,\avg{p|\psi}\,\dd p \;=\; \int p\,\abs{\hat\psi(p)}^{2}\,\dd p.

Check it against §6.3, which gives ψ,p^ψ=kψ~(k)2dk\avg{\psi,\hat p\psi}=\int\hbar k\abs{\tilde\psi(k)}^{2}\dd k in the wavenumber variable. Substituting p=kp=\hbar k gives dk=dp/\dd k=\dd p/\hbar, and (4.5.46) gives ψ^(p)=ψ~(k)/\hat\psi(p)=\tilde\psi(k)/\sqrt\hbar, so the integrand picks up \hbar from ψ~2\abs{\tilde\psi}^{2} and loses it again to dk\dd k. The two agree, and the agreement is a check on the \hbar bookkeeping rather than on the physics.

(d) Because xx\avg{x|x} would be the squared norm of x\ket x, and setting x=xx'=x in the first relation gives δ(0)\delta(0), which is not a number. The same objection applies to p\ket p. Both relations are true in §7.2's sense, as statements about what the symbols do under an integral, and neither is a statement about the length of a vector. There is also a counting objection that is independent of any limit: Chapter 4.3 §7.4 proved every orthonormal basis of this space is countable, and neither family is.

11 · Your turn

Problem 1 — the spectrum of a multiplication operator, both inclusions

(a) Let M^\hat M be multiplication by a real measurable gg on L2(X,μ)L^{2}(X,\mu), with μ\mu σ\sigma-finite. Show that every λ\lambda in the essential range of gg is in the spectrum, by building unit vectors that M^λ\hat M-\lambda shrinks to nothing. (b) Show the converse, that every λ\lambda outside the essential range is in the resolvent set. (c) Apply both to x^\hat x on L2[0,1]L^{2}[0,1] and give the spectrum and its parts. (d) Compare the norms of x^\hat x on L2[0,1]L^{2}[0,1] and on L2(R)L^{2}(\R), and say which of §2's three demands is doing the work in each case.

Solution

(a) Let λ\lambda be in the essential range, so Sn={ξ:g(ξ)λ<1/n}S_n=\{\xi:\abs{g(\xi)-\lambda}\lt1/n\} has positive measure for every nn. Since μ\mu is σ\sigma-finite, each SnS_n contains a subset TnT_n of finite positive measure. Put Ψn=1Tn/μ(Tn)\Psi_n=\mathbf 1_{T_n}/\sqrt{\mu(T_n)}, which is a unit vector, and then (M^λ)Ψn2=μ(Tn)1Tngλ2dμ1/n2\norm{(\hat M-\lambda)\Psi_n}^{2}=\mu(T_n)^{-1}\int_{T_n}\abs{g-\lambda}^{2}\dd\mu \le 1/n^{2}. Unit vectors going to zero means no bounded inverse, exactly as in (4.5.6).

(b) If λ\lambda is not in the essential range there is ϵ>0\epsilon\gt0 with μ({gλ<ϵ})=0\mu(\{\abs{g-\lambda}\lt\epsilon\})=0, so gλϵ\abs{g-\lambda}\ge\epsilon almost everywhere. Then 1/(gλ)1/(g-\lambda) is bounded by 1/ϵ1/\epsilon, multiplication by it is defined on the whole space and bounded, and it inverts M^λ\hat M-\lambda on both sides. So λ\lambda is in the resolvent set, and with (a) this proves σ(M^)=ess ran(g)\sigma(\hat M)=\operatorname{ess\,ran}(g).

(c) Here g(x)=xg(x)=x on [0,1][0,1] with Lebesgue measure, whose essential range is [0,1][0,1], so σ(x^)=[0,1]\sigma(\hat x)=[0,1]. No value is taken on a set of positive measure, so there are no eigenvalues, and §2.7 rules out a residual part, leaving σc=[0,1]\sigma_c=[0,1]. The picture is the same as on the line with the spectrum cut down to the interval the particle is confined to, which is what it should be.

(d) On [0,1][0,1] the operator is bounded, with x^=1\norm{\hat x}=1, since 01x2ψ201ψ2\int_0^{1}x^{2}\abs\psi^{2}\le\int_0^{1}\abs\psi^{2}, and the supremum is approached by states concentrated near x=1x=1. On R\R it is unbounded, by Chapter 4.4's Worked example 1. So on the interval the operator is defined everywhere and the only failure at a spectral point is unboundedness of the inverse, while on the line the operator itself is unbounded and carries a restricted domain as well. The third demand of (4.5.2) is what fails in both cases, and boundedness of the operator has nothing to do with it.

Problem 2 — the Hermite ladder, and half the energy is potential

(a) Read H2H_2 off (4.5.19) and check it satisfies (4.5.21). (b) Show that ξhn=(n+1)/2hn+1+n/2hn1\xi h_n=\sqrt{(n+1)/2}\,h_{n+1}+\sqrt{n/2}\,h_{n-1}, and deduce hn,ξ2hn=n+12\avg{h_n,\xi^{2}h_n}=n+\half. (c) Deduce that in the state hnh_n the expected potential energy is exactly half the total, and say which classical statement that reproduces. (d) For the displaced ground state ψa\psi_a of §5.7, compute H^osc\avg{\hat H_{\text{osc}}} two ways, from the Poisson coefficients and from the classical displacement energy, and check they agree.

Solution

(a) Expand e2ξtt2=1+(2ξtt2)+12(2ξtt2)2+\ee^{2\xi t-t^{2}}=1+(2\xi t-t^{2})+\half(2\xi t-t^{2})^{2}+\cdots and collect the terms in t2t^{2}, which are t2-t^{2} from the second bracket and 2ξ2t22\xi^{2}t^{2} from the third. Their sum is (2ξ21)t2(2\xi^{2}-1)t^{2}, and (4.5.19) says that equals H2t2/2!H_2t^{2}/2!, so H2=4ξ22H_2=4\xi^{2}-2. Then H2=8ξH_2'=8\xi, H2=8H_2''=8, and 82ξ(8ξ)+4(4ξ22)=816ξ2+16ξ28=08-2\xi(8\xi)+4(4\xi^{2}-2)=8-16\xi^{2}+16\xi^{2}-8=0.

(b) The first recurrence in (4.5.21) rearranges to ξHn=12Hn+1+nHn1\xi H_n=\half H_{n+1}+nH_{n-1}. Divide by the normalising constant Nn=2nn!πN_n=\sqrt{2^{n}n!\sqrt\pi} of (4.5.20) and multiply by the Gaussian. Since Nn+1/Nn=2(n+1)N_{n+1}/N_n=\sqrt{2(n+1)} and Nn1/Nn=1/2nN_{n-1}/N_n=1/\sqrt{2n}, the two coefficients become 2(n+1)/2=(n+1)/2\sqrt{2(n+1)}/2=\sqrt{(n+1)/2} and n/2n=n/2n/\sqrt{2n}=\sqrt{n/2}. Then hn,ξ2hn\avg{h_n,\xi^{2}h_n} is ξhn2\norm{\xi h_n}^{2}, and orthonormality kills the cross term, leaving (n+1)/2+n/2=n+12(n+1)/2+n/2=n+\half.

(c) The potential energy is 12mω2x2=12ωξ2\half m\omega^{2}x^{2}=\half\hbar\omega\,\xi^{2} by the scaling in (4.5.18), so its expectation in hnh_n is 12ω(n+12)\half\hbar\omega(n+\half), which is half of (n+12)ω(n+\half)\hbar\omega. The kinetic part is therefore the other half. That is the virial theorem for a quadratic potential, which in classical mechanics says the time-averaged kinetic and potential energies of a harmonic oscillator are equal, and here it holds state by state rather than on average over a period.

(d) From the coefficients, H^osc=n(n+12)ωcn2=ω(nˉ+12)\avg{\hat H_{\text{osc}}}=\sum_n(n+\half)\hbar\omega\abs{c_n}^{2} =\hbar\omega(\bar n+\half), and nˉ\bar n is the mean of the Poisson distribution in (4.5.30), which is a2/2a^{2}/2. So the answer is ω(a2/2+12)\hbar\omega(a^{2}/2+\half). Classically, displacing the ground state by aa in the ξ\xi variable means displacing it by x=ax0x=ax_0, at a cost of 12mω2a2x02=12ωa2\half m\omega^{2}a^{2}x_0^{2}=\half\hbar\omega a^{2}, on top of the ground-state energy 12ω\half\hbar\omega. The two agree, and the agreement is why a displaced ground state is the quantum state that behaves most like a classical oscillation, which Chapter 4.10 takes up.

Problem 3 — Stone, run on translations, and the flow that cannot be reversed

(a) On L2(R)L^{2}(\R) define (U^(a)ψ)(x)=ψ(xa)(\hat U(a)\psi)(x)=\psi(x-a). Show it is unitary and obeys the group law. (b) Show it is strongly continuous, and say why the argument needs a density statement rather than a pointwise one. (c) Identify the generator, using (4.5.48) in the transformed picture rather than by differentiating ψ\psi. (d) Repeat (a) on L2[0,)L^{2}[0,\infty) and say which hypothesis of §9.3 fails, connecting it to Chapter 4.4 §5.5.

Solution

(a) Sliding a function does not change ψ2\int\abs\psi^{2}, so the map preserves the norm, and it is inverted by sliding back, so it is unitary. Sliding by aa and then by bb slides by a+ba+b, and sliding by zero does nothing, which is the group law.

(b) For a continuous ψ\psi vanishing outside a bounded set, ψ(a)ψ0\norm{\psi(\cdot-a)-\psi}\to0 as a0a\to0 because the difference is uniformly small on a fixed bounded set. A general ψL2\psi\in L^{2} has no pointwise regularity at all, so no such estimate is available directly. The route is the one §7.3 used. Choose a continuous ϕ\phi within ϵ\epsilon by Chapter 4.3 §8.1, note that U^(a)\hat U(a) has norm one so it moves ψϕ\psi-\phi by at most ϵ\epsilon, and bound the total by 3ϵ3\epsilon.

(c) Transform. By Chapter 0.9 §3, sliding by aa multiplies the transform by eika\ee^{-\ii ka}, so in the transformed picture U^(a)\hat U(a) is multiplication by eika\ee^{-\ii ka}. Comparing with eiG^a/\ee^{-\ii\hat Ga/\hbar}, which §9.1 says is multiplication by eig(k)a/\ee^{-\ii g(k)a/\hbar}, gives g(k)=kg(k)=\hbar k, and by (4.5.15) that is p^\hat p. So momentum generates translation, which Chapter 4.4's Worked example 3 obtained from a Taylor series and which is now a statement about two functions of kk agreeing.

(d) On the half-line the recipe makes sense only for a0a\ge0, since sliding right leaves vacated space to fill with zeros while sliding left would need values of ψ\psi at negative xx that do not exist. So the family is a semigroup: it has the group law for non-negative parameters and no inverses. The hypothesis that fails is the first, that each U^(a)\hat U(a) is unitary, since the forward maps are isometries that are not onto, exactly the shift of Chapter 4.4 §1.2 with a continuous parameter. Chapter 4.4 §5.5 found the matching fact on the other side, that momentum on the half-line has no self-adjoint extension, and Stone is what says those are one fact.

Problem 4 — the spectral measure does the statistics

(a) Show P(E)P(E) commutes with A^\hat A on dom(A^)\operatorname{dom}(\hat A), and that P(E)P(E) maps the domain into itself. (b) Show ψ,A^2ψ=λ2dμψ(λ)\avg{\psi,\hat A^{2}\psi}=\int\lambda^{2}\dd\mu_\psi(\lambda) for ψ\psi in the domain of A^2\hat A^{2}, and deduce that the variance of A^\hat A in the state ψ\psi is the variance of the distribution μψ\mu_\psi. (c) Show that if μψ\mu_\psi has an atom at λ\lambda then ψ\psi has a non-zero component in the eigenspace at λ\lambda, and conversely. (d) Use (b) to show that a state with μψ\mu_\psi concentrated at a single point is an eigenvector, and say why this does not contradict §1.2.

Solution

(a) In the multiplication picture P(E)P(E) is multiplication by an indicator and A^\hat A is multiplication by gg, and two multiplications commute pointwise. If gΨL2g\Psi\in L^{2} then g1ΨL2g\mathbf 1\Psi\in L^{2} as well, since multiplying by an indicator only decreases moduli, so the domain is preserved. Conjugating back by U^\hat U gives both statements on HH.

(b) Section 6.6 makes A^2\hat A^{2} multiplication by g2g^{2}, so ψ,A^2ψ=g2Ψ2dμ\avg{\psi,\hat A^{2}\psi}=\int g^{2}\abs\Psi^{2}\dd\mu, and the pushforward argument of §6.3 applied to the function λ2\lambda^{2} rather than λ\lambda turns that into λ2dμψ\int\lambda^{2}\dd\mu_\psi. Subtracting the square of (4.5.35) gives A^2A^2=λ2dμψ(λdμψ)2\avg{\hat A^{2}}-\avg{\hat A}^{2}=\int\lambda^{2}\dd\mu_\psi-\big(\int\lambda\,\dd\mu_\psi\big)^{2}, which is the variance of μψ\mu_\psi. So every moment of the measurement statistics is a moment of the same distribution, and no separate probabilistic assumption is being made.

(c) An atom at λ\lambda means μψ({λ})=P({λ})ψ2>0\mu_\psi(\{\lambda\})=\norm{P(\{\lambda\})\psi}^{2}\gt0, and §6.1 identified P({λ})P(\{\lambda\}) as the projection onto the eigenspace at λ\lambda. So the component of ψ\psi in that eigenspace is non-zero, and the converse is the same sentence read backwards, since the projection of a vector with a non-zero component there is non-zero.

(d) If μψ=δλ0\mu_\psi=\delta_{\lambda_0} then P({λ0})ψ2=1=ψ2\norm{P(\{\lambda_0\})\psi}^{2}=1=\norm\psi^{2}, so ψ\psi lies entirely in the eigenspace at λ0\lambda_0 and is an eigenvector. Equivalently, the variance in (b) is zero. There is no contradiction with §1.2, which said x^\hat x has no eigenvectors: what part (d) shows is that no state of x^\hat x has a spectral measure concentrated at a point, which is the same statement. The sequence (4.5.6) has μψn\mu_{\psi_n} concentrated on an interval of width 1/n1/n, narrowing without ever reaching a point.

The brick you just laid — the values, when there are no eigenvectors to attach them to

The word had to be widened before a theorem could be stated at all. Chapter 0.5 defined an eigenvalue by the existence of an eigenvector, and §1.2 showed that position on L2(R)L^{2}(\R) has none, since (xλ)ψ=0(x-\lambda)\psi=0 forces ψ\psi to vanish off a null set and the quotient of Chapter 4.3 §5.3 then makes it the zero vector. So the definition was rebuilt on invertibility instead: λ\lambda is in the spectrum when A^λ\hat A-\lambda fails to have a bounded inverse defined on the whole space. In finite dimensions that gives back the eigenvalues exactly, by rank–nullity, so nothing was renamed. In infinite dimensions it gives σ(x^)=σ(p^)=R\sigma(\hat x)=\sigma(\hat p)=\R with no eigenvectors anywhere, and the states that stand in for the missing ones are the unit vectors of (4.5.6), which x^λ\hat x-\lambda shrinks to nothing. Two facts about self-adjoint operators were proved rather than assumed and both are used everywhere afterwards: the spectrum is real, and the residual part is empty.

One theorem quoted, and checked three times. Every self-adjoint operator is unitarily equivalent to multiplication by a real function on some L2(μ)L^{2}(\mu), which is the infinite-dimensional reading of Chapter 0.5's A=UDUA=UDU^{\dagger} with the list of eigenvalues widened into a function and the eigenvectors dropped. That is the one substantial mathematical mark of this part, and §§4 and 5 discharge it into arithmetic. Position is already multiplication by xx. Momentum becomes multiplication by k\hbar k under the Fourier transform Chapter 0.9 proved unitary, verified numerically against a finite difference to 1.96×1041.96\times10^{-4} on a 40964096-point grid, with the discrepancy falling as the square of the spacing. And H^osc\hat H_{\text{osc}} becomes multiplication by (n+12)ω(n+\half)\hbar\omega on the non-negative integers with counting measure. Those three are what the theorem was verified on, and §5.6 says what the verification does and does not reach. Every later use is a function of one of the three under §6.6's functional calculus, or an operator whose own chapter exhibits a complete family of eigenfunctions, or p^2/2m+V(x^)\hat p^{2}/2m+V(\hat x), which is a fourth operator and is checked as one in Chapter 4.6 §4.4. Unitary equivalence to a multiplication operator is not inherited by sums, so nothing here is closed under addition and no Hamiltonian is certified by the company it is built from. The mark has a second clause as well, uniqueness of the projection-valued measure, which §6.1 needs before the definite article in the measure of A^\hat A is allowed and which §3.2 records alongside the first.

The Hermite functions were built, not quoted. Defined by a generating function, they satisfy two recurrences, hence the Hermite equation, hence H^oschn=(n+12)ωhn\hat H_{\text{osc}}h_n=(n+\half)\hbar\omega h_n, and they are orthonormal by squaring the same generating function and comparing coefficients. Completeness is the part that costs something, and the proof runs statement (d) of Chapter 4.3 §7.3: a vector orthogonal to all of them is orthogonal to every power against a Gaussian weight, so the Fourier transform of feξ2/2f\ee^{-\xi^{2}/2} vanishes identically, so ff is the zero vector. The interchange of sum and integral in the middle is dominated convergence with feξ2/2ekξ\abs f\ee^{-\xi^{2}/2}\ee^{\abs{k\xi}} as the dominating function, and Chapter 4.3's closing brick predicted exactly that shape. No complex analysis and no mark. Parseval in this basis was checked against a closed form: the coefficients of a displaced ground state are Poisson-distributed, n<20cn2=1.00000000\sum_{n\lt20}\abs{c_n}^{2}=1.00000000 at mean 22 against an exact tail of 6.44×10146.44\times10^{-14}, and the Gram matrix of the first forty functions differs from the identity by 4.0×10154.0\times10^{-15}, which is the rounding floor rather than a measurement of anything.

The sum became an integral, and Chapter 0.5's bill is settled. Indicator functions in the multiplication picture become projections P(E)P(E) on the original space, the family is a measure taking projections as values, and pairing it with a state gives an honest probability distribution μψ\mu_\psi on the real line, concentrated on the spectrum. Then ψ,A^ψ=λdμψ\avg{\psi,\hat A\psi}=\int\lambda\,\dd\mu_\psi, which is what (4.5.36) abbreviates, and the domain it comes with is the same one the multiplication form supplies rather than an extra stipulation. Chapter 0.5 §6.4 said the projection form "is the one that survives to infinite dimensions" and its warning box said the sum "becomes an integral λdP(λ)\int\lambda\,\dd P(\lambda) over projection-valued measures, and completeness must be re-proved rather than assumed". Both are now paid, and completeness was re-proved three times over: of the space in Chapter 4.3 §6.2, of the Fourier basis in Chapter 4.3 §8.3, of the Hermite basis in §5.4. The Born rule for a continuous variable falls out with no new postulate, and it explains rather than stipulates why the probability attaches to an interval.

x\ket x and p\ket p have a meaning, and it is a procedure. Box normalisation is the name of it. Put momentum in a periodic box and position on a lattice, where both objects are ordinary vectors and Chapter 4.3 §7.3 applies unaltered, then rescale so that every term carries the spacing and let the spacing go to zero. Sums become integrals, and a Kronecker delta over the spacing becomes a Dirac delta in Chapter 0.9 §5.1's sense of the word. That closes the gap Chapter 0.9 §5.3 named, and the word closes applies to that gap and to nothing wider: the general theory of distributions is Chapter 5.4's. Chapter 4.3 §8.5's warning stands unamended, and is worth carrying out of this chapter intact. These families are not orthonormal bases in Chapter 4.3 §7.3's sense, eikx\ee^{\ii kx} is not in L2(R)L^{2}(\R), and a state at a point is the zero vector.

Stone, half built and half quoted. A self-adjoint H^\hat H gives eiH^t/\ee^{-\ii\hat Ht/\hbar} unitary, obeying the group law, strongly continuous, and returning H^\hat H when differentiated at the origin. All four are one line each in the multiplication picture, and the last two are dominated convergence with 4Ψ24\abs\Psi^{2} and (2g/)2Ψ2(2\abs g/\hbar)^{2} \abs\Psi^{2} as the dominating functions. The converse is quoted with its hypotheses, and the hypothesis that matters is that continuity is strong rather than uniform, since a norm-continuous group has a bounded generator and every Hamiltonian in this book is unbounded. Together the two halves make time evolution is unitary and the Hamiltonian is self-adjoint the same statement, which is what Chapter 4.2 §7.3 was told to expect at this section number.

Three marks, and where completeness was spent. The spectral theorem in multiplication form, in §3.2, quoted with its hypotheses and then verified three times. That one mark carries two clauses, existence of the display and uniqueness of the projection-valued measure read off it, and the second clause is what §6 spends when it calls PP the measure of A^\hat A and builds the Born rule on it. The rigged Hilbert space in §7.5, quoted for orientation only, since every manipulation with x\ket x and k\ket k in this book abbreviates §7.4's procedure rather than leaning on that framework. The converse half of Stone in §9.3, which is one implication of an equivalence whose other implication was proved here. That is three, and nothing else in this chapter is asserted without being derived. Three marks standing elsewhere are leaned on and cited rather than re-raised: Chapter 4.3 §2.3's measure construction, wherever an integral appears, Chapter 4.4 §2.3's closed graph theorem, in §2.4, where it is what excludes a fourth way for (4.5.2) to fail, and Chapter 4.4 §6.2's classification, inside the quoted proof of Stone's converse. Completeness of the space was spent twice in the open. Once in §2.7's grind box, where a Cauchy sequence has to converge for the image of A^λ\hat A-\lambda to be closed. Once in §5.4, where Chapter 4.3 §7.3's equivalence turns statement (d) into statement (b), and that equivalence needs §7.2, which needs completeness. Dominated convergence was spent five times, in §2.5, §5.4, §6.1, §9.1 and §9.2, and naming the dominating function each time is what makes those uses checkable.

Where this gets spent. Chapter 4.6 takes three things from here and could not start without them. The first is Stone's theorem from §9, which turns a strongly continuous unitary flow into a self-adjoint generator and so produces itψ=H^ψ\ii\hbar\,\partial_t\ket\psi=\hat H\ket\psi rather than assuming it. The second is the multiplication form of §3, which is what lets an operator built as p^2/2m+V(x^)\hat p^{2}/2m+V(\hat x) be discussed before anyone has solved anything. The third is §7's meaning for x\ket x, which is what the position representation is. Chapter 4.7 takes Worked example 1, since a scattering state is a state in the continuous spectrum of p^2/2m\hat p^{2}/2m and §8.2 is explicit that an expansion in bound states alone misses them. Chapter 4.8 expands in the Hermite functions freely, which §5.4 is what earns. Chapter 4.9's uncertainty relation says that no preparation narrows both spreads below the floor the commutator sets, and for position and momentum that floor is /2\hbar/2 in every state. That no joint distribution lies beneath the two spectral measures at all is a separate statement, not derivable from that inequality, and §6.5 leaves it open rather than assuming it: it belongs to Chapters 4.11 and 4.20. Chapter 4.13 does not treat hydrogen the way §5.5 treats the oscillator, and §8.2 says why: the Coulomb spectrum carries the whole of [0,)[0,\infty) above its bound states, so those states span the point-spectrum subspace and are not a basis of L2(R3)L^{2}(\R^{3}). Self-adjointness there is a separate theorem that this book does not build, it has to be settled before an eigenfunction is worth looking for, and Chapter 4.13 carries a mark for it. Chapter 4.17 needs the density of states that Worked example 1 computed. What this chapter does not do, and says so in §7.6, is build the general theory of distributions. That is Chapter 5.4, and until then every appearance of x\ket x is an abbreviation with a finite procedure behind it.