Part IV · Quantum Mechanics — Chapter 4.4

Domains, and the Adjoint's Domain

In infinite dimensions an operator is not a formula. It is a formula together with a domain, the domain is forced rather than chosen, and the word "Hermitian" turns out not to have been enough.

Where we are

Chapter 4.3 built the space. Square-integrable functions, with two functions that agree almost everywhere counted as one vector, form a complete inner-product space, and the Fourier modes are a genuine orthonormal basis of it rather than an assumption. That is a space. It has no operators on it yet.

This chapter puts them there, and the first thing it has to do is admit that the finite-dimensional recipe does not survive the move. In Chapter 0.5 an operator was a matrix, a matrix acts on every vector there is, and a Hermitian matrix is self-adjoint by the same stroke of the pen. None of those three sentences is true here. The derivative cannot act on every square-integrable function, because differentiating one need not produce another. So it comes with a restricted set of functions it is allowed to act on, and once you allow that, the adjoint acquires a restricted set of its own, fixed by the definition, with no reason at all to be the same one.

Here is the route. Section 1 reads back the sentence Chapter 0.5 left standing and says which of its four items belongs to this chapter and which to Chapter 4.5. Section 2 measures how badly the derivative misbehaves and then proves, from one quoted theorem, that the misbehaviour cannot be designed away: a symmetric operator defined on the whole space is bounded, so an unbounded observable must be restricted. Section 3 defines the domain and the adjoint's domain carefully. Section 4 separates symmetric from self-adjoint and corrects the second postulate of Chapter 4.2 in the section that chapter names. Section 5 is the payoff, and it is arithmetic you can check by hand: momentum on the line, on an interval and on a half-line, with the boundary term from integration by parts kept rather than waved through. Only then does Section 6 quote the theorem that counts the answers. Section 7 does the same for the particle in a box, where the answer is richer and one widely repeated statement about it is wrong.

Conventions. The inner product is linear in its second slot, as Chapter 0.5 §1.1 chose it, and u,v=uv\avg{u,v}=\int\overline{u}\,v throughout. Every integral is Chapter 4.3's. Two results are quoted rather than proved, and both are marked where they arrive: the closed graph theorem in §2.3, and von Neumann's classification of self-adjoint extensions in §6.2. There are no others, and one mark standing elsewhere is leaned on once, Fubini's theorem from Chapter 0.2 §4.2, inside the grind box of §5.2. The closing brick says where completeness was spent and what each mark bought.

Tools you'll need  — Chapter 4.3 above all: §5 for L2L^{2} and the quotient that repaired the third inner-product axiom, §5.4 for Cauchy–Schwarz carried over intact, §6 for completeness, §7.4 for the dense set of step functions, and §8 for the Fourier basis on an interval. Chapter 0.5 §4 for the definition of the adjoint, which this chapter reads more carefully than that chapter needed to, §6 for the spectral theorem in finite dimensions, and its closing paragraph, which is the debt §1 collects. Chapter 0.2 §3.2 for integration by parts and for the sentence that sent its boundary term here. Chapter 0.6 §2 for the lemma that every linear map on Rn\R^{n} is bounded, and for the clause naming the derivative as the operator that breaks it. Chapter 4.2 §4 for the second postulate and the warning attached to it, §7 for time evolution, and §8 for the three lines that force infinite dimension. Chapter 0.9 §5 for what integration by parts means when the derivative is taken away from the function and given to the test function instead.

1 · The four places Chapter 0.5 used finite dimension

Here is where this section is going. Chapter 0.5 finished by naming four steps in which its proofs leaned on the dimension being finite, and it said the bill would come due in this chapter and the next. We read that sentence back and take the four in order. By the end you will know which one belongs here, which belong to Chapter 4.5, and why the repair could not be done in a single chapter. The fourth item is the one to watch, because it is the only one with no finite-dimensional shadow at all.

1.1 · The sentence, read back

Chapter 0.5's last paragraph is short enough to quote whole, and it is worth having in front of you rather than in memory:

"One thing remains outstanding, and it is worth naming so that you notice when it is paid. Every proof in this chapter used finite dimension: in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective. Quantum mechanics happens in infinite dimensions. Chapters 4.4 and 4.5 are where the bill comes due."

Chapter 4.2 §8 then made the second sentence unavoidable rather than atmospheric. If [x^,p^]=iI^[\hat x,\hat p]=\ii\hbar\hat I held on a space of finite dimension nn, taking the trace of both sides would give 00 on the left, because the trace of a commutator of matrices vanishes, and in\ii\hbar n on the right. Three lines, no physics beyond the commutator itself, and quantum mechanics is infinite-dimensional before a single measurement has been described. Chapter 4.3 built the space that follows from that. This chapter and the next have to rebuild the operators.

1.2 · The four, one at a time

Take them in Chapter 0.5's own order, because the order is roughly the order of increasing damage.

The induction. Chapter 0.5 §6.3 proved the spectral theorem by finding one eigenvector, restricting the operator to the orthogonal complement of that eigenvector, and repeating on a space one dimension smaller. The recursion halts because the dimension is a positive integer and cannot decrease forever. Strip out finiteness and two things break rather than one. The recursion no longer halts, which is the visible failure. The first step also fails, which is the serious one: there need not be any eigenvector to start from. The operator "multiply by xx" on L2(R)L^{2}(\R) has none at all, since xψ(x)=λψ(x)x\psi(x)=\lambda\psi(x) forces ψ\psi to vanish wherever xλx\ne\lambda, and a function supported at one point is the zero vector of Chapter 4.3 §5.3. That failure is Chapter 4.5's, and it is why that chapter has to widen the word spectrum before it can state a theorem at all.

Rank–nullity. The identity dimkerA^+dimimA^=dimV\dim\ker\hat A+\dim\operatorname{im}\hat A=\dim V is not so much false in infinite dimensions as empty, because both sides can be infinite and an equation between two infinities constrains nothing. What that identity was for was converting a statement about the kernel into a statement about the image, and the replacement is not a theorem about dimensions at all. It is the observation that an image can be dense without being closed, so that coming arbitrarily close to every vector and reaching every vector are two different achievements, and counting cannot tell them apart. Chapter 4.5 is where that becomes the difference between the parts of the spectrum, because A^λ\hat A-\lambda can be one-to-one with a dense image and still have no bounded inverse.

The interchange of sums. Chapter 0.5 rearranged double sums freely, which finitely many terms always permit. The tempting move is to hand this item to Chapter 4.3's convergence theorems and walk away, but those license the exchange of a limit with an integral, whereas rearranging jkajk\sum_j\sum_k a_{jk} exchanges two sums, and an exchange of two sums is an exchange of two integrals in disguise. Chapter 4.3 §4.1 is explicit about that one: exchanging two integrals is Fubini's theorem, and it stays where Chapter 0.2 §4.2 marked it. What pays this item without raising anything new is the corollary the same section proves for non-negative terms, khkdμ=khkdμ\int\sum_k h_k\,\dd\mu=\sum_k\int h_k\,\dd\mu. Apply it to the functions hk=jajk1[j,j+1)h_k=\sum_j a_{jk}\mathbf 1_{[j,j+1)}, whose integral is jajk\sum_j a_{jk}, and the two sides of the corollary are the two orders of summation. That settles it whenever every ajk0a_{jk}\ge0, and a double sum that converges absolutely follows by splitting it into positive and negative parts the way Chapter 4.3 §3 splits a function before integrating it. This item is paid, and the material that pays it was in place before this chapter opened.

Injective does not imply surjective. A one-to-one linear map of a finite-dimensional space to itself is onto, and Chapter 0.5's grind box on the converse of U^=eiA^\hat U=\ee^{\ii\hat A} used exactly that, in the words "being injective on a finite-dimensional space it is onto, hence unitary on WW". It is false here, and the counterexample is small enough to hold in one hand. Take an orthonormal basis e1,e2,e3,e_1,e_2,e_3,\ldots of the space, which Chapter 4.3 §7.4 says can always be listed, and let S^\hat S move every basis vector one place along that list:

S^en  =  en+1,soS^(n1cnen)  =  n1cnen+1. \hat S\,e_n \;=\; e_{n+1}, \qquad\text{so}\qquad \hat S\Big(\sum_{n\ge1}c_n e_n\Big) \;=\; \sum_{n\ge1}c_n e_{n+1}. (4.4.1)

Nothing non-zero is sent to zero, so S^\hat S is injective, and by Parseval it preserves the norm exactly, since the same coefficients appear in the answer. Yet e1e_1 is not in its image, because every vector S^ψ\hat S\psi has coefficient zero against e1e_1. So S^\hat S is one-to-one and not onto, and in finite dimensions no such map exists. Hold on to this one. It returns in §5.5 wearing physical clothes, as the reason a particle confined to a half-line has no momentum observable, and again in Worked example 3, where S^\hat S turns out to be a translation with nowhere to come from.

1.3 · Which chapter pays which

The division follows from the four items and is worth stating before you meet it rather than after. This chapter is about the first half of the word "self-adjoint": what an operator is allowed to act on, and what its adjoint is therefore allowed to act on. That is the machinery the induction's first step will need, and it is a prerequisite for the spectrum rather than a part of it. Chapter 4.5 is about the values. It widens spectrum so that an operator with no eigenvectors still has one, states the theorem that replaces A^=U^D^U^\hat A=\hat U\hat D\hat U^{\dagger}, and checks it in every case this book uses.

Neither half stands alone. You cannot ask what the spectrum of an operator is until you have said which operator, domain included, and Chapter 4.5's theorem has self-adjoint in its hypothesis, which is a statement about domains and nothing else. So the bill Chapter 0.5 named is settled across two chapters, in that order, for that reason.

In plain terms 4.4.1

A chapter of Part 0 built the whole theory of matrices and eigenvectors, and it ended with an honest confession: four of its arguments had used the fact that there were only finitely many directions to work with. Quantum mechanics does not have finitely many directions, and the previous chapter proved that in three lines rather than asserting it. So the confession has to be settled.

Of the four, one has already been paid. The freedom to rearrange a double sum was bought in the previous chapter: one of its convergence theorems has a corollary that lets a sum be moved through an integral whenever the terms are all positive, and rearranging a double sum is that same move in disguise. Two more belong to the next chapter, and they are both about the same thing: in a space of functions an operator may have no eigenvectors whatsoever, and the machinery built for matrices starts by finding one.

The item that belongs here is the strangest of the four, and it has no counterpart in the finite world at all. A map can be perfectly one-to-one and still miss part of the space it maps into. Nothing is lost and yet something is unreachable. In a room with finitely many directions that cannot happen, so the situation has no picture attached, and the way to get one is to watch it happen. It happens twice in this chapter, and the second time it is the reason a particle held on one side of a wall has no measurable momentum at all.

2 · Bounded, unbounded, and why the derivative cannot be tamed

Section 1 closed Chapter 0.5's account. The debt this section collects is a different one, owed to Chapter 0.6, which proved that every linear map on Rn\R^{n} is bounded and named the derivative as the map that breaks the proof once the dimension is infinite. The plan for this section is to make one word precise and then use it to prove something that sounds like an inconvenience and is really a structural fact. We define what it means for an operator to be bounded and show that it is the same thing as being continuous. We then measure the derivative and find that it is neither. Finally we quote one theorem and get from it, in two lines, the statement that decides the shape of this whole chapter: an unbounded operator that we want to be an observable cannot be defined on the whole space, no matter how cleverly we try.

2.1 · The operator norm, and what it measures

Chapter 0.6 §2 proved that every linear map on Rn\R^{n} satisfies A^vCv\abs{\hat A\vv v}\le C\abs{\vv v} for some constant depending only on A^\hat A, and the proof wrote v\vv v in a basis and used Cauchy–Schwarz on the finitely many coefficients. The property that proof establishes has a name and a number attached. For a linear A^\hat A defined on a subspace D\mathcal{D} of a Hilbert space, put

A^  =  supψD, ψ0 A^ψψ, \norm{\hat A} \;=\; \sup_{\psi\in\mathcal{D},\ \psi\ne0}\ \frac{\norm{\hat A\psi}}{\norm{\psi}}, (4.4.2)

which is the largest factor by which A^\hat A can stretch anything. Call A^\hat A bounded when this number is finite and unbounded when it is not. The supremum in (4.4.2) is over a set of non-negative real numbers, so it always exists in the extended sense, and the only question is whether it is a number.

Before using the word it is worth knowing what it means, because "bounded" sounds like a technical convenience and it is not one. It is continuity, under another name. One direction takes a line: if A^\norm{\hat A} is finite then

A^ψA^ϕ  =  A^(ψϕ)    A^ψϕ, \norm{\hat A\psi-\hat A\phi} \;=\; \norm{\hat A(\psi-\phi)} \;\le\; \norm{\hat A}\,\norm{\psi-\phi}, (4.4.3)

so nearby inputs have nearby outputs, with a uniform rate. The converse takes three lines and is worth having because it is what makes the word carry weight. Suppose A^\norm{\hat A} is infinite. Then for each nn there is a unit vector ψn\psi_n in the domain with A^ψnn\norm{\hat A\psi_n}\ge n. Scale it down by setting ϕn=ψn/n\phi_n=\psi_n/\sqrt n. Now ϕn=1/n0\norm{\phi_n}=1/\sqrt n\to0, so the inputs march into the origin, while A^ϕnn/n=n\norm{\hat A\phi_n}\ge n/\sqrt n=\sqrt n\to\infty, so the outputs run away. An unbounded operator is therefore discontinuous, and discontinuous at every point rather than at some awkward ones, since linearity carries the failure at the origin everywhere else.

That is the honest content of the word. An unbounded observable is one whose value cannot be controlled by controlling the state to within a small error in the norm of the space, and the next subsection shows that momentum is one.

2.2 · The derivative is unbounded, exhibited

Chapter 0.6 §2 named the offender in advance, saying that "in infinite dimensions (Chapter 4.4) linear maps can be unbounded, ddx\dv{}{x} being the standard offender, and a great deal of quantum mechanics' technical difficulty descends from exactly that." The exhibit takes one family of functions. Work on L2[0,L]L^{2}[0,L] and take the exponential modes, with kLkL a multiple of 2π2\pi, which are the basis Chapter 4.3 §8 built on [π,π][-\pi,\pi] with the interval rescaled:

ψk(x)=eikx,ψk2=0L ⁣1dx=L,p^ψk=idψkdx=kψk. \psi_k(x)=\ee^{\ii kx}, \qquad \norm{\psi_k}^{2}=\int_0^{L}\!1\,\dd x=L, \qquad \hat p\,\psi_k=-\ii\hbar\,\dv{\psi_k}{x}=\hbar k\,\psi_k . (4.4.4)

Read the three parts of (4.4.4) together and the conclusion is immediate. The norm of ψk\psi_k does not depend on kk at all, while the norm of p^ψk\hat p\psi_k is kL\hbar\abs{k}\sqrt L, so the ratio in (4.4.2) is k\hbar\abs k and kk runs over 2πZ/L2\pi\mathbb{Z}/L without limit. The operator norm of p^\hat p is infinite. Every one of these functions is smooth, periodic and as well behaved as a function can be, so the failure is not caused by anything pathological in the inputs.

The same conclusion holds on the whole line, and it needs one extra step because the modes themselves are not square-integrable there. Fix any smooth gg that vanishes outside a bounded interval and has g=1\norm g=1, and set ψk=eikxg(x)\psi_k=\ee^{\ii kx}g(x). The modulus of ψk\psi_k is the modulus of gg, so ψk=1\norm{\psi_k}=1 for every kk. Differentiating the product gives p^ψk=kψkieikxg\hat p\psi_k=\hbar k\psi_k-\ii\hbar\ee^{\ii kx}g', and the triangle inequality then bounds the second term away rather than adding to it:

p^ψk    kψkg  =  kg  k  . \norm{\hat p\,\psi_k} \;\ge\; \hbar\abs{k}\,\norm{\psi_k} - \hbar\,\norm{g'} \;=\; \hbar\abs k-\hbar\norm{g'} \;\xrightarrow[\abs k\to\infty]{}\;\infty . (4.4.5)

So p^\hat p is unbounded on L2(R)L^{2}(\R) too, and the reason is not subtle. The number k\hbar k in (4.4.4) is a momentum, the operator norm of p^\hat p would be the largest momentum a state can have, and there is no largest momentum. Unboundedness is what a physically unbounded quantity looks like when you write it as an operator. The same argument, run on multiplication by xx, says the same thing about position, and Worked example 1 runs it.

Familiar ground — why you smooth a curve before you differentiate it

Take a measured concentration–time curve and add to it a small ripple, εsin(kx)\varepsilon\sin(kx), with ε\varepsilon tiny and kk large. In the norm this book uses, the contaminated curve differs from the clean one by about ε\varepsilon, independently of kk, because the amplitude of the ripple is all that the norm sees. Differentiate both curves and the difference between the answers is εkcos(kx)\varepsilon k\cos(kx), whose size is εk\varepsilon k. The mathematics is identical to (4.4.4), with the same family of functions doing the same job.

That is the whole of why nobody differentiates raw data. The error in the input is bounded and the error in the output is not, so any high-frequency noise, however small in amplitude, arrives in the derivative multiplied by its frequency. Smoothing first is not cosmetic tidying. It is the restriction of the derivative to a set of inputs on which the operation is stable, which is exactly the move §3 is about to make and give a name to.

Two terms of the analogy have no counterpart, and both are worth saying out loud. The first is the freedom. In data analysis you choose the smoothing, and choosing more of it is always available. In quantum mechanics the operator is fixed by physics and the restriction has to be chosen so that the operator stays an observable, which turns out to leave very little freedom and sometimes none. That is §5.

The second break is the load-bearing one, because it is about what the restriction buys. Smoothing genuinely buys stability: the smoothed derivative has a worst case, so a small error going in stays a small error coming out. Restricting the domain buys nothing of the kind. Every mode in (4.4.4) lies in the domain §3.2 is about to write down, so the ratio in (4.4.2) is still k\hbar\abs k and p^\hat p is exactly as unbounded on its domain as it was on all of L2L^{2}. The restriction is not a cure for the instability, and §2.3 is where you find out that it is compulsory anyway.

2.3 · Unboundedness cannot be designed away

The natural response to §2.2 is to look for a better formulation. The hope would be that momentum has been written down clumsily, and that some equivalent operator, defined on every state, does the same physical job while staying continuous. This subsection closes that door, and closes it completely, using one theorem that this book quotes rather than proves.

⚑ Quoted, not derived — the closed graph theorem

We use, without proof, the following. Let HH be a complete inner-product space and let T^\hat T be a linear map defined on all of HH, with values in HH. Suppose T^\hat T has this property, which we will call the closed-graph condition:

whenever ψnψ\psi_n\to\psi in HH and the images T^ψn\hat T\psi_n converge to some ϕ\phi, the limit ϕ\phi is T^ψ\hat T\psi.

Then T^\hat T is bounded.

What the hypothesis is doing. Continuity would say that T^ψn\hat T\psi_n converges and converges to the right thing. The closed-graph condition assumes only the second half: it says nothing about whether T^ψn\hat T\psi_n converges, and only rules out its converging to the wrong place. The theorem is that on a complete space the weaker demand implies the stronger one, and completeness is not decoration in that sentence. The result is false without it, and this is the first of the two places this chapter spends what Chapter 4.3 §6 proved. The standard proof runs through the Baire category theorem in about a page, and Rudin's Functional Analysis is where to read it. This book uses the theorem exactly twice, in the two lines below and once more in Chapter 4.5 §2.4, where it is what rules out a fourth way for a point of the spectrum to fail, and never again.

Now the two lines. Let A^\hat A be defined on all of HH and satisfy A^u,v=u,A^v\avg{\hat Au,v}=\avg{u,\hat Av} for every pair of vectors, which is Chapter 0.5's Hermitian condition with nothing added. Call an operator satisfying that condition symmetric. The word is worth having now, before it is needed in anger: §4.1 states it again once an operator has a domain to state it on, and the whole force of this chapter is in the difference between the two statements. We check the closed-graph condition. Suppose ψnψ\psi_n\to\psi and A^ψnϕ\hat A\psi_n\to\phi. For an arbitrary vv in the space, move A^\hat A across, take the limit, and move it back:

v,ϕ  =  limnv,A^ψn  =  limnA^v,ψn  =  A^v,ψ  =  v,A^ψ, \avg{v,\phi}\;=\;\lim_{n\to\infty}\avg{v,\hat A\psi_n}\;=\;\lim_{n\to\infty}\avg{\hat Av,\psi_n}\;=\;\avg{\hat Av,\psi}\;=\;\avg{v,\hat A\psi}, (4.4.6)

where each limit passes through the inner product because Cauchy–Schwarz makes the inner product continuous in each slot, as Chapter 4.3 §5.4 transferred it. So v,ϕA^ψ=0\avg{v,\phi-\hat A\psi}=0 for every vv, and choosing v=ϕA^ψv=\phi-\hat A\psi gives ϕ=A^ψ\phi=\hat A\psi. The closed-graph condition holds, the quoted theorem applies, and A^\hat A is bounded. That is Hellinger–Toeplitz, and it is worth stating on its own line because the rest of the chapter is a consequence of it.

The reframing this chapter runs on

A symmetric operator defined on all of a Hilbert space is bounded. Contrapositive: an operator that is unbounded, and that we intend to be an observable, cannot be defined on all of the space. Not "is awkward to define". Cannot.

So the restricted domain of p^\hat p is not a piece of caution, a smoothness assumption, or a convenience that a more careful formulation would remove. It is forced, by a theorem, from two facts we already have: that momentum is unbounded, which §2.2 exhibited, and that an observable satisfies Chapter 4.2's postulate P2. Everything that follows in this chapter is bookkeeping about a restriction we are not free to decline.

In plain terms 4.4.2

Some operations stretch, and the useful question is whether there is a worst case. If there is a single number such that no input is ever stretched by more than that factor, the operation is stable: a small error going in produces a small error coming out, and the two smallnesses are tied together by that one number. If there is no such number, small errors can be amplified without limit, and stability is gone.

Differentiation has no such number, and the reason is worth keeping rather than the proof. Wiggles of any frequency have the same size as functions and get bigger the faster they wiggle, so the faster the wiggle the larger the derivative, with no ceiling. Anyone who has tried to take a derivative of noisy measurements has met this in person. It is also the correct statement about momentum, because there is no largest momentum a particle can have, and an operator representing an unlimited quantity had better be unlimited.

Then comes the move that decides everything else. You might hope to write momentum down some other way and recover stability. A theorem says no. Any operation that acts on every state whatever, and that has the symmetry a measurable quantity must have, is automatically stable. Turn that around: an unstable measurable quantity cannot act on every state. So the list of states that momentum is allowed to act on is not a hedge and not a technicality. It is compulsory, and the rest of the chapter is about what it costs.

a natural place to stop  ·  the restriction is compulsory; what follows is the vocabulary for it

3 · Domains, and the adjoint's domain

What follows is the vocabulary the rest of the chapter uses, and it is short. An operator is promoted from a formula to a pair, a formula and a set of vectors it acts on. We fix the domain of p^\hat p and check that it is dense, saying what "dense" is for. Then we re-read Chapter 0.5's definition of the adjoint with that in hand, and find that the definition hands the adjoint a domain of its own, determined but not chosen, with no reason to agree with the one we started from. The last subsection draws the consequence that makes the whole of §5 predictable in advance.

3.1 · An operator is a formula and a domain

From here on, an operator on a Hilbert space HH means two things given together: a subspace dom(A^)H\operatorname{dom}(\hat A)\subseteq H, called its domain, and a linear map from that subspace into HH. Two operators are the same operator when the formulae agree and the domains agree. If the domains differ, they are different operators, even when the formula is the same symbols in the same order.

That convention looks pedantic on first meeting and it is the entire subject of the chapter. In finite dimensions it costs nothing, because a matrix acts on everything and there is only one domain available, which is why Chapter 0.5 never mentioned the word. Here the choice is real, the choice changes the physics, and §7 will show two operators with the same formula, both perfectly legitimate, whose energy levels are different numbers.

When B^\hat B has a larger domain than A^\hat A and agrees with A^\hat A everywhere on dom(A^)\operatorname{dom}(\hat A), we say B^\hat B extends A^\hat A. That word does the work in §6, where the question is how many self-adjoint extensions a given operator has.

3.2 · The domain of p^\hat p, and what "dense" is for

Momentum needs a domain, and §2.3 says the domain cannot be everything. The natural choice is the largest set on which the formula makes sense and lands back in the space. Take

dom(p^)  =  {ψL2 : ψ is the integral of a function ψ,  and ψL2},p^ψ  =  iψ. \begin{aligned} \operatorname{dom}(\hat p) \;&=\; \big\{\,\psi\in L^{2}\ :\ \psi \text{ is the integral of a function } \psi',\ \text{ and } \psi'\in L^{2}\,\big\},\\[3pt] \hat p\,\psi \;&=\; -\ii\hbar\,\psi' . \end{aligned} (4.4.7)

The phrase "is the integral of a function" is the fundamental theorem of Chapter 0.2 §2 used as a definition rather than as a conclusion, and it is what lets ψ\psi be non-smooth while still having a derivative that Chapter 4.3's integral can handle. A function with a jump in it fails the condition, because no integral has a jump. The condition ψL2\psi'\in L^{2} is what makes p^ψ\hat p\psi a vector of the space rather than merely a function.

One property of this set matters more than any other, and it is the reason the next subsection works at all. The domain is dense: every vector of L2L^{2} is a limit of vectors in it. On an interval this is one sentence, since the modes of (4.4.4) all lie in the domain and Chapter 4.3 §8 proved that finite combinations of them reach everything. On the whole line the argument is Chapter 4.3's as well, and it is already written down there. That chapter's §7.4 showed that step functions are dense, and its §8.1 then replaced the vertical sides of an indicator by straight ramps of width δ\delta, at a cost of at most 2δ2\delta in squared distance. Those trapezoids are exactly what is wanted here: each is continuous, each is the integral of its own step-function derivative, and that derivative is bounded and supported on a bounded set, hence in L2L^{2}. So every trapezoid is in (4.4.7), and density transfers along Chapter 4.3's own chain.

Now the reason to care. Density is not a technical hygiene condition. It is what makes the adjoint exist as a function of uu at all, and the next subsection is where you can watch it do that job.

3.3 · The adjoint, with a domain of its own

Chapter 0.5 §4 defined the adjoint by the relation A^u,v=u,A^v\avg{\hat A^{\dagger}u,v}=\avg{u,\hat Av}, and in finite dimensions that relation determines a matrix, namely the conjugate transpose, with nothing further to say. Read the same relation here, where A^\hat A acts only on its own domain, and it stops being a formula for A^\hat A^{\dagger} and becomes a demand that a given uu may or may not be able to meet:

A^u,  v  what the adjoint must produce    =  u,  A^v  a number you can already compute  for every vdom(A^). \begin{gathered} \ann{\avg{\hat A^{\dagger}u,\;v}}{what the adjoint must produce} \;=\; \ann{\avg{u,\;\hat Av}}{a number you can already compute} \\[8pt] \text{for every } v\in\operatorname{dom}(\hat A). \end{gathered} (4.4.8)

The two labels under (4.4.8) are the whole of the difficulty. What is being demanded is a single vector A^u\hat A^{\dagger}u whose inner product against every vv in the domain reproduces those computable numbers, and sometimes no vector in the space does that job. So the definition sorts the vectors of HH into those for which the demand can be met and those for which it cannot, and admitting the first sort and refusing the second is what defines dom(A^)\operatorname{dom}(\hat A^{\dagger}):

udom(A^)some wH satisfiesw,v=u,A^v  for every vdom(A^), \begin{aligned} u\in\operatorname{dom}(\hat A^{\dagger}) \quad&\Longleftrightarrow\quad \text{some } w\in H \text{ satisfies}\\[3pt] &\qquad\quad \avg{w,v}=\avg{u,\hat Av}\ \text{ for every } v\in\operatorname{dom}(\hat A), \end{aligned} (4.4.9)

with A^u\hat A^{\dagger}u defined to be that ww. Here is where density earns its keep. If two vectors w1w_1 and w2w_2 both satisfied (4.4.9), then w1w2,v=0\avg{w_1-w_2,v}=0 for every vv in the domain, and because the domain is dense and the inner product is continuous, the same holds for every vv in the whole space. Taking v=w1w2v=w_1-w_2 gives w1=w2w_1=w_2. Without density there would be a direction the test vectors never probe, ww could be shifted along it freely, and A^u\hat A^{\dagger}u would not be a well-defined vector. That is the whole job of the word dense, and it is done here and nowhere else.

This is the sentence the chapter turns on, so it is worth saying without any equation attached. The definition of the adjoint determines its domain. Nobody chooses it, and there is no reason whatever for it to coincide with the domain we started from. In finite dimensions the two coincide because both are the whole space and there is nothing to compare. Here they are two different subspaces produced by two different considerations, and the rest of the chapter is the study of when they happen to agree.

3.4 · The two domains move in opposite directions

One structural fact makes everything in §§5 to 7 predictable before you compute anything, and it follows from (4.4.9) by reading the words rather than the symbols. Suppose B^\hat B extends A^\hat A, so that dom(A^)dom(B^)\operatorname{dom}(\hat A)\subseteq\operatorname{dom}(\hat B). The condition on uu in (4.4.9) then has to hold for more vectors vv, which is a harder test, so fewer vectors uu pass it. In symbols,

dom(A^)dom(B^)dom(B^)dom(A^). \operatorname{dom}(\hat A)\subseteq\operatorname{dom}(\hat B) \qquad\Longrightarrow\qquad \operatorname{dom}(\hat B^{\dagger})\subseteq\operatorname{dom}(\hat A^{\dagger}). (4.4.10)

Enlarging an operator shrinks its adjoint, and the two move towards each other. That single observation is why the family of self-adjoint operators built from a given formula is a family at all, rather than a single answer or none: you start with a domain too small and an adjoint domain too big, and you enlarge the first, which shrinks the second, until they meet. Whether they can be made to meet, and in how many ways, is the question §6 answers and §5 answers first by hand.

In plain terms 4.4.3

An operator is no longer a rule. It is a rule together with a list of the states it is allowed to be applied to, and changing the list changes the operator even though the rule is written with the same symbols. That sounds like fussiness and it is the opposite. The list is where the physics of a boundary lives, and two lists with the same rule will shortly give a box two different sets of energy levels.

The partner of an operator, its adjoint, was defined in Part 0 by a single requirement: moving the operator from one side of an inner product to the other must not change the answer. In a space of finitely many directions that requirement pins down a matrix and there is nothing further to discuss. In a space of functions the same requirement has to be read as a question asked of each state separately: is there a vector that does this job for this state? For some states there is and for others there is not, so the requirement quietly hands the partner its own list, which nobody chose and which has no reason to match the first.

One more thing, and it is the thing that makes the rest predictable. Lengthen the first list and the second list gets shorter, because there are now more conditions for a state to satisfy. The two lists move towards each other. An operator worthy of representing a measurement is one where they coincide, and the interesting question is how many ways there are to make that happen.

4 · Symmetric is not self-adjoint

The two conditions in the title are the same condition in finite dimensions, and Chapter 4.2 §4 used the finite-dimensional word. This section separates them, says what each one buys, corrects the second postulate as that chapter announced would happen here, and is honest about what the difference is not: it is invisible to every test you would naturally run, including the test that measured values come out real. Section 5 then makes the difference bite on an operator you can check by hand.

4.1 · Two conditions, and how far apart they are

Section 2.3 has met the first of the two words already, for an operator that acts on everything. With domains in hand it can be said in the form the rest of the chapter needs. Call A^\hat A symmetric when moving it across the inner product changes nothing, for vectors it is allowed to act on:

A^u,v  =  u,A^vfor all u,vdom(A^). \avg{\hat Au,\,v} \;=\; \avg{u,\,\hat Av} \qquad \text{for all } u,v\in\operatorname{dom}(\hat A). (4.4.11)

Compare (4.4.11) with the definition of the adjoint in (4.4.9) and read off what symmetry says in that language. It says that every uu in dom(A^)\operatorname{dom}(\hat A) passes the test defining dom(A^)\operatorname{dom}(\hat A^{\dagger}), with w=A^uw=\hat Au as the witness. So symmetry is exactly the statement that

dom(A^)dom(A^)andA^u=A^u  for udom(A^). \operatorname{dom}(\hat A)\subseteq\operatorname{dom}(\hat A^{\dagger}) \qquad\text{and}\qquad \hat A^{\dagger}u=\hat Au \ \text{ for } u\in\operatorname{dom}(\hat A). (4.4.12)

In the vocabulary of §3.1, a symmetric operator is one whose adjoint extends it. Call A^\hat A self-adjoint when the containment in (4.4.12) is an equality, so that A^=A^\hat A^{\dagger}=\hat A with the domains included and not merely the formulae. The gap between the two words is therefore a gap between two subspaces, and measuring it on one operator after another is what the rest of the chapter does. Each measurement is of a stated operator, since the gap belongs to the domain and not to the formula. Section 5.3 measures it for momentum on the line, on the largest domain the formula allows, and finds no gap at all. Section 5.4 measures it on a bounded interval, for momentum pinned to zero at both ends, and finds a gap of two complex dimensions. It then closes that gap in a circle of ways, and each of them is an operator with no gap left. Section 5.5 measures it on a half-line, for the largest symmetric domain available there, finds a gap of exactly one complex dimension, and shows that nothing closes it.

In finite dimensions this distinction cannot be drawn, because both domains are the whole space and (4.4.12) is an equality the moment it is a containment. Chapter 0.5 was therefore right to use one word, and Chapter 4.2 inherited it correctly for the finite-dimensional systems that chapter's examples used.

4.2 · The second postulate, corrected

Chapter 4.2 §4 stated P2 as an observable is a Hermitian operator, and attached a warning saying the word would be sharpened here, in this section, for a reason that is physical rather than pedantic. The correction is now available and it is one line.

⚠ P2, in the form it keeps for the rest of the book

An observable is a self-adjoint operator: A^=A^\hat A^{\dagger}=\hat A with dom(A^)=dom(A^)\operatorname{dom}(\hat A^{\dagger})=\operatorname{dom}(\hat A). Symmetric is not enough. Everywhere Chapter 4.2 wrote "Hermitian", read this instead, and nothing in that chapter's arguments changes, because every operator it worked with acted on a space of finite dimension where the two words agree.

The postulate is no more assumed than it was. What has changed is which mathematical condition the assumption is attached to, and Chapter 4.2's own box said this would happen and named this section. The physical reason is §5.5: momentum on a half-line is symmetric and is not self-adjoint, and there is no repair. If P2 said only "symmetric", that operator would qualify as an observable and the theory would be committed to a measurable quantity for which Chapter 4.5's spectral theorem gives nothing and Chapter 4.5 §9's Stone's theorem gives no time evolution.

4.3 · What the difference is not visible to

It is worth naming the difficulty out loud here, because the difference between the two words is invisible to every test a physicist would think to apply, and a reader who does not know that will reasonably suspect the distinction of being empty.

Consider the property that made Hermitian operators attractive in the first place. If A^\hat A is symmetric and A^ψ=λψ\hat A\psi=\lambda\psi for a non-zero ψ\psi in the domain, then (4.4.11) with u=v=ψu=v=\psi gives λψ2=λψ2\lambda\norm\psi^{2}=\overline\lambda\norm\psi^{2}, so λ\lambda is real. Chapter 0.5 §6.1's argument transfers with no change and needs only symmetry. Symmetric operators also have orthogonal eigenvectors for distinct eigenvalues, by Chapter 0.5 §6.2's argument, again unchanged. And every expectation value ψ,A^ψ\avg{\psi,\hat A\psi} is real for the same reason. Chapter 4.2 §4.1 collected three theorems as the payoff of P2, and the first two of them are already true for a merely symmetric operator.

The third one is where the words part company, and it is the one you cannot check by making a measurement. Chapter 4.2 §4.1's third theorem was every state is a superposition of outcomes, and its proof there was Chapter 0.5 §6.3's orthonormal basis of eigenvectors. That is the spectral theorem, and the spectral theorem is exactly what symmetry fails to buy. Two structures the rest of Part IV stands on go with it.

  • The spectral theorem. Chapter 4.5's hypothesis is self-adjointness, not symmetry, and there are symmetric operators for which the conclusion is false. Without it there is no resolution of the identity, so no way to write a state as a superposition of the outcomes of a measurement, so no Born rule for that quantity.
  • Stone's theorem. Chapter 4.5 §9 gives eiH^t/\ee^{-\ii\hat Ht/\hbar} as a unitary flow exactly when H^\hat H is self-adjoint. A merely symmetric Hamiltonian does not generate one, so the state after a finite time is undefined rather than wrong.

Notice the shape of that. The two facts a reader would test by hand, on a single measurement, survive the weaker condition. The one that fails is a statement about all states at once, and no single measurement can see it. That is why the next section works entirely by hand: the only way to believe the distinction is to watch it happen on an operator you can differentiate.

In plain terms 4.4.4

Two conditions have been separated that Part 0 could treat as one. The weaker says that moving the operator from one side of an inner product to the other changes nothing, so long as both states involved are on the allowed list. The stronger says that, and adds that the partner operator's allowed list is the same list, not a longer one.

The uncomfortable part, and it is worth sitting with, is that the weaker condition already delivers everything a working physicist would check on a single measurement. Measured values come out real. Different values belong to perpendicular states. Averages are real numbers. All of that is in hand without the stronger condition, so no calculation you would naturally perform can tell you which of the two you have.

What the stronger condition buys is not a reassurance but a structure, and it is a statement about every state at once rather than about any one of them. Only for the stronger one is there a guarantee that the possible outcomes of a measurement account for the whole space, and only for the stronger one does the operator generate a flow in time. Lose it and you do not get wrong answers. You get no answers, because the machinery that turns an operator into predictions has a hypothesis it does not meet. The next section shows this happening to the most familiar operator in the subject, on the most familiar interval.

a natural place to stop  ·  the vocabulary is built; what follows is one operator on three intervals

5 · Three intervals, one operator, worked by hand

This is the section the chapter exists for, and everything in it is integration by parts. We redo Chapter 0.2's calculation with the boundary term kept instead of waved through, identify what the adjoint of p^\hat p is when no boundary condition has been imposed at all, and then ask on which functions the boundary term vanishes. Three intervals give three different answers. On the line there is nothing to impose. On a bounded interval there is exactly a circle of possible impositions, and you will see the whole circle. On a half-line the boundary term cannot be killed without killing the operator, and the conclusion is that momentum there is not an observable. No classification theorem is used anywhere in this section; §6 supplies one afterwards and it will have to agree with what you find here.

5.1 · The boundary term, kept this time

Chapter 0.2 §3.2 listed the places integration by parts does structural work, and one of the entries was a promise about this chapter: "An operator gets an adjoint. In Chapter 4.4, ψxϕdx=(xψ)ϕdx\int \psi^{*}\,\partial_x\phi\,\dd x = -\int(\partial_x\psi)^{*}\phi\,\dd x when boundary terms vanish. The minus sign is why p^=ix\hat p = -\ii\hbar\,\partial_x is Hermitian and x\partial_x alone is not." That chapter needed the identity with the boundary term gone. Here the boundary term is the whole content, so we keep it.

Work on an interval (a,b)(a,b), where aa may be -\infty and bb may be ++\infty. For uu and vv with u,vL2u',v'\in L^{2}, Chapter 0.2's identity reads abuv=[uv]ababuv\int_a^b\overline u\,v'=[\overline u v]_a^b-\int_a^b\overline{u'}v. We want that in terms of p^\hat p rather than of d/dx\dd/\dd x, so multiply through by i-\ii\hbar. The left side becomes u,p^v\avg{u,\hat pv}. On the right the integral becomes +iuv+\ii\hbar\int\overline{u'}v, which is p^u,v\avg{\hat pu,v}, because conjugating the constant flips its sign: iu=+iu\overline{-\ii\hbar u'}=+\ii\hbar\,\overline{u'}. Moving that term across leaves

p^u,vu,p^v  =  i[u(x)v(x)]ab \avg{\hat pu,\,v}-\avg{u,\,\hat pv} \;=\; \ii\hbar\Big[\,\overline{u(x)}\,v(x)\,\Big]_{a}^{b} (4.4.13)

One point about the hypotheses is worth settling here rather than leaving it to be noticed. Chapter 0.2 proved integration by parts for continuously differentiable functions, and the members of (4.4.7) need not be differentiable everywhere. The extension is a single step: write each function as the integral of its derivative, substitute, and split the resulting square of integration into its two triangles, whereupon each side of the identity turns into the other. That step is Fubini's theorem, the standing mark of Chapter 0.2 §4.2, and it is the same interchange the grind box of §5.2 names below. Nothing else is added.

Read (4.4.13) as the definition of symmetry rather than as an identity about integrals, and the section is already organised. By (4.4.11), an operator with the formula id/dx-\ii\hbar\,\dd/\dd x is symmetric on a given domain precisely when the right-hand side vanishes for every pair u,vu,v drawn from that domain. Everything below is that one question, asked on three intervals.

5.2 · What the adjoint is before any condition is imposed

Before choosing domains it is worth knowing the answer to a question that does not depend on the choice: which uu satisfy (4.4.9), given that the domain of p^\hat p contains at least the functions that die away before either end of the interval is reached? Every domain considered below contains those, whatever it does at the ends, so the answer is the same in all three cases, and it is the most permissive one available.

Write C\mathcal{C} for the members of (4.4.7) that vanish outside a closed subinterval of the open interval (a,b)(a,b). If Cdom(p^)\mathcal{C}\subseteq\operatorname{dom}(\hat p), then every uu in dom(p^)\operatorname{dom}(\hat p^{\dagger}) is an integral of a function uL2u'\in L^{2}, and p^u=iu\hat p^{\dagger}u=-\ii\hbar u'. No condition at the endpoints comes out of this step. The grind box proves it, and the one part of it that is not integration by parts is the fact that a function whose derivative vanishes in this weak sense is constant, which is done there in full.

That leaves one further requirement, and it is the whole of the three subsections below. Testing against C\mathcal{C} alone cannot see the ends of the interval, so it cannot decide which of those uu are admitted. What decides it is that the defining relation must hold against every vv in the domain rather than only against those that die away first, and the difference between the two demands is precisely the boundary term of (4.4.13). So the adjoint's domain is the maximal set cut down by the condition that the boundary term vanish, and the three intervals cut it down by different amounts.

Grind box — identifying the adjoint of the derivative, once, for all three intervals

Recall C\mathcal{C}: the continuous functions that vanish outside a closed subinterval of (a,b)(a,b) and are integrals of a derivative lying in L2L^{2}. It contains §3.2's trapezoids, so it is dense in L2L^{2} over every bounded subinterval by the chain Chapter 4.3 §8.1 built, and it is contained in every domain this section considers, since its members vanish at both ends. Let udom(p^)u\in\operatorname{dom}(\hat p^{\dagger}) with w=p^uw=\hat p^{\dagger}u, so w,v=u,iv\avg{w,v}=\avg{u,-\ii\hbar v'} for all vdom(p^)v\in\operatorname{dom}(\hat p) and in particular for all vCv\in\mathcal{C}. Put h=iw/h=\ii w/\hbar, which lies in L2L^{2}. Dividing the relation by i-\ii\hbar and conjugating the constant gives

abuv  =  abhvfor all vC. \int_a^b \overline{u}\,v' \;=\; -\int_a^b\overline{h}\,v \qquad\text{for all } v\in\mathcal{C}.

Fix x0(a,b)x_0\in(a,b) and set H(x)=x0xhH(x)=\int_{x_0}^{x}h, which is finite for xx in (a,b)(a,b) because L2L^{2} functions are integrable on bounded subintervals by Cauchy–Schwarz. Interchanging the two integrations then gives Hv=hv\int\overline H v'=-\int\overline h v for the same vv. That interchange is Fubini's theorem, and it is worth naming rather than performing quietly. Chapter 4.3 §4.1 says in as many words that it settles the exchange of a sum with an integral and leaves the exchange of two integrals where Chapter 0.2 §4.2 marked it, so we are spending that standing mark here rather than raising a new one. The version needed is the undemanding one: the region is a bounded triangle on which the modulus of the integrand has finite integral, which is the non-negative case Chapter 0.2 noted holds with no extra hypotheses. Subtracting,

ab(uH)  v  =  0for all vC. \int_a^b \overline{(u-H)}\;v' \;=\;0 \qquad\text{for all } v\in\mathcal{C}.

The constancy step. Fix ηC\eta\in\mathcal{C} with η=1\int\eta=1. Given any φC\varphi\in\mathcal{C}, put c=φc=\int\varphi and χ(x)=ax(φcη)\chi(x)=\int_a^{x}(\varphi-c\,\eta). Then χ=φcη\chi'=\varphi-c\eta, and χ\chi vanishes outside a closed subinterval of (a,b)(a,b), because the total integral of φcη\varphi-c\eta is zero, so χ\chi returns to zero beyond both supports. Hence χC\chi\in\mathcal{C} and

(uH)φ  =  (uH)(cη+χ)  =  c(uH)η  =  γφ, \int\overline{(u-H)}\,\varphi \;=\; \int\overline{(u-H)}\,(c\,\eta+\chi') \;=\; c\int\overline{(u-H)}\,\eta \;=\; \gamma\int\varphi ,

where γ=(uH)η\gamma=\int\overline{(u-H)}\eta is one fixed number, the middle step being the displayed identity applied to v=χv=\chi. So (uH)φ=γφ\int\overline{(u-H)}\varphi=\int\gamma\varphi for every φC\varphi\in\mathcal{C}, and by the density noted at the top, uH=γ\overline{u-H}=\gamma almost everywhere. Therefore u=H+γu=H+\overline\gamma.

Conclusion. uu is an integral of hh plus a constant, so uu is the integral of a function, that function is hL2h\in L^{2}, and p^u=w=ih=iu\hat p^{\dagger}u=w=-\ii\hbar h=-\ii\hbar u'. No endpoint condition was used and none was produced, so nothing constrains uu at aa or bb beyond membership of L2L^{2}. Conversely every such uu is in dom(p^)\operatorname{dom}(\hat p^{\dagger}) when the boundary term of (4.4.13) vanishes against the domain, which is checked case by case below. \blacksquare

Two remarks. The argument used only test functions supported strictly inside the interval, which is why it is indifferent to what happens at the ends and why one proof covers all three cases. And it is the honest content of the phrase "integrate by parts and drop the boundary term": the dropping is legitimate against test functions that vanish there, and against nothing else without a reason.

5.3 · The line: nothing to impose

Take (a,b)=(,)(a,b)=(-\infty,\infty) and the domain (4.4.7), which is the largest one the formula admits. The boundary term in (4.4.13) is a limit at each end, and the claim is that both limits are zero for every pair in the domain, with no condition imposed at all.

Here is the argument, and after (4.4.13) it uses Cauchy–Schwarz twice and nothing else. Put F=uvF=\overline u v. Read (4.4.13) on the bounded interval [0,X][0,X] and divide by i\ii\hbar: it says exactly that F(X)F(0)=0X(uv+uv)F(X)-F(0)=\int_0^{X}\big(\overline{u'}v+\overline uv'\big). Now bound the two pieces. Both uu and vv are in L2L^{2}, so FF is in L1L^{1} by Cauchy–Schwarz, meaning F\int\abs F is finite. The integrand above is a sum of two products of L2L^{2} functions, so it is in L1L^{1} for the same reason, and its integral over [0,X][0,X] therefore converges as XX\to\infty. So F(X)F(X) has a finite limit. A function with a finite limit at infinity whose modulus has a finite integral must have that limit equal to zero, since a non-zero limit would make F\int\abs F diverge. So F(X)0F(X)\to0, and the same at -\infty.

The boundary term therefore vanishes identically and p^\hat p is symmetric on the domain (4.4.7). Now compare that domain with the adjoint's. Section 5.2 says every uu in the adjoint's domain is an integral of a function in L2L^{2}, which is membership of (4.4.7), and the further requirement it left open is that the boundary term vanish, which the paragraph above has just shown it does for every such uu. Nothing is cut down. The two domains are the same set, so

dom(p^)  =  dom(p^)andp^  =  p^on L2(R). \operatorname{dom}(\hat p^{\dagger})\;=\;\operatorname{dom}(\hat p) \qquad\text{and}\qquad \hat p^{\dagger}\;=\;\hat p \qquad\text{on }L^{2}(\R). (4.4.14)

There is nothing to choose, no parameter, and no family. The momentum of a particle on a line is an observable, uniquely, and the reason it looked as though no work was needed is that in this one case no work is needed. Keep the feeling of that, because the next two cases are the same operator and the same calculation with different endpoints, and they do not come out this way.

5.4 · The interval: a circle of momentum operators

Now take (a,b)=(0,L)(a,b)=(0,L) with LL finite. The boundary term in (4.4.13) is no longer a limit, it is two numbers, and symmetry on a domain D\mathcal D is the requirement

u(L)v(L)  =  u(0)v(0)for all u,vD. \overline{u(L)}\,v(L) \;=\; \overline{u(0)}\,v(0) \qquad \text{for all } u,v\in\mathcal D. (4.4.15)

Start with the choice almost everyone makes first, which is to pin the wavefunction to zero at both ends: Dmin={v:vL2, v(0)=v(L)=0}\mathcal D_{\min}=\{v : v'\in L^{2},\ v(0)=v(L)=0\}. Condition (4.4.15) holds, since both sides are zero, so this operator is symmetric. Now ask what its adjoint is. By §5.2 the adjoint's domain is every uu with uL2u'\in L^{2} and no endpoint condition, because the boundary term is already zero for every such uu once vv is pinned. So dom(p^)\operatorname{dom}(\hat p^{\dagger}) is strictly larger than Dmin\mathcal D_{\min}, by every function that fails to vanish at an end.

That is the promised example, and it is worth pausing on. Take the claims one at a time. The operator id/dx-\ii\hbar\,\dd/\dd x with the wavefunction pinned at both ends is symmetric, by the paragraph above. Its expectation values are real, by §4.3, and so would any eigenvalue be. It passes that eigenvalue test vacuously, because iψ=pψ-\ii\hbar\psi'=p\psi forces ψ=eipx/\psi=\ee^{\ii px/\hbar}, which never vanishes and so cannot be pinned at either end, leaving the operator with no eigenvalues at all. And it is not self-adjoint. The gap between the two domains is exactly two complex dimensions, because what an element of the larger domain may do, and an element of the smaller one may not, is take a non-zero value at each of the two ends. Chapter 4.2 §4's warning box was pointing at this operator.

So pinning both ends is too strong. The question is which domains are not, and (4.4.15) answers it in two lines. Set u=vu=v in it and take moduli:

v(L)  =  v(0)for every vD. \abs{v(L)} \;=\; \abs{v(0)} \qquad\text{for every } v\in\mathcal D. (4.4.16)

Two cases follow, and there are no others. Either every vv in D\mathcal D has v(0)=0v(0)=0, and then (4.4.16) forces v(L)=0v(L)=0 as well, so D\mathcal D is Dmin\mathcal D_{\min} or a subspace of it. The first was dealt with in the paragraph above, and a subspace of it has an adjoint at least as large again, by (4.4.10), so it fails for the same reason and worse. Or some v0Dv_0\in\mathcal D has v0(0)0v_0(0)\ne0, and we may scale it so that v0(0)=1v_0(0)=1, whereupon (4.4.16) makes v0(L)v_0(L) a number of modulus one, say eiθ\ee^{\ii\theta}. Put u=v0u=v_0 in (4.4.15) and every other vv in the domain is caught by it:

eiθv(L)  =  v(0)v(L)  =  eiθv(0). \ee^{-\ii\theta}v(L)\;=\;v(0) \qquad\Longleftrightarrow\qquad v(L)\;=\;\ee^{\ii\theta}\,v(0). (4.4.17)

So a single phase, fixed once for the whole domain, is the only symmetric possibility beyond the pinned one. Define p^θ\hat p_\theta to be id/dx-\ii\hbar\,\dd/\dd x on Dθ={v:vL2, v(L)=eiθv(0)}\mathcal D_\theta=\{v: v'\in L^{2},\ v(L)=\ee^{\ii\theta}v(0)\}.

Each of these is self-adjoint, and the check is the same equation read in the other direction. Let uu be in the adjoint's domain, which by §5.2 means uL2u'\in L^{2} with nothing at the ends, and impose (4.4.15) against every vDθv\in\mathcal D_\theta. Substituting v(L)=eiθv(0)v(L)=\ee^{\ii\theta}v(0) turns it into (u(L)eiθu(0))v(0)=0\big(\overline{u(L)}\ee^{\ii\theta}-\overline{u(0)}\big)v(0)=0, and there are functions in Dθ\mathcal D_\theta with v(0)0v(0)\ne0, for instance eiθx/L\ee^{\ii\theta x/L}. So the bracket vanishes, which conjugates to u(L)=eiθu(0)u(L)=\ee^{\ii\theta}u(0): the adjoint's domain is Dθ\mathcal D_\theta again. The two domains coincide and p^θ\hat p_\theta is self-adjoint.

One step of the case analysis is still open, and it is the step the first case took explicitly. Condition (4.4.17) puts D\mathcal D inside Dθ\mathcal D_\theta without saying the two are equal, and a subspace of a symmetric domain is symmetric as well, so containment on its own settles nothing. But the operator on a proper subspace cannot be self-adjoint. If D\mathcal D is strictly inside Dθ\mathcal D_\theta, then p^θ\hat p_\theta extends id/dx-\ii\hbar\,\dd/\dd x on D\mathcal D, so (4.4.10) makes the adjoint's domain contain dom(p^θ)=Dθ\operatorname{dom}(\hat p_\theta^{\dagger})=\mathcal D_\theta, which is strictly larger than D\mathcal D. So D=Dθ\mathcal D=\mathcal D_\theta exactly, and the phases are the whole list.

Since θ\theta and θ+2π\theta+2\pi give the same condition, the family is a circle. Momentum on a bounded interval is not one observable but a circle of them, and they are genuinely different operators, because their eigenvalues differ. Solving iψ=pψ-\ii\hbar\psi'=p\psi gives ψ=eipx/\psi=\ee^{\ii px/\hbar}, and the condition in (4.4.17) reads eipL/=eiθ\ee^{\ii pL/\hbar}=\ee^{\ii\theta}, so

pn  =  (θ+2πn)L,nZ. p_n \;=\; \frac{\hbar\,(\theta+2\pi n)}{L}, \qquad n\in\mathbb{Z} . (4.4.18)

A whole ladder of allowed momenta, sliding rigidly as θ\theta turns, and returning to itself after one full turn. Choosing θ\theta is choosing what happens at the wall, and §7 will show the same phenomenon for energy with more room in it.

5.5 · The half-line: the term that cannot be killed

Take (a,b)=(0,)(a,b)=(0,\infty), which is the case Chapter 4.2 §4 named twice. The end at infinity is handled by §5.3's argument word for word, so the boundary term of (4.4.13) is one number:

p^u,vu,p^v  =  iu(0)v(0). \avg{\hat pu,\,v}-\avg{u,\,\hat pv} \;=\; -\,\ii\hbar\,\overline{u(0)}\,v(0). (4.4.19)

Ask the same question as before: for which domains does this vanish on every pair? Set u=vu=v in (4.4.19) and the answer arrives at once, because v(0)2=0\abs{v(0)}^{2}=0 has only one solution. Every vv in a symmetric domain must satisfy v(0)=0v(0)=0. There is no phase to choose, because there is only one endpoint and a phase relates two.

So the largest symmetric domain available is D0={v:vL2, v(0)=0}\mathcal D_{0}=\{v:v'\in L^{2},\ v(0)=0\}, and we can ask whether that one is self-adjoint. Section 5.2 gives its adjoint's domain as every uu with uL2u'\in L^{2} and no condition at zero, since (4.4.19) already vanishes once v(0)=0v(0)=0. That is strictly larger: the function u(x)=ex/au(x)=\ee^{-x/a}, for any length aa, has uL2u'\in L^{2} and u(0)=1u(0)=1, so it sits in the adjoint's domain and not in D0\mathcal D_0. The operator is symmetric and not self-adjoint.

Now put the two paragraphs together, which is the step that makes this different from §5.4. In that section the repair was to enlarge the domain until the two met. Here enlarging is not available: any enlargement contains a function with v(0)0v(0)\ne0, and the first paragraph shows that such a domain is not symmetric, so it is not a candidate at all. The domain cannot grow and the adjoint's domain cannot shrink. They never meet.

⚠ The conclusion, which is a fact about nature and not about notation

id/dx-\ii\hbar\,\dd/\dd x on [0,)[0,\infty) is symmetric and has no self-adjoint extension whatever. By P2 as corrected in §4.2, the momentum of a particle confined to a half-line is not an observable. There is no operator to measure, so there is no spectrum of possible values, no probability distribution over them, and no expectation value beyond the formal integral.

Nothing was quoted to reach that. It is (4.4.13), which is integration by parts, together with the observation that v(0)2=0\abs{v(0)}^{2}=0 forces v(0)=0v(0)=0.

Where it bites. A radial coordinate runs over [0,)[0,\infty), so this is the statement that there is no radial momentum observable, and it is the reason Chapter 4.13 has to argue for the behaviour of the hydrogen wavefunction at the origin rather than assume it. It is also why Chapter 4.7's infinite well cannot be declared to have ψ=0\psi=0 at the walls without an argument: that is one choice among the family §7 counts, and the physics of the wall is what selects it.

And notice what did not go wrong. Momentum on a half-line has real expectation values, real would-be eigenvalues, and a perfectly sensible-looking formula. Every check §4.3 listed passes. The failure is invisible until you ask the one question this chapter taught you to ask.

The reason for the asymmetry deserves a sentence, because it is Chapter 0.5's fourth item returning under a different name. On a bounded interval, whatever leaves one end can be made to arrive at the other, and θ\theta records how much phase it picks up on the way. On a half-line there is a place to go and nowhere to come from. That is exactly the shift S^\hat S of (4.4.1): norm-preserving, one-to-one, and not onto. Worked example 3 makes the identification precise by exponentiating p^\hat p and finding that the flow it generates runs in one direction only.

In plain terms 4.4.5

Everything in this section is one line of calculus done honestly. When you move a derivative from one function to the other inside an integral, a leftover appears at the two ends of the interval. In Part 0 it was always thrown away, with the note that throwing it away is a physical claim. Here it is the only thing being looked at, because it is precisely the obstruction to a momentum operator being a measurable quantity.

On an infinite line the leftover disappears by itself. Anything with finite total probability and a finite spread of momentum has to die away at both ends, so there is nothing to arrange and nothing to choose, and momentum is a perfectly ordinary observable. On a segment the leftover is the difference between the two ends, and it can be cancelled by insisting that whatever leaves one end arrives at the other with a fixed change of phase. The size of that change is a real number, free to be anything, and every value of it gives a different genuine momentum operator with a different ladder of allowed values. The condition at the wall is not bookkeeping. It is a physical input, and the measurable momenta depend on it.

On a half-line the arithmetic gives a different kind of answer. The leftover involves one end only, and the sole way to cancel it is to demand that every state vanish there, which is a demand so strong that the partner operator is left with a larger list of states than the operator itself and no amount of adjustment closes the gap. The conclusion is not a technicality: a particle held on one side of an impenetrable wall has no momentum observable at all. Not one that is hard to compute. One that does not exist.

a natural place to stop  ·  the answers are in hand; what follows is the theorem that counts them

6 · Counting the extensions: deficiency indices

Section 5 produced three answers by hand: a unique operator, a circle of them, and none. This section quotes the theorem that predicts which of those three outcomes occurs, without solving a boundary-value problem, and reduces the prediction to counting the square-integrable solutions of one first-order equation. It has to reproduce §5 exactly three times, and it does.

6.1 · The question, in a form a theorem can answer

Section 3.4 said what the general situation looks like. A symmetric operator sits inside its adjoint, enlarging the operator shrinks the adjoint, and a self-adjoint extension is a place where the two meet. So the size of the gap between dom(A^)\operatorname{dom}(\hat A) and dom(A^)\operatorname{dom}(\hat A^{\dagger}) is what decides everything, and the useful way to measure that gap is to look at what A^\hat A^{\dagger} can do that A^\hat A cannot.

The measurement von Neumann found is this. A symmetric operator has real eigenvalues, by §4.3, so it can have no eigenvector with a strictly complex eigenvalue. Its adjoint has a larger domain and is under no such restriction, so it can. Counting the eigenvectors of A^\hat A^{\dagger} at one complex eigenvalue above the axis and at its mirror image below measures exactly how much room there is between the operator and its adjoint, and that count is what the theorem is stated in terms of.

6.2 · The classification, quoted

⚑ Quoted, not derived — von Neumann's classification of self-adjoint extensions

We use, without proof, the following. Let A^\hat A be symmetric with a dense domain in a Hilbert space. Fix any positive number λ\lambda carrying the units of A^\hat A, and define the two deficiency spaces and their dimensions. They take an N\mathcal N rather than the D\mathcal D this chapter has been using for domains, because they are not domains of anything: they are the sets of solutions of an eigenvalue equation, sitting inside a domain already fixed.

N±  =  {fdom(A^) : A^f=±iλf},n±=dimN±. \mathcal{N}_{\pm} \;=\; \big\{f\in\operatorname{dom}(\hat A^{\dagger})\ :\ \hat A^{\dagger}f=\pm\,\ii\lambda f\big\}, \qquad n_{\pm}=\dim\mathcal{N}_{\pm}.

The two numbers do not depend on which λ\lambda was chosen. They are the deficiency indices, written (n+,n)(n_+,n_-), and they settle the question completely.

  • (0,0)(0,0). Exactly one self-adjoint operator extends A^\hat A.
  • (n,n)(n,n) with n1n\ge1. The self-adjoint operators extending A^\hat A correspond one-to-one with the unitary maps of N+\mathcal{N}_{+} onto N\mathcal{N}_{-}. That set of maps is U(n)U(n), and it has n2n^{2} real parameters.
  • n+nn_+\ne n_-. There are none.

What the mark covers and what it does not. The proof runs through the Cayley transform, which turns a symmetric operator into an isometry and a self-adjoint one into a unitary, so that extending the operator becomes extending an isometry and the count becomes linear algebra on the two leftover spaces. It is about six pages and it is in Reed and Simon's Methods of Modern Mathematical Physics, volume II, and in Rudin. The mark also covers the technical point that the deficiency spaces are read off A^\hat A^{\dagger} alone, so nothing about A^\hat A beyond symmetry and density enters the count. What the mark does not cover is any of §5, which used none of this.

Why it is computable. The deficiency spaces are read off A^\hat A^{\dagger}, and for the operators of §5 that domain came out maximal, with no endpoint condition surviving. So finding the deficiency spaces means solving a differential equation with no boundary conditions at all and then asking which solutions are square integrable. For a first-order operator that is one line of calculus.

6.3 · The three cases, counted

Say first which operator is being extended in each case, since the theorem counts the extensions of a given one. On the line it is p^\hat p on (4.4.7). On the interval it is p^\hat p on Dmin\mathcal D_{\min}, the pinned domain, which is the smallest sensible starting point and the one whose adjoint §5.4 computed. On the half-line it is p^\hat p on D0\mathcal D_{0}. In all three the adjoint acts by the same formula on the maximal domain, so the calculation is the same one three times over.

Take λ\lambda to be a fixed positive momentum p0p_0, so that the units work. The equation p^f=±ip0f\hat p^{\dagger}f=\pm\ii p_0 f is if=±ip0f-\ii\hbar f'=\pm\ii p_0f, so f=p0f/f'=\mp p_0f/\hbar and

p^f=+ip0f  f=ep0x/,p^f=ip0f  f=e+p0x/, \hat p^{\dagger}f=+\ii p_0 f \ \Longrightarrow\ f=\ee^{-p_0x/\hbar}, \qquad \hat p^{\dagger}f=-\ii p_0 f \ \Longrightarrow\ f=\ee^{+p_0x/\hbar}, (4.4.20)

each unique up to a constant, since a first-order equation has a one-dimensional solution space. An exponential has its derivative proportional to itself, so being in the domain of p^\hat p^{\dagger} at all reduces to the single question of whether the exponential is square-integrable. Hence n±n_\pm is 11 when the corresponding exponential lies in L2L^{2} of the interval and 00 when it does not, and that is the entire computation. Take the three intervals in the order §5 took them.

  • The line. ep0x/\ee^{-p_0x/\hbar} blows up as xx\to-\infty and e+p0x/\ee^{+p_0x/\hbar} blows up as x+x\to+\infty, so neither is in L2(R)L^{2}(\R) and the indices are (0,0)(0,0). Exactly one self-adjoint extension. Section 5.3 found exactly one.
  • The interval. Both exponentials are continuous on a bounded interval, so both are in L2[0,L]L^{2}[0,L], and the indices are (1,1)(1,1). The unitary maps between two one-dimensional spaces are multiplication by a phase, so the family is U(1)U(1), a circle, with 12=11^{2}=1 real parameter. Section 5.4 found a circle, parametrised by one phase.
  • The half-line. ep0x/\ee^{-p_0x/\hbar} decays and is in L2[0,)L^{2}[0,\infty), while e+p0x/\ee^{+p_0x/\hbar} grows and is not. The indices are (1,0)(1,0), they are unequal, and there are no self-adjoint extensions. Section 5.5 found none.

Three cases, three agreements, and the agreement is not a coincidence of arithmetic. Look at the totals. The sum n++nn_++n_- comes out 00, 22 and 11, which is the number of independent pieces of boundary data the interval carries, because each piece is one thing the boundary term of (4.4.13) has to cancel. For a first derivative that data is one number at each finite end, the value of the function there, which is why the totals here also read as a count of finite endpoints. Take the count of data as the rule and the count of endpoints as a coincidence of first order, because §7 has the same two finite ends and gets (2,2)(2,2): a second derivative reads a value and a slope at each of them. Self-adjoint extensions exist when the data can be paired off against itself, which needs an even total split evenly, and that is n+=nn_+=n_-. The half-line has one datum and no partner for it. What the theorem adds is that this survives into situations where the hand calculation would be much longer, and §7 is the first of those.

In plain terms 4.4.6

The previous section answered a question three times by direct calculation, and got three different kinds of answer: one operator, a circle of operators, and no operator. A theorem exists that predicts which of the three you are about to get, and it does so by a count that takes a single line.

The recipe is this. Take the partner operator, which has the longer list of allowed states, and look for states it sends to themselves multiplied by a purely imaginary number. The original operator can never do this, because it always returns real multiples, so anything the partner manages here is a direct measurement of how much larger the partner is. Count these states above the axis and below it separately, and the pair of counts decides which of the three answers you are going to get. It gets all three right here, which is why it can be trusted in the next section, where the hand calculation would be considerably longer.

a natural place to stop  ·  the first derivative is finished; what follows is the second, with the machinery put to work

7 · The particle in a box has four parameters, not one

Everything so far has been about a first derivative. Energy is a second derivative, and this section runs the whole machine on it. The boundary term acquires a second pair of terms, the deficiency indices come out (2,2)(2,2) rather than (1,1)(1,1), and the family of self-adjoint Hamiltonians for a particle in a box is therefore a U(2)U(2): four real parameters, not one. We then take a one-parameter slice of that family, solve it, and find that different members have genuinely different energy levels. That is the point of the section. The boundary condition is not bookkeeping and it is not a choice of convenience. It is physics, and it is measurable.

7.1 · The boundary form for a second derivative

Take H^=22md2dx2\hat H=-\dfrac{\hbar^{2}}{2m}\,\dvn{2}{}{x} on L2[0,L]L^{2}[0,L], which is Chapter 4.2's free Hamiltonian confined to a box. Integrating by parts twice moves both derivatives across and leaves two boundary terms rather than one:

H^u,vu,H^v  =  22m[u(x)v(x)u(x)v(x)]0L. \avg{\hat Hu,\,v}-\avg{u,\,\hat Hv} \;=\; \frac{\hbar^{2}}{2m}\Big[\,\overline{u(x)}\,v'(x)-\overline{u'(x)}\,v(x)\,\Big]_{0}^{L} . (4.4.21)

The structure is the same as (4.4.13) and the arithmetic is one step longer. Symmetry on a domain is again the demand that the right-hand side of (4.4.21) vanish for every pair drawn from it, but now the data at each end is a pair of numbers, the value and the slope, so there is more to arrange and more ways to arrange it.

7.2 · The count, and the number it produces

Rather than classify by hand, use §6.2, which is exactly the situation it was quoted for. As §6.3 insisted, the theorem counts the extensions of a given operator, so the first job is to say which operator is being extended. It is the one whose domain kills every boundary number there is:

D00  =  {ψ : ψL2[0,L],  ψ(0)=ψ(L)=ψ(0)=ψ(L)=0}, \mathcal{D}_{00} \;=\; \big\{\psi\ :\ \psi''\in L^{2}[0,L],\ \ \psi(0)=\psi(L)=\psi'(0)=\psi'(L)=0\big\}, (4.4.22)

with H^00\hat H_{00} the formula 22md2dx2-\tfrac{\hbar^{2}}{2m}\dvn{2}{}{x} on that domain. All four numbers are killed and not two, which is more than §5.4's Dmin\mathcal D_{\min} asked: that domain pinned the two values and left the slopes free. The distinction decides what is being counted. Pinning the values alone, with ψ\psi'' in L2L^{2}, is the Dirichlet domain, and the list below puts Dirichlet at a single point of the family we are about to count, whereas the operator a family is counted from has to be one that every member of the family extends.

Two hypotheses of §6.2 have to be in place before it can be used. H^00\hat H_{00} is symmetric, because every term in the bracket of (4.4.21) carries a factor that vanishes at each end. Its domain is dense, by §3.2's chain with one adjustment: replace each straight ramp there by the ramp sin2(πt/2δ)\sin^{2}(\pi t/2\delta) across the same width δ\delta, which meets the flat parts on either side with zero slope and has a bounded second derivative, so the approximating functions lie in (4.4.22) and are no further away in L2L^{2} than the straight ones were. The adjoint of H^00\hat H_{00} is then §5.2's answer run twice: the formula on the maximal domain, with no endpoint condition surviving, because the test functions that argument uses vanish identically near both ends and so constrain nothing there. So we solve H^00f=±iE0f\hat H_{00}^{\dagger}f=\pm\ii E_0f for a fixed positive energy E0E_0, with no boundary conditions imposed:

22mf  =  ±iE0ff  =  μ±2f,μ±2=2miE02. -\frac{\hbar^{2}}{2m}f'' \;=\; \pm\,\ii E_0\,f \qquad\Longleftrightarrow\qquad f''\;=\;\mu_{\pm}^{2}f, \qquad \mu_{\pm}^{2}=\mp\,\frac{2m\,\ii E_0}{\hbar^{2}} . (4.4.23)

What matters about (4.4.23) is not the value of μ±\mu_\pm but the order of the equation. A second-order linear equation with constant coefficients has a two-dimensional solution space, spanned by eμx\ee^{\mu x} and eμx\ee^{-\mu x}, and every one of those solutions is continuous on the bounded interval [0,L][0,L] and therefore square-integrable on it, whatever complex number μ\mu happens to be. Both deficiency spaces are two-dimensional and the indices are (2,2)(2,2).

That is where §6.3's rule comes good, and the two versions of it are worth reading against each other here. Two finite ends, two pieces of boundary data at each, four pieces in all, split evenly into (2,2)(2,2). Counting data gives the right answer and counting endpoints gives two, which is why §6.3 asked you to keep the first version and treat the second as an accident of first order.

Feed that into §6.2 and the answer is a family of self-adjoint Hamiltonians in one-to-one correspondence with the unitary maps of one two-dimensional space onto another, which is U(2)U(2):

(n+,n)=(2,2)U(2) of self-adjoint extensions,dimRU(2)=4. (n_{+},n_{-})=(2,2) \quad\Longrightarrow\quad \text{a } U(2)\ \text{of self-adjoint extensions}, \quad \dim_{\R}U(2)=4 . (4.4.24)

The real dimension of U(n)U(n) is n2n^{2}, and §6.2 quoted that along with the theorem. At n=1n=1 you can check it on sight, a 1×11\times1 unitary being a phase and nothing else, but n=2n=2 is not visible that way, so here is the count. A 2×22\times2 unitary matrix is one whose two columns are orthonormal. The first column is a unit vector in C2\C^{2}: four real numbers with one real equation on them, so three real parameters. The second column has to be a unit vector orthogonal to the first, and the vectors orthogonal to a given non-zero vector in C2\C^{2} form a line of one complex dimension, on which the unit vectors differ by a phase. One more real parameter, and three and one make four.

Four real parameters. That is worth stating loudly, because the number one is easy to arrive at and wrong. Momentum on an interval gave a single phase in §5.4, the ladder in (4.4.18) slid rigidly as that phase turned, and it is natural to carry the picture across. Energy is a second derivative, the boundary data at each end is a value and a slope, and the family is four-dimensional. Some familiar conditions and where they sit in it:

  • Dirichlet, ψ(0)=ψ(L)=0\psi(0)=\psi(L)=0, which is the textbook infinite well. One point of U(2)U(2).
  • Neumann, ψ(0)=ψ(L)=0\psi'(0)=\psi'(L)=0. A different point.
  • Periodic, ψ(L)=ψ(0)\psi(L)=\psi(0) and ψ(L)=ψ(0)\psi'(L)=\psi'(0), which is a particle on a ring, and antiperiodic, the same with both signs reversed. Two more.
  • The Robin family, ψ(0)=αψ(0)\psi'(0)=\alpha\,\psi(0) and ψ(L)=αψ(L)\psi'(L)=-\alpha\,\psi(L) with α\alpha real, a curve through the first two: α=0\alpha=0 is Neumann and the limit of large α\abs\alpha is Dirichlet. This is the slice §7.3 works.

Each entry satisfies (4.4.21) directly, and you can check any of them in a line by substituting the two conditions into the bracket. The Robin case is done in full below, including the step from symmetric to self-adjoint, because it is the one whose spectrum we want.

7.3 · The Robin family, worked

Fix a real number α\alpha with the units of an inverse length and take

dom(H^α)  =  {ψ : ψL2[0,L],  ψ(0)=αψ(0),  ψ(L)=αψ(L)}. \operatorname{dom}(\hat H_\alpha) \;=\; \big\{\psi\ :\ \psi''\in L^{2}[0,L],\ \ \psi'(0)=\alpha\,\psi(0),\ \ \psi'(L)=-\alpha\,\psi(L)\big\}. (4.4.25)

Check symmetry first. Substitute the two conditions into the bracket of (4.4.21) and look at each end separately. At x=Lx=L the bracket is u(L)(αv(L))(αu(L))v(L)\overline{u(L)}\,(-\alpha v(L))-\overline{(-\alpha u(L))}\,v(L), which is αu(L)v(L)+αu(L)v(L)=0-\alpha\overline{u(L)}v(L)+\alpha\overline{u(L)}v(L)=0 because α\alpha is real and comes out of the conjugate unchanged. At x=0x=0 the same cancellation happens with the sign the other way. Each end contributes zero on its own, so H^α\hat H_\alpha is symmetric.

Self-adjointness is the same computation read backwards, and it starts from §7.2 rather than from nothing. A function with all four boundary numbers zero satisfies the two Robin conditions, so D00dom(H^α)\mathcal{D}_{00}\subseteq\operatorname{dom}(\hat H_\alpha), and (4.4.10) then puts dom(H^α)\operatorname{dom}(\hat H_\alpha^{\dagger}) inside dom(H^00)\operatorname{dom}(\hat H_{00}^{\dagger}), which §7.2 identified as the maximal domain. So let uu have uL2u''\in L^{2} and no endpoint condition, and demand that the bracket vanish against every vdom(H^α)v\in\operatorname{dom}(\hat H_\alpha). Substituting the conditions on vv collects the bracket into

(αu(L)+u(L))v(L)    (αu(0)u(0))v(0)  =  0. -\big(\alpha\,\overline{u(L)}+\overline{u'(L)}\big)\,v(L) \;-\; \big(\alpha\,\overline{u(0)}-\overline{u'(0)}\big)\,v(0) \;=\;0 . (4.4.26)

The two numbers v(0)v(0) and v(L)v(L) can be prescribed independently inside the domain, because for any pair of target values there is a cubic polynomial matching those values and the two slopes the conditions then demand, and a cubic has exactly four coefficients to do it with. So both brackets in (4.4.26) vanish separately, and conjugating them returns u(L)=αu(L)u'(L)=-\alpha u(L) and u(0)=αu(0)u'(0)=\alpha u(0). Those are the conditions defining (4.4.25), so the adjoint's domain is the domain, and H^α\hat H_\alpha is self-adjoint for every real α\alpha.

7.4 · Different extensions, different spectra

Now solve it, because the whole claim of this section is that the choice is physical, and the way to demonstrate that is to show the choice moving the energy levels. Write E=2k2/2mE=\hbar^{2}k^{2}/2m for a positive energy, so that ψ=k2ψ\psi''=-k^{2}\psi and ψ=Acoskx+Bsinkx\psi=A\cos kx+B\sin kx. The condition at x=0x=0 reads Bk=αABk=\alpha A, which is solved by A=kA=k and B=αB=\alpha up to scale. Imposing the condition at x=Lx=L on that function and collecting terms gives one equation in one unknown:

(α2k2)sinkL  +  2αkcoskL  =  0tankL  =  2αkk2α2. (\alpha^{2}-k^{2})\,\sin kL \;+\; 2\alpha k\,\cos kL \;=\;0 \qquad\Longleftrightarrow\qquad \tan kL \;=\; \frac{2\alpha k}{k^{2}-\alpha^{2}} . (4.4.27)

That is a transcendental equation and it has no closed-form solution, which is itself informative: the clean n2n^{2} spectrum of the textbook box is a property of one point of the family and not of the formula. Two limits are worth reading off before solving it numerically. At α=0\alpha=0 the equation collapses to sinkL=0\sin kL=0, so k=nπ/Lk=n\pi/L and En=n2π22/2mL2E_n=n^{2}\pi^{2}\hbar^{2}/2mL^{2}, which is the Neumann ladder. It runs from n=0n=0, and the n=0n=0 member is the constant function, which the scaling A=kA=k above discards and which the function Δ(E)\Delta(E) defined below puts back. As α\abs\alpha\to\infty the condition ψ(0)=αψ(0)\psi'(0)=\alpha\psi(0) forces ψ(0)0\psi(0)\to0 with the slope finite, so the limit is the Dirichlet ladder, the same numbers with n=0n=0 removed and the lowest level at π22/2mL2\pi^{2}\hbar^{2}/2mL^{2}.

Negative energies have to be looked for separately, and they are where the family does something the textbook box never does. Put E=2κ2/2mE=-\hbar^{2}\kappa^{2}/2m with κ>0\kappa\gt0, repeat the two steps with hyperbolic functions in place of trigonometric ones, and the condition becomes (α2+κ2)sinhκL+2ακcoshκL=0(\alpha^{2}+\kappa^{2})\sinh\kappa L+2\alpha\kappa\cosh\kappa L=0, which needs α<0\alpha\lt0 to have any solution at all. Both cases are captured by one function of the energy, the characteristic function of this boundary-value problem, which shares its name with Chapter 0.9's characteristic function of a distribution and nothing else. It is obtained by dividing (4.4.27) by kk and letting kk be imaginary:

Δ(E)  =  (α2q)s(q)+2αc(q),q=2mE2, \Delta(E) \;=\; (\alpha^{2}-q)\,s(q)+2\alpha\,c(q), \qquad q=\frac{2mE}{\hbar^{2}}, (4.4.28)

where s(q)=sin(qL)/qs(q)=\sin(\sqrt q\,L)/\sqrt q and c(q)=cos(qL)c(q)=\cos(\sqrt q\,L), both of which are power series in qq with infinite radius of convergence and are therefore perfectly well defined for qq negative, where they become sinh\sinh and cosh\cosh. The energies are the zeros of Δ\Delta, and because Δ\Delta depends smoothly on α\alpha as well as on EE, the levels slide continuously as α\alpha is turned.

One value of Δ\Delta can be read without any computation and it predicts the whole qualitative story. At q=0q=0 we have s=Ls=L and c=1c=1, so

Δ(0)  =  α2L+2α  =  α(αL+2), \Delta(0) \;=\; \alpha^{2}L+2\alpha \;=\; \alpha\,(\alpha L+2), (4.4.29)

which vanishes at α=0\alpha=0 and at α=2/L\alpha=-2/L and nowhere else. Those are the only two settings at which a level sits exactly at zero energy, so they are the only two at which a level can pass from one side of zero to the other. Solving the equation confirms that nothing else happens: below zero there are no states for α>0\alpha\gt0, one state for 2/L<α<0-2/L\lt\alpha\lt0, and two for α<2/L\alpha\lt-2/L. A negative α\alpha makes each wall attractive, and the two thresholds in (4.4.29) are where the first and then the second state bound to a wall drops below the bottom of the positive ladder.

Solving (4.4.28) by Newton's method gives the numbers. Energies are in units of 2/2mL2\hbar^{2}/2mL^{2}, and the last column is the Dirichlet limit for comparison.

levelαL=4\alpha L=-4αL=0\alpha L=0αL=+5\alpha L=+5αL\alpha L\to\infty
n=0n=017.06249-17.06249005.2187295.218729π2=9.869604\pi^{2}=9.869604
n=1n=114.66902-14.66902π2=9.869604\pi^{2}=9.86960422.6698722.669874π2=39.478424\pi^{2}=39.47842
n=2n=224.1811024.181104π2=39.478424\pi^{2}=39.4784255.7064655.706469π2=88.826449\pi^{2}=88.82644
n=3n=373.0950573.095059π2=88.826449\pi^{2}=88.82644106.6389106.638916π2=157.913716\pi^{2}=157.9137

Read the middle column first. At α=0\alpha=0 the roots of (4.4.28) come out at n2π2n^{2}\pi^{2} to every digit printed, which is the check that the numerical method is solving the equation we wrote down rather than a neighbour of it. Then read across. The three columns are three different sets of numbers. Same particle, same mass, same box, same formula for the energy, and three different spectra, because three different self-adjoint operators were selected out of the same U(2)U(2).

alpha L = 0.000
levels: E0 = 0.000000 E1 = 9.869604 E2 = 39.478418 E3 = 88.826440
n^2 pi^2: 0.000000 9.869604 39.478418 88.826440 largest relative departure from n^2 pi^2 = 0.000e+0 (at alpha = 0 the closed form is n^2 pi^2 hbar^2 / 2 m L^2 and this reads zero to the last stored digit)
levels strictly below E = 0: 0 Delta(0) = alpha (alpha L + 2) = 0.000000 -- it vanishes only at alpha L = 0 and alpha L = -2, and those are the only two settings at which the count above changes
One box, one formula for the energy, and a four-parameter family of legitimate operators. This is a one-parameter slice of it. Top: the four lowest levels of (4.4.28) as functions of αL\alpha L, in units of 2/2mL2\hbar^{2}/2mL^{2}, computed by bracketing and bisecting the zeros of Δ\Delta on a grid that is uniform in sgn(E)E\mathrm{sgn}(E)\sqrt{\abs E} so that high levels are resolved as well as low ones. The amber rule is E=0E=0 and the two grey uprights are the roots of (4.4.29) at αL=0\alpha L=0 and αL=2\alpha L=-2. Nothing about those two values was put into the root finder; they are where the computed levels cross zero. Bottom: the two lowest normalised eigenfunctions at the current setting, drawn across the box. Push α\alpha positive and they are pressed away from both walls towards Dirichlet behaviour. Pull it negative and the lowest one piles up against the walls instead, which is what a negative-energy state bound to the boundary looks like. The readouts are the test. The first prints the levels. The second prints n2π2n^{2}\pi^{2} beside them together with the largest relative departure, which reads exactly zero at α=0\alpha=0, the computed roots agreeing with the closed form in every digit a double holds, and grows smoothly away from it. The third counts the levels strictly below zero and prints Δ(0)\Delta(0), so you can watch the count change at exactly the two places (4.4.29) says it must. In this figure only, energies are quoted in units of 2/2mL2\hbar^{2}/2mL^{2} and α\alpha in units of 1/L1/L.

The physical reading is the point of the whole chapter and it is worth stating without hedging. The formula 2ψ/2m-\hbar^{2}\psi''/2m does not determine a Hamiltonian for a particle in a box. It determines a four-parameter family of them, all self-adjoint, all legitimate observables, with different energy levels. Choosing among them is not a mathematical formality performed before the physics starts. It is the physics of what the wall is made of, and any experiment that measures the levels is measuring which member of the family the wall implements.

In plain terms 4.4.7

Energy involves a second derivative rather than a first, so at each wall there are two numbers to worry about instead of one: the value of the wavefunction and its slope. Running the counting recipe gives two independent directions of freedom at each end rather than one, and the family of legitimate energy operators for a particle in a box turns out to have four adjustable real numbers in it. Four, not one. The single knob belonged to momentum, where each wall carries one number, and it does not carry over.

Every boundary condition anyone writes down for a box is one point of that family. The textbook condition that the wavefunction vanishes at both walls is one point. Demanding instead that the slope vanish is another. Gluing the two ends together into a ring is a third. A wall that is neither perfectly hard nor perfectly soft sits somewhere in between, on a curve through the family that can be solved and plotted.

And the levels move. Turning the knob slides every energy continuously, and past a certain setting a level drops below zero altogether and becomes a state stuck to the wall rather than rattling around inside. Same particle, same box, same formula, different numbers on the spectrometer. That is the sense in which the condition at the boundary is not bookkeeping. It is a physical property of the wall, it is one of the four numbers, and it is measurable.

8 · Worked examples

Worked example 1 — position, where the domain is forced and there is nothing to choose

On L2(R)L^{2}(\R) let x^\hat x be multiplication by xx. (a) Show that x^\hat x is unbounded. (b) Write down the largest domain on which the formula lands back in the space, and show it is dense. (c) Show that x^\hat x is symmetric on it. (d) Show that dom(x^)\operatorname{dom}(\hat x^{\dagger}) is the same set, so x^\hat x is self-adjoint. (e) Say why no analogue of §5.4's circle appears.

(a) Take ψn\psi_n to be 1[n,n+1]\mathbf 1_{[n,n+1]}, which has ψn=1\norm{\psi_n}=1. Then x^ψn\hat x\psi_n is xx on [n,n+1][n,n+1] and zero elsewhere, so x^ψn2=nn+1x2dxn2\norm{\hat x\psi_n}^{2}=\int_n^{n+1}x^{2}\dd x\ge n^{2} and the ratio in (4.4.2) is at least nn. The supremum is infinite. As with momentum, the reason is physical: there is no largest position.

(b) The largest domain is dom(x^)={ψL2:xψL2}\operatorname{dom}(\hat x)=\{\psi\in L^{2}:x\psi\in L^{2}\}. It is dense because the truncations ψ1[R,R]\psi\mathbf 1_{[-R,R]} all lie in it, since xx is bounded by RR there, and ψψ1[R,R]0\norm{\psi-\psi\mathbf 1_{[-R,R]}}\to0 as RR\to\infty by dominated convergence with ψ2\abs\psi^{2} as the dominating function, which is Chapter 4.3 §4.3.

(c) For u,vu,v in the domain, x^u,v=xuv=xuv\avg{\hat xu,v}=\int\overline{xu}\,v=\int x\,\overline u\,v since xx is real, and that is u,x^v\avg{u,\hat xv}. The integrals converge because xuxu and vv are both in L2L^{2}. There is no integration by parts and therefore no boundary term.

(d) Let udom(x^)u\in\operatorname{dom}(\hat x^{\dagger}) with witness ww, so wv=uxv\int\overline w v=\int\overline u\,xv for every vv in the domain. Taking vv supported in a bounded interval, this says w=xuw=xu almost everywhere on that interval, and letting the interval grow gives w=xuw=xu on the line. Since wL2w\in L^{2} by definition of the adjoint's domain, xuL2xu\in L^{2} and udom(x^)u\in\operatorname{dom}(\hat x). So the two domains coincide.

(e) Because there is no boundary term to cancel. The whole family of §5.4 came from (4.4.13), which exists because integration by parts moves a derivative and leaves a remainder at the ends. Multiplication moves nothing. Running §6.2 confirms it: solving xf=±ix0fxf=\pm\ii x_0f gives (xix0)f=0(x\mp\ii x_0)f=0, so ff vanishes wherever x±ix0x\ne\pm\ii x_0, which is everywhere on the real line, and ff is the zero vector. The indices are (0,0)(0,0) and the extension is unique. Position and momentum on the line are alike in being forced and unbounded, and unlike in where the difficulty sits: for x^\hat x the domain is the only issue, and for p^\hat p the domain is the beginning of the issue.

Worked example 2 — squaring a self-adjoint momentum does not give every box Hamiltonian

On [0,L][0,L] take p^θ\hat p_\theta from §5.4. (a) Write down the domain of p^θ2\hat p_\theta^{2}. (b) Find its spectrum. (c) Show that the Dirichlet Hamiltonian is not p^θ2\hat p_\theta^{2} for any θ\theta. (d) Say what this means about the sentence "kinetic energy is p2/2mp^{2}/2m".

(a) Applying an operator twice requires the intermediate vector to be in the domain again, so dom(p^θ2)={ψ:ψDθ and p^θψDθ}\operatorname{dom}(\hat p_\theta^{2})=\{\psi:\psi\in\mathcal D_\theta \text{ and } \hat p_\theta\psi\in\mathcal D_\theta\}. Since p^θψ\hat p_\theta\psi is ψ\psi' up to a constant, this is

ψL2,ψ(L)=eiθψ(0),ψ(L)=eiθψ(0). \psi''\in L^{2}, \qquad \psi(L)=\ee^{\ii\theta}\psi(0), \qquad \psi'(L)=\ee^{\ii\theta}\psi'(0).

Two conditions, which is the right number for a point of the U(2)U(2) of §7.2, and the family p^θ2/2m\hat p_\theta^{2}/2m is a one-parameter curve inside that four-parameter family.

(b) The eigenvalues are the squares of those in (4.4.18), divided by 2m2m:

En  =  pn22m  =  2(θ+2πn)22mL2,nZ. E_n \;=\; \frac{p_n^{2}}{2m} \;=\; \frac{\hbar^{2}(\theta+2\pi n)^{2}}{2mL^{2}}, \qquad n\in\mathbb{Z}.

At θ=0\theta=0 this is the ring, with a ground state at zero energy and every level above it doubly degenerate, since nn and n-n give the same energy.

(c) Suppose the Dirichlet operator, whose domain requires ψ(0)=ψ(L)=0\psi(0)=\psi(L)=0 and nothing about the slopes, equalled p^θ2\hat p_\theta^{2} for some θ\theta. Two operators are equal only when their domains agree, so compare the domains. The function ψ(x)=sin(πx/L)\psi(x)=\sin(\pi x/L) has ψ(0)=ψ(L)=0\psi(0)=\psi(L)=0, so it is in the Dirichlet domain. Its derivative is (π/L)cos(πx/L)(\pi/L)\cos(\pi x/L), giving ψ(0)=π/L\psi'(0)=\pi/L and ψ(L)=π/L\psi'(L)=-\pi/L, so ψ(L)=eiθψ(0)\psi'(L)=\ee^{\ii\theta}\psi'(0) would force eiθ=1\ee^{\ii\theta}=-1 and hence θ=π\theta=\pi. But then the first condition demands ψ(L)=ψ(0)\psi(L)=-\psi(0), and both are zero, which is consistent, so test a second function. Take ϕ(x)=sin(2πx/L)\phi(x)=\sin(2\pi x/L), also Dirichlet, with ϕ(0)=ϕ(L)=2π/L\phi'(0)=\phi'(L)=2\pi/L, which needs eiθ=+1\ee^{\ii\theta}=+1. No single θ\theta serves both, so the Dirichlet domain is not dom(p^θ2)\operatorname{dom}(\hat p_\theta^{2}) for any θ\theta. The spectra confirm it independently. The lowest level of p^θ2/2m\hat p_\theta^{2}/2m is 2θ2/2mL2\hbar^{2}\theta_{*}^{2}/2mL^{2} with θ=minnθ+2πn\theta_{*}=\min_{n}\abs{\theta+2\pi n}, which never exceeds π\pi, while Dirichlet's lowest is π22/2mL2\pi^{2}\hbar^{2}/2mL^{2}. The two agree only at θ=π\theta=\pi, and there the level is doubly degenerate, from n=0n=0 and n=1n=-1, while every Dirichlet level is simple.

(d) The sentence survives as a statement about formulae and fails as a statement about operators. On the line it is exact, because §5.3 gives a unique p^\hat p and squaring it gives the unique free Hamiltonian. In a box it is not, and part (c) is the whole reason. Momentum observables on [0,L][0,L] are not in short supply, but §5.4 showed that every one of them is some p^θ\hat p_\theta, and part (a) showed that squaring one drags the slope condition ψ(L)=eiθψ(0)\psi'(L)=\ee^{\ii\theta}\psi'(0) along with it. The Dirichlet domain imposes nothing on the slopes, and part (c) turned that difference into two of its own members, sin(πx/L)\sin(\pi x/L) and sin(2πx/L)\sin(2\pi x/L), which demand eiθ=1\ee^{\ii\theta}=-1 and eiθ=+1\ee^{\ii\theta}=+1 of the same θ\theta. So the textbook infinite well has a Hamiltonian that is not the square of any momentum observable, and not for want of momentum observables to square. This is worth carrying into Chapter 4.7. The infinite well is not a free particle with a restriction bolted on. It is a different operator, chosen from a family, and the choice is what the walls are.

Worked example 3 — the half-line flow that runs one way, and Chapter 0.5's fourth item

On L2(R)L^{2}(\R), Chapter 4.2 §7 writes the flow generated by an observable as eiG^a/\ee^{-\ii\hat Ga/\hbar}. (a) Show that the flow generated by p^\hat p on the line is translation. (b) Ask the same question on [0,)[0,\infty) with the domain D0\mathcal D_0 of §5.5, and say what happens for each sign of aa. (c) Identify the resulting family with (4.4.1). (d) Say how this shows up in the deficiency indices.

(a) Expand eip^a/ψ\ee^{-\ii\hat pa/\hbar}\psi formally, using p^=iddx\hat p=-\ii\hbar\dv{}{x}:

eip^a/ψ  =  n01n!(addx)nψ  =  n0(a)nn!ψ(n)(x)  =  ψ(xa), \ee^{-\ii\hat pa/\hbar}\psi \;=\; \sum_{n\ge0}\frac{1}{n!}\left(-a\,\dv{}{x}\right)^{n}\psi \;=\; \sum_{n\ge0}\frac{(-a)^{n}}{n!}\,\psi^{(n)}(x) \;=\; \psi(x-a),

which is the Taylor series of ψ\psi about xx, convergent for analytic ψ\psi and extending to all of L2L^{2} by continuity since each map ψψ(a)\psi\mapsto\psi(\cdot-a) preserves the norm. So p^\hat p generates translation by aa, in either direction, and each translation is unitary. That is the content of (0,0)(0,0): a full group of motions, invertible, with a self-adjoint generator.

(b) On [0,)[0,\infty) the map ψ(x)ψ(xa)\psi(x)\mapsto\psi(x-a) makes sense as a map into L2[0,)L^{2}[0,\infty) only for a0a\ge0. It slides the function to the right, and the vacated interval [0,a)[0,a) is filled with zeros, which joins on continuously precisely because every member of D0\mathcal D_0 vanishes at the origin. For a<0a\lt0 the recipe would need values of ψ\psi at negative xx, and there are none. So there is a flow forwards and no flow backwards.

(c) The forward maps preserve the norm exactly, since sliding does not change ψ2\int\abs\psi^{2}, and they are one-to-one, since sliding back recovers the function wherever it was defined. They are not onto: nothing in the image is non-zero on [0,a)[0,a). That is exactly (4.4.1) with a continuous parameter in place of a discrete one, and it is Chapter 0.5's "the claim that an injective map is surjective" failing in the flesh. Composition adds the parameters, so the family is a semigroup rather than a group: it has an associative law and no inverses.

(d) A self-adjoint generator would give a two-sided flow, by Chapter 4.5 §9's Stone's theorem, and §5.5 showed there is no self-adjoint generator here. The deficiency indices (1,0)(1,0) of §6.3 are the arithmetic of exactly that one-sidedness: one square-integrable solution at +ip0+\ii p_0 and none at ip0-\ii p_0, because the exponential that decays at ++\infty has no partner decaying at a left end that does not exist. The three descriptions are one fact: a boundary term that cannot be cancelled, an index pair that cannot be balanced, and a motion that cannot be run backwards.

9 · Your turn

Problem 1 — the domain is not everything, exhibited

(a) Show that 1[0,1]\mathbf 1_{[0,1]} is in L2(R)L^{2}(\R) and not in dom(p^)\operatorname{dom}(\hat p). (b) Show that ψ(x)=ex\psi(x)=\ee^{-\abs x} is in the domain, and compute p^ψ\hat p\psi. (c) Show that dom(p^)\operatorname{dom}(\hat p) is not closed as a subset of L2L^{2}, by exhibiting a sequence in it converging to the function of part (a). (d) Explain why (c) is required by §2.3 rather than merely permitted by it.

Solution

(a) 1[0,1]2=1\int\abs{\mathbf 1_{[0,1]}}^{2}=1, so it is in L2L^{2}. For membership of the domain it would have to be the integral of some function, and no integral has a jump: if F(x)=cxgF(x)=\int_{c}^{x}g with gg locally integrable, then F(x)F(y)yxg0\abs{F(x)-F(y)}\le\int_y^{x}\abs g\to0 as yxy\to x, so FF is continuous. The indicator is not continuous at 00 or at 11.

(b) On each half-line ex\ee^{-\abs x} is smooth, and at the origin it is continuous with a corner. It is the integral of g(x)=sgn(x)exg(x)=-\mathrm{sgn}(x)\ee^{-\abs x}, which one checks by evaluating 0xg\int_0^{x}g separately for x>0x\gt0 and x<0x\lt0. That gg is in L2L^{2}, with g2=e2x=1\norm g^{2}=\int\ee^{-2\abs x}=1. So ψ\psi is in the domain and p^ψ=isgn(x)ex\hat p\psi=\ii\hbar\,\mathrm{sgn}(x)\ee^{-\abs x}. A corner is allowed; a jump is not.

(c) Let ψn\psi_n be continuous, zero outside [1/n,1+1/n][-1/n,1+1/n], equal to 11 on [0,1][0,1], and linear on the two ramps of width 1/n1/n. Each is in the domain, being piecewise linear and continuous with a step-function derivative in L2L^{2}. The difference from 1[0,1]\mathbf 1_{[0,1]} is supported on the two ramps and bounded by 11 there, so ψn1[0,1]22/n0\norm{\psi_n-\mathbf 1_{[0,1]}}^{2}\le2/n\to0. The limit is not in the domain by (a), so the domain is not closed.

(d) Suppose the domain were closed. Being dense by §3.2, it would then be all of L2L^{2}, and p^\hat p would be a symmetric operator defined on the whole space. Hellinger–Toeplitz would make it bounded, contradicting §2.2. So the domain of any unbounded symmetric operator is necessarily dense and not closed, which is an uncomfortable combination and is the price of the subject. Notice also the sizes in (c). The ramp derivatives have ψn2=2n\norm{\psi_n'}^{2}=2n, so the inputs converge while the images run away, which is §2.1's discontinuity happening in front of you.

Problem 2 — the adjoint of an extension

(a) Prove (4.4.10) from (4.4.9) directly. (b) Deduce that if A^\hat A is symmetric and B^\hat B is a symmetric extension of it, then dom(A^)dom(B^)dom(B^)dom(A^)\operatorname{dom}(\hat A)\subseteq\operatorname{dom}(\hat B)\subseteq\operatorname{dom}(\hat B^{\dagger})\subseteq\operatorname{dom}(\hat A^{\dagger}). (c) Conclude that every self-adjoint extension of a symmetric A^\hat A has a domain lying between dom(A^)\operatorname{dom}(\hat A) and dom(A^)\operatorname{dom}(\hat A^{\dagger}). (d) Use (c) to give a second proof that momentum on [0,)[0,\infty) has no self-adjoint extension, using only the two domains computed in §5.5.

Solution

(a) Let udom(B^)u\in\operatorname{dom}(\hat B^{\dagger}), so there is ww with w,v=u,B^v\avg{w,v}=\avg{u,\hat Bv} for every vdom(B^)v\in\operatorname{dom}(\hat B). Every vdom(A^)v\in\operatorname{dom}(\hat A) is such a vv, and B^v=A^v\hat Bv=\hat Av there, so the same ww witnesses udom(A^)u\in\operatorname{dom}(\hat A^{\dagger}). The containment is that one substitution.

(b) The first containment is the hypothesis. The second is symmetry of B^\hat B in the form (4.4.12). The third is part (a).

(c) A self-adjoint extension B^\hat B is in particular a symmetric extension, and it has dom(B^)=dom(B^)\operatorname{dom}(\hat B)=\operatorname{dom}(\hat B^{\dagger}), so the chain in (b) puts that common domain between the two ends. This is the precise form of §3.4's picture: the search for a self-adjoint extension is a search inside a fixed interval of subspaces, which is why the answer is a finite-dimensional family rather than an open-ended construction.

(d) Section 5.5 computed dom(p^)={v(0)=0}\operatorname{dom}(\hat p)=\{v(0)=0\} and dom(p^)={\operatorname{dom}(\hat p^{\dagger})=\{no condition}\}. Any self-adjoint extension has a domain D\mathcal D with {v(0)=0}D{\{v(0)=0\}\subseteq\mathcal D\subseteq\{no condition}\} and D\mathcal D symmetric. If D\mathcal D is strictly larger than {v(0)=0}\{v(0)=0\} it contains some vv with v(0)0v(0)\ne0, and then (4.4.19) with u=vu=v is non-zero, so D\mathcal D is not symmetric. So D={v(0)=0}\mathcal D=\{v(0)=0\}, which is not self-adjoint. No candidate survives.

Problem 3 — the circle of momenta, and what turning it does

(a) For p^θ\hat p_\theta on [0,L][0,L], verify that the eigenfunctions ψn(x)=L1/2ei(θ+2πn)x/L\psi_n(x)=L^{-1/2}\ee^{\ii(\theta+2\pi n)x/L} are normalised and orthogonal. (b) Show that U^ψ=eiθx/Lψ\hat U\psi=\ee^{\ii\theta x/L}\psi is unitary on L2[0,L]L^{2}[0,L] and maps the periodic domain Dθ=0\mathcal D_{\theta=0} of §5.4 onto Dθ\mathcal D_\theta. (c) Compute U^p^θU^\hat U^{\dagger}\hat p_\theta\hat U and read off the relation between the spectra of p^θ\hat p_\theta and p^0\hat p_0. (d) Two operators related as in (c) are unitarily equivalent. Explain why they are nevertheless different observables, and give the experimental statement that distinguishes them.

Solution

(a) ψn2=1/L\abs{\psi_n}^{2}=1/L, so 0Lψn2=1\int_0^{L}\abs{\psi_n}^{2}=1. For mnm\ne n, ψm,ψn=L10Le2πi(nm)x/Ldx=0\avg{\psi_m,\psi_n}=L^{-1}\int_0^{L}\ee^{2\pi\ii(n-m)x/L}\dd x=0, the integrand being a full number of periods. Each satisfies ψn(L)=eiθψn(0)\psi_n(L)=\ee^{\ii\theta}\psi_n(0) since e2πin=1\ee^{2\pi\ii n}=1.

(b) Multiplication by a function of modulus one preserves ψ\abs\psi pointwise, hence the norm, and it is invertible by multiplication by the conjugate, so it is unitary. If ψDθ=0\psi\in\mathcal D_{\theta=0}, meaning ψ(L)=ψ(0)\psi(L)=\psi(0), then (U^ψ)(L)=eiθψ(L)(\hat U\psi)(L)=\ee^{\ii\theta}\psi(L) and (U^ψ)(0)=ψ(0)(\hat U\psi)(0)=\psi(0), so (U^ψ)(L)=eiθ(U^ψ)(0)(\hat U\psi)(L)=\ee^{\ii\theta}(\hat U\psi)(0). The map is onto Dθ\mathcal D_\theta because U^\hat U^{\dagger} reverses it.

(c) By the product rule, p^(eiθx/Lψ)=eiθx/L(p^ψ+θLψ)\hat p(\ee^{\ii\theta x/L}\psi)=\ee^{\ii\theta x/L}\big(\hat p\psi+\tfrac{\hbar\theta}{L}\psi\big), so U^p^θU^=p^0+θ/L\hat U^{\dagger}\hat p_\theta\hat U=\hat p_0+\hbar\theta/L. The spectrum of p^θ\hat p_\theta is the spectrum of p^0\hat p_0 shifted by θ/L\hbar\theta/L, which is (4.4.18) read again.

(d) Unitary equivalence says the two operators have the same abstract structure, not that they are the same observable of the same system. The measured quantity is momentum in both cases, the states are functions on the same interval, and the sets of possible readings differ by θ/L\hbar\theta/L, which is a number an experiment reports. The equivalence relabels the states by multiplying them by a position-dependent phase, and the physical statement is that this is not a relabelling one is free to perform, because it moves the very observable at issue. Position it leaves untouched. Both U^\hat U and x^\hat x are multiplication by a function, so they commute exactly, and U^ψ2=ψ2\abs{\hat U\psi}^{2}=\abs{\psi}^{2} at every point, which means the relabelled state says precisely what the original said about where the particle is. Momentum is what moves, and part (c) is the computation: conjugating by U^\hat U turns p^θ\hat p_\theta into p^0+θ/L\hat p_0+\hbar\theta/L, so the relabelling shifts every momentum reading by θ/L\hbar\theta/L. Two boxes with different θ\theta are told apart by measuring momentum in each and comparing the ladders.

Problem 4 — the Robin box at the two limits, and one number in between

(a) From (4.4.27), recover the Neumann and Dirichlet ladders as the limits α0\alpha\to0 and α\abs\alpha\to\infty, and say which levels are lost or gained. (b) Show that the negative-energy condition is (α2+κ2)sinhκL+2ακcoshκL=0(\alpha^{2}+\kappa^{2})\sinh\kappa L+2\alpha\kappa\cosh\kappa L=0 and that it has no solution for α>0\alpha\gt0. (c) Show that a level sits at exactly zero energy only for α=0\alpha=0 or α=2/L\alpha=-2/L, by solving ψ=0\psi''=0 subject to the Robin conditions. (d) For αL=1\alpha L=-1, show that there is exactly one negative level and bracket it between 3-3 and 2-2 in units of 2/2mL2\hbar^{2}/2mL^{2}.

Solution

(a) At α=0\alpha=0 the equation is k2sinkL=0-k^{2}\sin kL=0, so kL=nπkL=n\pi with n0n\ge0, and k=0k=0 is a genuine solution with the constant eigenfunction, giving a state at E=0E=0. Dividing (4.4.27) by α2\alpha^{2} and letting α\abs\alpha\to\infty leaves sinkL=0\sin kL=0 again, but now k=0k=0 fails, because the corresponding solution of the original problem is the constant, which cannot satisfy ψ=αψ\psi'=\alpha\psi with α\alpha infinite unless ψ0\psi\equiv0. So the Neumann ladder is n=0,1,2,n=0,1,2,\ldots and the Dirichlet ladder is the same numbers with n=0n=0 removed. The zero-energy state is what the two limits differ by.

(b) Put E=2κ2/2mE=-\hbar^{2}\kappa^{2}/2m, so ψ=κ2ψ\psi''=\kappa^{2}\psi and ψ=Acoshκx+Bsinhκx\psi=A\cosh\kappa x+B\sinh\kappa x. The condition at 00 gives Bκ=αAB\kappa=\alpha A, so take A=κA=\kappa, B=αB=\alpha. Imposing ψ(L)=αψ(L)\psi'(L)=-\alpha\psi(L) and collecting gives the stated identity. For α>0\alpha\gt0 every term is positive when κ>0\kappa\gt0, so the sum cannot vanish. Rearranged, the condition reads tanhκL=2ακ/(κ2+α2)\tanh\kappa L=-2\alpha\kappa/(\kappa^{2}+\alpha^{2}), whose left side is positive and whose right side needs α<0\alpha\lt0.

(c) ψ=0\psi''=0 gives ψ=A+Bx\psi=A+Bx. Then ψ(0)=B=αA\psi'(0)=B=\alpha A and ψ(L)=B=α(A+BL)\psi'(L)=B=-\alpha(A+BL). Substituting the first into the second gives αA=αAα2LA\alpha A=-\alpha A-\alpha^{2}LA, that is αA(2+αL)=0\alpha A(2+\alpha L)=0. A non-zero solution needs A0A\ne0, since A=0A=0 forces B=0B=0, so α=0\alpha=0 or αL=2\alpha L=-2. This is (4.4.29) derived from the eigenfunction instead of from the characteristic function, and the two agree.

(d) With L=1L=1 and α=1\alpha=-1 the condition of (b) is f(κ)=(1+κ2)sinhκ2κcoshκ=0f(\kappa)=(1+\kappa^{2})\sinh\kappa-2\kappa\cosh\kappa=0. For small κ\kappa, expanding gives fκf\approx-\kappa, so f<0f\lt0 near zero. At κ=2\kappa=2, f=5(3.626860)4(3.762196)=18.13430215.048783=3.085519>0f=5(3.626860)-4(3.762196)=18.134302-15.048783=3.085519\gt0, so a root lies between. Tighten the bracket. At κ=2\kappa=\sqrt2, f=3(1.935067)22(2.178184)=5.8052006.160833=0.355633<0f=3(1.935067)-2\sqrt2\,(2.178184)=5.805200-6.160833=-0.355633\lt0, and at κ=3\kappa=\sqrt3, f=4(2.737656)23(2.914577)=10.95062510.096392=0.854233>0f=4(2.737656)-2\sqrt3\,(2.914577)=10.950625-10.096392=0.854233\gt0. So κ2\kappa^{2} lies between 22 and 33 and E=κ2E=-\kappa^{2} lies between 3-3 and 2-2 in these units. The figure's readout gives 2.382098-2.382098, so κ=1.543405\kappa=1.543405. Only one root exists, because ff is negative near zero, positive beyond, and increasing once past its single turning point.

Problem 5 — where the four parameters are, and what they are not

(a) Count the real parameters in the boundary data of a box, and explain why a self-adjoint Hamiltonian corresponds to two complex conditions rather than one or three. (b) Verify directly from (4.4.21) that the periodic and antiperiodic conditions are symmetric, and that a mixed pair, ψ(L)=ψ(0)\psi(L)=\psi(0) with ψ(L)=ψ(0)\psi'(L)=-\psi'(0), is not. (c) The Robin slice of §7.3 uses the same α\alpha at both ends. Show that allowing different values α0\alpha_0 and αL\alpha_L still gives a self-adjoint operator, and say how many of the four parameters that family covers. (d) A student says that since all four parameters give self-adjoint operators, and self-adjointness is the whole of P2, the choice among them cannot matter. Say precisely what is wrong.

Solution

(a) The boundary data is (ψ(0),ψ(L),ψ(0),ψ(L))\big(\psi(0),\psi(L),\psi'(0),\psi'(L)\big), four complex numbers, so eight real ones. A domain is cut out of that data by complex linear conditions, and how many of them there must be is forced from both sides. One condition would leave the boundary term of (4.4.21) non-zero on some pair, so the operator would not be symmetric; three would leave the domain too small, so the adjoint's domain would be strictly larger by (4.4.10) and the operator symmetric but not self-adjoint. Two is the only count that can balance.

(b) With ψ(L)=ψ(0)\psi(L)=\psi(0) and ψ(L)=ψ(0)\psi'(L)=\psi'(0) for both uu and vv, the bracket at LL equals the bracket at 00 term by term, so the difference is zero. With both signs reversed, each factor in each product changes sign, so each product is unchanged and the difference is again zero. For the mixed pair, the bracket at LL is u(0)(v(0))(u(0))v(0)\overline{u(0)}(-v'(0))-\overline{(-u'(0))}v(0), which is u(0)v(0)+u(0)v(0)-\overline{u(0)}v'(0)+\overline{u'(0)}v(0), the negative of the bracket at 00. The difference across the interval is therefore minus twice the bracket at 00, which is not zero in general. Take u=vu=v with u(0)=1u(0)=1 and u(0)=iu'(0)=\ii, which a cubic can arrange along with the two conditions at LL, and the bracket at 00 is 2i2\ii.

(c) The verification in §7.3 treated the two ends separately and used only that the coefficient at each end is real, so it goes through unchanged with ψ(0)=α0ψ(0)\psi'(0)=\alpha_0\psi(0) and ψ(L)=αLψ(L)\psi'(L)=-\alpha_L\psi(L) for real α0,αL\alpha_0,\alpha_L. That is two of the four, once Dirichlet is counted as the limiting value at each end. The remaining two are the ones that couple the ends to each other, which the periodic and antiperiodic conditions use and the Robin family does not. A Robin condition never lets the wavefunction at one wall know anything about the other.

(d) Self-adjointness is a requirement on an operator, not a description of one, and P2 says every observable is self-adjoint rather than that every self-adjoint operator is the observable you want. The four parameters index four different physical situations, distinguished by their energy levels, and §7.4's table shows three of them disagreeing in the first digit. What decides which one describes a given box is the physics of its walls, and no amount of checking self-adjointness will produce that. The same confusion in Chapter 4.2's language would be reading P2 as a converse, which that chapter's §4.4 explicitly declined to assert.

The brick you just laid — an operator is a formula and a domain, and the domain is forced

The restriction is not a choice. Chapter 0.6 proved that every linear map on Rn\R^{n} is bounded and named the derivative as the operator that breaks the proof in infinite dimensions. Section 2.2 measured the break: on the modes eikx\ee^{\ii kx} the norm stays fixed and the norm of the derivative grows without limit, so momentum is unbounded, which is what an unbounded physical quantity has to look like. Section 2.3 then closed the escape route. A symmetric operator defined on the whole of a Hilbert space is bounded, by the closed graph theorem in two lines, so an unbounded observable cannot be defined on the whole space. The domain of p^\hat p is compulsory.

And the adjoint gets a domain nobody chose. Chapter 0.5's defining relation for A^\hat A^{\dagger}, read where A^\hat A acts on a subspace, becomes a test that some vectors pass and others fail, and the set of vectors that pass is dom(A^)\operatorname{dom}(\hat A^{\dagger}). Density is what makes the answer unique, and it is doing that job and no other. The two domains move in opposite directions: enlarging the operator shrinks its adjoint, so a self-adjoint operator is a place where two moving subspaces meet. Symmetric is the containment. Self-adjoint is the equality. Chapter 4.2's P2 was stated with the finite-dimensional word and now reads self-adjoint, which is the correction that chapter's §4 named this section for.

Three intervals, one operator, three different answers, and all of it integration by parts. Keeping the boundary term of Chapter 0.2 §3.2 rather than dropping it gives i[uv]ab\ii\hbar\big[\overline u v\big]_a^{b} as the exact obstruction to symmetry. On R\R it vanishes by itself, both domains are the same set, and momentum is a unique observable. On [0,L][0,L] it vanishes exactly when v(L)=eiθv(0)v(L)=\ee^{\ii\theta}v(0) for one fixed phase, giving a circle of momentum operators whose ladders of allowed values slide as the phase turns, and showing that the natural choice of pinning both ends is symmetric and not self-adjoint. On [0,)[0,\infty) it forces v(0)=0v(0)=0, that domain is not self-adjoint, and nothing larger is symmetric. The momentum of a particle confined to a half-line is not an observable at all, and the argument used nothing beyond v(0)2=0\abs{v(0)}^{2}=0.

The count, quoted after the answers were already in hand. Von Neumann's classification reduces the question to solving A^f=±iλf\hat A^{\dagger}f=\pm\ii\lambda f with no boundary conditions and counting square-integrable solutions. For momentum those solutions are ep0x/\ee^{\mp p_0x/\hbar}, and the indices come out (0,0)(0,0) on the line, (1,1)(1,1) on the interval and (1,0)(1,0) on the half-line, matching §5 three times out of three. Placing the mark after the hand calculation rather than before it was deliberate: the theorem is checked against arithmetic you own.

The number this chapter owed, and a correction with it. The same count run on 2d2dx2/2m-\hbar^{2}\dvn{2}{}{x}/2m over [0,L][0,L] gives (2,2)(2,2), because a second-order equation has a two-dimensional solution space and every solution is square-integrable on a bounded interval. So the particle in a box has a U(2)U(2) of self-adjoint Hamiltonians: four real parameters, not one. The single parameter belongs to momentum, and carrying it across is the mistake §7 exists to prevent. Dirichlet, Neumann, periodic, antiperiodic and the Robin family are all points of that U(2)U(2). Solving the Robin slice gives the transcendental condition tankL=2αk/(k2α2)\tan kL=2\alpha k/(k^{2}-\alpha^{2}), whose roots at α=0\alpha=0 reproduce n2π22/2mL2n^{2}\pi^{2}\hbar^{2}/2mL^{2} to every digit printed and at other settings do not, which is the measurement the table and the figure report.

Two marks, and where completeness was spent. The closed graph theorem in §2.3, used once, for Hellinger–Toeplitz and nothing else. Von Neumann's classification in §6.2, quoted with its hypotheses and then checked against three cases already worked. Nothing else here is asserted without being derived, and in particular the whole of §5, which is the section two sentences of Chapter 4.2 point at, uses neither mark. One mark standing elsewhere is leaned on, Fubini's theorem from Chapter 0.2 §4.2, in the grind box of §5.2, cited there rather than re-raised, which is the same treatment Chapter 4.3 gave Heine–Cantor. Chapter 4.3's closing brick said this chapter would need completeness at every step, and the exact answer is that it was spent twice, both times inside one of those two quoted theorems. The closed graph theorem is false on an incomplete space, and von Neumann's carries a Hilbert space in its hypothesis for the same kind of reason. Everything in between runs on two other things Chapter 4.3 supplied, density and Cauchy–Schwarz, and saying which is which is better than crediting the whole chapter with everything.

Where this gets spent. Chapter 4.5 is the third instalment of the bill Chapter 0.5 named, and it needs this chapter's output as its hypothesis: its spectral theorem is stated for self-adjoint operators, which is a condition about domains and nothing else, and its Stone's theorem in §9 is what makes "time evolution is unitary" and "the Hamiltonian is self-adjoint" the same statement. The shape has now repeated twice and it is worth saying in the same words: Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it. Chapter 4.7 takes §7's U(2)U(2) and spends one paragraph, not one clause, on why the infinite well selects the Dirichlet point of it. Chapter 4.12 needs the domain language to state honestly that single-valuedness of a wavefunction in the azimuthal angle is an assumption about the domain of L^z\hat L_z rather than a theorem. Chapter 4.13 needs §5.5, because the radial coordinate lives on a half-line and the behaviour of the hydrogen wavefunction at the origin has to be argued rather than assumed. What you do not yet have is the values: an operator with no eigenvectors in the space still has a set of possible readings, and naming that set is the whole of Chapter 4.5.