Part IV · Quantum Mechanics — Chapter 4.4
Domains, and the Adjoint's Domain
In infinite dimensions an operator is not a formula. It is a formula together with a domain, the domain is forced rather than chosen, and the word "Hermitian" turns out not to have been enough.
Chapter 4.3 built the space. Square-integrable functions, with two functions that agree almost everywhere counted as one vector, form a complete inner-product space, and the Fourier modes are a genuine orthonormal basis of it rather than an assumption. That is a space. It has no operators on it yet.
This chapter puts them there, and the first thing it has to do is admit that the finite-dimensional recipe does not survive the move. In Chapter 0.5 an operator was a matrix, a matrix acts on every vector there is, and a Hermitian matrix is self-adjoint by the same stroke of the pen. None of those three sentences is true here. The derivative cannot act on every square-integrable function, because differentiating one need not produce another. So it comes with a restricted set of functions it is allowed to act on, and once you allow that, the adjoint acquires a restricted set of its own, fixed by the definition, with no reason at all to be the same one.
Here is the route. Section 1 reads back the sentence Chapter 0.5 left standing and says which of its four items belongs to this chapter and which to Chapter 4.5. Section 2 measures how badly the derivative misbehaves and then proves, from one quoted theorem, that the misbehaviour cannot be designed away: a symmetric operator defined on the whole space is bounded, so an unbounded observable must be restricted. Section 3 defines the domain and the adjoint's domain carefully. Section 4 separates symmetric from self-adjoint and corrects the second postulate of Chapter 4.2 in the section that chapter names. Section 5 is the payoff, and it is arithmetic you can check by hand: momentum on the line, on an interval and on a half-line, with the boundary term from integration by parts kept rather than waved through. Only then does Section 6 quote the theorem that counts the answers. Section 7 does the same for the particle in a box, where the answer is richer and one widely repeated statement about it is wrong.
Conventions. The inner product is linear in its second slot, as Chapter 0.5 §1.1 chose it, and throughout. Every integral is Chapter 4.3's. Two results are quoted rather than proved, and both are marked where they arrive: the closed graph theorem in §2.3, and von Neumann's classification of self-adjoint extensions in §6.2. There are no others, and one mark standing elsewhere is leaned on once, Fubini's theorem from Chapter 0.2 §4.2, inside the grind box of §5.2. The closing brick says where completeness was spent and what each mark bought.
Tools you'll need — Chapter 4.3 above all: §5 for and the quotient that repaired the third inner-product axiom, §5.4 for Cauchy–Schwarz carried over intact, §6 for completeness, §7.4 for the dense set of step functions, and §8 for the Fourier basis on an interval. Chapter 0.5 §4 for the definition of the adjoint, which this chapter reads more carefully than that chapter needed to, §6 for the spectral theorem in finite dimensions, and its closing paragraph, which is the debt §1 collects. Chapter 0.2 §3.2 for integration by parts and for the sentence that sent its boundary term here. Chapter 0.6 §2 for the lemma that every linear map on is bounded, and for the clause naming the derivative as the operator that breaks it. Chapter 4.2 §4 for the second postulate and the warning attached to it, §7 for time evolution, and §8 for the three lines that force infinite dimension. Chapter 0.9 §5 for what integration by parts means when the derivative is taken away from the function and given to the test function instead.
1 · The four places Chapter 0.5 used finite dimension
Here is where this section is going. Chapter 0.5 finished by naming four steps in which its proofs leaned on the dimension being finite, and it said the bill would come due in this chapter and the next. We read that sentence back and take the four in order. By the end you will know which one belongs here, which belong to Chapter 4.5, and why the repair could not be done in a single chapter. The fourth item is the one to watch, because it is the only one with no finite-dimensional shadow at all.
1.1 · The sentence, read back
Chapter 0.5's last paragraph is short enough to quote whole, and it is worth having in front of you rather than in memory:
"One thing remains outstanding, and it is worth naming so that you notice when it is paid. Every proof in this chapter used finite dimension: in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective. Quantum mechanics happens in infinite dimensions. Chapters 4.4 and 4.5 are where the bill comes due."
Chapter 4.2 §8 then made the second sentence unavoidable rather than atmospheric. If held on a space of finite dimension , taking the trace of both sides would give on the left, because the trace of a commutator of matrices vanishes, and on the right. Three lines, no physics beyond the commutator itself, and quantum mechanics is infinite-dimensional before a single measurement has been described. Chapter 4.3 built the space that follows from that. This chapter and the next have to rebuild the operators.
1.2 · The four, one at a time
Take them in Chapter 0.5's own order, because the order is roughly the order of increasing damage.
The induction. Chapter 0.5 §6.3 proved the spectral theorem by finding one eigenvector, restricting the operator to the orthogonal complement of that eigenvector, and repeating on a space one dimension smaller. The recursion halts because the dimension is a positive integer and cannot decrease forever. Strip out finiteness and two things break rather than one. The recursion no longer halts, which is the visible failure. The first step also fails, which is the serious one: there need not be any eigenvector to start from. The operator "multiply by " on has none at all, since forces to vanish wherever , and a function supported at one point is the zero vector of Chapter 4.3 §5.3. That failure is Chapter 4.5's, and it is why that chapter has to widen the word spectrum before it can state a theorem at all.
Rank–nullity. The identity is not so much false in infinite dimensions as empty, because both sides can be infinite and an equation between two infinities constrains nothing. What that identity was for was converting a statement about the kernel into a statement about the image, and the replacement is not a theorem about dimensions at all. It is the observation that an image can be dense without being closed, so that coming arbitrarily close to every vector and reaching every vector are two different achievements, and counting cannot tell them apart. Chapter 4.5 is where that becomes the difference between the parts of the spectrum, because can be one-to-one with a dense image and still have no bounded inverse.
The interchange of sums. Chapter 0.5 rearranged double sums freely, which finitely many terms always permit. The tempting move is to hand this item to Chapter 4.3's convergence theorems and walk away, but those license the exchange of a limit with an integral, whereas rearranging exchanges two sums, and an exchange of two sums is an exchange of two integrals in disguise. Chapter 4.3 §4.1 is explicit about that one: exchanging two integrals is Fubini's theorem, and it stays where Chapter 0.2 §4.2 marked it. What pays this item without raising anything new is the corollary the same section proves for non-negative terms, . Apply it to the functions , whose integral is , and the two sides of the corollary are the two orders of summation. That settles it whenever every , and a double sum that converges absolutely follows by splitting it into positive and negative parts the way Chapter 4.3 §3 splits a function before integrating it. This item is paid, and the material that pays it was in place before this chapter opened.
Injective does not imply surjective. A one-to-one linear map of a finite-dimensional space to itself is onto, and Chapter 0.5's grind box on the converse of used exactly that, in the words "being injective on a finite-dimensional space it is onto, hence unitary on ". It is false here, and the counterexample is small enough to hold in one hand. Take an orthonormal basis of the space, which Chapter 4.3 §7.4 says can always be listed, and let move every basis vector one place along that list:
Nothing non-zero is sent to zero, so is injective, and by Parseval it preserves the norm exactly, since the same coefficients appear in the answer. Yet is not in its image, because every vector has coefficient zero against . So is one-to-one and not onto, and in finite dimensions no such map exists. Hold on to this one. It returns in §5.5 wearing physical clothes, as the reason a particle confined to a half-line has no momentum observable, and again in Worked example 3, where turns out to be a translation with nowhere to come from.
1.3 · Which chapter pays which
The division follows from the four items and is worth stating before you meet it rather than after. This chapter is about the first half of the word "self-adjoint": what an operator is allowed to act on, and what its adjoint is therefore allowed to act on. That is the machinery the induction's first step will need, and it is a prerequisite for the spectrum rather than a part of it. Chapter 4.5 is about the values. It widens spectrum so that an operator with no eigenvectors still has one, states the theorem that replaces , and checks it in every case this book uses.
Neither half stands alone. You cannot ask what the spectrum of an operator is until you have said which operator, domain included, and Chapter 4.5's theorem has self-adjoint in its hypothesis, which is a statement about domains and nothing else. So the bill Chapter 0.5 named is settled across two chapters, in that order, for that reason.
A chapter of Part 0 built the whole theory of matrices and eigenvectors, and it ended with an honest confession: four of its arguments had used the fact that there were only finitely many directions to work with. Quantum mechanics does not have finitely many directions, and the previous chapter proved that in three lines rather than asserting it. So the confession has to be settled.
Of the four, one has already been paid. The freedom to rearrange a double sum was bought in the previous chapter: one of its convergence theorems has a corollary that lets a sum be moved through an integral whenever the terms are all positive, and rearranging a double sum is that same move in disguise. Two more belong to the next chapter, and they are both about the same thing: in a space of functions an operator may have no eigenvectors whatsoever, and the machinery built for matrices starts by finding one.
The item that belongs here is the strangest of the four, and it has no counterpart in the finite world at all. A map can be perfectly one-to-one and still miss part of the space it maps into. Nothing is lost and yet something is unreachable. In a room with finitely many directions that cannot happen, so the situation has no picture attached, and the way to get one is to watch it happen. It happens twice in this chapter, and the second time it is the reason a particle held on one side of a wall has no measurable momentum at all.
2 · Bounded, unbounded, and why the derivative cannot be tamed
Section 1 closed Chapter 0.5's account. The debt this section collects is a different one, owed to Chapter 0.6, which proved that every linear map on is bounded and named the derivative as the map that breaks the proof once the dimension is infinite. The plan for this section is to make one word precise and then use it to prove something that sounds like an inconvenience and is really a structural fact. We define what it means for an operator to be bounded and show that it is the same thing as being continuous. We then measure the derivative and find that it is neither. Finally we quote one theorem and get from it, in two lines, the statement that decides the shape of this whole chapter: an unbounded operator that we want to be an observable cannot be defined on the whole space, no matter how cleverly we try.
2.1 · The operator norm, and what it measures
Chapter 0.6 §2 proved that every linear map on satisfies for some constant depending only on , and the proof wrote in a basis and used Cauchy–Schwarz on the finitely many coefficients. The property that proof establishes has a name and a number attached. For a linear defined on a subspace of a Hilbert space, put
which is the largest factor by which can stretch anything. Call bounded when this number is finite and unbounded when it is not. The supremum in (4.4.2) is over a set of non-negative real numbers, so it always exists in the extended sense, and the only question is whether it is a number.
Before using the word it is worth knowing what it means, because "bounded" sounds like a technical convenience and it is not one. It is continuity, under another name. One direction takes a line: if is finite then
so nearby inputs have nearby outputs, with a uniform rate. The converse takes three lines and is worth having because it is what makes the word carry weight. Suppose is infinite. Then for each there is a unit vector in the domain with . Scale it down by setting . Now , so the inputs march into the origin, while , so the outputs run away. An unbounded operator is therefore discontinuous, and discontinuous at every point rather than at some awkward ones, since linearity carries the failure at the origin everywhere else.
That is the honest content of the word. An unbounded observable is one whose value cannot be controlled by controlling the state to within a small error in the norm of the space, and the next subsection shows that momentum is one.
2.2 · The derivative is unbounded, exhibited
Chapter 0.6 §2 named the offender in advance, saying that "in infinite dimensions (Chapter 4.4) linear maps can be unbounded, being the standard offender, and a great deal of quantum mechanics' technical difficulty descends from exactly that." The exhibit takes one family of functions. Work on and take the exponential modes, with a multiple of , which are the basis Chapter 4.3 §8 built on with the interval rescaled:
Read the three parts of (4.4.4) together and the conclusion is immediate. The norm of does not depend on at all, while the norm of is , so the ratio in (4.4.2) is and runs over without limit. The operator norm of is infinite. Every one of these functions is smooth, periodic and as well behaved as a function can be, so the failure is not caused by anything pathological in the inputs.
The same conclusion holds on the whole line, and it needs one extra step because the modes themselves are not square-integrable there. Fix any smooth that vanishes outside a bounded interval and has , and set . The modulus of is the modulus of , so for every . Differentiating the product gives , and the triangle inequality then bounds the second term away rather than adding to it:
So is unbounded on too, and the reason is not subtle. The number in (4.4.4) is a momentum, the operator norm of would be the largest momentum a state can have, and there is no largest momentum. Unboundedness is what a physically unbounded quantity looks like when you write it as an operator. The same argument, run on multiplication by , says the same thing about position, and Worked example 1 runs it.
Take a measured concentration–time curve and add to it a small ripple, , with tiny and large. In the norm this book uses, the contaminated curve differs from the clean one by about , independently of , because the amplitude of the ripple is all that the norm sees. Differentiate both curves and the difference between the answers is , whose size is . The mathematics is identical to (4.4.4), with the same family of functions doing the same job.
That is the whole of why nobody differentiates raw data. The error in the input is bounded and the error in the output is not, so any high-frequency noise, however small in amplitude, arrives in the derivative multiplied by its frequency. Smoothing first is not cosmetic tidying. It is the restriction of the derivative to a set of inputs on which the operation is stable, which is exactly the move §3 is about to make and give a name to.
Two terms of the analogy have no counterpart, and both are worth saying out loud. The first is the freedom. In data analysis you choose the smoothing, and choosing more of it is always available. In quantum mechanics the operator is fixed by physics and the restriction has to be chosen so that the operator stays an observable, which turns out to leave very little freedom and sometimes none. That is §5.
The second break is the load-bearing one, because it is about what the restriction buys. Smoothing genuinely buys stability: the smoothed derivative has a worst case, so a small error going in stays a small error coming out. Restricting the domain buys nothing of the kind. Every mode in (4.4.4) lies in the domain §3.2 is about to write down, so the ratio in (4.4.2) is still and is exactly as unbounded on its domain as it was on all of . The restriction is not a cure for the instability, and §2.3 is where you find out that it is compulsory anyway.
2.3 · Unboundedness cannot be designed away
The natural response to §2.2 is to look for a better formulation. The hope would be that momentum has been written down clumsily, and that some equivalent operator, defined on every state, does the same physical job while staying continuous. This subsection closes that door, and closes it completely, using one theorem that this book quotes rather than proves.
We use, without proof, the following. Let be a complete inner-product space and let be a linear map defined on all of , with values in . Suppose has this property, which we will call the closed-graph condition:
whenever in and the images converge to some , the limit is .
Then is bounded.
What the hypothesis is doing. Continuity would say that converges and converges to the right thing. The closed-graph condition assumes only the second half: it says nothing about whether converges, and only rules out its converging to the wrong place. The theorem is that on a complete space the weaker demand implies the stronger one, and completeness is not decoration in that sentence. The result is false without it, and this is the first of the two places this chapter spends what Chapter 4.3 §6 proved. The standard proof runs through the Baire category theorem in about a page, and Rudin's Functional Analysis is where to read it. This book uses the theorem exactly twice, in the two lines below and once more in Chapter 4.5 §2.4, where it is what rules out a fourth way for a point of the spectrum to fail, and never again.
Now the two lines. Let be defined on all of and satisfy for every pair of vectors, which is Chapter 0.5's Hermitian condition with nothing added. Call an operator satisfying that condition symmetric. The word is worth having now, before it is needed in anger: §4.1 states it again once an operator has a domain to state it on, and the whole force of this chapter is in the difference between the two statements. We check the closed-graph condition. Suppose and . For an arbitrary in the space, move across, take the limit, and move it back:
where each limit passes through the inner product because Cauchy–Schwarz makes the inner product continuous in each slot, as Chapter 4.3 §5.4 transferred it. So for every , and choosing gives . The closed-graph condition holds, the quoted theorem applies, and is bounded. That is Hellinger–Toeplitz, and it is worth stating on its own line because the rest of the chapter is a consequence of it.
A symmetric operator defined on all of a Hilbert space is bounded. Contrapositive: an operator that is unbounded, and that we intend to be an observable, cannot be defined on all of the space. Not "is awkward to define". Cannot.
So the restricted domain of is not a piece of caution, a smoothness assumption, or a convenience that a more careful formulation would remove. It is forced, by a theorem, from two facts we already have: that momentum is unbounded, which §2.2 exhibited, and that an observable satisfies Chapter 4.2's postulate P2. Everything that follows in this chapter is bookkeeping about a restriction we are not free to decline.
Some operations stretch, and the useful question is whether there is a worst case. If there is a single number such that no input is ever stretched by more than that factor, the operation is stable: a small error going in produces a small error coming out, and the two smallnesses are tied together by that one number. If there is no such number, small errors can be amplified without limit, and stability is gone.
Differentiation has no such number, and the reason is worth keeping rather than the proof. Wiggles of any frequency have the same size as functions and get bigger the faster they wiggle, so the faster the wiggle the larger the derivative, with no ceiling. Anyone who has tried to take a derivative of noisy measurements has met this in person. It is also the correct statement about momentum, because there is no largest momentum a particle can have, and an operator representing an unlimited quantity had better be unlimited.
Then comes the move that decides everything else. You might hope to write momentum down some other way and recover stability. A theorem says no. Any operation that acts on every state whatever, and that has the symmetry a measurable quantity must have, is automatically stable. Turn that around: an unstable measurable quantity cannot act on every state. So the list of states that momentum is allowed to act on is not a hedge and not a technicality. It is compulsory, and the rest of the chapter is about what it costs.
a natural place to stop · the restriction is compulsory; what follows is the vocabulary for it
3 · Domains, and the adjoint's domain
What follows is the vocabulary the rest of the chapter uses, and it is short. An operator is promoted from a formula to a pair, a formula and a set of vectors it acts on. We fix the domain of and check that it is dense, saying what "dense" is for. Then we re-read Chapter 0.5's definition of the adjoint with that in hand, and find that the definition hands the adjoint a domain of its own, determined but not chosen, with no reason to agree with the one we started from. The last subsection draws the consequence that makes the whole of §5 predictable in advance.
3.1 · An operator is a formula and a domain
From here on, an operator on a Hilbert space means two things given together: a subspace , called its domain, and a linear map from that subspace into . Two operators are the same operator when the formulae agree and the domains agree. If the domains differ, they are different operators, even when the formula is the same symbols in the same order.
That convention looks pedantic on first meeting and it is the entire subject of the chapter. In finite dimensions it costs nothing, because a matrix acts on everything and there is only one domain available, which is why Chapter 0.5 never mentioned the word. Here the choice is real, the choice changes the physics, and §7 will show two operators with the same formula, both perfectly legitimate, whose energy levels are different numbers.
When has a larger domain than and agrees with everywhere on , we say extends . That word does the work in §6, where the question is how many self-adjoint extensions a given operator has.
3.2 · The domain of , and what "dense" is for
Momentum needs a domain, and §2.3 says the domain cannot be everything. The natural choice is the largest set on which the formula makes sense and lands back in the space. Take
The phrase "is the integral of a function" is the fundamental theorem of Chapter 0.2 §2 used as a definition rather than as a conclusion, and it is what lets be non-smooth while still having a derivative that Chapter 4.3's integral can handle. A function with a jump in it fails the condition, because no integral has a jump. The condition is what makes a vector of the space rather than merely a function.
One property of this set matters more than any other, and it is the reason the next subsection works at all. The domain is dense: every vector of is a limit of vectors in it. On an interval this is one sentence, since the modes of (4.4.4) all lie in the domain and Chapter 4.3 §8 proved that finite combinations of them reach everything. On the whole line the argument is Chapter 4.3's as well, and it is already written down there. That chapter's §7.4 showed that step functions are dense, and its §8.1 then replaced the vertical sides of an indicator by straight ramps of width , at a cost of at most in squared distance. Those trapezoids are exactly what is wanted here: each is continuous, each is the integral of its own step-function derivative, and that derivative is bounded and supported on a bounded set, hence in . So every trapezoid is in (4.4.7), and density transfers along Chapter 4.3's own chain.
Now the reason to care. Density is not a technical hygiene condition. It is what makes the adjoint exist as a function of at all, and the next subsection is where you can watch it do that job.
3.3 · The adjoint, with a domain of its own
Chapter 0.5 §4 defined the adjoint by the relation , and in finite dimensions that relation determines a matrix, namely the conjugate transpose, with nothing further to say. Read the same relation here, where acts only on its own domain, and it stops being a formula for and becomes a demand that a given may or may not be able to meet:
The two labels under (4.4.8) are the whole of the difficulty. What is being demanded is a single vector whose inner product against every in the domain reproduces those computable numbers, and sometimes no vector in the space does that job. So the definition sorts the vectors of into those for which the demand can be met and those for which it cannot, and admitting the first sort and refusing the second is what defines :
with defined to be that . Here is where density earns its keep. If two vectors and both satisfied (4.4.9), then for every in the domain, and because the domain is dense and the inner product is continuous, the same holds for every in the whole space. Taking gives . Without density there would be a direction the test vectors never probe, could be shifted along it freely, and would not be a well-defined vector. That is the whole job of the word dense, and it is done here and nowhere else.
This is the sentence the chapter turns on, so it is worth saying without any equation attached. The definition of the adjoint determines its domain. Nobody chooses it, and there is no reason whatever for it to coincide with the domain we started from. In finite dimensions the two coincide because both are the whole space and there is nothing to compare. Here they are two different subspaces produced by two different considerations, and the rest of the chapter is the study of when they happen to agree.
3.4 · The two domains move in opposite directions
One structural fact makes everything in §§5 to 7 predictable before you compute anything, and it follows from (4.4.9) by reading the words rather than the symbols. Suppose extends , so that . The condition on in (4.4.9) then has to hold for more vectors , which is a harder test, so fewer vectors pass it. In symbols,
Enlarging an operator shrinks its adjoint, and the two move towards each other. That single observation is why the family of self-adjoint operators built from a given formula is a family at all, rather than a single answer or none: you start with a domain too small and an adjoint domain too big, and you enlarge the first, which shrinks the second, until they meet. Whether they can be made to meet, and in how many ways, is the question §6 answers and §5 answers first by hand.
An operator is no longer a rule. It is a rule together with a list of the states it is allowed to be applied to, and changing the list changes the operator even though the rule is written with the same symbols. That sounds like fussiness and it is the opposite. The list is where the physics of a boundary lives, and two lists with the same rule will shortly give a box two different sets of energy levels.
The partner of an operator, its adjoint, was defined in Part 0 by a single requirement: moving the operator from one side of an inner product to the other must not change the answer. In a space of finitely many directions that requirement pins down a matrix and there is nothing further to discuss. In a space of functions the same requirement has to be read as a question asked of each state separately: is there a vector that does this job for this state? For some states there is and for others there is not, so the requirement quietly hands the partner its own list, which nobody chose and which has no reason to match the first.
One more thing, and it is the thing that makes the rest predictable. Lengthen the first list and the second list gets shorter, because there are now more conditions for a state to satisfy. The two lists move towards each other. An operator worthy of representing a measurement is one where they coincide, and the interesting question is how many ways there are to make that happen.
4 · Symmetric is not self-adjoint
The two conditions in the title are the same condition in finite dimensions, and Chapter 4.2 §4 used the finite-dimensional word. This section separates them, says what each one buys, corrects the second postulate as that chapter announced would happen here, and is honest about what the difference is not: it is invisible to every test you would naturally run, including the test that measured values come out real. Section 5 then makes the difference bite on an operator you can check by hand.
4.1 · Two conditions, and how far apart they are
Section 2.3 has met the first of the two words already, for an operator that acts on everything. With domains in hand it can be said in the form the rest of the chapter needs. Call symmetric when moving it across the inner product changes nothing, for vectors it is allowed to act on:
Compare (4.4.11) with the definition of the adjoint in (4.4.9) and read off what symmetry says in that language. It says that every in passes the test defining , with as the witness. So symmetry is exactly the statement that
In the vocabulary of §3.1, a symmetric operator is one whose adjoint extends it. Call self-adjoint when the containment in (4.4.12) is an equality, so that with the domains included and not merely the formulae. The gap between the two words is therefore a gap between two subspaces, and measuring it on one operator after another is what the rest of the chapter does. Each measurement is of a stated operator, since the gap belongs to the domain and not to the formula. Section 5.3 measures it for momentum on the line, on the largest domain the formula allows, and finds no gap at all. Section 5.4 measures it on a bounded interval, for momentum pinned to zero at both ends, and finds a gap of two complex dimensions. It then closes that gap in a circle of ways, and each of them is an operator with no gap left. Section 5.5 measures it on a half-line, for the largest symmetric domain available there, finds a gap of exactly one complex dimension, and shows that nothing closes it.
In finite dimensions this distinction cannot be drawn, because both domains are the whole space and (4.4.12) is an equality the moment it is a containment. Chapter 0.5 was therefore right to use one word, and Chapter 4.2 inherited it correctly for the finite-dimensional systems that chapter's examples used.
4.2 · The second postulate, corrected
Chapter 4.2 §4 stated P2 as an observable is a Hermitian operator, and attached a warning saying the word would be sharpened here, in this section, for a reason that is physical rather than pedantic. The correction is now available and it is one line.
An observable is a self-adjoint operator: with . Symmetric is not enough. Everywhere Chapter 4.2 wrote "Hermitian", read this instead, and nothing in that chapter's arguments changes, because every operator it worked with acted on a space of finite dimension where the two words agree.
The postulate is no more assumed than it was. What has changed is which mathematical condition the assumption is attached to, and Chapter 4.2's own box said this would happen and named this section. The physical reason is §5.5: momentum on a half-line is symmetric and is not self-adjoint, and there is no repair. If P2 said only "symmetric", that operator would qualify as an observable and the theory would be committed to a measurable quantity for which Chapter 4.5's spectral theorem gives nothing and Chapter 4.5 §9's Stone's theorem gives no time evolution.
4.3 · What the difference is not visible to
It is worth naming the difficulty out loud here, because the difference between the two words is invisible to every test a physicist would think to apply, and a reader who does not know that will reasonably suspect the distinction of being empty.
Consider the property that made Hermitian operators attractive in the first place. If is symmetric and for a non-zero in the domain, then (4.4.11) with gives , so is real. Chapter 0.5 §6.1's argument transfers with no change and needs only symmetry. Symmetric operators also have orthogonal eigenvectors for distinct eigenvalues, by Chapter 0.5 §6.2's argument, again unchanged. And every expectation value is real for the same reason. Chapter 4.2 §4.1 collected three theorems as the payoff of P2, and the first two of them are already true for a merely symmetric operator.
The third one is where the words part company, and it is the one you cannot check by making a measurement. Chapter 4.2 §4.1's third theorem was every state is a superposition of outcomes, and its proof there was Chapter 0.5 §6.3's orthonormal basis of eigenvectors. That is the spectral theorem, and the spectral theorem is exactly what symmetry fails to buy. Two structures the rest of Part IV stands on go with it.
- The spectral theorem. Chapter 4.5's hypothesis is self-adjointness, not symmetry, and there are symmetric operators for which the conclusion is false. Without it there is no resolution of the identity, so no way to write a state as a superposition of the outcomes of a measurement, so no Born rule for that quantity.
- Stone's theorem. Chapter 4.5 §9 gives as a unitary flow exactly when is self-adjoint. A merely symmetric Hamiltonian does not generate one, so the state after a finite time is undefined rather than wrong.
Notice the shape of that. The two facts a reader would test by hand, on a single measurement, survive the weaker condition. The one that fails is a statement about all states at once, and no single measurement can see it. That is why the next section works entirely by hand: the only way to believe the distinction is to watch it happen on an operator you can differentiate.
Two conditions have been separated that Part 0 could treat as one. The weaker says that moving the operator from one side of an inner product to the other changes nothing, so long as both states involved are on the allowed list. The stronger says that, and adds that the partner operator's allowed list is the same list, not a longer one.
The uncomfortable part, and it is worth sitting with, is that the weaker condition already delivers everything a working physicist would check on a single measurement. Measured values come out real. Different values belong to perpendicular states. Averages are real numbers. All of that is in hand without the stronger condition, so no calculation you would naturally perform can tell you which of the two you have.
What the stronger condition buys is not a reassurance but a structure, and it is a statement about every state at once rather than about any one of them. Only for the stronger one is there a guarantee that the possible outcomes of a measurement account for the whole space, and only for the stronger one does the operator generate a flow in time. Lose it and you do not get wrong answers. You get no answers, because the machinery that turns an operator into predictions has a hypothesis it does not meet. The next section shows this happening to the most familiar operator in the subject, on the most familiar interval.
a natural place to stop · the vocabulary is built; what follows is one operator on three intervals
5 · Three intervals, one operator, worked by hand
This is the section the chapter exists for, and everything in it is integration by parts. We redo Chapter 0.2's calculation with the boundary term kept instead of waved through, identify what the adjoint of is when no boundary condition has been imposed at all, and then ask on which functions the boundary term vanishes. Three intervals give three different answers. On the line there is nothing to impose. On a bounded interval there is exactly a circle of possible impositions, and you will see the whole circle. On a half-line the boundary term cannot be killed without killing the operator, and the conclusion is that momentum there is not an observable. No classification theorem is used anywhere in this section; §6 supplies one afterwards and it will have to agree with what you find here.
5.1 · The boundary term, kept this time
Chapter 0.2 §3.2 listed the places integration by parts does structural work, and one of the entries was a promise about this chapter: "An operator gets an adjoint. In Chapter 4.4, when boundary terms vanish. The minus sign is why is Hermitian and alone is not." That chapter needed the identity with the boundary term gone. Here the boundary term is the whole content, so we keep it.
Work on an interval , where may be and may be . For and with , Chapter 0.2's identity reads . We want that in terms of rather than of , so multiply through by . The left side becomes . On the right the integral becomes , which is , because conjugating the constant flips its sign: . Moving that term across leaves
One point about the hypotheses is worth settling here rather than leaving it to be noticed. Chapter 0.2 proved integration by parts for continuously differentiable functions, and the members of (4.4.7) need not be differentiable everywhere. The extension is a single step: write each function as the integral of its derivative, substitute, and split the resulting square of integration into its two triangles, whereupon each side of the identity turns into the other. That step is Fubini's theorem, the standing mark of Chapter 0.2 §4.2, and it is the same interchange the grind box of §5.2 names below. Nothing else is added.
Read (4.4.13) as the definition of symmetry rather than as an identity about integrals, and the section is already organised. By (4.4.11), an operator with the formula is symmetric on a given domain precisely when the right-hand side vanishes for every pair drawn from that domain. Everything below is that one question, asked on three intervals.
5.2 · What the adjoint is before any condition is imposed
Before choosing domains it is worth knowing the answer to a question that does not depend on the choice: which satisfy (4.4.9), given that the domain of contains at least the functions that die away before either end of the interval is reached? Every domain considered below contains those, whatever it does at the ends, so the answer is the same in all three cases, and it is the most permissive one available.
Write for the members of (4.4.7) that vanish outside a closed subinterval of the open interval . If , then every in is an integral of a function , and . No condition at the endpoints comes out of this step. The grind box proves it, and the one part of it that is not integration by parts is the fact that a function whose derivative vanishes in this weak sense is constant, which is done there in full.
That leaves one further requirement, and it is the whole of the three subsections below. Testing against alone cannot see the ends of the interval, so it cannot decide which of those are admitted. What decides it is that the defining relation must hold against every in the domain rather than only against those that die away first, and the difference between the two demands is precisely the boundary term of (4.4.13). So the adjoint's domain is the maximal set cut down by the condition that the boundary term vanish, and the three intervals cut it down by different amounts.
Grind box — identifying the adjoint of the derivative, once, for all three intervals
Recall : the continuous functions that vanish outside a closed subinterval of and are integrals of a derivative lying in . It contains §3.2's trapezoids, so it is dense in over every bounded subinterval by the chain Chapter 4.3 §8.1 built, and it is contained in every domain this section considers, since its members vanish at both ends. Let with , so for all and in particular for all . Put , which lies in . Dividing the relation by and conjugating the constant gives
Fix and set , which is finite for in because functions are integrable on bounded subintervals by Cauchy–Schwarz. Interchanging the two integrations then gives for the same . That interchange is Fubini's theorem, and it is worth naming rather than performing quietly. Chapter 4.3 §4.1 says in as many words that it settles the exchange of a sum with an integral and leaves the exchange of two integrals where Chapter 0.2 §4.2 marked it, so we are spending that standing mark here rather than raising a new one. The version needed is the undemanding one: the region is a bounded triangle on which the modulus of the integrand has finite integral, which is the non-negative case Chapter 0.2 noted holds with no extra hypotheses. Subtracting,
The constancy step. Fix with . Given any , put and . Then , and vanishes outside a closed subinterval of , because the total integral of is zero, so returns to zero beyond both supports. Hence and
where is one fixed number, the middle step being the displayed identity applied to . So for every , and by the density noted at the top, almost everywhere. Therefore .
Conclusion. is an integral of plus a constant, so is the integral of a function, that function is , and . No endpoint condition was used and none was produced, so nothing constrains at or beyond membership of . Conversely every such is in when the boundary term of (4.4.13) vanishes against the domain, which is checked case by case below.
Two remarks. The argument used only test functions supported strictly inside the interval, which is why it is indifferent to what happens at the ends and why one proof covers all three cases. And it is the honest content of the phrase "integrate by parts and drop the boundary term": the dropping is legitimate against test functions that vanish there, and against nothing else without a reason.
5.3 · The line: nothing to impose
Take and the domain (4.4.7), which is the largest one the formula admits. The boundary term in (4.4.13) is a limit at each end, and the claim is that both limits are zero for every pair in the domain, with no condition imposed at all.
Here is the argument, and after (4.4.13) it uses Cauchy–Schwarz twice and nothing else. Put . Read (4.4.13) on the bounded interval and divide by : it says exactly that . Now bound the two pieces. Both and are in , so is in by Cauchy–Schwarz, meaning is finite. The integrand above is a sum of two products of functions, so it is in for the same reason, and its integral over therefore converges as . So has a finite limit. A function with a finite limit at infinity whose modulus has a finite integral must have that limit equal to zero, since a non-zero limit would make diverge. So , and the same at .
The boundary term therefore vanishes identically and is symmetric on the domain (4.4.7). Now compare that domain with the adjoint's. Section 5.2 says every in the adjoint's domain is an integral of a function in , which is membership of (4.4.7), and the further requirement it left open is that the boundary term vanish, which the paragraph above has just shown it does for every such . Nothing is cut down. The two domains are the same set, so
There is nothing to choose, no parameter, and no family. The momentum of a particle on a line is an observable, uniquely, and the reason it looked as though no work was needed is that in this one case no work is needed. Keep the feeling of that, because the next two cases are the same operator and the same calculation with different endpoints, and they do not come out this way.
5.4 · The interval: a circle of momentum operators
Now take with finite. The boundary term in (4.4.13) is no longer a limit, it is two numbers, and symmetry on a domain is the requirement
Start with the choice almost everyone makes first, which is to pin the wavefunction to zero at both ends: . Condition (4.4.15) holds, since both sides are zero, so this operator is symmetric. Now ask what its adjoint is. By §5.2 the adjoint's domain is every with and no endpoint condition, because the boundary term is already zero for every such once is pinned. So is strictly larger than , by every function that fails to vanish at an end.
That is the promised example, and it is worth pausing on. Take the claims one at a time. The operator with the wavefunction pinned at both ends is symmetric, by the paragraph above. Its expectation values are real, by §4.3, and so would any eigenvalue be. It passes that eigenvalue test vacuously, because forces , which never vanishes and so cannot be pinned at either end, leaving the operator with no eigenvalues at all. And it is not self-adjoint. The gap between the two domains is exactly two complex dimensions, because what an element of the larger domain may do, and an element of the smaller one may not, is take a non-zero value at each of the two ends. Chapter 4.2 §4's warning box was pointing at this operator.
So pinning both ends is too strong. The question is which domains are not, and (4.4.15) answers it in two lines. Set in it and take moduli:
Two cases follow, and there are no others. Either every in has , and then (4.4.16) forces as well, so is or a subspace of it. The first was dealt with in the paragraph above, and a subspace of it has an adjoint at least as large again, by (4.4.10), so it fails for the same reason and worse. Or some has , and we may scale it so that , whereupon (4.4.16) makes a number of modulus one, say . Put in (4.4.15) and every other in the domain is caught by it:
So a single phase, fixed once for the whole domain, is the only symmetric possibility beyond the pinned one. Define to be on .
Each of these is self-adjoint, and the check is the same equation read in the other direction. Let be in the adjoint's domain, which by §5.2 means with nothing at the ends, and impose (4.4.15) against every . Substituting turns it into , and there are functions in with , for instance . So the bracket vanishes, which conjugates to : the adjoint's domain is again. The two domains coincide and is self-adjoint.
One step of the case analysis is still open, and it is the step the first case took explicitly. Condition (4.4.17) puts inside without saying the two are equal, and a subspace of a symmetric domain is symmetric as well, so containment on its own settles nothing. But the operator on a proper subspace cannot be self-adjoint. If is strictly inside , then extends on , so (4.4.10) makes the adjoint's domain contain , which is strictly larger than . So exactly, and the phases are the whole list.
Since and give the same condition, the family is a circle. Momentum on a bounded interval is not one observable but a circle of them, and they are genuinely different operators, because their eigenvalues differ. Solving gives , and the condition in (4.4.17) reads , so
A whole ladder of allowed momenta, sliding rigidly as turns, and returning to itself after one full turn. Choosing is choosing what happens at the wall, and §7 will show the same phenomenon for energy with more room in it.
5.5 · The half-line: the term that cannot be killed
Take , which is the case Chapter 4.2 §4 named twice. The end at infinity is handled by §5.3's argument word for word, so the boundary term of (4.4.13) is one number:
Ask the same question as before: for which domains does this vanish on every pair? Set in (4.4.19) and the answer arrives at once, because has only one solution. Every in a symmetric domain must satisfy . There is no phase to choose, because there is only one endpoint and a phase relates two.
So the largest symmetric domain available is , and we can ask whether that one is self-adjoint. Section 5.2 gives its adjoint's domain as every with and no condition at zero, since (4.4.19) already vanishes once . That is strictly larger: the function , for any length , has and , so it sits in the adjoint's domain and not in . The operator is symmetric and not self-adjoint.
Now put the two paragraphs together, which is the step that makes this different from §5.4. In that section the repair was to enlarge the domain until the two met. Here enlarging is not available: any enlargement contains a function with , and the first paragraph shows that such a domain is not symmetric, so it is not a candidate at all. The domain cannot grow and the adjoint's domain cannot shrink. They never meet.
on is symmetric and has no self-adjoint extension whatever. By P2 as corrected in §4.2, the momentum of a particle confined to a half-line is not an observable. There is no operator to measure, so there is no spectrum of possible values, no probability distribution over them, and no expectation value beyond the formal integral.
Nothing was quoted to reach that. It is (4.4.13), which is integration by parts, together with the observation that forces .
Where it bites. A radial coordinate runs over , so this is the statement that there is no radial momentum observable, and it is the reason Chapter 4.13 has to argue for the behaviour of the hydrogen wavefunction at the origin rather than assume it. It is also why Chapter 4.7's infinite well cannot be declared to have at the walls without an argument: that is one choice among the family §7 counts, and the physics of the wall is what selects it.
And notice what did not go wrong. Momentum on a half-line has real expectation values, real would-be eigenvalues, and a perfectly sensible-looking formula. Every check §4.3 listed passes. The failure is invisible until you ask the one question this chapter taught you to ask.
The reason for the asymmetry deserves a sentence, because it is Chapter 0.5's fourth item returning under a different name. On a bounded interval, whatever leaves one end can be made to arrive at the other, and records how much phase it picks up on the way. On a half-line there is a place to go and nowhere to come from. That is exactly the shift of (4.4.1): norm-preserving, one-to-one, and not onto. Worked example 3 makes the identification precise by exponentiating and finding that the flow it generates runs in one direction only.
Everything in this section is one line of calculus done honestly. When you move a derivative from one function to the other inside an integral, a leftover appears at the two ends of the interval. In Part 0 it was always thrown away, with the note that throwing it away is a physical claim. Here it is the only thing being looked at, because it is precisely the obstruction to a momentum operator being a measurable quantity.
On an infinite line the leftover disappears by itself. Anything with finite total probability and a finite spread of momentum has to die away at both ends, so there is nothing to arrange and nothing to choose, and momentum is a perfectly ordinary observable. On a segment the leftover is the difference between the two ends, and it can be cancelled by insisting that whatever leaves one end arrives at the other with a fixed change of phase. The size of that change is a real number, free to be anything, and every value of it gives a different genuine momentum operator with a different ladder of allowed values. The condition at the wall is not bookkeeping. It is a physical input, and the measurable momenta depend on it.
On a half-line the arithmetic gives a different kind of answer. The leftover involves one end only, and the sole way to cancel it is to demand that every state vanish there, which is a demand so strong that the partner operator is left with a larger list of states than the operator itself and no amount of adjustment closes the gap. The conclusion is not a technicality: a particle held on one side of an impenetrable wall has no momentum observable at all. Not one that is hard to compute. One that does not exist.
a natural place to stop · the answers are in hand; what follows is the theorem that counts them
6 · Counting the extensions: deficiency indices
Section 5 produced three answers by hand: a unique operator, a circle of them, and none. This section quotes the theorem that predicts which of those three outcomes occurs, without solving a boundary-value problem, and reduces the prediction to counting the square-integrable solutions of one first-order equation. It has to reproduce §5 exactly three times, and it does.
6.1 · The question, in a form a theorem can answer
Section 3.4 said what the general situation looks like. A symmetric operator sits inside its adjoint, enlarging the operator shrinks the adjoint, and a self-adjoint extension is a place where the two meet. So the size of the gap between and is what decides everything, and the useful way to measure that gap is to look at what can do that cannot.
The measurement von Neumann found is this. A symmetric operator has real eigenvalues, by §4.3, so it can have no eigenvector with a strictly complex eigenvalue. Its adjoint has a larger domain and is under no such restriction, so it can. Counting the eigenvectors of at one complex eigenvalue above the axis and at its mirror image below measures exactly how much room there is between the operator and its adjoint, and that count is what the theorem is stated in terms of.
6.2 · The classification, quoted
We use, without proof, the following. Let be symmetric with a dense domain in a Hilbert space. Fix any positive number carrying the units of , and define the two deficiency spaces and their dimensions. They take an rather than the this chapter has been using for domains, because they are not domains of anything: they are the sets of solutions of an eigenvalue equation, sitting inside a domain already fixed.
The two numbers do not depend on which was chosen. They are the deficiency indices, written , and they settle the question completely.
- . Exactly one self-adjoint operator extends .
- with . The self-adjoint operators extending correspond one-to-one with the unitary maps of onto . That set of maps is , and it has real parameters.
- . There are none.
What the mark covers and what it does not. The proof runs through the Cayley transform, which turns a symmetric operator into an isometry and a self-adjoint one into a unitary, so that extending the operator becomes extending an isometry and the count becomes linear algebra on the two leftover spaces. It is about six pages and it is in Reed and Simon's Methods of Modern Mathematical Physics, volume II, and in Rudin. The mark also covers the technical point that the deficiency spaces are read off alone, so nothing about beyond symmetry and density enters the count. What the mark does not cover is any of §5, which used none of this.
Why it is computable. The deficiency spaces are read off , and for the operators of §5 that domain came out maximal, with no endpoint condition surviving. So finding the deficiency spaces means solving a differential equation with no boundary conditions at all and then asking which solutions are square integrable. For a first-order operator that is one line of calculus.
6.3 · The three cases, counted
Say first which operator is being extended in each case, since the theorem counts the extensions of a given one. On the line it is on (4.4.7). On the interval it is on , the pinned domain, which is the smallest sensible starting point and the one whose adjoint §5.4 computed. On the half-line it is on . In all three the adjoint acts by the same formula on the maximal domain, so the calculation is the same one three times over.
Take to be a fixed positive momentum , so that the units work. The equation is , so and
each unique up to a constant, since a first-order equation has a one-dimensional solution space. An exponential has its derivative proportional to itself, so being in the domain of at all reduces to the single question of whether the exponential is square-integrable. Hence is when the corresponding exponential lies in of the interval and when it does not, and that is the entire computation. Take the three intervals in the order §5 took them.
- The line. blows up as and blows up as , so neither is in and the indices are . Exactly one self-adjoint extension. Section 5.3 found exactly one.
- The interval. Both exponentials are continuous on a bounded interval, so both are in , and the indices are . The unitary maps between two one-dimensional spaces are multiplication by a phase, so the family is , a circle, with real parameter. Section 5.4 found a circle, parametrised by one phase.
- The half-line. decays and is in , while grows and is not. The indices are , they are unequal, and there are no self-adjoint extensions. Section 5.5 found none.
Three cases, three agreements, and the agreement is not a coincidence of arithmetic. Look at the totals. The sum comes out , and , which is the number of independent pieces of boundary data the interval carries, because each piece is one thing the boundary term of (4.4.13) has to cancel. For a first derivative that data is one number at each finite end, the value of the function there, which is why the totals here also read as a count of finite endpoints. Take the count of data as the rule and the count of endpoints as a coincidence of first order, because §7 has the same two finite ends and gets : a second derivative reads a value and a slope at each of them. Self-adjoint extensions exist when the data can be paired off against itself, which needs an even total split evenly, and that is . The half-line has one datum and no partner for it. What the theorem adds is that this survives into situations where the hand calculation would be much longer, and §7 is the first of those.
The previous section answered a question three times by direct calculation, and got three different kinds of answer: one operator, a circle of operators, and no operator. A theorem exists that predicts which of the three you are about to get, and it does so by a count that takes a single line.
The recipe is this. Take the partner operator, which has the longer list of allowed states, and look for states it sends to themselves multiplied by a purely imaginary number. The original operator can never do this, because it always returns real multiples, so anything the partner manages here is a direct measurement of how much larger the partner is. Count these states above the axis and below it separately, and the pair of counts decides which of the three answers you are going to get. It gets all three right here, which is why it can be trusted in the next section, where the hand calculation would be considerably longer.
a natural place to stop · the first derivative is finished; what follows is the second, with the machinery put to work
7 · The particle in a box has four parameters, not one
Everything so far has been about a first derivative. Energy is a second derivative, and this section runs the whole machine on it. The boundary term acquires a second pair of terms, the deficiency indices come out rather than , and the family of self-adjoint Hamiltonians for a particle in a box is therefore a : four real parameters, not one. We then take a one-parameter slice of that family, solve it, and find that different members have genuinely different energy levels. That is the point of the section. The boundary condition is not bookkeeping and it is not a choice of convenience. It is physics, and it is measurable.
7.1 · The boundary form for a second derivative
Take on , which is Chapter 4.2's free Hamiltonian confined to a box. Integrating by parts twice moves both derivatives across and leaves two boundary terms rather than one:
The structure is the same as (4.4.13) and the arithmetic is one step longer. Symmetry on a domain is again the demand that the right-hand side of (4.4.21) vanish for every pair drawn from it, but now the data at each end is a pair of numbers, the value and the slope, so there is more to arrange and more ways to arrange it.
7.2 · The count, and the number it produces
Rather than classify by hand, use §6.2, which is exactly the situation it was quoted for. As §6.3 insisted, the theorem counts the extensions of a given operator, so the first job is to say which operator is being extended. It is the one whose domain kills every boundary number there is:
with the formula on that domain. All four numbers are killed and not two, which is more than §5.4's asked: that domain pinned the two values and left the slopes free. The distinction decides what is being counted. Pinning the values alone, with in , is the Dirichlet domain, and the list below puts Dirichlet at a single point of the family we are about to count, whereas the operator a family is counted from has to be one that every member of the family extends.
Two hypotheses of §6.2 have to be in place before it can be used. is symmetric, because every term in the bracket of (4.4.21) carries a factor that vanishes at each end. Its domain is dense, by §3.2's chain with one adjustment: replace each straight ramp there by the ramp across the same width , which meets the flat parts on either side with zero slope and has a bounded second derivative, so the approximating functions lie in (4.4.22) and are no further away in than the straight ones were. The adjoint of is then §5.2's answer run twice: the formula on the maximal domain, with no endpoint condition surviving, because the test functions that argument uses vanish identically near both ends and so constrain nothing there. So we solve for a fixed positive energy , with no boundary conditions imposed:
What matters about (4.4.23) is not the value of but the order of the equation. A second-order linear equation with constant coefficients has a two-dimensional solution space, spanned by and , and every one of those solutions is continuous on the bounded interval and therefore square-integrable on it, whatever complex number happens to be. Both deficiency spaces are two-dimensional and the indices are .
That is where §6.3's rule comes good, and the two versions of it are worth reading against each other here. Two finite ends, two pieces of boundary data at each, four pieces in all, split evenly into . Counting data gives the right answer and counting endpoints gives two, which is why §6.3 asked you to keep the first version and treat the second as an accident of first order.
Feed that into §6.2 and the answer is a family of self-adjoint Hamiltonians in one-to-one correspondence with the unitary maps of one two-dimensional space onto another, which is :
The real dimension of is , and §6.2 quoted that along with the theorem. At you can check it on sight, a unitary being a phase and nothing else, but is not visible that way, so here is the count. A unitary matrix is one whose two columns are orthonormal. The first column is a unit vector in : four real numbers with one real equation on them, so three real parameters. The second column has to be a unit vector orthogonal to the first, and the vectors orthogonal to a given non-zero vector in form a line of one complex dimension, on which the unit vectors differ by a phase. One more real parameter, and three and one make four.
Four real parameters. That is worth stating loudly, because the number one is easy to arrive at and wrong. Momentum on an interval gave a single phase in §5.4, the ladder in (4.4.18) slid rigidly as that phase turned, and it is natural to carry the picture across. Energy is a second derivative, the boundary data at each end is a value and a slope, and the family is four-dimensional. Some familiar conditions and where they sit in it:
- Dirichlet, , which is the textbook infinite well. One point of .
- Neumann, . A different point.
- Periodic, and , which is a particle on a ring, and antiperiodic, the same with both signs reversed. Two more.
- The Robin family, and with real, a curve through the first two: is Neumann and the limit of large is Dirichlet. This is the slice §7.3 works.
Each entry satisfies (4.4.21) directly, and you can check any of them in a line by substituting the two conditions into the bracket. The Robin case is done in full below, including the step from symmetric to self-adjoint, because it is the one whose spectrum we want.
7.3 · The Robin family, worked
Fix a real number with the units of an inverse length and take
Check symmetry first. Substitute the two conditions into the bracket of (4.4.21) and look at each end separately. At the bracket is , which is because is real and comes out of the conjugate unchanged. At the same cancellation happens with the sign the other way. Each end contributes zero on its own, so is symmetric.
Self-adjointness is the same computation read backwards, and it starts from §7.2 rather than from nothing. A function with all four boundary numbers zero satisfies the two Robin conditions, so , and (4.4.10) then puts inside , which §7.2 identified as the maximal domain. So let have and no endpoint condition, and demand that the bracket vanish against every . Substituting the conditions on collects the bracket into
The two numbers and can be prescribed independently inside the domain, because for any pair of target values there is a cubic polynomial matching those values and the two slopes the conditions then demand, and a cubic has exactly four coefficients to do it with. So both brackets in (4.4.26) vanish separately, and conjugating them returns and . Those are the conditions defining (4.4.25), so the adjoint's domain is the domain, and is self-adjoint for every real .
7.4 · Different extensions, different spectra
Now solve it, because the whole claim of this section is that the choice is physical, and the way to demonstrate that is to show the choice moving the energy levels. Write for a positive energy, so that and . The condition at reads , which is solved by and up to scale. Imposing the condition at on that function and collecting terms gives one equation in one unknown:
That is a transcendental equation and it has no closed-form solution, which is itself informative: the clean spectrum of the textbook box is a property of one point of the family and not of the formula. Two limits are worth reading off before solving it numerically. At the equation collapses to , so and , which is the Neumann ladder. It runs from , and the member is the constant function, which the scaling above discards and which the function defined below puts back. As the condition forces with the slope finite, so the limit is the Dirichlet ladder, the same numbers with removed and the lowest level at .
Negative energies have to be looked for separately, and they are where the family does something the textbook box never does. Put with , repeat the two steps with hyperbolic functions in place of trigonometric ones, and the condition becomes , which needs to have any solution at all. Both cases are captured by one function of the energy, the characteristic function of this boundary-value problem, which shares its name with Chapter 0.9's characteristic function of a distribution and nothing else. It is obtained by dividing (4.4.27) by and letting be imaginary:
where and , both of which are power series in with infinite radius of convergence and are therefore perfectly well defined for negative, where they become and . The energies are the zeros of , and because depends smoothly on as well as on , the levels slide continuously as is turned.
One value of can be read without any computation and it predicts the whole qualitative story. At we have and , so
which vanishes at and at and nowhere else. Those are the only two settings at which a level sits exactly at zero energy, so they are the only two at which a level can pass from one side of zero to the other. Solving the equation confirms that nothing else happens: below zero there are no states for , one state for , and two for . A negative makes each wall attractive, and the two thresholds in (4.4.29) are where the first and then the second state bound to a wall drops below the bottom of the positive ladder.
Solving (4.4.28) by Newton's method gives the numbers. Energies are in units of , and the last column is the Dirichlet limit for comparison.
| level | ||||
|---|---|---|---|---|
Read the middle column first. At the roots of (4.4.28) come out at to every digit printed, which is the check that the numerical method is solving the equation we wrote down rather than a neighbour of it. Then read across. The three columns are three different sets of numbers. Same particle, same mass, same box, same formula for the energy, and three different spectra, because three different self-adjoint operators were selected out of the same .
The physical reading is the point of the whole chapter and it is worth stating without hedging. The formula does not determine a Hamiltonian for a particle in a box. It determines a four-parameter family of them, all self-adjoint, all legitimate observables, with different energy levels. Choosing among them is not a mathematical formality performed before the physics starts. It is the physics of what the wall is made of, and any experiment that measures the levels is measuring which member of the family the wall implements.
Energy involves a second derivative rather than a first, so at each wall there are two numbers to worry about instead of one: the value of the wavefunction and its slope. Running the counting recipe gives two independent directions of freedom at each end rather than one, and the family of legitimate energy operators for a particle in a box turns out to have four adjustable real numbers in it. Four, not one. The single knob belonged to momentum, where each wall carries one number, and it does not carry over.
Every boundary condition anyone writes down for a box is one point of that family. The textbook condition that the wavefunction vanishes at both walls is one point. Demanding instead that the slope vanish is another. Gluing the two ends together into a ring is a third. A wall that is neither perfectly hard nor perfectly soft sits somewhere in between, on a curve through the family that can be solved and plotted.
And the levels move. Turning the knob slides every energy continuously, and past a certain setting a level drops below zero altogether and becomes a state stuck to the wall rather than rattling around inside. Same particle, same box, same formula, different numbers on the spectrometer. That is the sense in which the condition at the boundary is not bookkeeping. It is a physical property of the wall, it is one of the four numbers, and it is measurable.
8 · Worked examples
On let be multiplication by . (a) Show that is unbounded. (b) Write down the largest domain on which the formula lands back in the space, and show it is dense. (c) Show that is symmetric on it. (d) Show that is the same set, so is self-adjoint. (e) Say why no analogue of §5.4's circle appears.
(a) Take to be , which has . Then is on and zero elsewhere, so and the ratio in (4.4.2) is at least . The supremum is infinite. As with momentum, the reason is physical: there is no largest position.
(b) The largest domain is . It is dense because the truncations all lie in it, since is bounded by there, and as by dominated convergence with as the dominating function, which is Chapter 4.3 §4.3.
(c) For in the domain, since is real, and that is . The integrals converge because and are both in . There is no integration by parts and therefore no boundary term.
(d) Let with witness , so for every in the domain. Taking supported in a bounded interval, this says almost everywhere on that interval, and letting the interval grow gives on the line. Since by definition of the adjoint's domain, and . So the two domains coincide.
(e) Because there is no boundary term to cancel. The whole family of §5.4 came from (4.4.13), which exists because integration by parts moves a derivative and leaves a remainder at the ends. Multiplication moves nothing. Running §6.2 confirms it: solving gives , so vanishes wherever , which is everywhere on the real line, and is the zero vector. The indices are and the extension is unique. Position and momentum on the line are alike in being forced and unbounded, and unlike in where the difficulty sits: for the domain is the only issue, and for the domain is the beginning of the issue.
On take from §5.4. (a) Write down the domain of . (b) Find its spectrum. (c) Show that the Dirichlet Hamiltonian is not for any . (d) Say what this means about the sentence "kinetic energy is ".
(a) Applying an operator twice requires the intermediate vector to be in the domain again, so . Since is up to a constant, this is
Two conditions, which is the right number for a point of the of §7.2, and the family is a one-parameter curve inside that four-parameter family.
(b) The eigenvalues are the squares of those in (4.4.18), divided by :
At this is the ring, with a ground state at zero energy and every level above it doubly degenerate, since and give the same energy.
(c) Suppose the Dirichlet operator, whose domain requires and nothing about the slopes, equalled for some . Two operators are equal only when their domains agree, so compare the domains. The function has , so it is in the Dirichlet domain. Its derivative is , giving and , so would force and hence . But then the first condition demands , and both are zero, which is consistent, so test a second function. Take , also Dirichlet, with , which needs . No single serves both, so the Dirichlet domain is not for any . The spectra confirm it independently. The lowest level of is with , which never exceeds , while Dirichlet's lowest is . The two agree only at , and there the level is doubly degenerate, from and , while every Dirichlet level is simple.
(d) The sentence survives as a statement about formulae and fails as a statement about operators. On the line it is exact, because §5.3 gives a unique and squaring it gives the unique free Hamiltonian. In a box it is not, and part (c) is the whole reason. Momentum observables on are not in short supply, but §5.4 showed that every one of them is some , and part (a) showed that squaring one drags the slope condition along with it. The Dirichlet domain imposes nothing on the slopes, and part (c) turned that difference into two of its own members, and , which demand and of the same . So the textbook infinite well has a Hamiltonian that is not the square of any momentum observable, and not for want of momentum observables to square. This is worth carrying into Chapter 4.7. The infinite well is not a free particle with a restriction bolted on. It is a different operator, chosen from a family, and the choice is what the walls are.
On , Chapter 4.2 §7 writes the flow generated by an observable as . (a) Show that the flow generated by on the line is translation. (b) Ask the same question on with the domain of §5.5, and say what happens for each sign of . (c) Identify the resulting family with (4.4.1). (d) Say how this shows up in the deficiency indices.
(a) Expand formally, using :
which is the Taylor series of about , convergent for analytic and extending to all of by continuity since each map preserves the norm. So generates translation by , in either direction, and each translation is unitary. That is the content of : a full group of motions, invertible, with a self-adjoint generator.
(b) On the map makes sense as a map into only for . It slides the function to the right, and the vacated interval is filled with zeros, which joins on continuously precisely because every member of vanishes at the origin. For the recipe would need values of at negative , and there are none. So there is a flow forwards and no flow backwards.
(c) The forward maps preserve the norm exactly, since sliding does not change , and they are one-to-one, since sliding back recovers the function wherever it was defined. They are not onto: nothing in the image is non-zero on . That is exactly (4.4.1) with a continuous parameter in place of a discrete one, and it is Chapter 0.5's "the claim that an injective map is surjective" failing in the flesh. Composition adds the parameters, so the family is a semigroup rather than a group: it has an associative law and no inverses.
(d) A self-adjoint generator would give a two-sided flow, by Chapter 4.5 §9's Stone's theorem, and §5.5 showed there is no self-adjoint generator here. The deficiency indices of §6.3 are the arithmetic of exactly that one-sidedness: one square-integrable solution at and none at , because the exponential that decays at has no partner decaying at a left end that does not exist. The three descriptions are one fact: a boundary term that cannot be cancelled, an index pair that cannot be balanced, and a motion that cannot be run backwards.
9 · Your turn
Problem 1 — the domain is not everything, exhibited
(a) Show that is in and not in . (b) Show that is in the domain, and compute . (c) Show that is not closed as a subset of , by exhibiting a sequence in it converging to the function of part (a). (d) Explain why (c) is required by §2.3 rather than merely permitted by it.
Solution
(a) , so it is in . For membership of the domain it would have to be the integral of some function, and no integral has a jump: if with locally integrable, then as , so is continuous. The indicator is not continuous at or at .
(b) On each half-line is smooth, and at the origin it is continuous with a corner. It is the integral of , which one checks by evaluating separately for and . That is in , with . So is in the domain and . A corner is allowed; a jump is not.
(c) Let be continuous, zero outside , equal to on , and linear on the two ramps of width . Each is in the domain, being piecewise linear and continuous with a step-function derivative in . The difference from is supported on the two ramps and bounded by there, so . The limit is not in the domain by (a), so the domain is not closed.
(d) Suppose the domain were closed. Being dense by §3.2, it would then be all of , and would be a symmetric operator defined on the whole space. Hellinger–Toeplitz would make it bounded, contradicting §2.2. So the domain of any unbounded symmetric operator is necessarily dense and not closed, which is an uncomfortable combination and is the price of the subject. Notice also the sizes in (c). The ramp derivatives have , so the inputs converge while the images run away, which is §2.1's discontinuity happening in front of you.
Problem 2 — the adjoint of an extension
(a) Prove (4.4.10) from (4.4.9) directly. (b) Deduce that if is symmetric and is a symmetric extension of it, then . (c) Conclude that every self-adjoint extension of a symmetric has a domain lying between and . (d) Use (c) to give a second proof that momentum on has no self-adjoint extension, using only the two domains computed in §5.5.
Solution
(a) Let , so there is with for every . Every is such a , and there, so the same witnesses . The containment is that one substitution.
(b) The first containment is the hypothesis. The second is symmetry of in the form (4.4.12). The third is part (a).
(c) A self-adjoint extension is in particular a symmetric extension, and it has , so the chain in (b) puts that common domain between the two ends. This is the precise form of §3.4's picture: the search for a self-adjoint extension is a search inside a fixed interval of subspaces, which is why the answer is a finite-dimensional family rather than an open-ended construction.
(d) Section 5.5 computed and no condition. Any self-adjoint extension has a domain with no condition and symmetric. If is strictly larger than it contains some with , and then (4.4.19) with is non-zero, so is not symmetric. So , which is not self-adjoint. No candidate survives.
Problem 3 — the circle of momenta, and what turning it does
(a) For on , verify that the eigenfunctions are normalised and orthogonal. (b) Show that is unitary on and maps the periodic domain of §5.4 onto . (c) Compute and read off the relation between the spectra of and . (d) Two operators related as in (c) are unitarily equivalent. Explain why they are nevertheless different observables, and give the experimental statement that distinguishes them.
Solution
(a) , so . For , , the integrand being a full number of periods. Each satisfies since .
(b) Multiplication by a function of modulus one preserves pointwise, hence the norm, and it is invertible by multiplication by the conjugate, so it is unitary. If , meaning , then and , so . The map is onto because reverses it.
(c) By the product rule, , so . The spectrum of is the spectrum of shifted by , which is (4.4.18) read again.
(d) Unitary equivalence says the two operators have the same abstract structure, not that they are the same observable of the same system. The measured quantity is momentum in both cases, the states are functions on the same interval, and the sets of possible readings differ by , which is a number an experiment reports. The equivalence relabels the states by multiplying them by a position-dependent phase, and the physical statement is that this is not a relabelling one is free to perform, because it moves the very observable at issue. Position it leaves untouched. Both and are multiplication by a function, so they commute exactly, and at every point, which means the relabelled state says precisely what the original said about where the particle is. Momentum is what moves, and part (c) is the computation: conjugating by turns into , so the relabelling shifts every momentum reading by . Two boxes with different are told apart by measuring momentum in each and comparing the ladders.
Problem 4 — the Robin box at the two limits, and one number in between
(a) From (4.4.27), recover the Neumann and Dirichlet ladders as the limits and , and say which levels are lost or gained. (b) Show that the negative-energy condition is and that it has no solution for . (c) Show that a level sits at exactly zero energy only for or , by solving subject to the Robin conditions. (d) For , show that there is exactly one negative level and bracket it between and in units of .
Solution
(a) At the equation is , so with , and is a genuine solution with the constant eigenfunction, giving a state at . Dividing (4.4.27) by and letting leaves again, but now fails, because the corresponding solution of the original problem is the constant, which cannot satisfy with infinite unless . So the Neumann ladder is and the Dirichlet ladder is the same numbers with removed. The zero-energy state is what the two limits differ by.
(b) Put , so and . The condition at gives , so take , . Imposing and collecting gives the stated identity. For every term is positive when , so the sum cannot vanish. Rearranged, the condition reads , whose left side is positive and whose right side needs .
(c) gives . Then and . Substituting the first into the second gives , that is . A non-zero solution needs , since forces , so or . This is (4.4.29) derived from the eigenfunction instead of from the characteristic function, and the two agree.
(d) With and the condition of (b) is . For small , expanding gives , so near zero. At , , so a root lies between. Tighten the bracket. At , , and at , . So lies between and and lies between and in these units. The figure's readout gives , so . Only one root exists, because is negative near zero, positive beyond, and increasing once past its single turning point.
Problem 5 — where the four parameters are, and what they are not
(a) Count the real parameters in the boundary data of a box, and explain why a self-adjoint Hamiltonian corresponds to two complex conditions rather than one or three. (b) Verify directly from (4.4.21) that the periodic and antiperiodic conditions are symmetric, and that a mixed pair, with , is not. (c) The Robin slice of §7.3 uses the same at both ends. Show that allowing different values and still gives a self-adjoint operator, and say how many of the four parameters that family covers. (d) A student says that since all four parameters give self-adjoint operators, and self-adjointness is the whole of P2, the choice among them cannot matter. Say precisely what is wrong.
Solution
(a) The boundary data is , four complex numbers, so eight real ones. A domain is cut out of that data by complex linear conditions, and how many of them there must be is forced from both sides. One condition would leave the boundary term of (4.4.21) non-zero on some pair, so the operator would not be symmetric; three would leave the domain too small, so the adjoint's domain would be strictly larger by (4.4.10) and the operator symmetric but not self-adjoint. Two is the only count that can balance.
(b) With and for both and , the bracket at equals the bracket at term by term, so the difference is zero. With both signs reversed, each factor in each product changes sign, so each product is unchanged and the difference is again zero. For the mixed pair, the bracket at is , which is , the negative of the bracket at . The difference across the interval is therefore minus twice the bracket at , which is not zero in general. Take with and , which a cubic can arrange along with the two conditions at , and the bracket at is .
(c) The verification in §7.3 treated the two ends separately and used only that the coefficient at each end is real, so it goes through unchanged with and for real . That is two of the four, once Dirichlet is counted as the limiting value at each end. The remaining two are the ones that couple the ends to each other, which the periodic and antiperiodic conditions use and the Robin family does not. A Robin condition never lets the wavefunction at one wall know anything about the other.
(d) Self-adjointness is a requirement on an operator, not a description of one, and P2 says every observable is self-adjoint rather than that every self-adjoint operator is the observable you want. The four parameters index four different physical situations, distinguished by their energy levels, and §7.4's table shows three of them disagreeing in the first digit. What decides which one describes a given box is the physics of its walls, and no amount of checking self-adjointness will produce that. The same confusion in Chapter 4.2's language would be reading P2 as a converse, which that chapter's §4.4 explicitly declined to assert.
The restriction is not a choice. Chapter 0.6 proved that every linear map on is bounded and named the derivative as the operator that breaks the proof in infinite dimensions. Section 2.2 measured the break: on the modes the norm stays fixed and the norm of the derivative grows without limit, so momentum is unbounded, which is what an unbounded physical quantity has to look like. Section 2.3 then closed the escape route. A symmetric operator defined on the whole of a Hilbert space is bounded, by the closed graph theorem in two lines, so an unbounded observable cannot be defined on the whole space. The domain of is compulsory.
And the adjoint gets a domain nobody chose. Chapter 0.5's defining relation for , read where acts on a subspace, becomes a test that some vectors pass and others fail, and the set of vectors that pass is . Density is what makes the answer unique, and it is doing that job and no other. The two domains move in opposite directions: enlarging the operator shrinks its adjoint, so a self-adjoint operator is a place where two moving subspaces meet. Symmetric is the containment. Self-adjoint is the equality. Chapter 4.2's P2 was stated with the finite-dimensional word and now reads self-adjoint, which is the correction that chapter's §4 named this section for.
Three intervals, one operator, three different answers, and all of it integration by parts. Keeping the boundary term of Chapter 0.2 §3.2 rather than dropping it gives as the exact obstruction to symmetry. On it vanishes by itself, both domains are the same set, and momentum is a unique observable. On it vanishes exactly when for one fixed phase, giving a circle of momentum operators whose ladders of allowed values slide as the phase turns, and showing that the natural choice of pinning both ends is symmetric and not self-adjoint. On it forces , that domain is not self-adjoint, and nothing larger is symmetric. The momentum of a particle confined to a half-line is not an observable at all, and the argument used nothing beyond .
The count, quoted after the answers were already in hand. Von Neumann's classification reduces the question to solving with no boundary conditions and counting square-integrable solutions. For momentum those solutions are , and the indices come out on the line, on the interval and on the half-line, matching §5 three times out of three. Placing the mark after the hand calculation rather than before it was deliberate: the theorem is checked against arithmetic you own.
The number this chapter owed, and a correction with it. The same count run on over gives , because a second-order equation has a two-dimensional solution space and every solution is square-integrable on a bounded interval. So the particle in a box has a of self-adjoint Hamiltonians: four real parameters, not one. The single parameter belongs to momentum, and carrying it across is the mistake §7 exists to prevent. Dirichlet, Neumann, periodic, antiperiodic and the Robin family are all points of that . Solving the Robin slice gives the transcendental condition , whose roots at reproduce to every digit printed and at other settings do not, which is the measurement the table and the figure report.
Two marks, and where completeness was spent. The closed graph theorem in §2.3, used once, for Hellinger–Toeplitz and nothing else. Von Neumann's classification in §6.2, quoted with its hypotheses and then checked against three cases already worked. Nothing else here is asserted without being derived, and in particular the whole of §5, which is the section two sentences of Chapter 4.2 point at, uses neither mark. One mark standing elsewhere is leaned on, Fubini's theorem from Chapter 0.2 §4.2, in the grind box of §5.2, cited there rather than re-raised, which is the same treatment Chapter 4.3 gave Heine–Cantor. Chapter 4.3's closing brick said this chapter would need completeness at every step, and the exact answer is that it was spent twice, both times inside one of those two quoted theorems. The closed graph theorem is false on an incomplete space, and von Neumann's carries a Hilbert space in its hypothesis for the same kind of reason. Everything in between runs on two other things Chapter 4.3 supplied, density and Cauchy–Schwarz, and saying which is which is better than crediting the whole chapter with everything.
Where this gets spent. Chapter 4.5 is the third instalment of the bill Chapter 0.5 named, and it needs this chapter's output as its hypothesis: its spectral theorem is stated for self-adjoint operators, which is a condition about domains and nothing else, and its Stone's theorem in §9 is what makes "time evolution is unitary" and "the Hamiltonian is self-adjoint" the same statement. The shape has now repeated twice and it is worth saying in the same words: Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it. Chapter 4.7 takes §7's and spends one paragraph, not one clause, on why the infinite well selects the Dirichlet point of it. Chapter 4.12 needs the domain language to state honestly that single-valuedness of a wavefunction in the azimuthal angle is an assumption about the domain of rather than a theorem. Chapter 4.13 needs §5.5, because the radial coordinate lives on a half-line and the behaviour of the hydrogen wavefunction at the origin has to be argued rather than assumed. What you do not yet have is the values: an operator with no eigenvectors in the space still has a set of possible readings, and naming that set is the whole of Chapter 4.5.