Part IV · Quantum Mechanics — Chapter 4.2

The Linear Algebra of Quantum States

A table of renamings, seven postulates, and nothing else.

Where we are

Chapter 4.1 ended with two statements. What has to change is what a state is, and the mathematics for that change was built nine chapters ago. This chapter makes good on both. It is the chapter Part 0 was a down-payment on. The claim it has to honour was written out in the opening paragraph of Chapter 0.5 in these words: "every postulate of quantum mechanics is a statement about Hermitian operators on an inner-product space. Not modelled by, not analogous to — is." That same chapter then told you what this one would look like: "Chapter 4.2 becomes a translation exercise: a table of renamings, not a new subject."

So the spine of what follows is a two-column table. On the left, a theorem of Chapter 0.5, with the number of the equation that states it. On the right, a sentence about measurement. The rows are not analogies, and the right-hand column is not an interpretation of the left. Each row is one statement written twice, in two vocabularies. Section 2 builds the table, and the rest of the chapter refers back to it.

What is genuinely new is short. That shortness is the reason this chapter is written differently from every chapter before it. Part I cornered the action principle out of Newton. Part II cornered Lorentz invariance out of two facts about light. Part III cornered the field equations out of free fall and one identity. Quantum mechanics cannot be cornered. The honest response is to say so out loud at each of the seven places where something is asserted rather than derived. Each of those seven gets its own box, before it is used, and each carries a ⚑ as well. The box says what is being claimed. The mark says it was not earned. This chapter therefore has the largest flag count in the book, and that is the point rather than a defect.

One of the seven matters more than the others. The Born rule, which is §5, is the first thing in twenty-nine chapters that is posited rather than cornered, and it is still open. It is not smuggled in as a definition. It is not derived from a plausible-sounding axiom, and it is not established by the theorem that comes closest. Section 5 says that plainly, at the moment it enters, and Chapter 4.20 returns to what it costs.

Conventions. SI throughout, with \hbar written explicitly everywhere. Part IV does not use natural units and does not adopt them until Chapter 5.1. Dirac notation is exactly as Chapter 0.5 §2.2 installed it: hats on operators, no hats on eigenvalues or classical quantities, I^\hat I for the identity. The inner product is linear in its second slot and conjugate-linear in its first, as Chapter 0.5 §1.1 chose and said it was choosing. Section 2.1 below restates why, because the Born rule is built on it. Everything proved in Chapter 0.5 is a finite-dimensional theorem. Where a statement here will need repair in infinite dimensions, the repair is named by chapter, and you are told which of the two mathematics chapters does it.

Tools you'll need  — Chapter 0.5, all of it, and this chapter spends nothing else of comparable weight. Section 1 has the inner-product axioms and the choice of linear slot. Section 1.4 has Cauchy–Schwarz, used in §8 here and cashed as the uncertainty principle in Chapter 4.9. Section 2 has orthonormal bases, Parseval and the resolution of the identity. Section 3 has orthogonal projection and P^2=P^\hat P^{2}=\hat P. Section 4 has the adjoint and the three species of operator. Section 6 has the spectral theorem in all three of its parts. Section 7 has functions of operators and eiA\ee^{\ii A}. Section 8 has commuting operators, simultaneous diagonalisation and its figure fig-joint. And its Worked example 1 has a 2×22\times2 matrix that §10 here uses three times. Chapter 0.4 §2 for basis and dimension, which §9 needs to build the tensor product. Its §6.1 for the cyclic property of the trace, which §8 turns into a proof that this subject cannot be finite-dimensional. Its §7 for the one theorem Part 0 imported rather than proved. Chapter 1.3 §6 for Poisson brackets and the fundamental brackets {qi,pj}=δji\{q^{i},p_{j}\}=\delta^{i}_{j}, and §7 for the statement that a conserved quantity is the generator of its symmetry. Sections 7 and 8 here repeat both with operators. Chapter 4.1 §6.5 for \hbar and pμ=kμp^{\mu}=\hbar k^{\mu}, and §6.6 for the tension this chapter is here to resolve. Chapter 2.5 §4.3 for E2=p2c2+m2c4E^{2}=p^{2}c^{2}+m^{2}c^{4}, which §10.3 expands to get the neutrino oscillation length. Chapter 0.3 for Euler's formula, used in §10 without comment.

1 · The claim, stated before it is earned

Announce the destination. This section does no mathematics. It states exactly what the chapter will assert and exactly what it will not. It gives the total number of assumptions quantum mechanics requires, and it sets out the two rules by which the rest of Part IV should be read. Doing this first costs two pages. What it buys is the ability to check, at every later line, whether something was derived or put in by hand.

1.1 · What Chapter 0.5 claimed

Chapter 0.5 proved a sequence of theorems about operators on a complex inner-product space, with no physics anywhere in the argument. Then, in an insight box at the end of its §6, it read three of them back with two words changed:

"(a) Eigenvalues are real becomes: measurement outcomes are real numbers. … (b) Eigenvectors are orthogonal becomes: distinct outcomes are perfectly distinguishable. … (c) The eigenvectors form a complete orthonormal basis becomes: any state whatsoever can be written as a superposition of outcomes… These are three of the postulates of quantum mechanics as they are usually taught, and they are not postulates. They are theorems about Hermitian matrices, proved above, in a chapter containing no physics. Chapter 4.2 will do exactly one thing that this chapter did not: assert that a physical state is a unit vector and a physical observable is a Hermitian operator. That is one postulate, not four. Everything else is renaming."

That is the claim. This chapter either honours it, or the design of the book fails retrospectively. So the chapter is organised to make the claim checkable. Section 2 lays out the renamings as a table, and every later section either points at a row of that table or announces an assumption in a box. If you find a place where something has been slipped in without one or the other, the claim has failed there.

1.2 · The complete bill, said out loud

Here is the number, because a subject that has to assume things should be willing to say how many. Quantum mechanics as this book develops it rests on eight postulates and one experimental measurement. Seven of the eight arrive in this chapter. The eighth is the symmetrisation of identical-particle states, which arrives in Chapter 4.18. The measurement is that the electron carries half-integer angular momentum, which is Stern–Gerlach's, quoted in Chapter 4.12. Nothing else in Part IV is asserted without being derived. Chapter 4.20 §9 puts the whole list back on one page, with an honest statement of what each item does and does not settle.

What is assertedWhere
P1A pure state is a unit vector, defined only up to an overall phase§3
P2An observable is a Hermitian operator§4
P3The Born rule. The probability of outcome λk\lambda_{k} is P^kψ2\norm{\hat P_{k}\ket\psi}^{2}§5
P4The state update. After outcome λk\lambda_{k} the state is P^kψ/P^kψ\hat P_{k}\ket\psi/\norm{\hat P_{k}\ket\psi}§6
P5The generator of time evolution is the energy§7
P6Canonical quantisation. [x^i,p^j]=iδij[\hat x_{i},\hat p_{j}]=\ii\hbar\delta_{ij}§8
P7A composite system's space is the tensor product of the parts'§9
P8Identical-particle states are totally symmetric or totally antisymmetricChapter 4.18
E1Measured, not postulated: the electron carries j=12j=\halfChapter 4.12

Now set that against the list of things which are not on it. Every one of them is a theorem of Chapter 0.5 that you have already proved:

  • that measured values are real numbers (0.5 §6.1).
  • that distinct measured values belong to exactly perpendicular states (0.5 §6.2).
  • that every state can be resolved into the possible outcomes of any observable (0.5 §6.3).
  • that the probabilities of the outcomes add to one (0.5 §2, Parseval).
  • that two quantities can be sharp together exactly when their operators commute (0.5 §8).
  • that a norm-preserving evolution preserves every amplitude, so that conserving total probability and being unitary are the same condition (0.5 §4.3).
  • that a Hermitian generator produces exactly such an evolution, and every such evolution has one (0.5 §7.1 and its grind box).

Seven theorems against seven assertions. The seven theorems are the ones usually presented as the mysterious part. That contrast is the payoff for nine chapters of mathematics, and it is the thing this chapter exists to make visible.

1.3 · Two rules for reading the rest of Part IV

Every postulate is announced in a box, before it is used, and carries a ⚑. Part III never had to do this, because general relativity was cornered rather than posited. There you had the equivalence principle plus one identity, and the field equations follow. Quantum mechanics has no such route. The response is not to hide that fact but to timestamp it. A box means this is being asserted. The ⚑ means the book uses this and does not derive it. Both, every time, in place.

Say each time whether a step is a derivation or an identification. These are different things, and Part IV moves between them from sentence to sentence. That eiH^t/\ee^{-\ii\hat Ht/\hbar} is unitary when H^\hat H is Hermitian is algebra, proved in Chapter 0.5 §7.1. That the operator in that exponent is the energy is a physical identification with experiments behind it. Chapter 4.1 already ran this distinction once. The form of the cavity spectrum was derived and its normalisation was matched, and the chapter said so rather than letting the two blur. You have been trained for twenty-nine chapters to ask which is which, and this is the part where the answer changes most often.

In plain terms 4.2.1

Something changes in this chapter, and it is worth naming before it happens rather than afterwards. Everything so far has been cornered. Geometry was not chosen; it was what survived once you insisted that no observer's description could be privileged. Even the field equations of gravity were forced, in the sense that the alternatives were eliminated one by one until a single form was left standing. What arrives now cannot be got that way, and no honest telling pretends otherwise.

So the method changes. Certain statements are going to be put down as assumptions, each in its own box, each marked, each stated before anything leans on it. That is not a weakness in the presentation. It is the difference between a subject that was derived and a subject that was discovered, and confusing the two is how people come away believing quantum mechanics was deduced from something more familiar. It was not. It was guessed, tested for a century, and never yet found wrong.

What makes the chapter worth reading anyway is how little has to be assumed. The mathematics is finished — it was finished nine chapters ago, in a chapter with no physics in it at all. What gets added here is short enough to list on one page, and the rest of the part is that short list being spent.

2 · The table of renamings

Announce the destination. We set out, in one place, every result of Chapter 0.5 that this part of the book uses, beside the physical statement it becomes. Nothing in this section is assumed and nothing is proved. The proofs are all in Chapter 0.5, and each one is cited by equation. The purpose is to make the size of what has to be added visible before any of it is added.

2.1 · One convention, restated because everything rests on it

Chapter 0.5 §1.1 defined the inner product to be linear in its second argument and conjugate-linear in its first, and gave the reason in advance: "in Chapter 4.2 the object ϕ,ψ\avg{\phi,\psi} becomes a probability amplitude, read right-to-left as 'amplitude to find ψ\psi in state ϕ\phi', and the linear slot had better be the one holding the state that is actually evolving." That reason can now be stated properly rather than promised.

The number ϕψ\avg{\phi|\psi} is going to be an amplitude, meaning a complex number whose squared modulus is a probability. The state that is evolving, being prepared, being superposed, is ψ\ket\psi, and superposition is a linear operation on it. Split a beam into two paths and recombine it, and what comes out is ψ1+ψ2\ket{\psi_{1}}+\ket{\psi_{2}}. The amplitude to find that in ϕ\ket\phi must be the sum of the two amplitudes, or interference could not be computed at all. So the slot holding ψ\ket\psi has to be the linear one.

The conjugation then lands on ϕ\ket\phi, and that is exactly what makes ϕψ=ψϕ\avg{\phi|\psi}=\overline{\avg{\psi|\phi}}, and hence ϕψ2=ψϕ2\abs{\avg{\phi|\psi}}^{2}=\abs{\avg{\psi|\phi}}^{2}. Read that last equality in words. The probability of finding ψ\psi in ϕ\phi equals the probability of finding ϕ\phi in ψ\psi, which is a symmetry a probability had better have.

Mathematics texts make the other choice, and then every formula below acquires its conjugates on the other side. Nothing physical depends on which convention you pick. What does depend on it is arithmetic. Mixing the two conventions inside one calculation is the single most common way to produce a wrong sign.

2.2 · The table

Read it once from left to right and once from right to left. Left to right it is a dictionary. Right to left it is a list of physical facts, with the place each one was proved.

Chapter 0.5 provedWhereWhich says, about a physical system
A complex inner-product space: vectors add and scale, and u,v\avg{u,v} is linear in vv §1 Superposition. If two states are possible, so is any combination of them, and ϕψ\avg{\phi|\psi} is the amplitude to find ψ\psi in ϕ\phi
v2=v,v>0\norm{v}^{2}=\avg{v,v}\gt0 for v0v\neq0 §1.1, axiom (iii) Two different states can always be told apart: no non-zero state has zero length
Cauchy–Schwarz, u,v2u,uv,v\abs{\avg{u,v}}^{2}\le\avg{u,u}\avg{v,v} §1.4 No probability exceeds one — and, in Chapter 4.9, the uncertainty relation, with nothing added but the meaning of the letters
Orthonormal basis {ei}\{e_{i}\}, ei,ej=δij\avg{e_{i},e_{j}}=\delta_{ij} §2 A complete set of mutually exclusive outcomes of one measurement
v=iei,veiv=\sum_{i}\avg{e_{i},v}\,e_{i} §2 Every state is a superposition of the outcomes, with the coefficients read off as single overlaps
Parseval, v2=iei,v2\norm{v}^{2}=\sum_{i}\abs{\avg{e_{i},v}}^{2} §2 The probabilities add to one, and they add to one automatically rather than by a normalisation imposed afterwards
Resolution of the identity, ieiei=I^\sum_{i}\ket{e_{i}}\bra{e_{i}}=\hat I §2.3 The set of outcomes has missed nothing: insert it anywhere and a hard amplitude becomes a sum of easy ones
Orthogonal projection P^\hat P, with P^2=P^\hat P^{2}=\hat P and P^=P^\hat P^{\dagger}=\hat P §3 An outcome. Idempotence becomes the statement that repeating a measurement immediately returns the same value
The adjoint, A^u,v=u,A^v\avg{\hat Au,v}=\avg{u,\hat A^{\dagger}v} §4 The rule for moving an operator across an amplitude — used on every page of Part IV
Hermitian, A^=A^\hat A^{\dagger}=\hat A §4.2 An observable
Unitary, U^U^=I^\hat U^{\dagger}\hat U=\hat I, preserves every inner product §4.3 Evolution, and every symmetry: exactly the maps under which total probability is conserved
Eigenvalue λ\lambda, eigenvector vv, A^v=λv\hat Av=\lambda v §5 A measured value, and the state that gives it with certainty
Over C\C every operator has an eigenvalue §5, from 0.4 §7 Every observable has possible measured values. Over R\R this is false, and that is why the space is complex
Hermitian \Rightarrow every eigenvalue is real §6.1 Measured values are real numbers. A dial cannot read 3+2i3+2\ii
Hermitian \Rightarrow eigenvectors with different eigenvalues are orthogonal §6.2 Distinct outcomes are perfectly distinguishable. The chance of confusing one for another is zero, not small
The spectral theorem: an orthonormal eigenbasis exists §6.3 Any state at all is a superposition of the outcomes of any observable
A^=kλkP^k\hat A=\sum_{k}\lambda_{k}\hat P_{k}, with P^kP^l=δklP^k\hat P_{k}\hat P_{l}=\delta_{kl}\hat P_{k} and kP^k=I^\sum_{k}\hat P_{k}=\hat I §6.4 An observable is its list of possible values together with the outcomes they belong to, and nothing else
Degeneracy: the eigenspace is unique, a basis inside it is not §6.4 Two states can share a measured value, and then that measurement does not determine the state
f(A^)=kf(λk)P^kf(\hat A)=\sum_{k}f(\lambda_{k})\hat P_{k} §7 Functions of observables, including the one that generates motion
A^\hat A Hermitian  eiA^\Rightarrow\ \ee^{\ii\hat A} unitary, and conversely §7.1 Observables generate symmetries and evolution, and every symmetry has an observable behind it
[A^,B^]=0    [\hat A,\hat B]=0\iff a common orthonormal eigenbasis exists §8 Compatible measurements, and the quantum numbers that label a state
[A^,B^]0[\hat A,\hat B]\neq0\Rightarrow no common eigenbasis §8 There is no state in which both quantities are sharp — the qualitative half of Chapter 4.9

2.3 · What has been added so far, exactly

Nothing. Not one line of the table is an assumption about the world. Every left-hand entry is a theorem with a proof you have read, and every right-hand entry is that theorem with two or three words changed. Chapter 0.5 §6.5 put the count in its own words: "That is one postulate, not four. Everything else is renaming". The table is that sentence, itemised.

The physics enters at exactly seven places, and the sections that follow take them one at a time.

  • Two of the seven, §3 and §4, are the pair Chapter 0.5 named: a state is a unit vector, an observable is a Hermitian operator.
  • Two more, §5 and §6, are the ones Chapter 0.5 never had to mention, because a chapter of linear algebra has no reason to talk about probability.
  • The last three, §§7–9, are dynamics, the relation between position and momentum, and how two systems combine.

That is the whole of it.

The inversion, and why it is worth noticing

If you have met a list of "the postulates of quantum mechanics" before, compare it with §1.2's. The usual list runs to five or six items. It includes, as postulates, the statements that measurement outcomes are real, that the eigenstates of an observable form a complete orthonormal set, and that probabilities add to one. None of those three is on §1.2's list, because none of them is an assumption. They are 0.5 §6.1, 0.5 §6.3 and Parseval, and they were proved four parts ago in a chapter with no physics in it.

Notice which way round that leaves things. The items a reader finds strange are the theorems: real outcomes, perfect distinguishability, superposition, probabilities that sum to one automatically, and the impossibility of knowing two incompatible quantities together. The items that are genuinely assumed are dull-looking by comparison: a state is a unit vector, an observable is a Hermitian operator, evolution is generated by the energy. So the strangeness and the assumption sit in different places. The whole reason Chapter 0.5 was written where it was, four parts before it was needed, was to make that visible rather than assertable.

One item on the usual list survives intact, and §5 is about it.

In plain terms 4.2.2

Here is the debt being settled. Nine chapters ago a chapter of pure linear algebra was written with the announcement that it was quantum mechanics with the physics taken out, and that when the physics came back the work would already be done. What that amounts to is a table with two columns. On the left sits a theorem about self-partnered maps on a space with an inner product, proved with no physics in the room. On the right sits a sentence about measurement. The rows are not analogies and the right-hand column is not an interpretation of the left; each row is one statement appearing twice, in two vocabularies.

Read the table and notice how little is added. Measured values are real because the multipliers of such a map are real. Different outcomes are never mistaken for one another because the special directions belonging to different multipliers are exactly perpendicular. Any state at all is a combination of outcomes because there are enough of those directions to span everything. Weights adding to one is the statement that the squared length of a vector is the sum of the squared sizes of its coordinates.

What has to be supplied on top is small enough to count: seven assertions in this chapter, one more in the last chapter of this part, and one experimental fact. Everything else in the table was earned long ago.

3 · States are unit rays

Announce the destination. The first assertion. A state is a unit vector, and two unit vectors differing by an overall complex factor of modulus one describe the same state. The first clause is bookkeeping, and we show why. The second clause is physics, and we show what it forbids and what it permits. The section ends with a parameter count that turns the abstraction into something you can hold: a two-state system has two real parameters, not four.

⚑ P1 — the state postulate

Asserted. The physical states of a system are in one-to-one correspondence with the unit vectors of a complex inner-product space. One understanding comes with that: ψ\ket\psi and eiαψ\ee^{\ii\alpha}\ket\psi for real α\alpha are the same state, not two states that agree. Such an equivalence class is called a ray.

Not derived, and nothing here derives it. The vector-space structure is the whole content. It says that if ψ1\ket{\psi_{1}} and ψ2\ket{\psi_{2}} are possible states, then so is c1ψ1+c2ψ2c_{1}\ket{\psi_{1}}+c_{2}\ket{\psi_{2}} for any complex c1,c2c_{1},c_{2} making the result a unit vector. That is the superposition principle, and it is an assertion about nature. Chapter 4.1 §6.6 described two descriptions of light, each supported by measurements the other cannot account for, and said the tension is resolved by changing what a state is. This is the change.

What is being deferred. Which inner-product space is not specified here, and cannot be. For a spin it is C2\C^{2}. For a particle on a line it is a space of functions that Chapter 4.3 builds and Chapters 4.4 and 4.5 supply operators for. Every theorem quoted below is Chapter 0.5's finite-dimensional one, and every use of it in infinite dimensions is a promissory note on those two chapters.

3.1 · Why the length is fixed at one, and why that is not a restriction

The reason to normalise is that the numbers we are about to call probabilities are squared overlaps, and Parseval says their sum is ψ2\norm\psi^{2}. So expand any state in an orthonormal basis {ei}\{\ket{e_{i}}\}, which Chapter 0.5 §2 permits:

ψ  =  iciei,ci=eiψ,ψ2  =  ici2. \ket\psi \;=\; \sum_{i}c_{i}\ket{e_{i}}, \qquad c_{i}=\avg{e_{i}|\psi}, \qquad \norm{\psi}^{2} \;=\; \sum_{i}\abs{c_{i}}^{2}. (4.2.1)

Setting ψ=1\norm\psi=1 makes the right-hand sum equal to one, which is what a complete list of probabilities has to do. Notice the direction of the logic. Normalisation is chosen so that Parseval's identity reads as a statement about probability. Probability is not being normalised afterwards, by hand.

Scaling ψ\ket\psi by any non-zero complex cc scales every cic_{i} by cc, and every ci2\abs{c_{i}}^{2} by c2\abs{c}^{2}. So it changes no ratio and describes nothing new. A state is therefore a direction, and the unit-length convention picks one representative from each direction.

3.2 · The overall phase, and why it is invisible

Fixing the length leaves one freedom untouched. If ψ=1\norm\psi=1 then eiαψ=1\norm{\ee^{\ii\alpha}\psi}=1 as well, since eiα=1\abs{\ee^{\ii\alpha}}=1. So unit length does not pick a unique vector. P1 asserts that this leftover freedom is physically empty. The reason can be given completely only once §5 has supplied the Born rule. But the computation itself is one line, and it does not depend on anything §5 adds beyond the form of the expression, so we can do it now. For any projection P^k\hat P_{k},

P^keiαψ2  =  eiαP^kψ, eiαP^kψ=  eiαeiαP^kψ,P^kψ  =  P^kψ2, \begin{aligned} \norm{\hat P_{k}\,\ee^{\ii\alpha}\ket\psi}^{2} \;&=\; \avg{\ee^{\ii\alpha}\hat P_{k}\psi,\ \ee^{\ii\alpha}\hat P_{k}\psi}\\[3pt] &=\; \overline{\ee^{\ii\alpha}}\,\ee^{\ii\alpha}\avg{\hat P_{k}\psi,\hat P_{k}\psi} \;=\; \norm{\hat P_{k}\ket\psi}^{2}, \end{aligned} (4.2.2)

because the first slot conjugates (Chapter 0.5 §1.1) and eiαeiα=1\ee^{-\ii\alpha}\ee^{\ii\alpha}=1. Every probability the theory produces is of this form, so every one of them is untouched. The same one-line argument kills the phase in every expectation value, since eiαψA^eiαψ=ψA^ψ\bra{\ee^{\ii\alpha}\psi}\hat A\ket{\ee^{\ii\alpha}\psi}=\bra\psi\hat A\ket\psi by the identical cancellation. There is no measurement in the theory that can see α\alpha.

Verified. Taking the Hermitian matrix A^=diag(2)(11/21/21)\hat A=\operatorname{diag}(2)\oplus\big(\begin{smallmatrix}1&1/2\\ 1/2&1\end{smallmatrix}\big) and the state ψ=13(1,2,2)\ket\psi=\tfrac13(1,2,2), symbolic algebra gives the three probabilities 00, 89\tfrac89, 19\tfrac19 both with and without a factor eiα\ee^{\ii\alpha}, identically in α\alpha, and they sum to 11.

3.3 · The relative phase, which is everything

Now the contrast that makes §3.2 content rather than triviality. Multiply one term of a superposition by a phase and nothing cancels. Take a two-state system, the two states being the two paths of an interferometer, and compare

ψ0=12(1+2)withψβ=12(1+eiβ2). \ket{\psi_{0}} = \tfrac{1}{\sqrt2}\big(\ket{1}+\ket{2}\big) \qquad\text{with}\qquad \ket{\psi_{\beta}} = \tfrac{1}{\sqrt2}\big(\ket{1}+\ee^{\ii\beta}\ket{2}\big). (4.2.3)

Both are unit vectors. Now we want the number a detector actually reports, so project onto the recombined state +=12(1+2)\ket{+}=\tfrac{1}{\sqrt2}(\ket1+\ket2), which is what a detector at one output port of an interferometer responds to. The amplitude is +ψβ=12(1+eiβ)\avg{+|\psi_{\beta}}=\tfrac12(1+\ee^{\ii\beta}), and squaring its modulus gives

+ψβ2  =  14(1+eiβ)(1+eiβ)  =  14(2+2cosβ)  =  cos2 ⁣β2, \abs{\avg{+|\psi_{\beta}}}^{2} \;=\; \tfrac14\big(1+\ee^{\ii\beta}\big)\overline{\big(1+\ee^{\ii\beta}\big)} \;=\; \tfrac14\big(2+2\cos\beta\big) \;=\; \cos^{2}\!\tfrac{\beta}{2}, (4.2.4)

which runs from 11 at β=0\beta=0 to 00 at β=π\beta=\pi. That is a fringe, and it is what an interferometer measures. So a phase applied to every component is unobservable, and a phase applied to one component is the difference between full transmission and complete extinction.

This is not a subtlety to be filed away. It is the whole reason the theory needs complex numbers, and the whole content of the two-slit experiment. Chapter 4.1 §6.6 insisted that light's wave behaviour was untouched by the evidence for parcels of energy, and (4.2.4) is where the wave behaviour now lives: in the relative phase between the terms of a superposition.

3.4 · How many parameters a state actually has

Count, because the count is the cleanest way to feel the difference between a vector and a ray. A unit vector in Cn\C^{n} has 2n2n real components, subject to one real constraint ψ=1\norm\psi=1, so the unit sphere has 2n12n-1 real dimensions. Identifying vectors that differ by a phase removes one more. So

#{states of an n-level system}  =  2n2  real parameters. \#\{\text{states of an $n$-level system}\} \;=\; 2n-2 \ \text{ real parameters}. (4.2.5)

For n=2n=2 that is two, not four. Two real parameters is a sphere, and we can see the sphere directly by writing the general unit vector as

ψ  =  cosθ21  +  eiφsinθ22,0θπ,0φ<2π, \ket\psi \;=\; \cos\tfrac\theta2\,\ket{1} \;+\; \ee^{\ii\varphi}\sin\tfrac\theta2\,\ket{2}, \qquad 0\le\theta\le\pi,\quad 0\le\varphi\lt2\pi, (4.2.6)

which uses the phase freedom to make the first coefficient real and non-negative, and that exhausts the freedom. Every state of a two-level system is (4.2.6) for exactly one (θ,φ)(\theta,\varphi), and (θ,φ)(\theta,\varphi) are polar coordinates on a sphere. That sphere is the Bloch sphere. It is not a picture drawn for intuition. It is the state space, exactly, and §12's Problem 1 shows that the three expectation values σx,σy,σz\avg{\sigma_{x}},\avg{\sigma_{y}},\avg{\sigma_{z}} are the Cartesian coordinates of the corresponding point. The sphere is therefore something you can measure.

For n=3n=3 the count gives four real parameters, and there is no comparably nice picture. The two-level case is special, and its picture does not generalise. Better to say that out loud than to let the sphere become a mental model of quantum mechanics in general.

In plain terms 4.2.3

The first assertion is about what a state is, and it has two clauses that do different work. A state is a direction in the space, and its length is fixed at one. Fixing the length is not a restriction on nature but a choice of bookkeeping, because the weights are going to be squared lengths and they have to add to one; doubling every component would double every weight and describe nothing new.

The second clause is stranger and more consequential. Multiplying the whole state by a phase — a complex number of size one, applied to every component at once — changes nothing that can ever be measured, so states differing only by such a factor are the same state rather than two states that happen to agree. What is emphatically not invisible is a phase applied to one component and not another. That relative phase is the entire difference between two beams that reinforce and two that cancel, and it is what a wave description was carrying all along.

So the correct object is not a vector but a direction with its overall phase forgotten, and the counting bears this out: a two-state system has two real parameters to its name rather than four, which is why every account of one draws a sphere.

a natural place to stop  ·  the dictionary is laid out; what follows is one postulate at a time

4 · Observables are Hermitian operators

Announce the destination. The second assertion, and the one Chapter 0.5 named in advance as "the one postulate that has to be made". We state it, then collect the three theorems it immediately buys without re-proving any of them. Then we read the spectral decomposition as a physical statement about what an observable is. Then we use Chapter 0.5 §8 to define a term the rest of Part IV lives on: a quantum number. One warning is attached to the postulate itself and cannot wait, because the word "Hermitian" is not strong enough for the spaces Chapter 4.3 builds.

⚑ P2 — the observable postulate, with a correction announced in advance

Asserted. Every measurable quantity of a physical system is represented by a Hermitian operator A^=A^\hat A=\hat A^{\dagger} on the state space, and the possible results of measuring it are the eigenvalues of A^\hat A and nothing else.

Not derived. This is Chapter 0.5's "an observable is a Hermitian operator is the one postulate that has to be made", collected. Notice what is being assumed and what is not. It is not being assumed that outcomes are real, or that different outcomes are distinguishable, or that the possible outcomes span the state space. Those are consequences, proved in Chapter 0.5 §6 with no physics in the argument, and §4.1 below reads them off.

⚠ The word will be sharpened, and the difference is physical. In a finite-dimensional space, Hermitian (A^=A^\hat A^{\dagger}=\hat A) is the right condition and nothing more need be said. In a space of functions it is not enough. An operator like iddx-\ii\hbar\,\dv{}{x} cannot act on every vector of the space, because differentiating a square-integrable function need not give a square-integrable function. So it carries a domain, and its adjoint carries a domain of its own, fixed by the definition and with no reason to be the same one. The condition A^u,v=u,A^v\avg{\hat Au,v}=\avg{u,\hat Av} holding on a domain is called symmetric, and that is strictly weaker than self-adjoint, which requires the domains to match as well. Chapter 4.4 §4 makes the correction, and P2 should be read as saying self-adjoint from that point on.

Here is why that is not pedantry. Chapter 4.4 §5 shows that iddx-\ii\hbar\,\dv{}{x} on the half-line [0,)[0,\infty) is symmetric and has no self-adjoint extension whatever. So "the momentum of a particle confined to a half-line" is not an observable at all. A physical conclusion, from a domain.

4.1 · Three theorems, collected rather than re-proved

With P2 in place, the following are facts about nature that were established in a chapter containing no physics. Each is cited at the section that states it, and none is proved again here. Proving them again would be a misreading of what this chapter is.

Measured values are real. Chapter 0.5 §6.1 showed that a Hermitian operator's eigenvalues satisfy λˉ=λ\bar\lambda=\lambda. A dial reads a real number, and now it has to.

Distinct measured values belong to orthogonal states. Chapter 0.5 §6.2 showed that eigenvectors with λμ\lambda\neq\mu satisfy v,w=0\avg{v,w}=0. Section 5 will make v,w2\abs{\avg{v,w}}^{2} the probability of finding one where the other is, so that probability is exactly zero. Not small, and not zero to within experimental error. Two states with different definite values of the same observable are perfectly distinguishable in a single measurement.

Every state is a superposition of outcomes. Chapter 0.5 §6.3 produced an orthonormal basis of eigenvectors. So for any state ψ\ket\psi and any observable A^\hat A whatever, ψ=iciei\ket\psi=\sum_{i}c_{i}\ket{e_{i}} with ici2=1\sum_{i}\abs{c_{i}}^{2}=1 by Parseval. The probabilities of the outcomes add to one automatically, and no separate assumption is needed to make them.

4.2 · What an observable is, read off the spectral decomposition

We want the form of the decomposition that will still be true when the space becomes infinite dimensional, so group Chapter 0.5 §6.4's decomposition by distinct eigenvalue rather than by basis vector. That is the form Part IV uses throughout:

A^  =  kλkP^k,P^kP^l=δklP^k,kP^k=I^, \hat A \;=\; \sum_{k}\lambda_{k}\hat P_{k}, \qquad \hat P_{k}\hat P_{l}=\delta_{kl}\hat P_{k}, \qquad \sum_{k}\hat P_{k}=\hat I, (4.2.7)

with the λk\lambda_{k} real and distinct. Let's read (4.2.7) as a definition of the object rather than as a factorisation of it. An observable is a list of real numbers, together with a set of mutually orthogonal outcomes that exhaust the space. It is a labelled partition of the state space into perpendicular pieces. The operator is not something extra that the list has. The list is the operator. Two observables are the same observable exactly when they have the same list, which is why in practice one specifies a measurement by saying what its possible answers are and which states give each of them with certainty.

The degeneracy clause carries physics. If P^k\hat P_{k} has rank greater than one, several independent states share the value λk\lambda_{k}, and Chapter 0.5 §6.4 was careful that the eigenspace is unique but a basis inside it is not, and no theorem prefers one. Physically: that measurement, on its own, does not determine the state. Something else has to be measured, and the next subsection says exactly what "something else" is allowed to be.

4.3 · Compatible observables, and the definition of a quantum number

Chapter 0.5 §8 proved that two Hermitian operators admit a common orthonormal eigenbasis if and only if they commute. Rename both halves.

If [A^,B^]=0[\hat A,\hat B]=0 there are states in which both quantities have definite values at once, and enough of them to span the space. Call two such observables compatible. Chapter 0.5 §8.2's Step 4 is where the content sits. Inside a degenerate eigenspace of A^\hat A, where A^\hat A has said all it can say, the restriction of B^\hat B is still Hermitian. So the spectral theorem applies inside that eigenspace, and B^\hat B chooses the basis A^\hat A could not.

If [A^,B^]0[\hat A,\hat B]\neq0 there is no common eigenbasis, so there is no state at all in which both quantities are sharp. That is the qualitative content of the uncertainty principle. Chapter 4.9 supplies the quantitative version by applying Cauchy–Schwarz to two particular vectors. That is a calculation Chapter 0.5 §1.4 already set up, and it adds nothing but the meaning of the symbols.

Now the definition this chapter owes the rest of the part. A set {A^1,,A^m}\{\hat A_{1},\dots,\hat A_{m}\} of mutually commuting observables is a complete set of commuting observables when their common eigenspaces are all one-dimensional. Said the other way round, the list of eigenvalues (a1,,am)(a_{1},\dots,a_{m}) determines the state uniquely up to phase. That list is called the quantum numbers of the state. Writing a hydrogen state as n,,m\ket{n,\ell,m_{\ell}} is naming the eigenvalues of three commuting operators and nothing more. Chapter 0.5's insight box promised the term to Chapter 4.11. The term is defined here instead, where the machinery arrives. Chapter 4.9 §3 asks how one knows a set is complete, and Chapters 4.11 and 4.13 spend it on angular momentum and on hydrogen.

One caution, and Chapter 0.5's figure fig-joint is where to see it rather than have it described again. The theorem promises that a common eigenbasis exists. It does not promise that the basis a calculation happens to produce is that one. Drive the commutator of two operators to exactly zero, and a standard eigensolver will still hand back states in which the second observable has a spread. Inside a degenerate eigenspace it returned an eigenbasis of the first, rather than the one the second prefers. Go and move the sliders. The point is worth the minute, and it is exactly what Chapter 4.13's degeneracies will need.

4.4 · Which operators are observables, and which are not

P2 says every observable is Hermitian. It does not say every Hermitian operator corresponds to something an apparatus can measure, and that converse is genuinely open. For a system with superselection rules, or for gauge-dependent quantities, some Hermitian operators correspond to nothing measurable. This book does not need the converse and does not assert it. What it does need is the direction P2 states, plus one immediate negative consequence worth recording now.

A unitary operator is not, in general, an observable. Chapter 0.5's Problem 2 showed that its eigenvalues have modulus one, so they are complex unless they are ±1\pm1, and P2 requires real outcomes. Unitary operators are how the state changes. Hermitian operators are what can be read. Chapter 0.5 §7.1's correspondence U^=eiA^\hat U=\ee^{\ii\hat A} says the two classes carry the same information, and §7 below is the section that spends it.

In plain terms 4.2.4

The second assertion is the one the whole design was aimed at. A measurable quantity is represented by a map that is its own partner, and that single sentence is the last thing needing to be assumed about measurement outcomes. Everything usually presented alongside it as a separate postulate is a theorem already proved: that the readings are real numbers, that distinct readings are perfectly distinguishable rather than merely different, that any state can be resolved into the possible readings, that the weights add to one.

One honest qualification belongs here rather than later. The word used is Hermitian, and in a space of functions that word is not quite strong enough — the stronger property involves specifying which functions the map is allowed to act on, and the difference is physical rather than pedantic. A particle confined to a half-line turns out to have no momentum observable at all. The repair has its own chapter.

The other dividend is labels. Two quantities can be sharp at once exactly when the order of the two maps does not matter, and when there are enough such quantities to leave no ambiguity, the list of their values names the state uniquely. Those numbers are what atoms are catalogued by.

5 · The Born rule

Announce the destination. This is the section the book's honesty rests on. We state the rule connecting a state to the probabilities of outcomes. We say without hedging that it is not derivable, and that nothing in this book or outside it has derived it. We deal properly with the theorem that comes closest. Then we derive from the rule everything that is derivable from it: that probabilities are non-negative and sum to one, that overall phase is invisible, and that the average and the spread of a measurement are the two short expressions everyone uses.

5.1 · Why nothing so far has produced a probability

Look at what P1 and P2 have supplied. There is a state, which is a unit ray. There is an observable, which by (4.2.7) is a set of real values with orthogonal outcomes summing to the identity. The state resolves into those outcomes, ψ=kP^kψ\ket\psi=\sum_{k}\hat P_{k}\ket\psi, since kP^k=I^\sum_{k}\hat P_{k}=\hat I. The pieces are mutually orthogonal, so Chapter 0.5 §1.3's Pythagoras applies with no cross terms, and their squared lengths add to one:

1  =  ψ2  =  kP^kψ2  =  kP^kψ2, 1 \;=\; \norm{\psi}^{2} \;=\; \Big\|\sum_{k}\hat P_{k}\ket\psi\Big\|^{2} \;=\; \sum_{k}\norm{\hat P_{k}\ket\psi}^{2}, (4.2.8)

the cross terms vanishing because P^kψ,P^lψ=ψ,P^kP^lψ=δklP^kψ2\avg{\hat P_{k}\psi,\hat P_{l}\psi}=\avg{\psi,\hat P_{k}\hat P_{l}\psi}=\delta_{kl}\norm{\hat P_{k}\psi}^{2} by Hermiticity and orthogonality of the projections. So there is a list of non-negative numbers, one per possible outcome, adding to one, sitting in the formalism already.

And that is as far as mathematics goes. A list of non-negative numbers summing to one is consistent with being a probability distribution. It is not a demonstration that it is one. Nothing in Chapter 0.5, and nothing in P1 or P2, connects any number in the theory to the relative frequency of an outcome in a long run of experiments. That connection has to be asserted, and here it is.

⚑ P3 — the Born rule

Asserted. A system in state ψ\ket\psi, measured for the observable A^=kλkP^k\hat A=\sum_{k}\lambda_{k}\hat P_{k}, yields the value λk\lambda_{k} with probability

Pr(λk)  =  P^kψ2  =  ψP^kψ. \Pr(\lambda_{k}) \;=\; \norm{\hat P_{k}\ket\psi}^{2} \;=\; \bra{\psi}\hat P_{k}\ket{\psi}.

When λk\lambda_{k} is non-degenerate this is ekψ2\abs{\avg{e_{k}|\psi}}^{2}, the squared modulus of a single amplitude. When it is degenerate it is the sum of the squared moduli over any orthonormal basis of the eigenspace, and the projection form is what makes that independent of which basis you pick.

This is the first thing in this book that is posited rather than cornered, and it is permanently open. Every earlier ⚑ in twenty-nine chapters marked something established elsewhere that this book chose not to establish, whether a theorem of mathematics, a measurement, or a result belonging to a subject not built here. P3 is none of those. It cannot be deduced from P1 and P2. It does not follow from the linearity of the theory. And no argument in physics has established it: the literature contains many attempted derivations, and each assumes something equivalent to what it concludes. The mark on this box is therefore of a different kind from most in this book. Elsewhere ⚑ means this book chose not to prove this. Here it means nobody has proved this. Chapter 4.20 §9 returns to what that costs, and states precisely which part of the problem decoherence addresses and which part it does not touch.

⚑ Gleason's theorem, with its hypotheses, because it is the thing that comes closest. Statement: work on a complex Hilbert space of dimension at least three. Take any function μ\mu assigning a number in [0,1][0,1] to every orthogonal projection, subject to μ(I^)=1\mu(\hat I)=1 and μ(kP^k)=kμ(P^k)\mu(\sum_{k}\hat P_{k})=\sum_{k}\mu(\hat P_{k}) for any countable family of mutually orthogonal projections. Then μ\mu has the form μ(P^)=tr(ρ^P^)\mu(\hat P)=\operatorname{tr}(\hat\rho\hat P) for a unique positive operator ρ^\hat\rho of unit trace. For a pure state that is exactly P3. It is a genuine theorem, and it is not proved here.

What it does and does not do: look at how much its hypotheses already assume. They assume that probabilities are assigned to projections. They assume those probabilities are additive over orthogonal projections. And they assume the dimension is at least three. That last condition is not a technicality. The two-dimensional case is a genuine counterexample, and spin-12\half is two-dimensional. Assume all of that and the squared length is forced. But assuming all of that is assuming most of what one wanted explained. It narrows the freedom inside a framework rather than deriving the framework. It is not a derivation of the Born rule, and this book does not present it as one.

5.2 · What follows immediately, and what does not

The probabilities are legitimate. Each is a squared norm, hence real and non-negative, and (4.2.8) says they add to one. Neither fact was assumed in P3. Both are Chapter 0.5's, and P3 was written in the form P^kψ2\norm{\hat P_{k}\psi}^{2} precisely so that they would be automatic.

The overall phase is invisible. Proved in (4.2.2), and it is worth noticing that §3.2's computation was written before P3 arrived and needed only the form P^kψ2\norm{\hat P_{k}\psi}^{2}. P1 and P3 fit together exactly. Had the rule been anything not quadratic in ψ\ket\psi, P1's identification of ψ\ket\psi with eiαψ\ee^{\ii\alpha}\ket\psi would have been inconsistent.

What does not follow: which outcome occurs. P3 gives the distribution and says nothing about the individual run. This is not a gap the chapter will close.

5.3 · The expectation value and the variance, derived

These are the two quantities actually reported by an experiment, and both come out of P3 in two lines each. We want the mean of the measured value over many repetitions. By P3 that mean is the weighted average kλkPr(λk)\sum_{k}\lambda_{k}\Pr(\lambda_{k}). So substitute P3 into it, and then run (4.2.7) backwards:

A^ψ    kλkP^kψ2  =  kλkψP^kψ=  ψ(kλkP^k)ψ  =    ψA^ψ   \begin{aligned} \avg{\hat A}_{\psi} \;&\equiv\; \sum_{k}\lambda_{k}\,\norm{\hat P_{k}\ket\psi}^{2} \;=\; \sum_{k}\lambda_{k}\bra\psi\hat P_{k}\ket\psi\\[4pt] &=\; \bra\psi\Big(\sum_{k}\lambda_{k}\hat P_{k}\Big)\ket\psi \;=\; \boxed{\;\bra{\psi}\hat A\ket{\psi}\;} \end{aligned} (4.2.9)

The middle step used P^kψ2=P^kψ,P^kψ=ψ,P^kP^kψ=ψ,P^kψ\norm{\hat P_{k}\psi}^{2}=\avg{\hat P_{k}\psi,\hat P_{k}\psi} =\avg{\psi,\hat P_{k}^{\dagger}\hat P_{k}\psi}=\avg{\psi,\hat P_{k}\psi}, which is P^k=P^k\hat P_{k}^{\dagger}=\hat P_{k} and P^k2=P^k\hat P_{k}^{2}=\hat P_{k} from Chapter 0.5 §3. The last step is the spectral decomposition read from right to left. So the ubiquitous formula A^=ψA^ψ\avg{\hat A}=\bra\psi\hat A\ket\psi is derived from P3. It is not a separate definition of what "average" means.

The spread comes the same way. The variance of the measured value is k(λkA^)2Pr(λk)\sum_{k}(\lambda_{k}-\avg{\hat A})^{2}\Pr(\lambda_{k}). We want that in operator form, so expand the square, then use kλk2P^k=A^2\sum_{k}\lambda_{k}^{2}\hat P_{k}=\hat A^{2} (which is Chapter 0.5 §7's function-of-an-operator with f(λ)=λ2f(\lambda)=\lambda^{2}) together with kP^k=I^\sum_{k}\hat P_{k}=\hat I:

(ΔA)2  =  A^2A^2  =  (A^A^I^)ψ2. (\Delta A)^{2} \;=\; \avg{\hat A^{2}} - \avg{\hat A}^{2} \;=\; \norm{\big(\hat A-\avg{\hat A}\hat I\big)\ket\psi}^{2}. (4.2.10)

The second form is the one Chapter 4.9 needs, because it exhibits ΔA\Delta A as the length of a vector, and Cauchy–Schwarz is a statement about lengths. That is the whole reason the uncertainty principle is one line there rather than a new idea.

Verified. With A^=(212121)\hat A=\big(\begin{smallmatrix}2&\half\\ \half&1\end{smallmatrix}\big) and ψ=(35,45)\ket\psi=(\tfrac35,\tfrac45), computing the two projections and summing gives kPr=1\sum_{k}\Pr=1, kλkPr(λk)=4625\sum_{k}\lambda_{k}\Pr(\lambda_{k})=\tfrac{46}{25} against ψA^ψ=4625\bra\psi\hat A\ket\psi=\tfrac{46}{25}, and k(λkA^)2Pr(λk)=9612500\sum_{k}(\lambda_{k}-\avg{\hat A})^{2}\Pr(\lambda_{k})=\tfrac{961}{2500} against A^2A^2=9612500\avg{\hat A^{2}}-\avg{\hat A}^{2}=\tfrac{961}{2500}. Exact rationals both ways.

5.4 · Amplitudes are not probabilities

⚠ The belief this section exists to prevent

A superposition is not a mixture. The state 12(1+2)\tfrac{1}{\sqrt2}\big(\ket1+\ket2\big) is not a description of a system that is really in 1\ket1 or really in 2\ket2 with equal chance and we do not know which. Those two situations give identical predictions for a measurement in the {1,2}\{\ket1,\ket2\} basis, where both give 12\half and 12\half. They give different predictions for every other measurement, and that is what makes the difference a fact rather than a preference.

Take the observable whose outcomes are ±=12(1±2)\ket{\pm}=\tfrac{1}{\sqrt2}(\ket1\pm\ket2). For the superposition, +ψ2=12(1+1)2=1\abs{\avg{+|\psi}}^{2}=\abs{\tfrac12(1+1)}^{2}=1: the outcome ++ occurs every time. For the mixture, the answer is 12+12+12+22=1212+1212=12\half\abs{\avg{+|1}}^{2} +\half\abs{\avg{+|2}}^{2}=\half\cdot\half+\half\cdot\half=\half. Certainty against a coin flip. The arithmetic of the two differs because in the first case the amplitudes are added and then squared, and in the second the squares are added. Amplitudes are complex numbers, so adding them first allows cancellation that adding squares can never produce.

This is why the theory needs a vector space over C\C and not a probability distribution over a list of possibilities, and it is the single most common place for a reader to install a picture that will have to be removed later. The formal object describing a genuine mixture exists, and it is the density operator. It is built in Chapter 4.19, and Chapter 4.20 spends it on decoherence, which is where this distinction does its work.

In plain terms 4.2.5

This is where the book stops cornering and starts assuming, and the moment deserves to be marked rather than slipped past. Everything so far has been forced: the action principle out of Newton, the geometry of spacetime out of two facts about light, the field equations out of free fall and one identity. The rule connecting a state to the odds of each outcome is not like that. It is put in by hand, because nothing available implies it and nothing since has managed to.

The rule itself is short. Project the state onto the directions belonging to a given reading, and the squared length of what survives is the probability of getting that reading. Non-negative because it is a squared length; adding to one because the projections reconstruct the whole vector with nothing left over; unaffected by an overall phase because that phase has size one and squaring removes it.

There is a celebrated theorem which comes close to deriving it, and its hypotheses are worth naming because they are where the content hides: it assumes that probabilities are already assigned to projections, additively, and it needs at least three dimensions. Assume that much and the squared length is forced. Which is to say the theorem explains the formula given the framework, and the framework is most of what one wanted explained.

6 · Measurement, and the state after it

Announce the destination. A fourth assertion is needed, and most books fold it into the third. We keep them apart, because they are logically independent, and because the separation is what lets Chapter 4.20 say precisely which of the two decoherence addresses. We then derive the one consequence that makes the postulate testable, which is that an immediately repeated measurement returns the same value with certainty. Next we set the algebra beside the conditioning you do every day, and say exactly where the two part company. We end by naming the one imported theorem that all four postulates rest on.

6.1 · Why P3 is not enough

P3 says what the odds are. It says nothing whatever about the system afterwards, and something has to be said, because measurements are performed in sequence. An apparatus is calibrated by measuring the same thing twice. A Stern–Gerlach beam is sent through a second magnet. A qubit is read out and then used. Without a rule for the post-measurement state, the theory cannot predict the second measurement at all.

⚑ P4 — the state update, a separate postulate

Asserted. If a measurement of A^=kλkP^k\hat A=\sum_{k}\lambda_{k}\hat P_{k} on the state ψ\ket\psi yields the value λk\lambda_{k}, then immediately afterwards the state of the system is

ψ  =  P^kψP^kψ. \ket{\psi'} \;=\; \frac{\hat P_{k}\ket\psi}{\norm{\hat P_{k}\ket\psi}}.

Not derived, and independent of P3. One could consistently write down a theory with P3 and a different update rule, or with P3 and no update rule at all for measurements never repeated. The division of labour matters: P3 is about frequencies in an ensemble. P4 is about an individual system after an individual outcome. Most treatments merge them into one "measurement postulate", and thereby lose the ability to say which half is in difficulty. Chapter 4.20 §9 needs the separation. Decoherence gives a good account of why the alternatives stop interfering, which is a statement about P3's probabilities becoming classical. It gives no account at all of why one of them happens, and that is P4's clause.

The denominator is not a second assumption. P^kψ\norm{\hat P_{k}\ket\psi} is exactly the square root of P3's probability, so the rescaling is forced by P1's requirement that a state be a unit vector. And the update is undefined precisely when P^kψ=0\norm{\hat P_{k}\ket\psi}=0, which is when the outcome has probability zero and therefore does not occur.

6.2 · Repeatability, derived

The postulate has one immediate consequence that is a genuine prediction, and it costs three lines. Measure A^\hat A again straight away, on the updated state ψ\ket{\psi'}. By P3 the probability of getting λl\lambda_{l} is P^lψ2\norm{\hat P_{l}\ket{\psi'}}^{2}, and

P^lψ  =  P^lP^kψP^kψ  =  δlkP^kψP^kψ, \hat P_{l}\ket{\psi'} \;=\; \frac{\hat P_{l}\hat P_{k}\ket\psi}{\norm{\hat P_{k}\ket\psi}} \;=\; \delta_{lk}\,\frac{\hat P_{k}\ket\psi}{\norm{\hat P_{k}\ket\psi}}, (4.2.11)

using P^lP^k=δlkP^k\hat P_{l}\hat P_{k}=\delta_{lk}\hat P_{k} from (4.2.7). So the probability is 11 for l=kl=k and 00 otherwise: an immediately repeated measurement returns the same value with certainty. That is a fact about laboratories, and it is what makes the word "measurement" mean anything. It comes from P^2=P^\hat P^{2}=\hat P, which is Chapter 0.5 §3's idempotence, proved there as the algebraic statement that projecting twice is projecting once.

Notice also which vector ψ\ket{\psi'} is when λk\lambda_{k} is non-degenerate. Then P^k\hat P_{k} has rank one, so P^kψ=ekekψ\hat P_{k}\ket\psi=\ket{e_{k}}\avg{e_{k}|\psi}, and the update gives ek\ket{e_{k}} up to a phase, which by P1 is ek\ket{e_{k}}. The state after the measurement is the eigenstate, and every memory of ψ\ket\psi except the fact of the outcome is gone. When λk\lambda_{k} is degenerate the update keeps more. It keeps the component of ψ\ket\psi inside the eigenspace, direction and all. That distinction is not decorative. It is what makes a measurement of a degenerate observable weaker than a measurement of a complete set, and it is the reason Chapter 4.13's hydrogen states need three labels.

Familiar ground — the update rule is conditioning, and where the identity ends

The algebra of P4 is one you perform daily. Given a joint distribution and an observation, you restrict to the cases consistent with what was seen and divide by the total weight of what survived, so that the remaining probabilities add to one again: p(xE)=p(x)1E(x)/p(E)p(x\mid E)=p(x)\mathbf{1}_{E}(x)/p(E). Restrict, then renormalise. P4 is the same two operations in the same order. Here P^k\hat P_{k} is the restriction, and P^kψ\norm{\hat P_{k}\ket\psi} is the weight of what survived. The correspondence is exact, symbol for symbol, including the fact that both are undefined when the observed event had probability zero. When you compute a positive predictive value from a prior and a test result, you are doing P4 with a diagonal projection.

Where it stops, and this is the whole of the difference. Conditioning restricts a probability. P4 restricts an amplitude. Probabilities are non-negative, so discarding branches can only remove weight, and a branch that survives conditioning contributes what it contributes. Amplitudes are complex, so two surviving branches can cancel. That is why the ordinary law of total probability fails here. For a classical variable p(B)=kp(BAk)p(Ak)p(B)=\sum_{k}p(B\mid A_{k})p(A_{k}) always holds. Worked example 1 exhibits a case where the quantum analogue gives 12\tfrac12 if the intermediate quantity is measured and exactly 00 if it is not. Nothing about the second run differs. The first apparatus was switched on, and that changed the answer from certainty to a coin flip. No amount of care with priors reproduces that, because the object being conditioned is not a probability.

The second difference is more subtle, and worth naming because it is the one people carry the wrong intuition about. Conditioning tells you something about a variable that already had a value. P4 makes no such claim, and §5.4 showed why: a superposition is not an unknown value. So the update is a change of state and not merely a change of information, and every attempt to read it as the latter runs into Chapter 4.20's inequalities.

6.3 · What is not settled, named

P3 and P4 together let you predict every laboratory result in this book. They do not explain anything about the process they describe. Three questions are left completely open, and it is better to list them than to let them accumulate as unease.

What counts as a measurement? P4 refers to "a measurement" as a primitive. The apparatus is itself made of atoms and should be describable by P1 and P5 like anything else, in which case nothing in the joint evolution ever produces a single outcome. This is the measurement problem and this book does not solve it.

Why one outcome rather than another? P3 supplies the distribution and P4 the conditional state. Neither says what selects the individual result.

When does the update happen? P4 says "immediately afterwards", which is not a time. Chapter 4.20 §9 states which of these decoherence answers, essentially the first and only in part, and which it leaves untouched.

6.4 · What all four postulates rest on

One structural remark, and it is owed. Every measurement postulate above is a statement about the projections P^k\hat P_{k} of an observable, and those exist because of Chapter 0.5's spectral theorem. Look at Step 1 of that theorem's proof: "Because VV is complex, [Chapter 0.4] supplies an eigenvalue." Chapter 0.4 §7 supplied it from the fundamental theorem of algebra, which says that every non-constant complex polynomial has a root. That chapter flagged the theorem as the one result Part 0 imported rather than proved ⚑.

So the whole apparatus of measurement in this book descends from one imported theorem. That is a small debt, and it is worth knowing exactly where it sits, because it is not a hidden one. Without it, an observable might have no eigenvalues at all, and there would be no possible measured values to assign probabilities to. Chapter 0.4 §7 already made the point in those words. The debt is paid in Chapter 5.4, where complex analysis is built and the fundamental theorem of algebra becomes a three-line corollary of Liouville's theorem. Until then it is quoted, and it is the only thing in this chapter that is quoted rather than either derived or postulated.

In plain terms 4.2.6

A second assertion is needed and most accounts fold it into the first, which costs them the ability to say later which of the two is in trouble. The first says what the odds are. The second says what is left afterwards: whatever part of the state pointed along the directions belonging to the reading obtained, rescaled back to length one. The rest is gone.

The mathematics of that is something done every day. Restrict attention to the cases consistent with what was observed, then renormalise so the surviving weights add to one again — this is conditioning, and the algebra is identical, line for line. The break comes from what is being restricted. Here it is amplitudes rather than probabilities, and amplitudes can cancel, so a component that survives conditioning may still be annulled by another that also survived. Send a beam through a filter, and whether a second filter passes anything depends on whether the first one was looked at.

What none of this settles is why anything is left at all — why one reading occurs rather than the others, and what physical process the rescaling describes. That question is not answered in this chapter or in this book, and the last chapter of the part says exactly which part of it decoherence addresses and which part it leaves alone.

a natural place to stop  ·  states, observables and measurement are in place; what follows is motion

7 · Evolution has a Hermitian generator

Announce the destination. Here almost everything is derived, and the assertion is a single clause. In four steps we show that time evolution must be unitary, that a unitary flow must be the exponential of a Hermitian operator, and that the operator therefore exists and has the dimensions of an energy. Then we state the one thing that does not follow, which is that it is the energy. Differentiating the result gives the Schrödinger equation, which Chapter 4.6 takes seriously. The section closes by collecting Chapter 1.3's promise that observables generate symmetries, with operators in place of functions and every word unchanged.

7.1 · Step one: evolution preserves the norm

Write U^(t)\hat U(t) for the map taking the state at time 00 to the state at time tt, and require two things of it.

The first is linearity. If a system can be prepared in ψ1\ket{\psi_{1}} or ψ2\ket{\psi_{2}}, P1 says it can be prepared in any superposition, and evolution must take that superposition to the corresponding superposition of the evolved states. Otherwise the two beams of an interferometer, prepared together and travelling together, could not be treated as they are.

The second is conservation of total probability. At every time, the outcomes of any measurement must have probabilities summing to one, which by §5.2 is the statement ψ(t)=1\norm{\psi(t)}=1. Put the two together, and U^(t)\hat U(t) is a linear map preserving the norm of every state.

7.2 · Step two: norm-preserving and linear means unitary

Preserving lengths sounds weaker than preserving inner products. It is not. The missing inner-product information can be extracted from lengths alone, and that extraction is worth doing on the page, because it is three lines and it is the hinge of the section. Apply Chapter 0.5 §1.3's expansion u+v2=u2+v2+2Reu,v\norm{u+v}^{2}=\norm{u}^{2}+\norm{v}^{2}+2\operatorname{Re}\avg{u,v} to U^u+U^v\hat Uu+\hat Uv, which is U^(u+v)\hat U(u+v) by linearity, and use U^w=w\norm{\hat Uw}=\norm{w} three times:

u2+v2+2ReU^u,U^v  =  U^(u+v)2  =  u+v2=  u2+v2+2Reu,v, \begin{aligned} \norm{u}^{2}+\norm{v}^{2}+2\operatorname{Re}\avg{\hat Uu,\hat Uv} \;&=\; \norm{\hat U(u+v)}^{2} \;=\; \norm{u+v}^{2}\\[3pt] &=\; \norm{u}^{2}+\norm{v}^{2}+2\operatorname{Re}\avg{u,v}, \end{aligned} (4.2.12)

so ReU^u,U^v=Reu,v\operatorname{Re}\avg{\hat Uu,\hat Uv}=\operatorname{Re}\avg{u,v}. That is the real part. Now we want the imaginary part, so run the same line with vv replaced by iv\ii v. The second slot is the linear one, so U^u,U^(iv)=iU^u,U^v\avg{\hat Uu,\hat U(\ii v)}=\ii\avg{\hat Uu,\hat Uv}, and Re(iz)=Imz\operatorname{Re}(\ii z)=-\operatorname{Im}z, which gives ImU^u,U^v=Imu,v\operatorname{Im}\avg{\hat Uu,\hat Uv}=\operatorname{Im}\avg{u,v}. Both parts agree, so

U^u,U^v  =  u,vfor all u,vU^U^=I^. \avg{\hat Uu,\hat Uv} \;=\; \avg{u,v}\quad\text{for all }u,v \qquad\Longleftrightarrow\qquad \hat U^{\dagger}\hat U=\hat I. (4.2.13)

The equivalence on the right is Chapter 0.5 §4.3 read backwards. We have U^u,U^v=u,U^U^v\avg{\hat Uu,\hat Uv} =\avg{u,\hat U^{\dagger}\hat Uv} by the definition of the adjoint, and an operator whose inner products with everything agree with the identity's is the identity. So time evolution is unitary, and this is a consequence rather than an assumption. Chapter 0.5 §4.3 said the same thing forward, in the words "U(t)U(t) must be unitary because ψ2=1\norm{\psi}^{2}=1 is a total probability". Here is the argument it was pointing at.

7.3 · Step three: a unitary flow has a Hermitian generator

Now bring in time. Take a system whose physical arrangement is not itself changing, with no field being switched on and no apparatus being moved. For such a system, evolving for ss and then for tt must be the same as evolving for s+ts+t, since there is nothing to distinguish the intermediate instant. So

U^(s+t)  =  U^(t)U^(s),U^(0)=I^. \hat U(s+t) \;=\; \hat U(t)\,\hat U(s), \qquad \hat U(0)=\hat I. (4.2.14)

We want a differential equation rather than a functional one, so differentiate (4.2.14) with respect to ss at s=0s=0. The left side gives U^(t)\hat U'(t) and the right gives U^(t)U^(0)\hat U(t)\hat U'(0), so writing K^U^(0)\hat K\equiv\hat U'(0),

dU^dt  =  U^(t)K^U^(t)  =  etK^, \dv{\hat U}{t} \;=\; \hat U(t)\,\hat K \qquad\Longrightarrow\qquad \hat U(t) \;=\; \ee^{t\hat K}, (4.2.15)

the solution being the matrix exponential mtmK^m/m!\sum_{m}t^{m}\hat K^{m}/m!, which converges for every operator and is the unique solution with U^(0)=I^\hat U(0)=\hat I. Once K^\hat K is identified below as iG^/-\ii\hat G/\hbar with G^\hat G Hermitian, that same exponential is Chapter 0.5 §7's function of an operator with f(λ)=etλf(\lambda)=\ee^{t\lambda}.

What kind of operator is K^\hat K? Unitarity is the condition we have not yet spent, so differentiate U^(t)U^(t)=I^\hat U^{\dagger}(t)\hat U(t)=\hat I at t=0t=0, using (U^)=(U^)(\hat U')^{\dagger}=(\hat U^{\dagger})':

K^+K^  =  0, \hat K^{\dagger}+\hat K \;=\; 0, (4.2.16)

so K^\hat K is anti-Hermitian. Then iK^\ii\hat K is Hermitian, because (iK^)=iK^=iK^(\ii\hat K)^{\dagger}=-\ii\hat K^{\dagger}=\ii\hat K using (αA^)=αˉA^(\alpha\hat A)^{\dagger}=\bar\alpha\hat A^{\dagger} from Chapter 0.5 §4.1. Every anti-Hermitian operator is i-\ii times a Hermitian one and conversely, which is the same correspondence as Chapter 0.5 §7.1's U^=eiA^\hat U=\ee^{\ii\hat A}, seen at the level of generators rather than group elements.

Now the dimensions. K^\hat K has the dimensions of one over time, since tK^t\hat K has to be dimensionless. Multiply and divide by the constant with the dimensions of action that Chapter 4.1 produced, writing K^=iG^/\hat K=-\ii\hat G/\hbar. Then G^\hat G is Hermitian, and it has the dimensions of /time\hbar/\text{time}, which is energy. Nothing has been assumed to get here. The existence of a Hermitian operator with the dimensions of energy generating the evolution is forced by unitarity, the group law and dimensional analysis.

U^(t)  =  eiG^t/,G^=G^,[G^]=energy. \hat U(t) \;=\; \ee^{-\ii\hat Gt/\hbar}, \qquad \hat G^{\dagger}=\hat G, \qquad [\hat G]=\text{energy}. (4.2.17)

7.4 · Step four: the one thing that does not follow

⚑ P5 — the generator of time evolution is the energy

Asserted. The Hermitian operator G^\hat G of (4.2.17) is the observable representing the system's energy. It is written H^\hat H and called the Hamiltonian, and time evolution is

ψ(t)  =  eiH^t/ψ(0). \ket{\psi(t)} \;=\; \ee^{-\ii\hat Ht/\hbar}\ket{\psi(0)}.

Exactly half of this was derived. That evolution has the form eiG^t/\ee^{-\ii\hat Gt/\hbar} for some Hermitian G^\hat G of energy dimensions is §7.3, with no physics beyond linearity and conservation of probability. That G^\hat G is the energy is the other half. The energy here means the same quantity a calorimeter measures, the same quantity Chapter 1.3 built as the Legendre transform of the Lagrangian, the same quantity Chapter 4.1's hνh\nu counted in parcels. Saying G^\hat G is that quantity is an identification with experiments behind it, and it is the postulate. Chapter 4.6 §2 states the sign convention loudly and shows what the identification buys. The rigorous infinite-dimensional version of §7.3, where "differentiate the group law" needs care, is Stone's theorem, quoted in Chapter 4.5 §9.

The hypothesis that was used, named. (4.2.14) assumed the physical arrangement does not change during the evolution. Think of a system being driven, such as an atom in a laser pulse or a spin in a swept magnetic field. There H^(t)\hat H(t) depends on time, the group law fails, and eiH^t/\ee^{-\ii\hat Ht/\hbar} is wrong. Chapter 4.17 §3 handles that case, where the exponential is replaced by a time-ordered series.

Differentiating (4.2.17) with G^=H^\hat G=\hat H and multiplying by i\ii\hbar gives the equation the next chapter is named after:

  iddtψ(t)  =  H^ψ(t).   \boxed{\;\ii\hbar\,\dv{}{t}\ket{\psi(t)} \;=\; \hat H\ket{\psi(t)}.\;} (4.2.18)

Chapter 0.5's closing insight box named this exact equation as where its eiA\ee^{\ii A} would be spent, and said the sentence "energy is observable, therefore probability is conserved" is not a physical argument but a line of algebra. That is now literally true. (4.2.13) is the algebra, and P5 is the only physics in the paragraph.

7.5 · Observables generate symmetries: Chapter 1.3, with operators

Nothing in §7.3 used the fact that the parameter was time. Any Hermitian G^\hat G generates a one-parameter family of unitaries U^ϵ=eiϵG^/\hat U_{\epsilon}=\ee^{-\ii\epsilon\hat G/\hbar}, and the statement that this family is a symmetry of the system is the statement that it leaves H^\hat H alone. So let's work out what such a family does to an arbitrary observable, to first order in ϵ\epsilon. Expand the exponential, keeping terms to first order:

U^ϵf^U^ϵ  =  (I^+iϵG^)f^(I^iϵG^)+O(ϵ2)  =  f^  +  iϵ[G^,f^]+O(ϵ2), \hat U_{\epsilon}^{\dagger}\,\hat f\,\hat U_{\epsilon} \;=\; \Big(\hat I+\tfrac{\ii\epsilon}{\hbar}\hat G\Big)\hat f\Big(\hat I-\tfrac{\ii\epsilon}{\hbar}\hat G\Big) + O(\epsilon^{2}) \;=\; \hat f \;+\; \frac{\ii\epsilon}{\hbar}\big[\hat G,\hat f\big] + O(\epsilon^{2}), (4.2.19)

the two cross terms combining into the commutator because G^f^f^G^=[G^,f^]\hat G\hat f-\hat f\hat G=[\hat G,\hat f] by definition. Subtracting f^\hat f leaves the change itself, which is the object we are after:

δf^  =  iϵ[G^,f^]  =  ϵ1i[f^,G^]. \delta\hat f \;=\; \frac{\ii\epsilon}{\hbar}\big[\hat G,\hat f\big] \;=\; \epsilon\cdot\frac{1}{\ii\hbar}\big[\hat f,\hat G\big]. (4.2.20)

Set that beside Chapter 1.3 §7, which proved for classical mechanics that the change in ff under the flow generated by GG is δf=ϵ{f,G}\delta f=\epsilon\{f,G\}. The two are the same statement under the single replacement

{f,g}    1i[f^,g^], \{f,g\} \;\longmapsto\; \frac{1}{\ii\hbar}\big[\hat f,\hat g\big], (4.2.21)

which is precisely the substitution Chapter 1.3 §6.4 announced in advance. Taking G^=H^\hat G=\hat H and ϵ=t\epsilon=t turns (4.2.20) into the equation of motion for an observable, df^dt=1i[f^,H^]\dv{\hat f}{t}=\frac{1}{\ii\hbar}[\hat f,\hat H], which is Chapter 1.3's dfdt={f,H}\dv{f}{t}=\{f,H\} with the same replacement.

Now read the same line the other way. The condition [f^,H^]=0[\hat f,\hat H]=0 says two things at once: that f^\hat f is conserved, and that H^\hat H is unchanged by the symmetry f^\hat f generates. A conserved quantity is the generator of its symmetry, in operator form, which is what Chapter 1.3 said this chapter would repeat word for word.

Two things are being collected there, and it is worth separating them. The algebraic statement (4.2.20) is derived, from §7.3 and nothing else. The claim that (4.2.21) is the right correspondence for every pair of classical observables is a different animal. It is not derived, it is not assumed here either, and it is false. Section 8 postulates it for the single pair it needs, and Chapter 4.10 §8 proves that it cannot be extended consistently to all of them.

Familiar ground — you already exponentiate a matrix to evolve a system, and one factor of i\ii changes everything

A two-compartment pharmacokinetic model, or a two-state Markov model of treatment response, is a vector of amounts p(t)\vv p(t) obeying p˙=Kp\dot{\vv p}=K\vv p, with KK a real matrix whose columns sum to zero. Its solution is p(t)=etKp(0)\vv p(t)=\ee^{tK}\vv p(0), a matrix exponential, evaluated exactly as Chapter 0.5 §7 evaluates one, by diagonalising KK and exponentiating the eigenvalues. Everything structural in §7.3 is already in that calculation: a one-parameter family, a group law e(s+t)K=esKetK\ee^{(s+t)K}=\ee^{sK}\ee^{tK}, a generator read off as the derivative at zero, and a conserved total because the columns sum to zero. (4.2.18) is that equation with KK replaced by iH^/-\ii\hat H/\hbar, and nothing else about the machinery changes.

Where the two part company is exactly the i\ii, and the consequence is total. A rate matrix has eigenvalues with non-positive real part, and real ones for the two-state case here. So etK\ee^{tK} has factors et/τ\ee^{-t/\tau}. The compartments relax, the transients die, and the system settles into a steady state that does not remember how it started. An anti-Hermitian generator has purely imaginary eigenvalues, so the factors are eiEkt/\ee^{-\ii E_{k}t/\hbar}, of modulus one at every time. Nothing decays, nothing relaxes, and there is no steady state at all. The components hold their sizes forever, and only their relative phases move.

That is why §10's two-state system oscillates undamped instead of equilibrating at fifty-fifty as the corresponding Markov chain does, and it is the entire content of requiring probability to be conserved rather than merely bounded. Suppose you want relaxation back, and real atoms do relax. It has to come from coupling to something else that has been left out of the state. That is Chapter 4.20's subject, and it is not available by adjusting H^\hat H.

In plain terms 4.2.7

Now motion, and here almost everything is derived rather than assumed. Total probability must remain one for as long as the system exists, so whatever moves a state forward in time cannot change its length, and a map that preserves all lengths preserves all overlaps and is a rotation of the space. That much is not a physical hypothesis; it is the bookkeeping of the state postulate followed to its conclusion.

Then the correspondence from the toolkit chapter takes over. Rotations of this kind are exactly the exponentials of self-partnered maps, so there is a generator, and it is an observable. The whole question is which one, and the answer is the only genuinely new thing here: the generator is the energy. Nothing forces that. It is an identification with an experiment behind it, in exactly the sense that the constant relating colour to energy was an identification in the last chapter, and it is worth marking as such because it is the single input.

Differentiate the rotation and an equation of motion appears, which the next chapter takes seriously. The pattern underneath is one already met in the classical setting: a conserved quantity and the motion it generates are one object, and putting operators in place of functions repeats every word.

8 · Canonical quantisation

Announce the destination. Position and momentum have to be related to each other somehow, and the classical theory already supplies the relation. Chapter 1.3 computed the Poisson bracket of a coordinate with its conjugate momentum and got one. We postulate the operator version, under §7.5's substitution, for that pair alone. Then we spend three lines proving something that decides the shape of the next two chapters. No finite-dimensional space can carry that relation. Quantum mechanics is infinite-dimensional before any physics is done, and you can prove it here with Chapter 0.4's trace.

8.1 · The classical relation, and the operator version

Chapter 1.3 §6 built the Poisson bracket and computed the fundamental brackets

{qi,pj}=δji,{qi,qj}=0,{pi,pj}=0, \{q^{i},p_{j}\}=\delta^{i}_{j}, \qquad \{q^{i},q^{j}\}=0, \qquad \{p_{i},p_{j}\}=0, (4.2.22)

which say that momentum generates translation of position, and that neither generates any motion of its own kind. Section 7.5 showed that the operator analogue of a Poisson bracket is 1i[f^,g^]\frac{1}{\ii\hbar}[\hat f,\hat g], derived for a generator acting on an observable. Applying it to (4.2.22) gives the relation this part runs on. It is an assumption, and here is its box.

⚑ P6 — canonical quantisation, for position and momentum

Asserted. The observables representing the Cartesian coordinates of a particle and their conjugate momenta satisfy

[x^i,p^j]=iδijI^,[x^i,x^j]=0,[p^i,p^j]=0. \big[\hat x_{i},\hat p_{j}\big]=\ii\hbar\,\delta_{ij}\,\hat I, \qquad \big[\hat x_{i},\hat x_{j}\big]=0, \qquad \big[\hat p_{i},\hat p_{j}\big]=0.

Not derived. Chapter 1.3 §6.4 announced the substitution {f,g}1i[f^,g^]\{f,g\}\mapsto\frac{1}{\ii\hbar}[\hat f,\hat g] and named Chapter 4.9 as where it would be taken seriously. The commutator itself is postulated here, because Chapters 4.6, 4.8 and 4.11 all need it before Chapter 4.9 arrives. What Chapter 4.10 §8 supplies is the sharper and more interesting statement, and that statement is negative. The substitution cannot be extended consistently to every classical observable at once. Assign operators to all polynomials in xx and pp obeying it, and you obtain a contradiction. So P6 is a postulate about this one pair and not a general dictionary, and it should not be read as one.

What it does not fix. P6 says nothing about which space, which functions, or what x^\hat x and p^\hat p look like. Chapter 4.6 supplies the standard realisation p^=i\hat p=-\ii\hbar\nabla acting on functions of position, and §8.4 below says why the realisation cannot be dodged. Notice also that the relation is dimensionally forced once it is assumed to be a multiple of the identity. The bracket [x^,p^][\hat x,\hat p] has the dimensions of position times momentum, which is action. Chapter 4.1 §5.7 already observed that the new constant has exactly those dimensions, and that it is exactly the phase-space unit classical statistical mechanics was missing.

8.2 · Three lines that decide the next two chapters

Suppose the state space were finite-dimensional, of dimension nn, so that x^\hat x and p^\hat p are n×nn\times n matrices. Take the trace of both sides of P6's first relation. Chapter 0.4 §6.1 proved that the trace is cyclic, tr(AB)=tr(BA)\operatorname{tr}(AB)=\operatorname{tr}(BA), for any two square matrices of matching size. Hence the trace of any commutator vanishes:

tr[x^,p^]  =  tr(x^p^)tr(p^x^)  =  0. \operatorname{tr}\big[\hat x,\hat p\big] \;=\; \operatorname{tr}(\hat x\hat p) - \operatorname{tr}(\hat p\hat x) \;=\; 0. (4.2.23)

Now take the trace of the right-hand side. The identity matrix in dimension nn has nn ones on its diagonal, so

tr(iI^n)  =  in    0for every n1. \operatorname{tr}\big(\ii\hbar\,\hat I_{n}\big) \;=\; \ii\hbar\,n \;\neq\; 0 \qquad\text{for every }n\ge1. (4.2.24)

Two numbers that must be equal are 00 and in\ii\hbar n. There are no n×nn\times n matrices satisfying P6, for any nn whatever.

Stop and look at what that argument is. It uses one fact about the trace, proved in Chapter 0.4 in a chapter about determinants and change of basis, with no analysis, no limits and no physics. It is complete. And it says that the state space of a particle with a position cannot be Cn\C^{n} for any nn. The reason is not that Cn\C^{n} is too coarse an approximation, and not that a continuum is more elegant. It is that the arithmetic is impossible. The infinite-dimensionality of quantum mechanics is forced by one postulate and one line of linear algebra, before any physics has been done at all.

8.3 · How badly finite matrices fail, quantified

"Impossible" invites the response that some large matrix must come close, and the question deserves a number rather than a reassurance. It does not come close, and the bound is exact.

Use the inner product on matrices from Chapter 0.5 §1.2's table, A,B=tr(AB)\avg{A,B}=\operatorname{tr}(A^{\dagger}B), whose norm is AF=(ijAij2)1/2\norm{A}_{F}=\big(\sum_{ij}\abs{A_{ij}}^{2}\big)^{1/2}. Apply Chapter 0.5 §1.4's Cauchy–Schwarz inequality in that space, to I^\hat I and any matrix MM:

trM2  =  I^,M2    I^F2MF2  =  nMF2. \abs{\operatorname{tr}M}^{2} \;=\; \abs{\avg{\hat I,M}}^{2} \;\le\; \norm{\hat I}_{F}^{2}\,\norm{M}_{F}^{2} \;=\; n\,\norm{M}_{F}^{2}. (4.2.25)

Now we want to apply that to the miss itself, so set M=[x^,p^]iI^M=[\hat x,\hat p]-\ii\hbar\hat I, the amount by which a candidate pair fails. By (4.2.23) its trace is 0in0-\ii\hbar n, so trM=n\abs{\operatorname{tr}M}=\hbar n, and (4.2.25) gives 2n2nMF2\hbar^{2}n^{2}\le n\norm{M}_{F}^{2}, that is

  [x^,p^]iI^F    n  =  iI^F   \boxed{\;\norm{\big[\hat x,\hat p\big]-\ii\hbar\hat I}_{F} \;\ge\; \hbar\sqrt{n} \;=\; \norm{\ii\hbar\hat I}_{F}\;} (4.2.26)

for every pair of n×nn\times n matrices and every nn. Read the right-hand equality. The error is at least as large as the thing being reproduced. In relative terms the miss is at least 100%100\%, at every size, so a finite matrix pair cannot get the canonical commutator even half right.

Verified numerically, and the bound is attained. Minimising [x^,p^]iI^F/iI^F\norm{[\hat x,\hat p]-\ii\hbar\hat I}_{F}/\norm{\ii\hbar\hat I}_{F} directly over all complex n×nn\times n pairs by Nelder–Mead from twelve random starts bottoms out at 1.0000001.000000 for n=2,3,4,5n=2,3,4,5, so the bound is right and it is attained. Problem 3 identifies the equality case, and it is worth having in advance. Equality needs MM parallel to I^\hat I, which forces [x^,p^]=0[\hat x,\hat p]=0. So the best a finite-dimensional model can do is to commute, that is to reproduce none of the relation at all, and every pair that makes the commutator non-zero does strictly worse. For the natural truncation x^=diag(0,1,,n1)\hat x=\operatorname{diag}(0,1,\dots,n-1) with a nearest-neighbour p^\hat p, the trace of the commutator is 00 to machine precision at n=4,16,64n=4,16,64 and 256256, exactly as (4.2.23) requires.

8.4 · Why spin is not a counterexample, and what has to happen next

Here is the objection to raise. Chapter 4.11 will describe electron spin with 2×22\times2 matrices, and that is finite-dimensional. It is, and no contradiction arises, because a spin has no position operator. The three spin components satisfy [S^i,S^j]=iϵijkS^k[\hat S_{i},\hat S_{j}]=\ii\hbar\epsilon_{ijk}\hat S_{k}, whose right-hand side is not a multiple of the identity. Its trace therefore can vanish, and does, since each S^k\hat S_{k} is traceless. Section 8.2 rules out finite dimensions only for a system carrying a canonically conjugate pair. That it rules them out so cleanly while leaving spin alone is a check on the argument rather than a limitation of it.

What §8.2 forces, then, is this. The space of states for a particle has to be infinite-dimensional, and every theorem quoted in this chapter was proved in Chapter 0.5 in finite dimensions. Chapter 0.5 was explicit about the four places its proofs used that hypothesis, naming them as "in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective", and it said the bill would come due. It comes due in two instalments, and the division is worth stating precisely rather than gesturing at:

  • Chapter 4.3 builds the space. The Riemann integral is discarded and rebuilt as the Lebesgue integral. That is what makes L2L^{2} complete, so that a limit of states is a state, and what makes the Fourier modes genuinely an orthonormal basis rather than an assumption. Everything in §§1–3 of this chapter that used "an orthonormal basis exists" is on credit until then.
  • Chapters 4.4 and 4.5 build the operators on it. Unbounded operators, domains, the difference between symmetric and self-adjoint that P2's box already warned about, spectra with no eigenvectors, the spectral theorem in the form that survives, and the meaning of x\ket x and p\ket p. Everything in §§4–7 that used "the spectral theorem" is on credit until then.

That shape is not new. Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it, in the same order and for the same reason.

In plain terms 4.2.8

Position and momentum need a relation to each other, and the classical theory supplies one already: the bracket that measured how a quantity changes under the flow another one generates gave a particular answer for those two. The assertion here is that the same relation holds with the bracket replaced by the failure of two maps to commute, times a constant. That is a real assumption, and whether it can be extended to every quantity at once is a question with a sharp and negative answer several chapters ahead.

What the assertion costs is immediate and enormous, and the argument is three lines long. Adding up the diagonal entries of a product does not care about the order of multiplication, so the diagonal sum of the failure-to-commute is zero for any two square arrays whatever. The relation demands that it equal a fixed non-zero number on every diagonal entry, whose sum is therefore not zero. No finite list of numbers can do this. Not approximately either: the best possible attempt at any size misses by as much as the target itself.

So the space cannot be finite-dimensional, and this is settled before any physics is done. Two chapters of mathematics follow, and they were not chosen for thoroughness. They were forced here.

9 · Two systems: the tensor product

Announce the destination. Everything so far has described one system. Two systems require a rule for combining their state spaces, and there are only two candidates: the dimensions add, or the dimensions multiply. We build the object in which they multiply out of Chapter 0.4's basis and dimension, and it turns out to be a basis of pairs and nothing more. We postulate that it is the right one, and then we count. The counting is what makes entanglement a fact about arithmetic rather than a mystery, and it is the whole of what Chapter 4.19 needs from this chapter.

9.1 · A basis of pairs

Let system AA have state space VV with orthonormal basis {a1,,am}\{\ket{a_{1}},\dots,\ket{a_{m}}\} and system BB have WW with {b1,,bn}\{\ket{b_{1}},\dots,\ket{b_{n}}\}. Whatever the joint space is, it must contain a state for each way of specifying both systems separately: AA in ai\ket{a_{i}} and BB in bj\ket{b_{j}}. Write that state aibj\ket{a_{i}}\otimes\ket{b_{j}}, abbreviated aibj\ket{a_{i}b_{j}}, and there are mnmn such pairs.

Now use P1. If those mnmn states are all possible states of the pair, then by the superposition principle so is every complex combination of them, and the joint space contains their span. Declare them to be an orthonormal basis of that span. That declaration is the definition of the tensor product VWV\otimes W, and by Chapter 0.4 §2.2's definition of dimension it gives

dim(VW)  =  (dimV)(dimW)  =  mn. \dim\big(V\otimes W\big) \;=\; \big(\dim V\big)\big(\dim W\big) \;=\; mn. (4.2.27)

The inner product comes with it, and nothing new has to be chosen. Declare aibjakbl=δikδjl\avg{a_{i}b_{j}\,|\,a_{k}b_{l}}=\delta_{ik}\delta_{jl} on the basis and extend by conjugate-linearity in the first slot and linearity in the second, which is Chapter 0.5 §1's construction unchanged. For a general pair of states ϕ=iϕiai\ket\phi=\sum_{i}\phi_{i}\ket{a_{i}} and χ=jχjbj\ket\chi=\sum_{j}\chi_{j}\ket{b_{j}} this gives

ϕχ  =  i,jϕiχjaibj,ϕχϕχ  =  ϕϕχχ, \ket{\phi}\otimes\ket{\chi} \;=\; \sum_{i,j}\phi_{i}\chi_{j}\,\ket{a_{i}b_{j}}, \qquad \avg{\phi\otimes\chi\,|\,\phi'\otimes\chi'} \;=\; \avg{\phi|\phi'}\avg{\chi|\chi'}, (4.2.28)

the second identity following by expanding both sides on the basis and collapsing with the deltas. Amplitudes for independent systems multiply, which is what they had better do.

⚑ P7 — the composition postulate

Asserted. The state space of a system composed of two parts is the tensor product of the parts' state spaces, so its dimension is the product of theirs. Observables belonging to part AA alone act as A^I^\hat A\otimes\hat I, and any two such observables commute with any two belonging to BB alone.

Not derived. There is a competing rule, which says that describing two systems means listing two descriptions, so that the dimensions add. That is what classical mechanics does with configuration spaces, and nothing in P1 to P6 rules it out. Which one nature uses is a physical question, and its answer is the source of every effect in Chapters 4.18 to 4.20.

What is not being claimed. Nothing here says a state of the whole is built from states of the parts, and §9.2 shows almost none of them are. Nothing here handles identical parts either, where a further restriction applies. That restriction is P8, the symmetrisation postulate, and it is stated in Chapter 4.18 §3.

9.2 · Entanglement, as a dimension count

Ask how many of the mnmn dimensions are reached by states of the form ϕχ\ket\phi\otimes\ket\chi. Count parameters. Choosing ϕ\ket\phi takes mm complex numbers and χ\ket\chi takes nn, but rescaling ϕ\ket\phi by cc and χ\ket\chi by 1/c1/c gives the same product, so one complex parameter is shared. The set of product vectors therefore has complex dimension m+n1m+n-1, inside a space of complex dimension mnmn. For two two-level systems that is 33 inside 44. For two ten-level systems it is 1919 inside 100100.

So almost every state of a composite system is not a pair of states of its parts. Such states are called entangled, and the word names a counting fact rather than an influence: the joint space has more directions in it than the product construction reaches.

One explicit case, checkable by hand in three lines. Take two two-level systems and the state 12(00+11)\tfrac{1}{\sqrt2}\big(\ket{00}+\ket{11}\big). If it were (a0+b1)(c0+d1)=ac00+ad01+bc10+bd11\big(a\ket0+b\ket1\big)\otimes\big(c\ket0+d\ket1\big)=ac\ket{00}+ad\ket{01}+bc\ket{10}+bd\ket{11}, then matching coefficients requires

ac=12,bd=12,ad=0,bc=0. ac=\tfrac{1}{\sqrt2}, \qquad bd=\tfrac{1}{\sqrt2}, \qquad ad=0, \qquad bc=0. (4.2.29)

From ad=0ad=0 either a=0a=0, which contradicts ac=1/2ac=1/\sqrt2, or d=0d=0, which contradicts bd=1/2bd=1/\sqrt2. No solution exists. The state assigns no state at all to either half on its own, and the question "what is the first system's state?" has no answer of the kind P1 provides. That is why Chapter 4.19 has to build a different object, the density operator, to answer it.

9.3 · The size of the space, with a number

Iterating (4.2.27) over NN two-level systems gives a state space of dimension 2N2^{N}, so specifying a general state takes 2N2^{N} complex amplitudes, while specifying a state of each system separately takes 2N2N real angles by (4.2.5). At N=300N=300:

2300  =  2.04×1090amplitudes,against600angles. 2^{300} \;=\; 2.04\times10^{90} \quad\text{amplitudes}, \qquad\text{against}\qquad 600 \quad\text{angles}. (4.2.30)

The observable universe contains of order 108010^{80} atoms. Three hundred two-level systems means three hundred atoms, which is a small molecule. They have more amplitudes in their joint description than there are atoms available to record them, and the excess is entirely the gap between mnmn and m+n1m+n-1, compounded three hundred times. That gap is what makes simulating quantum systems on ordinary computers hard, what makes building quantum ones interesting, and what Chapter 4.20 measures with Bell's inequality. It follows from P7 and nothing else.

In plain terms 4.2.9

Put two systems side by side and ask what describes the pair. The answer is the one place where quantum mechanics departs from ordinary intuition by an amount that can be written as a number. Classically, describing two things means describing each and listing both, so the descriptions add. Here a basis for the pair is a list of pairs of basis states, one drawn from each, so the dimensions multiply.

That difference is entirely responsible for the phenomenon everyone finds strange. States describing each system separately and then pairing them do exist, but they form a vanishingly thin subset — counted properly, of dimension roughly the sum where the whole space has dimension the product. Almost every state of the pair is therefore not of that form, which means it assigns no state at all to either half on its own. A concrete two-by-two example can be checked by hand in three lines: four conditions on four numbers, and no solution.

The practical scale of this is worth a number. Three hundred two-state systems require more amplitudes to specify than there are atoms in the observable universe, while a description of each one separately needs six hundred angles. That gap is where the interest in building such machines comes from, and the last chapter of this part is about what fills it.

a natural place to stop  ·  the framework is complete; what follows is three systems worked in full

10 · Three two-state systems, in full

Announce the destination. Chapter 0.5's Worked example 1 diagonalised one 2×22\times2 Hermitian matrix completely and then said what it was: "It is the ammonia molecule… it is neutrino oscillation… it is the qubit. Chapter 4.2 will do all three, and the linear algebra will already be finished." Here it is done. We solve the matrix once with §7's evolution operator, and obtain a single formula. Then we read that formula three times, in three sets of units: picoseconds, kilometres, nanoseconds. The linear algebra takes half a page because it was finished nine chapters ago. The rest is identification and arithmetic.

10.1 · One matrix, solved once

Two states, called 1\ket1 and 2\ket2, whatever they turn out to be physically. The most general Hermitian operator on their span has four real parameters. One of them, the phase of the off-diagonal entry, can be absorbed into the definition of 2\ket2, so take it real. Write the remaining three as a mean, a difference and a coupling:

H^  =  (Eˉ+δΔΔEˉδ),Eˉ, δ, ΔR,Δ>0. \hat H \;=\; \begin{pmatrix} \bar E+\delta & \Delta\\[2pt] \Delta & \bar E-\delta\end{pmatrix}, \qquad \bar E,\ \delta,\ \Delta\in\R,\quad \Delta\gt0. (4.2.31)

Chapter 0.5's Worked example 1 is the case δ=0\delta=0. The eigenvalues come from the characteristic polynomial, (Eˉ+δλ)(Eˉδλ)Δ2=0(\bar E+\delta-\lambda)(\bar E-\delta-\lambda)-\Delta^{2}=0, giving

E±  =  Eˉ±R,R    δ2+Δ2,ΔE    E+E  =  2R. E_{\pm} \;=\; \bar E \pm R, \qquad R \;\equiv\; \sqrt{\delta^{2}+\Delta^{2}}, \qquad \Delta E \;\equiv\; E_{+}-E_{-} \;=\; 2R. (4.2.32)

Real, as Chapter 0.5 §6.1 requires, and the level splitting is ΔE=2R\Delta E=2R. Notice the shape. The coupling Δ\Delta pushes the two levels apart whatever the diagonal difference is, and the splitting is never smaller than 2Δ2\Delta. Two levels that would cross as δ\delta passes through zero instead approach to 2Δ2\Delta and separate again. That is an avoided crossing, and the phenomenon is entirely (4.2.32).

Define the mixing angle θ\theta by

cos2θ  =  δR,sin2θ  =  ΔR,equivalentlytan2θ=Δδ, \cos2\theta \;=\; \frac{\delta}{R}, \qquad \sin2\theta \;=\; \frac{\Delta}{R}, \qquad\text{equivalently}\qquad \tan2\theta=\frac{\Delta}{\delta}, (4.2.33)

which is legitimate because cos22θ+sin22θ=(δ2+Δ2)/R2=1\cos^{2}2\theta+\sin^{2}2\theta=(\delta^{2}+\Delta^{2})/R^{2}=1. Then the normalised eigenvectors are

+  =  cosθ1+sinθ2,  =  sinθ1+cosθ2, \ket{+} \;=\; \cos\theta\,\ket1+\sin\theta\,\ket2, \qquad \ket{-} \;=\; -\sin\theta\,\ket1+\cos\theta\,\ket2, (4.2.34)

orthogonal by inspection, as Chapter 0.5 §6.2 guarantees without being consulted. The grind box verifies them and runs §7's evolution operator on the state that starts as 1\ket1. The result is one formula, and it is the whole of this section:

  P12(t)  =  sin22θ  sin2 ⁣(ΔEt2)  =  Δ2δ2+Δ2sin2 ⁣(δ2+Δ2  t)   \boxed{\;P_{1\to2}(t) \;=\; \sin^{2}2\theta\;\sin^{2}\!\left(\frac{\Delta E\,t}{2\hbar}\right) \;=\; \frac{\Delta^{2}}{\delta^{2}+\Delta^{2}}\,\sin^{2}\!\left(\frac{\sqrt{\delta^{2}+\Delta^{2}}\;t}{\hbar}\right)\;} (4.2.35)

with P11=1P12P_{1\to1}=1-P_{1\to2} exactly. Four features are worth naming before the physics starts, because each is one of this chapter's postulates doing visible work.

The mean energy Eˉ\bar E does not appear. It cancelled, because it contributes an overall factor eiEˉt/\ee^{-\ii\bar Et/\hbar} to the state. That is a global phase, which P1 says is not there. Only energy differences are observable, and that is P1 rather than a separate principle.

The oscillation never stops and never damps. By §7's familiar box, the eigenvalues of the generator are purely imaginary, so nothing decays. A classical two-state rate process would relax to fifty-fifty and stay there.

Depth and speed move in opposite directions. Increasing the detuning δ\delta at fixed coupling makes sin22θ\sin^{2}2\theta smaller, which is a shallower oscillation, and RR larger, which is a faster one.

And at short times the two effects cancel exactly. Expanding (4.2.35) for small tt, the sin2\sin^{2} contributes R2t2/2R^{2}t^{2}/\hbar^{2} and the prefactor contributes Δ2/R2\Delta^{2}/R^{2}, so

P12(t)  =  Δ2t22    Δ2(Δ2+δ2)t434  +   P_{1\to2}(t) \;=\; \frac{\Delta^{2}t^{2}}{\hbar^{2}} \;-\; \frac{\Delta^{2}\big(\Delta^{2}+\delta^{2}\big)t^{4}}{3\hbar^{4}} \;+\; \cdots (4.2.36)

and the leading term does not contain δ\delta at all. However far off resonance the system is driven, the transition probability leaves zero along the same parabola. Detuning changes when the curve turns over, not how fast it starts. That is the seed of the transition-rate formula Chapter 4.17 derives, and the figure below makes it visible.

Grind box — the eigenvectors, the evolution, and P12P_{1\to2}

Step 1 · check the eigenvectors. Subtract EˉI^\bar E\hat I, which shifts both eigenvalues and moves no eigenvector. Using (4.2.33),

(H^EˉI^)+=(δΔΔδ)(cosθsinθ)=(δcosθ+ΔsinθΔcosθδsinθ)=R(cosθsinθ), \big(\hat H-\bar E\hat I\big)\ket{+} = \begin{pmatrix}\delta&\Delta\\ \Delta&-\delta\end{pmatrix}\begin{pmatrix}\cos\theta\\ \sin\theta\end{pmatrix} = \begin{pmatrix}\delta\cos\theta+\Delta\sin\theta\\ \Delta\cos\theta-\delta\sin\theta\end{pmatrix} = R\begin{pmatrix}\cos\theta\\ \sin\theta\end{pmatrix},

the last equality being the two identities cos2θcosθ+sin2θsinθ=cosθ\cos2\theta\cos\theta+\sin2\theta\sin\theta =\cos\theta and sin2θcosθcos2θsinθ=sinθ\sin2\theta\cos\theta-\cos2\theta\sin\theta=\sin\theta, both instances of cos(2θθ)\cos(2\theta-\theta) and sin(2θθ)\sin(2\theta-\theta). The same computation with \ket- gives R-R.

Step 2 · invert. (4.2.34) is a rotation by θ\theta, so its inverse is a rotation by θ-\theta:

1=cosθ+    sinθ,2=sinθ+  +  cosθ. \ket1 = \cos\theta\,\ket+ \;-\; \sin\theta\,\ket-, \qquad \ket2 = \sin\theta\,\ket+ \;+\; \cos\theta\,\ket-.

Step 3 · evolve. By §7, each eigenstate picks up its own phase and nothing else:

ψ(t)=eiEˉt/(cosθeiRt/+    sinθe+iRt/). \ket{\psi(t)} = \ee^{-\ii\bar Et/\hbar}\Big(\cos\theta\,\ee^{-\ii Rt/\hbar}\ket+ \;-\; \sin\theta\,\ee^{+\ii Rt/\hbar}\ket-\Big).

Step 4 · project. By P3 the amplitude to find 2\ket2 is 2ψ(t)\avg{2|\psi(t)}, and Step 2 gives 2+=sinθ\avg{2|+}=\sin\theta, 2=cosθ\avg{2|-}=\cos\theta:

2ψ(t)=eiEˉt/sinθcosθ(eiRt/e+iRt/)=ieiEˉt/sin2θsin ⁣Rt, \avg{2|\psi(t)} = \ee^{-\ii\bar Et/\hbar}\sin\theta\cos\theta\Big(\ee^{-\ii Rt/\hbar}-\ee^{+\ii Rt/\hbar}\Big) = -\ii\,\ee^{-\ii\bar Et/\hbar}\,\sin2\theta\,\sin\!\frac{Rt}{\hbar},

using eiueiu=2isinu\ee^{-\ii u}-\ee^{\ii u}=-2\ii\sin u from Euler's formula and 2sinθcosθ=sin2θ2\sin\theta\cos\theta=\sin2\theta. Squaring the modulus kills both the global phase and the i-\ii, leaving (4.2.35). \blacksquare

Verified symbolically. Forming eiH^t/\ee^{-\ii\hat Ht/\hbar} directly from (4.2.31) and taking the modulus squared of its entries returns Δ2sin2(Rt/)/R2\Delta^{2}\sin^{2}(Rt/\hbar)/R^{2} for the off-diagonal transition, and gives P11+P121=0P_{1\to1}+P_{1\to2}-1=0 identically in δ,Δ,t\delta,\Delta,t. Not to some precision. Identically. The series expansion of the same expression returns (4.2.36) term by term.

⚑ The measured parameters used below

Everything from here to the end of §10 is (4.2.35) with numbers put in. The numbers are measurements and are quoted, not derived. The formula is derived, and it is the same one in all three cases.

Ammonia. The inversion line of NH3\mathrm{NH}_{3} in the (J,K)=(3,3)(J,K)=(3,3) rotational state is at ν0=23.8701 GHz\nu_{0}=23.8701\ \mathrm{GHz}, which is the line the first maser ran on in 1954. The corresponding splitting for the non-rotating molecule is 0.7934 cm1=23.7855 GHz0.7934\ \mathrm{cm^{-1}}=23.7855\ \mathrm{GHz}. The barrier to inversion is about 2020 cm12020\ \mathrm{cm^{-1}}.

Neutrinos. Δm322=2.45×103 eV2\Delta m^{2}_{32}=2.45\times10^{-3}\ \mathrm{eV^{2}} with sin22θ23\sin^{2}2\theta_{23} consistent with 11. Then sin22θ13=0.085\sin^{2}2\theta_{13}=0.085 with Δm312=2.5×103 eV2\Delta m^{2}_{31}=2.5\times10^{-3}\ \mathrm{eV^{2}}. Baselines and energies as quoted per experiment.

Qubit. A superconducting two-level circuit with level splitting ω0/2π=5 GHz\omega_{0}/2\pi=5\ \mathrm{GHz}, driven at a Rabi frequency ΩR/2π=25 MHz\Omega_{R}/2\pi=25\ \mathrm{MHz}. These are representative numbers for a transmon, chosen because they are round.

The two-state description itself is a model in every case.

  • Ammonia is a molecule with many vibrational and rotational levels, and the two-state treatment is the restriction to the lowest inversion doublet.
  • The neutrino has three flavours, and the two-flavour treatment is exact only when one mass splitting dominates.
  • The qubit is a weakly anharmonic ladder truncated to its bottom two rungs.

Each restriction is a good approximation for the reason its own field gives, and none of them is being derived here.

10.2 · The ammonia molecule

Ammonia is a nitrogen atom and three hydrogens arranged as a shallow pyramid. The nitrogen can sit above the plane of the hydrogens or below it, and these two arrangements are mirror images with identical energy. Call them L\ket L and R\ket R. Two facts fix (4.2.31) completely.

The detuning is zero, by symmetry. Reflecting the molecule exchanges L\ket L and R\ket R and changes no energy, so the two diagonal entries of H^\hat H are equal and δ=0\delta=0. This is not an approximation or a convenient choice. It is a symmetry of the Coulomb interaction between the same four nuclei and the same ten electrons. Hence by (4.2.33), θ=π/4\theta=\pi/4 and sin22θ=1\sin^{2}2\theta=1, and the oscillation goes all the way from one configuration to the other and back.

The coupling is not zero. Classically the nitrogen cannot get from one side to the other, because passing through the plane of the hydrogens costs about 2020 cm12020\ \mathrm{cm^{-1}}, which is 0.2504 eV0.2504\ \mathrm{eV}, far more than the molecule has. The off-diagonal entry Δ\Delta of a Hermitian operator connecting L\ket L and R\ket R is under no such prohibition. Nothing in P1 to P5 requires a state to get from one place to another through the intervening places, and Chapter 4.7 computes Δ\Delta for a barrier. Here it is read off a measurement.

So the energy eigenstates are the symmetric and antisymmetric combinations, ±=(L±R)/2\ket\pm=(\ket L\pm\ket R)/\sqrt2, split by ΔE=2Δ\Delta E=2\Delta, and

PLR(t)  =  sin2 ⁣(ΔEt2)  =  sin2(πν0t),ΔE  =  hν0. P_{L\to R}(t) \;=\; \sin^{2}\!\left(\frac{\Delta E\,t}{2\hbar}\right) \;=\; \sin^{2}\big(\pi\nu_{0}t\big), \qquad \Delta E \;=\; h\nu_{0}. (4.2.37)

We want a number out of that rather than a shape, so put the measured ν0=23.8701 GHz\nu_{0}=23.8701\ \mathrm{GHz} into it. The splitting is

ΔE  =  hν0  =  6.62607015×1034×2.38701×1010=  1.58165×1023 J  =  98.72 μeV, \begin{aligned} \Delta E \;=\; h\nu_{0} \;&=\; 6.62607015\times10^{-34}\times2.38701\times10^{10}\\[3pt] &=\; 1.58165\times10^{-23}\ \mathrm{J} \;=\; 98.72\ \mu\mathrm{eV}, \end{aligned} (4.2.38)

and the time for the molecule to turn itself inside out and back is 1/ν0=41.89 ps1/\nu_{0}=41.89\ \mathrm{ps}, with the halfway point, nitrogen fully on the other side, at 20.95 ps20.95\ \mathrm{ps}. Three consequences follow, and each of them is a postulate of this chapter made visible.

A molecule in its ground state has no shape. The stationary states are ±\ket\pm, not L\ket L or R\ket R. A state of definite shape is a superposition of the two energy levels, and therefore not stationary. This is §4.3 in the most concrete form available. Shape and energy are represented by operators that do not commute, so there is no state with both sharp, and the molecule in its lowest energy level is not in either configuration.

The splitting is tiny because the barrier is large. The ratio of barrier to splitting is 0.2504 eV/98.72 μeV=25360.2504\ \mathrm{eV}/98.72\ \mu\mathrm{eV}=2536. That the coupling is suppressed by three orders of magnitude rather than being zero is the quantitative content of tunnelling, and Chapter 4.10's WKB approximation computes the suppression as an exponential in the barrier's width and height.

The frequency lands in the microwave band, which is why the maser came first. The wavelength is c/ν0=1.2559 cmc/\nu_{0}=1.2559\ \mathrm{cm}. Chapter 4.1's Worked example 3 derived the ratio of stimulated to spontaneous emission for an atom in thermal radiation as 1/(ehν/kBT1)1/(\ee^{h\nu/k_{B}T}-1). At 23.8701 GHz23.8701\ \mathrm{GHz} and 300 K300\ \mathrm{K}, hν/kBT=3.819×103h\nu/k_{B}T=3.819\times10^{-3}, and the ratio is

RstimRspont  =  1e0.0038191  =  261.4, \frac{R^{\text{stim}}}{R^{\text{spont}}} \;=\; \frac{1}{\ee^{0.003819}-1} \;=\; 261.4, (4.2.39)

so stimulated emission beats spontaneous emission by a factor of 260260 at room temperature at this frequency, against 103310^{-33} for an optical transition. Now sort the two energy eigenstates with an inhomogeneous electric field, which works because ±\ket\pm have opposite parity and so respond oppositely. Feed the upper one into a cavity tuned to 1.26 cm1.26\ \mathrm{cm}, and the third process Einstein was forced to invent in Chapter 4.1 §4.5 amplifies. That is the ammonia maser, built in 1954, and every number in its design is in this paragraph.

10.3 · Neutrino oscillation

The same matrix, and the identification is where all the work is. A neutrino is produced by a weak interaction in a definite flavour, with a muon or with an electron, and it is detected the same way. But flavour is not what propagation cares about. Propagation is generated by H^\hat H, whose eigenstates are the states of definite mass.

Call the flavour states νμ\ket{\nu_{\mu}} and ντ\ket{\nu_{\tau}}, and the mass states ν1\ket{\nu_{1}} and ν2\ket{\nu_{2}}. The assertion tested by every oscillation experiment is that these are two different orthonormal bases of the same two-dimensional space, related by (4.2.34) for some angle θ\theta. Nothing about that is unusual. It is Chapter 0.4's change of basis, and §4.3 of this chapter says the two observables do not commute.

Now the energies, and this is the one step needing Part II. A neutrino of definite mass mim_{i} and momentum pp has, by Chapter 2.5 §4.3, energy Ei=p2c2+mi2c4E_{i}=\sqrt{p^{2}c^{2}+m_{i}^{2}c^{4}}. Neutrinos in these experiments carry energies in the MeV to GeV range, while their masses are below an electronvolt, so mic2pcm_{i}c^{2}\ll pc and the square root can be expanded. We want the difference of two nearly equal energies, and the leading terms will cancel, which is exactly why expanding is the right move:

Ei  =  pc1+mi2c2p2  =  pc  +  mi2c32p  +  O ⁣(m4c5p3), E_{i} \;=\; pc\sqrt{1+\frac{m_{i}^{2}c^{2}}{p^{2}}} \;=\; pc \;+\; \frac{m_{i}^{2}c^{3}}{2p} \;+\; O\!\left(\frac{m^{4}c^{5}}{p^{3}}\right), (4.2.40)

using 1+u=1+u/2+O(u2)\sqrt{1+u}=1+u/2+O(u^{2}) from Chapter 0.3. The common pcpc cancels in the difference, and with EpcE\approx pc for the beam energy,

ΔE  =  E1E2  =  (m12m22)c42E    Δm2c42E. \Delta E \;=\; E_{1}-E_{2} \;=\; \frac{\big(m_{1}^{2}-m_{2}^{2}\big)c^{4}}{2E} \;\equiv\; \frac{\Delta m^{2}c^{4}}{2E}. (4.2.41)

That is the level splitting in (4.2.35). One more identification is needed. A neutrino travels at essentially cc, so the proper substitution for the elapsed time is t=L/ct=L/c, where LL is the distance from source to detector. That distance is the quantity an experiment actually controls. Substituting both into (4.2.35):

  Pνμντ(L)  =  sin22θ sin2 ⁣(Δm2c4L4cE)   \boxed{\;P_{\nu_{\mu}\to\nu_{\tau}}(L) \;=\; \sin^{2}2\theta\ \sin^{2}\!\left(\frac{\Delta m^{2}c^{4}L}{4\hbar c E}\right)\;} (4.2.42)

which is the formula every oscillation experiment is analysed with, obtained from (4.2.35) by two substitutions and no new physics. Now convert the phase to the units experiments use, with Δm2\Delta m^{2} in eV2\mathrm{eV^{2}}, LL in kilometres and EE in GeV. Using c=1.973269804×107 eVm\hbar c=1.973269804\times10^{-7}\ \mathrm{eV\,m},

Δm2c4L4cE  =  Δm2×L×1034×1.973269804×107×E×109  =  1.26693Δm2LE, \frac{\Delta m^{2}c^{4}L}{4\hbar cE} \;=\; \frac{\Delta m^{2}\times L\times10^{3}}{4\times1.973269804\times10^{-7}\times E\times10^{9}} \;=\; 1.26693\,\frac{\Delta m^{2}\,L}{E}, (4.2.43)

the familiar coefficient, derived rather than quoted. Next we want the distance over which the pattern repeats. That is the oscillation length, the analogue of ammonia's 41.89 ps41.89\ \mathrm{ps}, and it is the LL making the phase advance by π\pi:

Losc  =  4πcEΔm2c4  =  2.4797  E[GeV]Δm2[eV2] km. L_{\text{osc}} \;=\; \frac{4\pi\hbar cE}{\Delta m^{2}c^{4}} \;=\; 2.4797\;\frac{E\,[\mathrm{GeV}]}{\Delta m^{2}\,[\mathrm{eV^{2}}]}\ \mathrm{km}. (4.2.44)

Numbers, for the T2K experiment. A muon-neutrino beam is made at Tokai and detected at Kamioka, L=295 kmL=295\ \mathrm{km} away, with the beam tuned to E0.600 GeVE\approx0.600\ \mathrm{GeV}. With Δm322=2.45×103 eV2\Delta m^{2}_{32}=2.45\times10^{-3}\ \mathrm{eV^{2}} the phase is 1.26693×2.45×103×295/0.600=1.5261 rad1.26693\times2.45\times10^{-3}\times295/0.600=1.5261\ \mathrm{rad}, which is 0.486π0.486\pi, within three per cent of π/2\pi/2, the first oscillation maximum. So

sin2(1.5261)=0.9980,Pνμνμ  =  1sin22θ23×0.9980  =  0.0020 \sin^{2}(1.5261)=0.9980, \qquad P_{\nu_{\mu}\to\nu_{\mu}} \;=\; 1-\sin^{2}2\theta_{23}\times0.9980 \;=\; 0.0020 (4.2.45)

for maximal mixing: essentially every muon neutrino in the beam has become something else by the time it arrives. The oscillation length is Losc=2.4797×0.600/2.45×103=607.3 kmL_{\text{osc}}=2.4797\times0.600/2.45\times10^{-3}=607.3\ \mathrm{km}, so the first minimum sits at 303.6 km303.6\ \mathrm{km} against a baseline of 295 km295\ \mathrm{km}. Equivalently, 295 km295\ \mathrm{km} is exactly the first minimum for a beam energy of 0.583 GeV0.583\ \mathrm{GeV}. The beam energy was chosen to put the detector at the minimum, and (4.2.42) is how it was chosen.

The same formula with a small mixing angle gives the other kind of experiment. At the Daya Bay reactor, electron antineutrinos of about 4 MeV4\ \mathrm{MeV} are counted at L=1.65 kmL=1.65\ \mathrm{km}. The phase is 1.26693×2.5×103×1.65/0.004=1.307 rad1.26693\times2.5\times10^{-3}\times1.65/0.004=1.307\ \mathrm{rad}, so with sin22θ13=0.085\sin^{2}2\theta_{13}=0.085 the predicted disappearance is 0.085×sin2(1.307)=0.0790.085\times\sin^{2}(1.307)=0.079. That is a deficit of about eight per cent, which is what is measured, and it is how θ13\theta_{13} is known.

One honesty note about the derivation. Treating the two mass states as plane waves with a common momentum, and setting t=L/ct=L/c, is a shortcut. A neutrino is produced as a localised wave packet, the two mass components travel at slightly different speeds, and the careful treatment follows the packets and asks when they still overlap at the detector. It gives (4.2.42) in every regime these experiments operate in, and the corrections are suppressed by the ratio of the packet width to LoscL_{\text{osc}}. The shortcut is standard, and it is a shortcut. The machinery to do it properly is Chapter 4.6's wave packets.

10.4 · The qubit

The third reading inverts the relationship between the formula and the apparatus. In ammonia the parameters are whatever chemistry supplies. In a neutrino they are whatever the mass matrix supplies. In an engineered two-level system both are set by the designer, and (4.2.35) becomes a specification rather than a prediction.

Take a superconducting circuit with two levels split by ω0\hbar\omega_{0}, with ω0/2π=5 GHz\omega_{0}/2\pi=5\ \mathrm{GHz}. That is a splitting of 20.68 μeV20.68\ \mu\mathrm{eV}, which is twelve times kBT=1.72 μeVk_{B}T=1.72\ \mu\mathrm{eV} at the 20 mK20\ \mathrm{mK} such circuits run at. So Chapter 0.6's Boltzmann factor eΔE/kBT=6×106\ee^{-\Delta E/k_{B}T}=6\times10^{-6} leaves the system in its lower level rather than thermally stirred.

Now drive it with a microwave field at exactly ω0\omega_{0}. Chapter 4.17 shows that in a frame rotating with the drive the problem becomes (4.2.31) with δ=0\delta=0 and Δ=ΩR/2\Delta=\hbar\Omega_{R}/2, where ΩR\Omega_{R} is proportional to the drive amplitude. That reduction is Chapter 4.17's and is not derived here. With it granted, (4.2.35) reads

P01(t)  =  sin2 ⁣(ΩRt2), P_{0\to1}(t) \;=\; \sin^{2}\!\left(\frac{\Omega_{R}t}{2}\right), (4.2.46)

the Rabi formula. With ΩR/2π=25 MHz\Omega_{R}/2\pi=25\ \mathrm{MHz} the full period is 40 ns40\ \mathrm{ns}, so a pulse lasting

tπ  =  πΩR  =  12×25 MHz  =  20 ns t_{\pi} \;=\; \frac{\pi}{\Omega_{R}} \;=\; \frac{1}{2\times25\ \mathrm{MHz}} \;=\; 20\ \mathrm{ns} (4.2.47)

takes 0\ket0 to 1\ket1 with probability 11. That is a NOT gate, performed by leaving the drive on for exactly half an oscillation. Half that, 10 ns10\ \mathrm{ns}, gives probability 12\half and produces the operator whose square is NOT. Worked example 3 builds it and shows that no classical stochastic process has such a square root. Detuning the drive off resonance is δ0\delta\neq0, and (4.2.35) says the gate then fails to reach probability 11 no matter how long it is left on, with the shortfall 1sin22θ=δ2/(δ2+Δ2)1-\sin^{2}2\theta=\delta^{2}/(\delta^{2}+\Delta^{2}). That is where the tolerance on a control line comes from.

Set the three side by side. Same matrix, same formula, three sets of units:

1,2\ket1,\ket2 areδ\deltaSplitting ΔE\Delta EFull period of P12P_{1\to2}
Ammonianitrogen above / below the hydrogens00, by mirror symmetryh×23.8701 GHz=98.72 μeVh\times23.8701\ \mathrm{GHz}=98.72\ \mu\mathrm{eV}41.89 ps41.89\ \mathrm{ps}
Neutrinothe two detectable flavoursset by the mass matrixΔm2c4/2E=2.04 peV\Delta m^{2}c^{4}/2E=2.04\ \mathrm{peV}607.3 km607.3\ \mathrm{km} of flight
Qubitthe two engineered levelsdrive detuning, chosenΩR=0.1034 μeV\hbar\Omega_{R}=0.1034\ \mu\mathrm{eV}40.0 ns40.0\ \mathrm{ns}

The splittings span a factor of 4.8×1074.8\times10^{7}, from 98.72 μeV98.72\ \mu\mathrm{eV} down to 2.04 peV2.04\ \mathrm{peV}, and the periods span exactly the same factor the other way, because (4.2.35) makes period and splitting reciprocal and nothing else enters. The linear algebra is one 2×22\times2 matrix that Chapter 0.5 diagonalised before any of this was mentioned.

10.5 · One computed evolution, read in three sets of units

The figure below integrates (4.2.18) numerically for (4.2.31) and plots what comes out. Nothing in it uses (4.2.35). The closed form appears only in the readout, as the thing being checked against. Two of the readouts are the numerical confirmation this chapter owes.

0.00
2.00
0.000 pi
normalisation: max | |c1|^2 + |c2|^2 - 1 | = 5.22e-15 integration against the closed form: max difference = 8.88e-15 (nothing plotted used the closed form)
measured from the curve: period = 1.0000000 pi, amplitude = 1.0000000 formula: pi/sqrt(1+d^2) = 1.0000000 pi, sin^2 2theta = 1/(1+d^2) = 1.0000000
the same curve, three sets of units: ammonia 41.893 ps neutrinos 607.27 km of flight qubit 40.00 ns (half of each is a complete transfer)
global phase alpha = 0.000 pi: the plotted curve differs from the alpha = 0 run by at most 0.00e+0 -- P1, measured
One matrix, integrated, and read off three ways. The Schrödinger equation idc/dτ=(d11d)c\ii\,\dd c/\dd\tau=\big(\begin{smallmatrix}d&1\\ 1&-d\end{smallmatrix}\big)c is integrated by fourth-order Runge–Kutta in the dimensionless time τ=Δt/\tau=\Delta t/\hbar with d=δ/Δd=\delta/\Delta on the slider, starting from eiα(1,0)\ee^{\ii\alpha}(1,0). Blue is c22\abs{c_{2}}^{2}, green is c12\abs{c_{1}}^{2}, and the flat amber line is their sum. The faint curve is the d=0d=0 case, held fixed for comparison; the purple dashed parabola is τ2\tau^{2}. Move the detuning slider and watch two things happen at once. The oscillation gets shallower, settling on the amplitude sin22θ=1/(1+d2)\sin^{2}2\theta=1/(1+d^{2}), and it gets faster, by the factor 1+d2\sqrt{1+d^{2}} — and the readout confirms both against (4.2.35) to seven figures although the plotted curve never used it. Then look at where the curves leave the origin. Every one of them, at every detuning, departs along the same parabola τ2\tau^{2}; the detuning changes when the curve turns over and not how fast it starts, which is (4.2.36) and is what Chapter 4.17's transition rate is built on. Pull the cycles slider down to 0.50.5 to give that corner of the plot the whole width, then sweep the detuning: the departure does not move. The amber line reads 11 with a worst deviation of about 101410^{-14} over the whole integration, which is the arithmetic of the machine rather than of the physics — that is P1 and P3 agreeing, and it holds at every slider setting. The global-phase slider does nothing, and the readout says by how much nothing: moving α\alpha through a full turn changes the plotted curve by about 101410^{-14}, which is (4.2.2) measured rather than asserted. The last line converts the same dimensionless curve into each system's own units. At δ=0\delta=0 it reads 41.8941.89 ps, 607.3607.3 km and 40.040.0 ns; away from δ=0\delta=0 the ammonia column is no longer ammonia, because §10.2's mirror symmetry forces δ=0\delta=0 for that molecule, while the neutrino and the qubit columns remain physical — θ23\theta_{23} is not exactly 4545^{\circ} and a drive is never exactly on resonance.
In plain terms 4.2.10

The point of this section is that there is one calculation, and it is finished. Take two states of nearly the same energy with something connecting them, which is the smallest interesting arrangement there is, and solve it once. What comes out is a single expression: the chance of finding the system in the other state swings back and forth forever, at a rate set by the separation of the two energy levels, with a depth set by how evenly the connection mixes them.

Now read the answer three times. In a molecule of ammonia the two states are the nitrogen atom sitting on either side of its three hydrogens, the connection is its ability to pass through them, and the swing takes forty-two trillionths of a second — which corresponds to a microwave line that was used to build the first device of its kind. In a neutrino the two states are the two identities it can be detected with, the connection is the mismatch between those and the two definite masses, and the swing takes six hundred kilometres of flight. In an engineered two-level circuit both numbers are chosen by the designer, and half a swing at twenty nanoseconds is what turns one state into the other.

Three subjects, three sets of units, one matrix. That is what the chapter has been claiming.

11 · Worked examples

Worked example 1 — three magnets in a row, and what P4 does that P3 cannot

A beam of spin-12\half particles is prepared with a definite value of S^z\hat S_{z}, passed through a second apparatus measuring S^x\hat S_{x}, and then through a third measuring S^z\hat S_{z} again. Using only P3 and P4: (a) compute the probability of the third magnet reporting spin-down. (b) Compute the same probability with the middle magnet removed. (c) Compute it with the middle magnet present but not read, so that the two xx-paths are recombined coherently. (d) Say which postulate each answer used, and what the comparison establishes.

Set up the two-dimensional space with z+,z\ket{z{+}},\ket{z{-}} an orthonormal basis. The observable S^x\hat S_{x} has, by Chapter 0.5's Worked example 1 with E0=0E_{0}=0, the eigenstates

x±  =  12(z+±z), \ket{x{\pm}} \;=\; \tfrac{1}{\sqrt2}\big(\ket{z{+}}\pm\ket{z{-}}\big),

which are orthogonal, as §4.1 requires without being asked. All four overlaps between the two bases have modulus 1/21/\sqrt2.

(a) Measure, keep the ++ beam, measure again. Starting from z+\ket{z{+}}, P3 gives Pr(x+)=x+z+2=12\Pr(x{+})=\abs{\avg{x{+}|z{+}}}^{2}=\half, and P4 replaces the state by x+\ket{x{+}}. Then P3 again: Pr(z)=zx+2=12\Pr(z{-})=\abs{\avg{z{-}|x{+}}}^{2}=\half. The two are independent because P4 wiped out all memory of the first state, so

Pr(z+x+z)  =  12×12  =  14. \Pr(z{+}\to x{+}\to z{-}) \;=\; \tfrac12\times\tfrac12 \;=\; \tfrac14.

(b) Middle magnet removed. Nothing happens between the two zz measurements, so the state is still z+\ket{z{+}} and P3 gives Pr(z)=zz+2=0\Pr(z{-})=\abs{\avg{z{-}|z{+}}}^{2}=0, exactly. The two eigenstates of S^z\hat S_{z} are orthogonal, and §4.1's second theorem says the confusion probability is zero rather than small.

(c) Middle magnet present, nothing read, beams recombined. Now no measurement occurred, so P4 does not apply and P3 is used once at the end. The state is unchanged and can be rewritten using the resolution of the identity in the xx basis, which is Chapter 0.5 §2.3's eiei=I^\sum\ket{e_{i}}\bra{e_{i}}=\hat I inserted and nothing more:

zz+=zI^z+=zx+x+z++12+zxxz+12  =  0. \avg{z{-}|z{+}} = \avg{z{-}|\hat I|z{+}} = \underbrace{\avg{z{-}|x{+}}\avg{x{+}|z{+}}}_{+\frac12} + \underbrace{\avg{z{-}|x{-}}\avg{x{-}|z{+}}}_{-\frac12} \;=\; 0.

The two amplitudes are equal in magnitude and opposite in sign, and they cancel exactly. So the answer is 00, agreeing with (b) as it must, since inserting I^\hat I changes nothing.

(d) What the three answers establish. Compare (a) with (c). The same apparatus is in the beam line in both, the same two paths are travelled, and the same final measurement is made. In (c) the two amplitudes are added and then squared, giving 12122=0\abs{\tfrac12-\tfrac12}^{2}=0. In (a) they are squared and then added, giving 14+14\tfrac14+\tfrac14, of which one branch was kept, so 14\tfrac14. The only difference in the beam line is that in (a) the intermediate value was read. Keeping one branch afterwards is what turns the 12\tfrac12 into 14\tfrac14, and the reading is what destroyed the cancellation.

Certainty becomes a coin flip because somebody looked. That is P4 doing work that P3 alone cannot do. P3 assigns probabilities to the outcomes of a measurement, and it is P4 that destroys the coherence between the branches so that the third magnet sees no interference. It also shows why §6.2's familiar box insisted that conditioning is not the whole story. The classical law of total probability would give Pr(z)=kPr(zxk)Pr(xk)=14+14=12\Pr(z{-})=\sum_{k}\Pr(z{-}\mid x_{k})\Pr(x_{k}) =\tfrac14+\tfrac14=\tfrac12 whether anyone looked or not, and the measured answer without looking is 00.

Worked example 2 — a degenerate observable, done completely

A three-level system has the observable

A^=(010100001),ψ=13(212). \hat A = \begin{pmatrix}0&1&0\\ 1&0&0\\ 0&0&1\end{pmatrix}, \qquad \ket\psi = \tfrac13\begin{pmatrix}2\\ 1\\ 2\end{pmatrix}.

(a) Find the eigenvalues and the projections. (b) Apply P3 and check the probabilities sum to one. (c) Compute A^\avg{\hat A} and (ΔA)2(\Delta A)^{2} two ways. (d) Apply P4 and say precisely what the state becomes. (e) Find a second observable that completes the set, and give the quantum numbers of the three joint eigenstates.

(a) A^\hat A is real symmetric, hence Hermitian. Its top-left 2×22\times2 block is Chapter 0.5's σx\sigma_{x} with eigenvalues ±1\pm1 and eigenvectors 12(1,±1,0)\tfrac{1}{\sqrt2}(1,\pm1,0). The third basis vector (0,0,1)(0,0,1) is an eigenvector with eigenvalue +1+1. So the distinct eigenvalues are +1+1, twice degenerate, and 1-1, once. Building P^k\hat P_{k} as the sum of ee\ket{e}\bra{e} over an orthonormal basis of each eigenspace:

P^+=12(110110002),P^=12(110110000). \hat P_{+} = \tfrac12\begin{pmatrix}1&1&0\\ 1&1&0\\ 0&0&2\end{pmatrix}, \qquad \hat P_{-} = \tfrac12\begin{pmatrix}1&-1&0\\ -1&1&0\\ 0&0&0\end{pmatrix}.

Checks: P^++P^=I^\hat P_{+}+\hat P_{-}=\hat I ✓, P^+P^=0\hat P_{+}\hat P_{-}=0 ✓, each is idempotent and Hermitian ✓, and P^+P^=A^\hat P_{+}-\hat P_{-}=\hat A ✓, which is (4.2.7).

(b) P^+ψ=16(3,3,4)\hat P_{+}\ket\psi=\tfrac16(3,3,4) and P^ψ=16(1,1,0)\hat P_{-}\ket\psi=\tfrac16(1,-1,0), so

Pr(+1)=P^+ψ2=136(9+9+16)=3436=1718,Pr(1)=136(1+1)=118, \Pr(+1)=\norm{\hat P_{+}\psi}^{2}=\tfrac{1}{36}(9+9+16)=\tfrac{34}{36}=\tfrac{17}{18}, \qquad \Pr(-1)=\tfrac{1}{36}(1+1)=\tfrac{1}{18},

summing to 11. Note what would have gone wrong with the careless version of P3. Writing "the probability is ekψ2\abs{\avg{e_{k}|\psi}}^{2}" and picking the eigenvector 12(1,1,0)\tfrac{1}{\sqrt2}(1,1,0) alone gives 1299=12\tfrac{1}{2}\cdot\tfrac{9}{9}=\tfrac12, and picking (0,0,1)(0,0,1) alone gives 49\tfrac49. Neither of those is the probability of the outcome +1+1, which is their sum 1718\tfrac{17}{18}. The projection form is written the way it is precisely so that a degenerate eigenvalue is handled correctly and independently of which basis was chosen inside the eigenspace.

(c) From P3, A^=(+1)1718+(1)118=1618=89\avg{\hat A}=(+1)\tfrac{17}{18}+(-1)\tfrac{1}{18}=\tfrac{16}{18}=\tfrac89. From (4.2.9), ψA^ψ=19(21+12+22)=89\bra\psi\hat A\ket\psi=\tfrac19(2\cdot1+1\cdot2+2\cdot2) =\tfrac89 ✓. For the spread, A^2=I^\hat A^{2}=\hat I because both eigenvalues square to 11, so A^2=1\avg{\hat A^{2}}=1 and (ΔA)2=16481=1781(\Delta A)^{2}=1-\tfrac{64}{81}=\tfrac{17}{81}. Directly, (189)21718+(189)2118=171458+2891458=3061458=1781\big(1-\tfrac89\big)^{2}\tfrac{17}{18}+\big(-1-\tfrac89\big)^{2}\tfrac{1}{18} =\tfrac{17}{1458}+\tfrac{289}{1458}=\tfrac{306}{1458}=\tfrac{17}{81} ✓.

(d) If +1+1 is obtained, P4 gives

ψ  =  P^+ψP^+ψ  =  134(334). \ket{\psi'} \;=\; \frac{\hat P_{+}\ket\psi}{\norm{\hat P_{+}\ket\psi}} \;=\; \frac{1}{\sqrt{34}}\begin{pmatrix}3\\ 3\\ 4\end{pmatrix}.

This is the point of the example. The state after the measurement is not 12(1,1,0)\tfrac{1}{\sqrt2}(1,1,0) and not (0,0,1)(0,0,1). It is the particular direction inside the eigenspace that ψ\ket\psi already pointed along, rescaled. A degenerate measurement removes the component outside the eigenspace and touches nothing inside it, which is exactly §6.2's remark that such a measurement is weaker than a complete one.

(e) Any B^\hat B commuting with A^\hat A and distinguishing the two directions inside P^+\hat P_{+}'s eigenspace will do. Take B^=diag(1,1,1)\hat B=\operatorname{diag}(1,1,-1). Its commutator with A^\hat A is the zero matrix by direct multiplication, since B^\hat B is a multiple of the identity on the block where A^\hat A is non-diagonal. Inside the +1+1 eigenspace of A^\hat A, the operator B^\hat B has eigenvalue +1+1 on 12(1,1,0)\tfrac{1}{\sqrt2}(1,1,0) and 1-1 on (0,0,1)(0,0,1), so it chooses the basis A^\hat A could not. That is Chapter 0.5 §8.2's Step 4, in three dimensions. The joint labels are

(a,b)=(+1,+1)  12(1,1,0),(+1,1)  (0,0,1),(1,+1)  12(1,1,0), \big(a,b\big) = (+1,+1)\ \to\ \tfrac{1}{\sqrt2}(1,1,0), \qquad (+1,-1)\ \to\ (0,0,1), \qquad (-1,+1)\ \to\ \tfrac{1}{\sqrt2}(1,-1,0),

each one-dimensional, so {A^,B^}\{\hat A,\hat B\} is a complete set of commuting observables and those pairs are the quantum numbers of §4.3. Measuring B^\hat B on ψ\ket{\psi'} gives Pr(+1)=12(3+3)/342=917\Pr(+1)=\abs{\tfrac{1}{\sqrt2}(3+3)/\sqrt{34}}^{2}=\tfrac{9}{17} and Pr(1)=4/342=817\Pr(-1)=\abs{4/\sqrt{34}}^{2}=\tfrac{8}{17}, summing to one, and after it the state carries a complete label.

Worked example 3 — the square root of NOT, and why no classical process has one

Section 10.4 said that half a Rabi period turns 0\ket0 into 1\ket1 with certainty, and that a quarter period produces an operator whose square is that gate. (a) Build the quarter-period operator from §7 and check it is unitary. (b) Square it. (c) Show that the classical analogue does not exist, meaning a two-state stochastic process whose square is a certain flip. (d) Say which postulate the difference belongs to.

(a) The resonant drive of §10.4 is (4.2.31) with δ=0\delta=0 and Eˉ=0\bar E=0, so H^=Δσx\hat H=\Delta\sigma_{x} and by (4.2.17) the evolution operator is U^(t)=eiΔσxt/\hat U(t)=\ee^{-\ii\Delta\sigma_{x}t/\hbar}. Chapter 0.5's Worked example 2 evaluated exactly this exponential using σx2=I^\sigma_{x}^{2}=\hat I and got cosϑI^isinϑσx\cos\vartheta\,\hat I-\ii\sin\vartheta\,\sigma_{x} with ϑ=Δt/\vartheta=\Delta t/\hbar. The quarter period is ϑ=π/4\vartheta=\pi/4, giving

V^    eiπσx/4  =  12(1ii1). \hat V \;\equiv\; \ee^{-\ii\pi\sigma_{x}/4} \;=\; \frac{1}{\sqrt2}\begin{pmatrix}1 & -\ii\\ -\ii & 1\end{pmatrix}.

Unitary, since V^V^=12(1ii1)(1ii1)=I^\hat V^{\dagger}\hat V=\tfrac12\big(\begin{smallmatrix}1&\ii\\ \ii&1\end{smallmatrix}\big)\big(\begin{smallmatrix}1&-\ii\\ -\ii&1\end{smallmatrix}\big)=\hat I ✓, as §7.2 guarantees for any eiH^t/\ee^{-\ii\hat Ht/\hbar} with H^\hat H Hermitian. Applied to 0\ket0 it gives 12(1,i)\tfrac{1}{\sqrt2}(1,-\ii), whose two probabilities are 12\half and 12\half.

(b) Squaring,

V^2  =  12(1ii1)(1ii1)  =  12(02i2i0)  =  iσx. \hat V^{2} \;=\; \tfrac12\begin{pmatrix}1&-\ii\\ -\ii&1\end{pmatrix}\begin{pmatrix}1&-\ii\\ -\ii&1\end{pmatrix} \;=\; \tfrac12\begin{pmatrix}0&-2\ii\\ -2\ii&0\end{pmatrix} \;=\; -\ii\,\sigma_{x}.

That is the NOT gate multiplied by i-\ii, and by P1 an overall factor of modulus one is not there. So V^\hat V applied twice takes 0\ket0 to 1\ket1 with probability 11. It is a genuine square root of NOT, and it is what the pulse of §10.4 performs after 10 ns10\ \mathrm{ns}.

(c) Now the classical version. A stochastic process on two states that treats them symmetrically is described by the matrix S=(a1a1aa)S=\big(\begin{smallmatrix}a&1-a\\ 1-a&a\end{smallmatrix}\big) with 0a10\le a\le1, meaning "stay with probability aa, flip with probability 1a1-a". Two applications give

S2=(a2+(1a)22a(1a)2a(1a)a2+(1a)2), S^{2} = \begin{pmatrix}a^{2}+(1-a)^{2} & 2a(1-a)\\ 2a(1-a) & a^{2}+(1-a)^{2}\end{pmatrix},

and requiring S2S^{2} to be the certain flip means requiring the diagonal entry a2+(1a)2a^{2}+(1-a)^{2} to vanish. It is a sum of two squares of real numbers, so it vanishes only if both vanish, which needs a=0a=0 and a=1a=1 at once. Its minimum over all real aa is 12\half, attained at a=12a=\half. No classical two-state process, applied twice, produces a certain flip. The best possible leaves a half chance of being where it started.

(d) The difference is P1's, not P3's. The quantum operator succeeds because the two intermediate amplitudes are 1/21/\sqrt2 and i/2-\ii/\sqrt2, and on the second application the contributions to "stay" are 12\tfrac12 and (i)212=12(-\ii)^{2}\tfrac12=-\tfrac12, which cancel. Cancellation requires the entries to be complex numbers of opposite sign, which requires the state to be a vector in a complex space rather than a list of probabilities. Section 5.4 named that as the single place a reader is most likely to install the wrong picture. Half-way through a NOT gate the qubit is not "probably still 0\ket0". It is in a definite state whose phase is what makes the second half of the gate work.

12 · Your turn

Problem 1 — rays, not vectors, and the sphere that follows

(a) Show that the set of states of an nn-level system has 2n22n-2 real parameters, and say which constraint removes each of the two. (b) For n=2n=2, show that every state can be written as (4.2.6) for exactly one pair (θ,φ)(\theta,\varphi) with 0θπ0\le\theta\le\pi and 0φ<2π0\le\varphi\lt2\pi, treating θ=0\theta=0 and θ=π\theta=\pi as the two exceptions and saying what happens there. (c) Compute σx,σy,σz\avg{\sigma_{x}},\avg{\sigma_{y}},\avg{\sigma_{z}} in the state (4.2.6), with σx=(0110)\sigma_{x}=\big(\begin{smallmatrix}0&1\\ 1&0\end{smallmatrix}\big), σy=(0ii0)\sigma_{y}=\big(\begin{smallmatrix}0&-\ii\\ \ii&0\end{smallmatrix}\big), σz=(1001)\sigma_{z}=\big(\begin{smallmatrix}1&0\\ 0&-1\end{smallmatrix}\big), and identify what the three numbers are. (d) Deduce that two states are the same state exactly when those three expectation values agree, and say why that makes the Bloch sphere a measurement rather than a picture.

Solution

(a) A vector in Cn\C^{n} is 2n2n real numbers. The condition ψ=1\norm\psi=1 is one real equation, leaving 2n12n-1. P1 identifies ψ\ket\psi with eiαψ\ee^{\ii\alpha}\ket\psi, and α\alpha is one real parameter, leaving 2n22n-2. Normalisation removes one, and the global phase removes the other.

(b) Write ψ=c11+c22\ket\psi=c_{1}\ket1+c_{2}\ket2. Choose the phase freedom to make c1c_{1} real and non-negative, which fixes α\alpha uniquely unless c1=0c_{1}=0. Then c1=cos(θ/2)c_{1}=\cos(\theta/2) for a unique θ[0,π]\theta\in[0,\pi], since cos(θ/2)\cos(\theta/2) runs monotonically from 11 to 00 on that range, and c2=1c12=sin(θ/2)\abs{c_{2}}=\sqrt{1-c_{1}^{2}}=\sin(\theta/2), so c2=eiφsin(θ/2)c_{2}=\ee^{\ii\varphi}\sin(\theta/2) with φ\varphi unique modulo 2π2\pi. The exceptions are the poles. At θ=0\theta=0 we have c2=0c_{2}=0 and φ\varphi is undefined, and at θ=π\theta=\pi we have c1=0c_{1}=0 and φ\varphi is the leftover phase, which P1 discards. Both are single states, and they are exactly the two points where a polar coordinate system on a sphere degenerates. The parameter count is right, and the coordinates are the ones misbehaving.

(c) With c1=cos(θ/2)c_{1}=\cos(\theta/2), c2=eiφsin(θ/2)c_{2}=\ee^{\ii\varphi}\sin(\theta/2) and (4.2.9),

σz=c12c22=cos2θ2sin2θ2=cosθ, \avg{\sigma_{z}} = \abs{c_{1}}^{2}-\abs{c_{2}}^{2} = \cos^{2}\tfrac\theta2-\sin^{2}\tfrac\theta2 = \cos\theta, σx=cˉ1c2+cˉ2c1=2Re(cˉ1c2)=2cosθ2sinθ2cosφ=sinθcosφ, \avg{\sigma_{x}} = \bar c_{1}c_{2}+\bar c_{2}c_{1} = 2\operatorname{Re}\big(\bar c_{1}c_{2}\big) = 2\cos\tfrac\theta2\sin\tfrac\theta2\cos\varphi = \sin\theta\cos\varphi,

and the same computation with σy\sigma_{y} gives 2Im(cˉ1c2)-2\operatorname{Im}(\bar c_{1}c_{2}) with the sign convention above, which is sinθsinφ\sin\theta\sin\varphi. So the three expectation values are the Cartesian coordinates of the point at polar angle θ\theta and azimuth φ\varphi on the unit sphere. Check the length: sin2θcos2φ+sin2θsin2φ+cos2θ=1\sin^{2}\theta\cos^{2}\varphi+\sin^{2}\theta\sin^{2}\varphi +\cos^{2}\theta=1, so every state sits exactly on the sphere and none inside it. (States inside it exist and are mixtures, which is Chapter 4.19's density operator.)

(d) The map from (θ,φ)(\theta,\varphi) to the point is a bijection onto the sphere by (b) and (c), so equal expectation values force equal (θ,φ)(\theta,\varphi) and hence the same state. Each of the three numbers is the average of a measurement, and Chapter 4.12 measures exactly these with three Stern–Gerlach magnets in three orientations. So the position of a state on the sphere is determined by an experiment, and the sphere is where the states are rather than a way of imagining them. Note that this fails for n3n\ge3. The parameter count gives four real parameters, there is no two-sphere, and no comparable picture exists.

Problem 2 — how the ammonia maser sorts its molecules

Put the molecule of §10.2 in a uniform electric field E\mathcal{E}. The two configurations L\ket L and R\ket R have electric dipole moments of magnitude μ\mu pointing in opposite directions along the symmetry axis, so the field shifts their energies by μE\mp\mu\mathcal{E}: the matrix is (4.2.31) with δ=μE\delta=\mu\mathcal{E} and Δ\Delta unchanged. (a) Write down the two energies as functions of E\mathcal{E} and sketch their behaviour at small and large field. (b) Find the field at which the two régimes cross over, and evaluate it for μ=1.4718 D=4.909×1030 Cm\mu=1.4718\ \mathrm{D}=4.909\times10^{-30}\ \mathrm{C\,m}. (c) Explain, from (a), why passing a beam through an inhomogeneous field separates the two energy eigenstates, and which way each is deflected. (d) What happens to the inversion oscillation of (4.2.37) as the field is raised?

Solution

(a) Straight from (4.2.32),

E±(E)  =  Eˉ±Δ2+μ2E2. E_{\pm}(\mathcal{E}) \;=\; \bar E \pm\sqrt{\Delta^{2}+\mu^{2}\mathcal{E}^{2}}.

At small field, expand: Δ2+μ2E2=Δ(1+μ2E2/2Δ2+)\sqrt{\Delta^{2}+\mu^{2}\mathcal{E}^{2}} =\Delta\big(1+\mu^{2}\mathcal{E}^{2}/2\Delta^{2}+\cdots\big), so the shift is quadratic in E\mathcal{E} with coefficient ±μ2/2Δ\pm\mu^{2}/2\Delta. At large field the square root is dominated by μE\mu\mathcal{E} and the shift is linear, ±μE\pm\mu\mathcal{E}, which is the classical answer for a dipole in a field. The two levels never cross. They repel to a minimum separation of 2Δ2\Delta at zero field, which is the avoided crossing of §10.1.

(b) The crossover is where the two terms under the root are equal, μE=Δ\mu\mathcal{E}=\Delta. With Δ=12hν0=7.908×1024 J\Delta=\half h\nu_{0}=7.908\times10^{-24}\ \mathrm{J},

E×=Δμ=7.908×10244.909×1030=1.611×106 Vm1=16.1 kVcm1. \mathcal{E}_{\times} = \frac{\Delta}{\mu} = \frac{7.908\times10^{-24}}{4.909\times10^{-30}} = 1.611\times10^{6}\ \mathrm{V\,m^{-1}} = 16.1\ \mathrm{kV\,cm^{-1}}.

At 1 kVcm11\ \mathrm{kV\,cm^{-1}} the splitting has grown by only 0.19%0.19\%. At 100 kVcm1100\ \mathrm{kV\,cm^{-1}} it has grown by a factor of 6.296.29, and the response is essentially linear.

(c) An inhomogeneous field exerts a force E±(E(x))-\nabla E_{\pm}(\mathcal{E}(\vv x)) on a molecule in the corresponding state. Since E+E_{+} increases with field strength and EE_{-} decreases, the upper state is pushed towards weak field and the lower state towards strong field. They go opposite ways, and a quadrupole field with a minimum on the axis focuses the upper state down the axis while throwing the lower state out. That is the maser's state selector, and what emerges is a beam of molecules in the upper energy state only. That is a population inversion, which Chapter 4.1's Worked example 3 proved cannot occur in thermal equilibrium at any temperature, and which therefore has to be manufactured.

(d) By (4.2.35) the oscillation amplitude is sin22θ=Δ2/(Δ2+μ2E2)\sin^{2}2\theta=\Delta^{2}/(\Delta^{2}+\mu^{2}\mathcal{E}^{2}), which falls as the field rises, while the frequency 2R/h2R/h rises. At EE×\mathcal{E}\gg\mathcal{E}_{\times} the amplitude goes as (Δ/μE)2(\Delta/\mu\mathcal{E})^{2} and the molecule is essentially frozen into whichever configuration it started in, because the field has made the two configurations energetically distinguishable by far more than the tunnelling can bridge. A strong field switches the inversion off, and this is the same mechanism by which a large molecule's environment freezes it into one enantiomer.

Problem 3 — how badly a finite matrix misses the canonical commutator

Section 8.3 asserted a bound. Prove it and interpret it. (a) Using the inner product A,B=tr(AB)\avg{A,B}=\operatorname{tr}(A^{\dagger}B) on n×nn\times n complex matrices, which is Chapter 0.5 §1.2's fifth row, verify the three axioms hold, and compute I^F\norm{\hat I}_{F}. (b) Prove trMnMF\abs{\operatorname{tr}M}\le\sqrt n\,\norm{M}_{F} for every MM, and state which theorem of Chapter 0.5 you used. (c) Deduce [X^,P^]iI^Fn\norm{[\hat X,\hat P]-\ii\hbar\hat I}_{F}\ge\hbar\sqrt n for all n×nn\times n matrices X^,P^\hat X,\hat P, and say what the right-hand side is the norm of. (d) Determine exactly when equality holds, using Chapter 0.5 §1.4's equality condition, and say what that means about how close a finite-dimensional model can get. Check your answer against the pair X^=(0100)\hat X=\big(\begin{smallmatrix}0&1\\ 0&0\end{smallmatrix}\big), P^=(00i0)\hat P=\hbar\big(\begin{smallmatrix}0&0\\ \ii&0\end{smallmatrix}\big).

Solution

(a) Conjugate symmetry: tr(AB)=tr((AB))=tr(BA)\overline{\operatorname{tr}(A^{\dagger}B)} =\operatorname{tr}\big((A^{\dagger}B)^{\dagger}\big)=\operatorname{tr}(B^{\dagger}A), since the trace of the conjugate transpose is the conjugate of the trace. Linearity in the second slot is linearity of the trace. Positive definiteness: tr(AA)=i(AA)ii=i,jAjiAji=ijAij20\operatorname{tr}(A^{\dagger}A)=\sum_{i}(A^{\dagger}A)_{ii} =\sum_{i,j}\overline{A_{ji}}A_{ji}=\sum_{ij}\abs{A_{ij}}^{2}\ge0, vanishing only if every entry does. And I^F2=trI^=n\norm{\hat I}_{F}^{2}=\operatorname{tr}\hat I=n, so I^F=n\norm{\hat I}_{F}=\sqrt n.

(b) trM=tr(I^M)=I^,M\operatorname{tr}M=\operatorname{tr}(\hat I^{\dagger}M)=\avg{\hat I,M}, so Cauchy–Schwarz (Chapter 0.5 §1.4) applied in this inner-product space gives I^,M2I^,I^M,M=nMF2\abs{\avg{\hat I,M}}^{2}\le\avg{\hat I,\hat I}\avg{M,M}=n\norm{M}_{F}^{2}. Take square roots. The point worth noticing is that Chapter 0.5 proved Cauchy–Schwarz from the three axioms alone and never mentioned what the vectors were, so it applies verbatim to a space whose vectors are matrices.

(c) Put M=[X^,P^]iI^M=[\hat X,\hat P]-\ii\hbar\hat I. By cyclicity of the trace (Chapter 0.4 §6.1) the commutator is traceless, so trM=in\operatorname{tr}M=-\ii\hbar n and trM=n\abs{\operatorname{tr}M}=\hbar n. Then (b) gives nnMF\hbar n\le\sqrt n\,\norm{M}_{F}, that is MFn\norm{M}_{F}\ge\hbar\sqrt n. The right-hand side is exactly iI^F=I^F=n\norm{\ii\hbar\hat I}_{F}=\hbar\norm{\hat I}_{F}=\hbar\sqrt n: the error is at least as big as the target.

(d) Chapter 0.5 §1.4 proved that Cauchy–Schwarz is an equality exactly when the two vectors are parallel, so equality here needs M=cI^M=c\hat I for some complex cc. Taking the trace, cn=trM=incn=\operatorname{tr}M=-\ii\hbar n, so c=ic=-\ii\hbar and

[X^,P^]iI^=iI^[X^,P^]=0. \big[\hat X,\hat P\big]-\ii\hbar\hat I = -\ii\hbar\hat I \qquad\Longleftrightarrow\qquad \big[\hat X,\hat P\big]=0.

The bound is attained exactly when the two matrices commute, which is to say when they reproduce none of the canonical relation whatever. Every pair that actually tries, by having a non-zero commutator, does strictly worse. Check it on the suggested pair. We have X^P^=diag(i,0)\hat X\hat P=\hbar\operatorname{diag}(\ii,0) and P^X^=diag(0,i)\hat P\hat X=\hbar\operatorname{diag}(0,\ii), so [X^,P^]=idiag(1,1)[\hat X,\hat P]=\ii\hbar\operatorname{diag}(1,-1), which has half of it exactly right. Then M=idiag(0,2)M=\ii\hbar\operatorname{diag}(0,-2) with MF=2\norm{M}_{F}=2\hbar, a relative error of 2/(2)=1.4142\hbar/(\hbar\sqrt2)=1.414 against the commuting pair's 1.0001.000. Getting one diagonal entry exactly right cost more than giving up did.

So there is no sequence of finite-dimensional models converging to the canonical commutator, not even slowly: the relative error is bounded below by 11 at every dimension, and the minimum is achieved by doing nothing. Numerical minimisation over all complex n×nn\times n pairs from twelve random starts returns 1.0000001.000000 at n=2,3,4,5n=2,3,4,5, which is the bound and its equality case, found without being told about either.

Problem 4 — design a neutrino experiment

The other mass splitting is Δm212=7.53×105 eV2\Delta m^{2}_{21}=7.53\times10^{-5}\ \mathrm{eV^{2}} with sin22θ12=0.851\sin^{2}2\theta_{12}=0.851. Reactors emit electron antineutrinos with energies of a few MeV; take E=4 MeVE=4\ \mathrm{MeV}. (a) Using (4.2.43), find the baseline at which the disappearance is greatest, and the oscillation length. (b) The KamLAND experiment used reactors at an average of about 180 km180\ \mathrm{km}. Compute the survival probability there at 4 MeV4\ \mathrm{MeV} and say where on the oscillation curve that sits. (c) Daya Bay, measuring the other angle, is at 1.65 km1.65\ \mathrm{km}. Show that this is the right baseline for Δm312\Delta m^{2}_{31} and the wrong one for Δm212\Delta m^{2}_{21}, with numbers. (d) A colleague proposes measuring Δm212\Delta m^{2}_{21} with a detector 100 m100\ \mathrm{m} from a reactor, arguing that being closer means more neutrinos. Say quantitatively why this fails.

Solution

(a) Maximum disappearance is the first maximum of sin2\sin^{2}, at phase π/2\pi/2:

1.26693Δm2LE=π2L=πE2×1.26693Δm2=π×0.0042×1.26693×7.53×105=65.9 km, 1.26693\,\frac{\Delta m^{2}L}{E}=\frac{\pi}{2} \quad\Longrightarrow\quad L=\frac{\pi E}{2\times1.26693\,\Delta m^{2}} = \frac{\pi\times0.004}{2\times1.26693\times7.53\times10^{-5}} = 65.9\ \mathrm{km},

with EE in GeV. The oscillation length from (4.2.44) is Losc=2.4797×0.004/7.53×105=131.7 kmL_{\text{osc}}=2.4797\times0.004/7.53\times10^{-5}=131.7\ \mathrm{km}, twice the first maximum as it must be.

(b) At L=180 kmL=180\ \mathrm{km} the phase is 1.26693×7.53×105×180/0.004=4.293 rad=1.366π1.26693\times7.53\times10^{-5}\times180/0.004=4.293\ \mathrm{rad}=1.366\pi, so sin2=0.834\sin^{2}=0.834 and

Psurvival=10.851×0.834=0.29. P_{\text{survival}} = 1-0.851\times0.834 = 0.29.

That is past the first minimum and on the way back up the second oscillation. Now notice what this implies about the analysis. The phase depends on EE, the reactor spectrum spans roughly 22 to 8 MeV8\ \mathrm{MeV}, and at 180 km180\ \mathrm{km} the phase varies from 2.12.1 to 8.6 rad8.6\ \mathrm{rad} across that band, which is more than a full oscillation. The measured survival is therefore an average over the spectrum and is closer to 0.60.6. What makes the measurement powerful is that the shape of the distortion as a function of 1/E1/E fixes Δm2\Delta m^{2} far better than any single rate could.

(c) For Δm312=2.5×103\Delta m^{2}_{31}=2.5\times10^{-3} at 4 MeV4\ \mathrm{MeV}, the first maximum is at L=π×0.004/(2×1.26693×2.5×103)=1.98 kmL=\pi\times0.004/(2\times1.26693\times2.5\times10^{-3})=1.98\ \mathrm{km}, so 1.65 km1.65\ \mathrm{km} sits at phase 1.31 rad1.31\ \mathrm{rad}, giving sin2=0.93\sin^{2}=0.93. That is near enough to the maximum, and §10.3 computed the resulting 8%8\% deficit. For Δm212\Delta m^{2}_{21} at the same baseline the phase is 1.26693×7.53×105×1.65/0.004=0.0394 rad1.26693\times7.53\times10^{-5}\times1.65/0.004=0.0394\ \mathrm{rad}, so sin2=1.55×103\sin^{2}=1.55\times10^{-3} and the disappearance from that term is 0.851×1.55×103=0.13%0.851\times1.55\times10^{-3}=0.13\%, invisible under the systematics. The two splittings differ by a factor of 3333, so the two baselines differ by the same factor, and one experiment cannot do both.

(d) The flux does rise, as 1/L21/L^{2}, so at 100 m100\ \mathrm{m} instead of 65.9 km65.9\ \mathrm{km} there are (65900/100)2=4.3×105(65900/100)^{2}=4.3\times10^{5} times as many events. But the signal is not the number of events. It is the fraction missing, and at L=0.1 kmL=0.1\ \mathrm{km} the phase is 2.4×103 rad2.4\times10^{-3}\ \mathrm{rad}, so sin2(phase)=5.7×106\sin^{2}(\text{phase})=5.7\times10^{-6} and the disappearance is 4.8×1064.8\times10^{-6}. That is five parts in a million, against a reactor flux normalisation known to a per cent at best. Increasing the count rate by 4.3×1054.3\times10^{5} improves the statistical error by 4.3×105=656\sqrt{4.3\times10^{5}}=656, while the effect being sought has shrunk by 0.851/4.8×1061.8×1050.851/4.8\times10^{-6}\approx1.8\times10^{5}. The trade is losing by a factor of about 270270, and that is before the systematic error, which does not improve with the count rate at all and which is already far larger than the signal. Oscillation experiments are placed by (4.2.44) and not by flux.

Problem 5 — entanglement, by counting

(a) Show that the four-dimensional space of two two-level systems contains states that are not of the form ϕχ\ket\phi\otimes\ket\chi, by exhibiting one and proving no factorisation exists. (b) A general two-qubit state is c0000+c0101+c1010+c1111c_{00}\ket{00}+c_{01}\ket{01}+c_{10}\ket{10}+c_{11}\ket{11}. Show that it factorises if and only if c00c11=c01c10c_{00}c_{11}=c_{01}c_{10}, and interpret that as one complex condition on four complex numbers. (c) Confirm §9.2's parameter count for m=n=2m=n=2 using (b), remembering P1. (d) For Ψ=12(00+11)\ket\Psi=\tfrac{1}{\sqrt2}(\ket{00}+\ket{11}), compute the probabilities of the four outcomes of measuring σz\sigma_{z} on both systems, and then the two outcomes of measuring σz\sigma_{z} on the first system alone. Say what is strange about the answer and which chapter handles it.

Solution

(a) Take Ψ=12(00+11)\ket\Psi=\tfrac{1}{\sqrt2}(\ket{00}+\ket{11}) and suppose Ψ=(a0+b1)(c0+d1)\ket\Psi=(a\ket0+b\ket1)\otimes(c\ket0+d\ket1). Expanding by (4.2.28) and matching the four coefficients gives (4.2.29): ac=bd=1/2ac=bd=1/\sqrt2 and ad=bc=0ad=bc=0. From ad=0ad=0, either a=0a=0, which contradicts ac=1/2ac=1/\sqrt2, or d=0d=0, which contradicts bd=1/2bd=1/\sqrt2. No factorisation exists.

(b) If the state factorises then cij=ϕiχjc_{ij}=\phi_{i}\chi_{j} and c00c11=ϕ0χ0ϕ1χ1=c01c10c_{00}c_{11}=\phi_{0}\chi_{0}\phi_{1}\chi_{1}=c_{01}c_{10}, so the condition is necessary. Conversely suppose c00c11=c01c10c_{00}c_{11}=c_{01}c_{10}. If c000c_{00}\neq0, set ϕ0=1\phi_{0}=1, ϕ1=c10/c00\phi_{1}=c_{10}/c_{00}, χ0=c00\chi_{0}=c_{00}, χ1=c01\chi_{1}=c_{01}. Then ϕ0χ0=c00\phi_{0}\chi_{0}=c_{00} ✓, ϕ0χ1=c01\phi_{0}\chi_{1}=c_{01} ✓, ϕ1χ0=c10\phi_{1}\chi_{0}=c_{10} ✓, and ϕ1χ1=c10c01/c00=c11\phi_{1}\chi_{1}=c_{10}c_{01}/c_{00}=c_{11} by hypothesis ✓. If c00=0c_{00}=0 the same argument runs from whichever coefficient is non-zero. So factorising is exactly the vanishing of the 2×22\times2 determinant c00c11c01c10c_{00}c_{11}-c_{01}c_{10}. That is one complex equation on four complex numbers, which is a two-real-dimensional condition, and it confirms that product states are a thin set rather than a large one.

(c) Product states satisfy one complex equation inside C4\C^{4}, so they form a set of complex dimension 3=m+n13=m+n-1 with m=n=2m=n=2, exactly §9.2's count. Now count them as states. Imposing Ψ=1\norm\Psi=1 and discarding the global phase removes two real dimensions, by Problem 1(a). So the whole space has 2×42=62\times4-2=6 real parameters and the product states have 2×32=42\times3-2=4, which is also 2+22+2: two real parameters for each qubit's own Bloch sphere, as it must be.

(d) The joint measurement projects onto the four basis states, and by P3 the probabilities are cij2\abs{c_{ij}}^{2}: Pr(00)=Pr(11)=12\Pr(00)=\Pr(11)=\half and Pr(01)=Pr(10)=0\Pr(01)=\Pr(10)=0. Measuring only the first system means measuring σzI^\sigma_{z}\otimes\hat I, whose eigenvalue +1+1 has the two-dimensional eigenspace spanned by 00,01\ket{00},\ket{01}. By P3 in its projection form, Pr(+1)=P^+Ψ2=12\Pr(+1)=\norm{\hat P_{+}\ket\Psi}^{2}=\half and likewise Pr(1)=12\Pr(-1)=\half.

What is strange is the combination. Each system on its own is completely unpredictable, a fair coin, and yet the two always agree. There is no state of the first system alone that reproduces this. Assigning it 0\ket0 or 1\ket1 contradicts the observed randomness, and assigning it any superposition contradicts (a). The description "each has a definite but unknown value, correlated at preparation" reproduces these particular numbers and is ruled out by measuring other observables, which is the content of Bell's inequality. Chapter 4.19 builds the density operator to answer "what is the first system's state", and Chapter 4.20 measures the inequality.

The brick you just laid — and the two chapters it forces

The table was the chapter. Twenty-two rows, every left-hand entry a theorem of Chapter 0.5 cited at the section that states it, every right-hand entry that theorem with two or three words changed. Real measured values are 0.5 §6.1. Perfectly distinguishable outcomes are §6.2, and the confusion probability is exactly zero rather than small. Any state being a superposition of the outcomes of any observable is §6.3. Probabilities adding to one is Parseval, §2. Compatible measurements and the quantum numbers that label a state are §8. Symmetries and evolution generated by observables are §7.1. None of that was assumed here and none of it was re-proved here, which was the instruction.

Seven assertions, boxed and flagged, and they are the whole of the addition.

  • A state is a unit ray, so that normalisation makes Parseval read as probability and the overall phase drops out of every prediction by one line of conjugation. The relative phase between terms is the opposite case. It produces cos2(β/2)\cos^{2}(\beta/2), and it is the entire wave behaviour of Chapter 4.1 §6.6, now living somewhere definite.
  • An observable is Hermitian, which is the only postulate the measurement outcomes need. Chapter 4.4 §4 will sharpen it to self-adjoint for a physical reason: momentum on a half-line has no self-adjoint extension and is therefore not an observable.
  • The Born rule, which is not derivable, is still open, and was marked as such at the moment it entered rather than afterwards. Gleason's theorem was quoted with its hypotheses and explicitly not presented as a derivation.
  • The state update is a separate postulate, and Worked example 1 shows why the separation is not bookkeeping. Three magnets give 12\tfrac12 when the middle one is read, and exactly 00 when it is not.
  • The generator of time evolution is the energy, which is the only new content in §7. Unitarity followed from probability conservation in three lines, and the existence of a Hermitian generator followed from unitarity in three more.
  • The canonical commutator is postulated here rather than in Chapter 4.9, because Chapters 4.6, 4.8 and 4.11 all need it first.
  • And a composite system's space is the tensor product, so dimensions multiply.

Three lines that settle the shape of the book. The trace of a commutator is zero by cyclicity, from Chapter 0.4 §6.1. The trace of iI^n\ii\hbar\hat I_{n} is in\ii\hbar n. So no finite matrices satisfy [x^,p^]=iI^[\hat x,\hat p]=\ii\hbar\hat I. Cauchy–Schwarz in the matrix inner product sharpens that: [x^,p^]iI^Fn=iI^F\norm{[\hat x,\hat p]-\ii\hbar\hat I}_{F}\ge\hbar\sqrt n=\norm{\ii\hbar\hat I}_{F} at every dimension, with equality exactly for commuting pairs. So a finite model's best strategy is to reproduce none of the relation, and direct numerical minimisation confirms the ratio 1.0000001.000000 at n=2,3,4,5n=2,3,4,5. Quantum mechanics is infinite-dimensional before a single physical question has been asked, and the proof needs nothing but a trace.

And one matrix did all three systems. With H^\hat H having mean Eˉ\bar E, difference δ\delta and coupling Δ\Delta, the splitting is 2δ2+Δ22\sqrt{\delta^{2}+\Delta^{2}} and P12=sin22θsin2(ΔEt/2)P_{1\to2}=\sin^{2}2\theta\,\sin^{2}(\Delta E\,t/2\hbar), with Eˉ\bar E absent because it is a global phase. Read once: ammonia, δ=0\delta=0 by mirror symmetry, splitting h×23.8701 GHz=98.72 μeVh\times23.8701\ \mathrm{GHz} =98.72\ \mu\mathrm{eV}, inversion period 41.89 ps41.89\ \mathrm{ps}, wavelength 1.256 cm1.256\ \mathrm{cm}, stimulated emission favoured over spontaneous by 261261 at room temperature. That is the maser, with every design number on the page. Read again with tL/ct\to L/c and ΔE=Δm2c4/2E\Delta E=\Delta m^{2}c^{4}/2E expanded from Chapter 2.5's dispersion relation: the oscillation formula sin22θsin2(Δm2c4L/4cE)\sin^{2}2\theta\sin^{2}(\Delta m^{2}c^{4}L/4\hbar cE), its coefficient 1.266931.26693 derived rather than quoted, and T2K's 295 km295\ \mathrm{km} at 0.600 GeV0.600\ \mathrm{GeV} sitting at 0.486π0.486\pi against an oscillation length of 607.3 km607.3\ \mathrm{km}. Read a third time with both parameters chosen by an engineer: a 20 ns20\ \mathrm{ns} NOT gate and a 10 ns10\ \mathrm{ns} square root of it, which Worked example 3 shows no classical two-state process possesses. The figure integrates the equation numerically and holds c12+c22=1\abs{c_{1}}^{2}+\abs{c_{2}}^{2}=1 to about 101410^{-14} through the whole cycle at every setting, moves not at all under a full turn of the global phase, and shows every detuned curve leaving the origin along the same parabola τ2\tau^{2}, which is the fact Chapter 4.17's transition rate is built on.

Where this gets spent. Ten flags. Seven are the postulates. One is Gleason with its hypotheses. One is the fundamental theorem of algebra inherited from Chapter 0.4 §7, which is the single thing this chapter quotes rather than derives or postulates, and the thing the whole measurement apparatus descends from, paid off in Chapter 5.4. And one is the measured parameters of §10. That is the largest count in the book and it is correct, because this is the only chapter whose subject is the postulates.

What is owed is now precise rather than vague. Chapter 0.5 named the four places its proofs used finite dimension, calling them "in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective", and §8 showed that infinite dimension is forced. The bill is paid in three instalments, and the shape repeats one you have already seen: Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it.

Chapter 4.3 discards the Riemann integral, builds the Lebesgue integral, proves L2L^{2} complete so that a limit of states is a state, and shows the Fourier modes really are an orthonormal basis. Everything here that said "expand in a basis" is on credit until then. Chapter 4.4 supplies domains and the difference between symmetric and self-adjoint that P2's box warned about. Chapter 4.5 supplies spectra with no eigenvectors in the space, the spectral theorem in the form that survives, Stone's theorem for §7.3, and the meaning of x\ket x and p\ket p. Everything here that said "the spectral theorem" is on credit until then.

After that, the list gets spent one chapter at a time.

  • Chapter 4.6 turns (4.2.18) into a differential equation and solves it.
  • Chapter 4.8 diagonalises the oscillator, which is §4.2's "solving a system means diagonalising its Hamiltonian".
  • Chapter 4.9 applies Cauchy–Schwarz to (A^A^)ψ(\hat A-\avg{\hat A})\ket\psi and (B^B^)ψ(\hat B-\avg{\hat B})\ket\psi and gets the uncertainty principle with nothing added, and Chapter 4.10 proves that P6 cannot be extended to every observable.
  • Chapter 4.11 spends §4.3's quantum numbers on angular momentum.
  • Chapter 4.20 returns to P3, P4 and P7 together, which is where the three of them turn out to have been the interesting ones all along.