Photon polarization experiments
To understand how quantum mechanics works, we look at the outcome of a number of Gedanken experiments involving polarized light beams. Typically, a monochromatic plane wave solution or Maxwell equations (see Problem 9.23) has electric field
where and are the unit vectors in the - and -directions and and are complex numbers such that
When is real we have a linearly polarized wave, as for example
If the wave is circularly polarized; the sign is said to be right cir cularly polarized, and the sign is left circularly polarized. In all other cases it is said to be elliptically polarized. If we pass a polarized beam through a polarizer with axis of polarization , then the beam is reduced in intensity by the factor and the emergen beam is -polarized. Thus, if the resultant beam is passed through another -polarizer it will be 100% transmitted, while if it is passed through an it will be totally absorbed and nothing will come through. This is the classical situation.
As was discovered by Planck and Einstein at the turn of the twentieth century, light beams come in discrete packets called photons, having energy where is Planck’s constant and . What happens if we send the beams through the polarizers one photon at a time? Since the frequency of each photon is unchanged it emerges with the same energy, and since the intensity of the beam is related to the energy, it must mean that the number of photons is reduced. However, the most obvious conclusion that the beam consists of a mixture of photons consisting of a fraction polarized in the -direction and in the -direction will not stand up to scrutiny. For, if a beam with were passed through a polarizer designed to only transmit waves linearly polarized in the direction, then it should be 100% transmitted. However, on the mixture hypothesis only half the -polarized photons should get through, and half the photons, leaving a total fraction being transmitted.
In quantum mechanics it is proposed that each photon is a ‘complex superposition of the two polarization states and . The probability of transmission by an -polarizer is given by , while the probability of transmission by an -polarizer is The effect of the -polarizer is essentially to ‘collapse’ the photon into an -polarized state. The polarizer can be regarded both as a measuring device or equally as a device fo preparing photons in an state. If used as a measuring device it returns the value 1 ifthe photon is transmitted, or 0 if not – in either case the act of measurement has changed the state of the photon being measured.
An interesting arrangement to illustrate the second point of view is shown in Fig. 14.1. Consider a beam of photons incident on an followed by an -polarizer; the net result is that no photons come out of the second polarizer. Now introduce a polarizer for the direction – in other words a device that should block some of the photons between the two initial polarizers. If the mixture theory were correct, it is inconceivable that this could increase transmission. Yet the reality is that half the photons emerge from this intermediary polarizer with polarization ), and a further half of these, namely a quarter in all, are now transmitted by the
How can we find the transmission probability of a polarized state with respect to an arbitrary polarization direction ? The following argument is designed to be motivational rather than rigorous. Let , where and are complex numbers subject to . We write , called the amplitude for -transmission of an

Figure 14.1 Photon polarization experiment
-polarized photon. It is a complex number having no obvious physical interpretation of itself, but its magnitude square is the probability of transmission by an -polarizer. Similarly is the amplitude for -transmission, and the probability of transmission by an . What is the polarization such that a polarizer of this type allows for no transmission, For linearly polarized waves with and both real we expect it to be geometrically orthogonal, . For circularly polarized waves, the orthogonal ‘direction’ is the opposite circular sense. Hence
since phase factors such as are irrelevant. In the general elliptical case we might guess tha , since it reduces to the correct answer for linear and circular polarization. Solving for and we have
Let be any other polarization, then substituting for and gives
Setting gives the normalization condition . Hence, since (transmission probability of 1),
Other systems such as the Stern–Gerlach experiment, in which an electron of magnetic moment is always deflected in a magnetic field in just two directions, exhibit a completely analogous formalism. The conclusion is that the quantum mechanical states of a system form a complex vector space with inner product satisfying the usua
conditions
The probability of obtaining a value corresponding to in a measurement is
As will be seen, states are in fact normalized to , so that only linear combinations with are permitted.
The Hilbert space of states
We will now assume that every physical system corresponds to a separable Hilbert space , representing all possible states of the system. The Hilbert space may be finite dimensional, as for example the states of polarization of a photon or electron, but often it is infinite dimensional. A state of the system is represented by a non-zero vector , but this correspondence is not one-to-one, as any two vectors and that are proportional through a non-zero complex factor, where , will be assumed to represent identical states. In other words, a state is an equivalence class or ray of vectors all related by proportionality. A state may be represented by any vector from the class, and it is standard to select a representative having unit norm . Even this restriction does not uniquely define a vector to represent the state, as any other vector with will also satisfy the unit norm condition. The angular freedom, , is sometimes referred to as the phase of the state vector. Phase is only significant in a relative sense; for example, is in general a different state to , but is not.
In this chapter we will adopt Dirac’s bra-ket notation which, though slightly quirky, has largely become the convention of choice among physicists. Vectors are written as kets and one makes the identification . By the Riesz representation theorem 13.10, to each linear functional there corresponds a unique vector such that . In Dirac’s terminology the linear functional is referred to as a bra, written . The relation between bras and kets is antilinear,
In Dirac’s notation it is common to think of a linear operator as acting to the lef on kets (vectors), while acting to the right on bras (linear functionals):
and if
The following notational usages for the matrix elements of an operator between two vectors are all equivalent:
is an o.n. basis of kets in a separable Hilbert space then we may write
Observables
In classical mechanics,physical observables refer to quantities such as position, momentum, energy or angular momentum, which are real numbers or real multicomponented objects. In quantum mechanics observables are represented by self-adjoint operators on the Hilbert space of states. We first consider the case where is a hermitian operator (bounded and continuous). Such an observable is said to be complete if the corresponding hermitian operator is complete, so that there is an orthonormal basis made up of eigenvectors . such that
The result of measuring a complete observable is always one of the eigenvalues , and the fact that these are real numbers provides a connection with classical physics. By Theorem 13.2 every state can be written uniquely in the form
or, since the vector is arbitrary, we can write
Show that the operator can be written in the form
The matrix element of the identity operator between two states and is
Its physical interpretation is that is the probability of realizing a state when the system is in the state . Since both state vectors are unit vectors, the Cauchy–Schwarz inequality ensures that the probability is less than one,
If is a complete hermitian operator with eigenstates satisfying Eq. (14.1) then, according to this assumption, the probability that the eigenstate is realized when the system is in the state is given by where the are the coefficients in the expansion Eq. (14.2). Thus is the probability that the value be obtained on measuring the observable when the system is in the state . By Parseval’s identity (13.7) we have
and the expectation value of the observable in a given state is given by
The act of measuring the observable ‘collapses’ the system into one of the eigenstates , with probability . This feature of quantum mechanics, that the result ofa measurement can only be known to within a probability, and that the system is no longer in the same state after a measurement as before, is one of the key differences between quantum and classical physics, where a measurement is always made delicately enough so as to minimally disturb the system. Quantum mechanics asserts that this is impossible, even in principle.
The root mean square deviation of an observable in a state is defined by
The quantity under the square root is positive, for
since is hermitian and is real. A useful formula for the RMS deviation is
Hence is an eigenstate of if and only if it is dispersion-free, . For, by Eq. (14.5), if then and immediately results in , and conversely then , which is only possible if . Dispersion-free states are sometimes referred to as pure states with respect to the observable .
Theorem 14.1 · Heisenberg uncertainty theorem
(Heisenberg) Let and be two hermitian operators, thenfor any state
where is the commutator of the two operators
Proof
Let
so that and . Using the Cauchy–Schwarz inequality,
Now
Hence
Show that for any two hermitian operators and , the operator is hermitian.
Show that for any state that is an eigenvector of either or .
A particularly interesting case of Theorem 14.1 occurs when and satisfy the canonical commutation relations,
where is Planck’s constant divided by 2. Such a pair of observables are said to be complementary. With some restriction on admissible domains they hold for the position operator and the momentum operator discussed in Examples 13.19 and 13.20. For, let be a function in the intersection of their domains, then
whence
Theorem 14.1 results in the classic Heisenberg uncertainty relation
Sometimes it is claimed that this relation has no effect at a macroscopic level because Planck’s constant is so ‘small ). Little could be further from the truth. The fact that we are supported by a solid Earth, and not collapse in towards its centre, can be traced to this and similar relations
Show that Eq. (14.7) cannot possibly hold in a finite dimensional space. [Hint: Take the trace of both sides.]
Theorem 14.2 · Compatible observables and a common eigenbasis
A pair of complete hermitian observables and commute, if and only ifthere exists a complete set of common eigenvectors. Such observables are said to be compatible.
Proof
If there exists a basis of common eigenvectors . such that
then for each . Hence for arbitrary vectors we have from Eq. (14.2)
Conversely, suppose that and commute. Let be an eigenvalue of with eigenspace , and set to be the projection operator into this subspace. If then , since
For any we therefore have . Hence
and since is an arbitrary vector,
Taking the adjoint of this equation, and using , gives
and it follows that , the operator commutes with all projection operators If is any eigenvalue of with projection map , then since is a hermitian operator that commutes with the above argument shows that it commutes with
Hence, the operator is hermitian and idempotent, and using Theorem 13.14 it is a projection operator. The space it projects into is . Two such spaces and are clearly orthogonal unless and . Choose an orthonorma basis for each . The collection of these vectors is a complete o.n. set consisting entirely of common eigenvectors to and . For, if is any non-zero vector orthogonal to all , then for all , . Since is complete this implies for all eigenvalues of , and since is complete we must have -
Consider spin electrons in a Stern–Gerlach device for measuring spin in the -direction. Let be the operator for the observable ‘spin in the -direction’. It can only take on two values – up or down. This results in two eigenvalues , and the eigenvectors are written
Thus
and setting results in the matrix components
Every state of the system can be written
The operator representing spin in an arbitrary direction
has expectation values in different directions given by the classical values
where refers to the pure states in the direction, , etc.
Since is hermitian with eigenvalues its matrix with respect to any o.n. basis has the form
where
Hence and . The expectation value of in the state is given by
so that where is a real number. For and we have
The states and are the eigenstates of and with normalized component
and as the expection values of in the orthogonal states vanish
Hence . Applying the unitary operator
results in , and the spin operators are given by the Pauli representation
For a spin operator in an arbitrary direction the expectation values are , etc., from which it is straightforward to verify that
Find the eigenstates and of .
Unbounded operators in quantum mechanics
An important part of the framework of quantum mechanics is the correspondence princi , which asserts that to every classical dynamical variable there corresponds a quantum mechanical observable. This is at best a sort of guide – for example, as there is no natural way of defining general functions for a pair of non-commuting operators such as and , it is not clear what operators correspond to generalized position and momentum in classical canonical coordinates. For rectangular cartesian coordinates , , and mo menta , etc. experience has taught that the Hilbert space of states corresponding to a one-dimensional dynamical system is , and the position and momentum operators are given by
These operators are unbounded operators and have been discussed in Examples 13.17, 13.19 and 13.20 of Chapter 13.
As these operators are not defined on all of it is most common to take domains
These domains are dense in since the basis of functions constructed from hermite polynomials in Example 13.7 (see Eq. (13.6)) belong to both of them. As shown in Example 13.19 the operator is self-adjoint, but is a symmetric operato that is not self-adjoint (see Example 13.20). To make it self-adjoint it must be extended to the domain of absolutely continuous functions.
The position operator has no eigenvalues and eigenfunctions in (see Example 13.14). For the momentum operator the eigenvalue equation reads
and even when is a real number the function does not belong to
For each real number set , and
is a linear functional on in fact, it is the Fourier transform ofSection 12.3. This linear functional can be thought of as a tempered distribution on the space of test functions of rapid decrease . It is a bra that corresponds to no ket vector (this does not violate the Riesz representation theorem 13.10 since the domain is not a closed subspace . In quantum theory it is common to write equations that may be interpreted as
which hold in the distributional sense,
Its integral version holds if we permit integration by parts, as for distributions,
Similarly, for each real define the linear functiona by
for all kets . These too can be thought of as distributions on a set of tes functions of rapid decrease. They behave as ‘eigenbras’ of the position operator
since
for all . While there is no function in having , the Dirac delta function can be thought of as fulfilling this role in a distributional sense (see Chapter 12).
For a self-adjoint operator we may apply Theorem 13.25. Let be the spectral family of increasing projection operators defined by , such that
The latter relation follows from
for all
Prove Eq. (14.12).
If is any measurable subset of then the probability of the measured value of lying in , when the system is in a state , is given by
The expectation value and RMS deviation are given by
and
The spectral family for the position operator is defined as the ‘cut-off operators
Firstly, these operators are projection operators since they are idempotent and hermitian:
for all . They are an increasing family since the image spaces are clearly increasing, and . The function is absolutely continuous, since
and has generalized derivative with respect to given by
Hence
which is equivalent to the required spectral decomposition
Show that for any for the spectral family of the previous example
Problems
Verify for each direction
the spin operator
has eigenvalues . Show that up to phase, the eigenvectors can be expressed as
and compute the expectation values for spin in the direction of the various axes
For a beam of particles in a pure state show that after a measurement of spin in the direction the probability that the spin is in this direction is
If and are vector observables that commute with the Pauli spin matrices, in general) show that
where
Prove the following commutator identities:
Using the identities of Problem 14.3 show the following identities:
where are the angular momentum operators.
Consider a one-dimensional wave packet
where
Show that is a Gaussian normal distribution whose peak moves with velocity and whose spread increases with time, always satisfying
If an electron is initially within an atomic radius , after how long will be , (b) the size of the solar system (about