FOUNDATIONS OF STATISTICAL MECHANICS, Part II
D. Ter Haar
Submitted 1956 | SovietRxiv: ru-195601.48954 | Translated from Russian

Full Text

FOUNDATIONS OF STATISTICAL MECHANICS, Part II

D. ter Haar

IV. The H-Theorem in Quantum Statistics*)

IV.1. The H-Theorem in an Elementary Treatment

It is known that the transition from classical mechanics to quantum mechanics had relatively little effect on statistical mechanics. The very same methods that had been successfully used in classical thermostatics proved capable of being transferred to the case of quantum thermostatics. Statistical methods, in comparison with others, turned out to be the most suitable for use under the new conditions. It is also known that quantum mechanics arose on the basis of statistical considerations—what is meant is Planck’s development of the theory of black-body radiation00P1, 00P2; see also43P, where Planck himself gives arguments in favor of introducing the quantum of action. It is therefore not surprising that questions connected with the application of quantum theory in statistical considerations were the subject of a number of articles even before the appearance of the old quantum theory of Bohr13B1, 13B2, not to mention modern quantum theory—the so-called wave or matrix mechanics25H, 25B, 26S1 ). Initially, the question of a priori weights was studied chiefly (see Section IV.3); however, after the appearance of wave mechanics many authors (for example,27N1, 27N2, 29D, 32N3, 35D, 36D, 40H) demonstrated how Gibbs ensembles can be used successfully in quantum-mechanical considerations. In many respects, thermostatics

) Conclusion. For the beginning see UFN 59, issue 4 (1956).
*) Reviews of the development of both the old and the modern quantum theory may be found in the literature (for example,19S, 33H, 38K).

more readily fits within the framework of quantum mechanics than of classical mechanics; this is connected with the fact that quantum-mechanical predictions have a statistical character. For example, even if we have obtained the maximum amount of information about the physical system under consideration, i.e. are dealing with what Neumann called a pure case, even then one can verify only the predictions made on the basis of the corresponding wave function by using an ensemble*).

In quantum statistics, just as in classical statistics, there are various ways of approaching the problem that we consider in the present work. We are again faced with questions (A) and (B), posed in the Introduction, and again we may try to consider first of all isolated systems either by using the quantum-mechanical analogue of the \(H\)-theorem, or by using the quantum-mechanical ergodic theorem. As an alternative, one may consider ensembles and the question of representative ensembles. In Part IV, at the beginning [section IV.1], the case of isolated systems is discussed by means of the so-called elementary method of consideration (see ESM, Ch. IV), and first of all it is shown how the equilibrium distribution is studied in this case, and then the \(H\)-theorem is considered. After this [section IV.2] the \(H\)-theorem in the theory of ensembles will be discussed. In conclusion we shall turn to the question of representative ensembles and the choice of a priori weights [section IV.3]. In the following Part V the quantum-mechanical ergodic theorem is considered.

Before presenting the elementary method, it is useful to become briefly acquainted with the work of Ornstein and Kramers \(^{27\mathrm{O}}\), devoted to the study of equilibrium in a system of Fermi–Dirac particles [cf. also the papers of Nordheim \(^{28\mathrm{N}}\) and Halpern and Derman \(^{39\mathrm{H}}\)]. In this work they showed how the Fermi distribution can be derived from purely kinetic considerations without invoking any (more precisely, almost any) other arguments. They reasoned as follows. Consider a system of independent particles, and let each particle be capable of occupying energy levels \(\varepsilon_k\), for simplicity nondegenerate. Let us now consider a “collision” between two particles, and let the energies before the collision be \(\varepsilon_i\) and \(\varepsilon_j\), and after the collision \(\varepsilon_{i'}\) and \(\varepsilon_{j'}\). The law of conservation of energy leads to the relation

\[ \varepsilon_i+\varepsilon_j=\varepsilon_{i'}+\varepsilon_{j'} . \tag{IV.1, 1} \]

Denote the probability of this transition by \(a\), the probability of the reverse transition by \(a'\), the number of transitions per unit time in one direction by \(A\), and the number of transitions in the reverse direction

*) This is, for example, analyzed in detail by Kemble \(^{37\mathrm{K}}\) (especially section 14b); see also section IV.3.

through \(A'\). The quantities \(A\) and \(A'\) satisfy the relations

\[ A = an_i n_j(1-n_{i'})(1-n_{j'}), \tag{IV.1,2} \]

\[ A' = a'n_{i'}n_{j'}(1-n_i)(1-n_j), \tag{IV.1,3} \]

where \(n_i, n_j, n_{i'}\), and \(n_{j'}\) are the mean numbers of particles in the energy states \(\varepsilon_i, \varepsilon_j, \varepsilon_{i'}\), and \(\varepsilon_{j'}\), respectively. Relations (IV.1,2) and (IV.1,3) follow from the fact that the number of transitions is proportional not only to the mean number of particles in the states \(i\) and \(j\), but also to the probability of finding the states \(i'\) and \(j'\) unoccupied, since otherwise the transition cannot take place because of the Pauli principle \(^{25P,\,27P}\).

If the system is in equilibrium, we have

\[ A = A'. \tag{IV.1,4} \]

Suppose now that

\[ a = a'; \tag{IV.1,5} \]

then from (IV.1,2)—(IV.1,5) we obtain

\[ m_i m_j = m_{i'}m_{j'}, \tag{IV.1,6} \]

where

\[ m_i = n_i/(1-n_i). \tag{IV.1,7} \]

Relation (IV.1,6) in equilibrium must be satisfied for any pair of states \(\varepsilon_i, \varepsilon_j\). Just as the Maxwell distribution follows from (I.1, 12) and the condition of conservation of energy, so in our case from (IV.1,6), under condition (IV.1,7), we obtain that in equilibrium

\[ m_i = \exp(\mu-\beta\varepsilon_i); \tag{IV.1,8} \]

or, using (IV.1,7),

\[ n_i = 1/[\exp(-\mu+\beta\varepsilon_i)+1], \tag{IV.1,9} \]

which is precisely the Fermi distribution \(^{26F1}\).

Here it is necessary to note (1) that relation (IV.1,2) is based on purely kinetic considerations and does not make use of any quantum-mechanical arguments, and (2) that relation (IV.1,5) is introduced as an assumption replacing the hypothesis on the number of collisions considered in Section I.1.

It may be mentioned that Ornstein and Kramers discussed the question of whether relation (IV.1,5) can be derived from fundamental principles and indicated that from the considerations of Heisenberg \(^{26H}\) and Jordan \(^{26J}\) it follows that, instead of (IV.1,5), proceeding from quantum mechanics, one can obtain only the much weaker relation

\[ \sum_{ij}\left(a_{ij;\,i'j'}-a_{i'j';\,ij}\right)=0. \tag{IV.1,10} \]

In (IV.1,10), \(a_{ij;\,i'j'}\) and \(a_{i'j';\,ij}\) denote the probabilities of transitions, to

hitherto denoted by \(a\) and \(a'\); the summation is carried out over all pairs of values \(\varepsilon_i,\varepsilon_j,\ldots\) for which \(\varepsilon_i+\varepsilon_j\) is the same. The question of the extent to which relation (IV.1.5) follows from (IV.1.10) is connected with questions considered in section (IV.3) and Appendix V, to which we shall refer the reader*).

We have not shown that a distribution different from the distribution specified by relation (IV.1.9) will tend to the latter. This is easy to see by using the expression for the entropy of a system of independent Fermi–Dirac particles**) [see, for example, ESM, p. 406, expression (I.3.6), and also (IV.1.28)]; we have

\[ H=-S/k=-\sum_i\bigl[(1-n_i)\ln(1-n_i)+n_i\ln n_i\bigr]. \tag{IV.1.11} \]

The rate of change of \(H\) is given by the expression

\[ dH/dt=\sum_i \ln [n_i/(1-n_i)]\,dn_i/dt = \]

\[ =\sum_{ij;i'j'} \ln [n_i/(1-n_i)] \{a_{i'j';ij}n_{i'}n_{j'}(1-n_i)(1-n_j)- \]

\[ -a_{ij;i'j'}n_i n_j(1-n_{i'})(1-n_{j'})\}. \tag{IV.1.12} \]

In (IV.1.12) the summation is performed first over all energy levels \(\varepsilon_i\), then over all energy levels \(\varepsilon_j\) with which collisions can occur, and further over all pairs of energy levels \(\varepsilon_{i'}\) and \(\varepsilon_{j'}\) which can be obtained as a result of collisions between particles in the states \(\varepsilon_i\) and \(\varepsilon_j\). In order to obtain a symmetric expression in which the summation would be carried out both over the pairs \(\varepsilon_i\) and \(\varepsilon_j\), and over the pairs \(\varepsilon_{i'}\) and \(\varepsilon_{j'}\), we write

\[ dH/dt=\sum_{ij,i'j'} a\,[n_{i'}n_{j'}(1-n_i)(1-n_j)-n_i n_j(1-n_{i'})(1-n_{j'})]\times \]

\[ \times \{\ln [n_i n_j/(1-n_i)(1-n_j)]\}, \tag{IV.1.13} \]

where relation (IV.1.5) has been used. If we now take the half-sum of the expression on the right-hand side of (IV.1.13) and of the expression obtained from it by interchanging \(ij\) and \(i'j'\) [cf. the discussion of expression (I.1.9)], we obtain

\[ dH/dt = \]

\[ =\frac12\sum_{ij,i'j'} a\{\ln [n_i n_j(1-n_{i'})(1-n_{j'})/(1-n_i)(1-n_j)n_{i'}n_{j'}]\}\times \]

\[ \times [n_{i'}n_{j'}(1-n_i)(1-n_j)-n_i n_j(1-n_{i'})(1-n_{j'})]\leqslant 0; \tag{IV.1.14} \]

*) It is interesting to note that relation (IV.1.10) is sufficient for deriving the Maxwell–Boltzmann distribution if, of course, in (IV.1.2) and (IV.1.3) the factors \((1-n_{j'})\), etc., are absent; cf. in this connection also \(^{54\mathrm{K}}\) section 14, \(^{53\mathrm{T}}\) and \(^{54\mathrm{L}}\).

) Such particles are sometimes called fermions, and Bose–Einstein particles bosons**.

here the equality sign occurs only if \(n_i\) satisfy (IV.1,9). From what has been said, it thus becomes clear once again that the approximation to the equilibrium distribution can be proved by using the assumption concerning the number of transitions. We shall discuss the derivation of expression (IV.114) after the elementary method of treatment has been presented.

In the elementary method the energy levels are taken in groups, the levels of one group corresponding to approximately the same energy. Let the \(i\)-th group contain \(Z_i\) levels and let there be \(N_i\) particles in the system with energies lying within this group; let, further, \(E_i\) be the approximate value of the energy for the levels of the group. First of all we wish to find the equilibrium distribution \(N_i\); we shall define it once more (see Section I.3) as the most probable distribution. We must then find the a priori probabilities \(W(N_i)\) for the given distribution, in accordance with \(W(Z)\) in expression (I.3,1). At this point it is necessary to distinguish between the statistics of Boltzmann (Bo), Fermi—Dirac (F.—D.), and Bose—Einstein (B.—E.). It is necessary first of all to remember that quantum mechanics involves taking into account two circumstances that have no place in classical mechanics, namely diffraction effects and symmetry effects\(^*\). The former already appear in the old quantum mechanics and are responsible for the existence of the quantization of energy, angular momentum, etc. The latter are connected with the necessity of taking into account that in nature there evidently exist only such systems of identical particles for which the wave functions are either completely symmetric in all particles (B.—E.), or completely antisymmetric (F.—D.)\(^ {**}\).

At first glance it may seem that, consequently, only the B.—E. or F.—D. statistics are admissible. However, there are cases, such as, for example, a system of particles of a crystal lattice, when identical particles may be distinguishable; say, in the case of a crystal, by the position occupied in the lattice [cf. \({}^{49\mathrm{R}}\), Ch. III, Sec. 1]. Then we are dealing with Boltzmann statistics—with one small difference, which will be indicated below. B.—E. statistics was introduced by Bose \({}^{24\mathrm{B}}\)\(^ {***}\) for light quanta; he used it to derive Planck’s radiation law. Einstein \({}^{24\mathrm{E2},\,25\mathrm{E1},\,25\mathrm{E2}}\) applied B.—E. statistics to an ideal gas and showed in what way

\(^*\) A discussion of the importance of this distinction is given in \({}^{54\mathrm{H3}}\).

\(^ {**}\) For an analysis of the reasons why some particles obey B.—E. statistics and others F.—D. statistics, we refer to the literature (for example, \({}^{39\mathrm{B2},\,40\mathrm{P1},\,40\mathrm{P2},\,49\mathrm{W2}}\)).

\(^ {***}\) This work was translated by Einstein, who added the following note: “Bose’s derivation of Planck’s formula is, in my opinion, an important step forward. The method used here can also be applied to the quantum theory of ideal gases, as I shall show elsewhere.”

there results the so-called Einstein condensation—a phenomenon that may be related to \(\lambda\)-transitions in liquid helium (cf. Ch. IX ESM). Fermi \(^{26F1}\) introduced the Pauli exclusion principle \(^{25P, 27P}\) into statistical considerations, and Dirac \(^{26D}\) analyzed the connection between statistics and wave mechanics. Ehrenfest and Uhlenbeck (\(^{27E}\), see also \(^{27U}\)) showed how Boltzmann statistics can be included in quantum statistics.

Let us now proceed to the calculation of \(W(N_i)\). Let \(n_j\) be the number of particles on the energy level \(\varepsilon_j\), and let \(W(n_j)\) be the probability for the distribution \(n_j\). We then have

\[ W(N_i)=\sum W(n_j); \tag{IV.1,15} \]

here the summation extends over all \(n_j\)-distributions, and for each group

\[ \sum n_j=N_i, \tag{IV.1,16} \]

where the \(\varepsilon_j\) over which the summation is carried out belong to the group \(Z_i\). Our problem has now been reduced to finding \(W(n_j)\), i.e., the probability that \(N\) particles are distributed among various energy levels, assumed to be nondegenerate, with \(n_j\) particles on the level \(\varepsilon_j\).

In the case of B.–E. statistics, for any given distribution there is only one wave function and, consequently,

\[ W_{\mathrm{B.-E.}}(n_j)=1. \tag{IV.1,17} \]

In the case of F.–D. statistics, only completely antisymmetric functions are possible. This means that each level can either be occupied by one particle or be free, but cannot be occupied by more than one particle. In this case, therefore, there is either one possible wave function or none:

\[ \left. \begin{aligned} W_{\mathrm{F.-D.}}(n_j)&=1,\quad \text{if all } n_j=0 \text{ or } 1,\\ W_{\mathrm{F.-D.}}(n_j)&=0,\quad \text{if at least one of the } n_j>1. \end{aligned} \right\} \tag{IV.1,18} \]

Finally, in the case of Boltzmann statistics, any one of the \(N!\) permutations of the variables of the \(N\) particles in the wave function again leads to an admissible wave function. Since a permutation of \(n_j\) particles on one and the same level does not change the wave function, one may put

\[ W'_{\mathrm{Bo}}=N!\Pi_j(1/n_j!). \tag{IV.1,19} \]

Here no account is taken of the fact that, although it is sometimes possible to renumber the particles, we are not able to trace the path of each individual atom, and thus the weight of each \(n_j\)-distribution is exaggerated by \(N!\) times, which is the number of different microstates corresponding to one and the same

macrostate*). Consequently, in the case of Boltzmann statistics it is necessary to use, instead of (IV.1,19), the expression

\[ W_{\mathrm{Bo}}(n_j)=\prod_j (1/n_j!). \tag{IV.1,20} \]

We can now use relations (IV.1,15)—(IV.1,18) and (IV.1,20) to compute \(W(N_i)\). In the case of Boltzmann statistics we easily obtain (cf. ESM, p. 74) that

\[ W_{\mathrm{Bo}}(N_i)=\prod_i \frac{1}{N_i!}\sum \frac{N_i!}{\prod n_j!} =\prod_i (Z_i^{N_i}/N_i!). \tag{IV.1,21} \]

In the case of B.—E. statistics we must find the number of different ways in which \(N_i\) particles can be distributed over \(Z_i\) levels; we obtain

\[ W_{\mathrm{B.-E.}}(N_i)=\prod_i [(N_i+Z_i-1)!/N_i!(Z_i-1)!]. \tag{IV.1,22} \]

Finally, in the case of F.—D. statistics, one must find the number of different ways in which \(N_i\) particles can be distributed over \(Z_i\) levels, with no more than one particle on each of the levels; we obtain

\[ W_{\mathrm{F.-D.}}(N_i)=\prod_i [Z_i!/N_i!(Z_i-N_i)!]. \tag{IV.1,23} \]

Suppose now that all \(Z_i\), \(N_i\), and \(Z_i-N_i\) (in the F.—D. case) are large numbers, so that expression (I.3, 3) may be used for the factorials. We obtain

\[ \ln W(N_i)=\sum_i \{(N_i+\alpha Z_i)\ln [(Z_i/N_i)+\alpha]-\alpha Z_i\ln (Z_i/N_i)\}, \tag{IV.1,24} \]

where

\[ \alpha_{\mathrm{Bo}}=0,\qquad \alpha_{\mathrm{B.-E.}}=1,\qquad \alpha_{\mathrm{F.-D.}}=-1; \tag{IV.1,25} \]

let us note that in the Boltzmann case the term \(\sum N_i(=N)\), which is inessential for our discussion, has been omitted. From (IV.1,24) it can be seen that in the limit \(N_i/Z_i\to 0\), i.e., for extremely dilute systems or for systems at very high temperatures, these three expressions coincide.

The equilibrium distribution is found by maximizing \(\ln W(N_i)\) under the conditions of constancy of the total energy and the total number of particles

\[ \sum N_i=N,\qquad \sum N_i E_i=E; \tag{IV.1,26} \]

as a result we find [cf. the discussion in Section I. 3]

\[ \ln [(Z_i+\alpha N_i)/N_i]=-\mu+\beta E_i, \tag{IV.1,27} \]

*) The importance of omitting the factor \(N!\) is also mentioned in Gibbs’s monograph \(^{02G}\) (ch. XV) and in Nordheim’s paper \(^{24N}\).

which leads to the well-known expressions for the Boltzmann, Fermi–Dirac, and Bose–Einstein distributions.

We shall now prove the approach to equilibrium*). Using the relation \(H=-S/k\) and the Boltzmann–Planck relation \(S=k\ln W\), we obtain for our \(H\)-function

\[ H=\sum_i \{\alpha Z_i \ln (Z_i/N_i)-(N_i+\alpha Z_i)\ln [(Z_i+\alpha N_i)/N_i]\}. \tag{IV.1,28} \]

Let \(N_{ij;i'j'}\) be the number of such transitions per unit time in which each of the groups \(Z_i\) and \(Z_j\) loses one particle, while each of the groups \(Z_{i'}\) and \(Z_{j'}\) acquires one particle. Such transitions are possible only in the case of conservation of energy, that is, if

\[ E_i+E_j=E_{i'}+E_{j'}. \tag{IV.1,29} \]

In Appendix V it will be shown that \(N_{ij;i'j'}\) satisfies the equation

\[ N_{ij;i'j'}=A_{ij;i'j'}N_iN_j(Z_{i'}+\alpha N_{i'})(Z_{j'}+\alpha N_{j'}), \tag{IV.1,30} \]

and we shall assume [cf. (IV.1,5)] that

\[ A_{ij;i'j'}=A_{i'j';ij}. \tag{IV.1,31} \]

Now, in exactly the same way as expression (IV.1,14) was obtained, we find**)

\[ \begin{aligned} dH/dt&=\sum_i (dN_i/dt)\ln [N_i/(Z_i+\alpha N_i)] = \\ &=\frac12 \sum_{ij;i'j'} A [N_iN_j(Z_{i'}+\alpha N_{i'})(Z_{j'}+\alpha N_{j'}) - \\ &\qquad\qquad - N_{i'}N_{j'}(Z_i+\alpha N_i)(Z_j+\alpha N_j)] \times \\ &\qquad\qquad \times \ln [N_iN_j(Z_{i'}+\alpha N_{i'})(Z_{j'}+\alpha N_{j'})/N_{i'}N_{j'} \times \\ &\qquad\qquad \times (Z_i+\alpha N_i)(Z_j+\alpha N_j)] \leqslant 0. \end{aligned} \tag{IV.1,32} \]

The equality sign occurs only in the case where relation (IV.1,27) is fulfilled.

We have once again obtained an equation from which it is clear that \(H\) decreases monotonically until equilibrium is established. Both (IV.1,32) and (IV.1,14) were derived using the hypothesis concerning the number of collisions and by considering transition probabilities. The discussion here can evidently be conducted in the spirit of Section 1.3; in particular, in the present case one may also consider

*) Here we follow Pauli’s treatment \(^{28P}\); cf. also Nordheim’s treatment \(^{28N}\). Nordheim was especially interested in the distribution of electrons in metals (cf. also \(^{29P}\)).

**) It should be noted that \(Z_i+\alpha N_i\) is never negative, since in the F.–D. case \(N_i\leq Z_i\).

the expected frequency of occurrence of fluctuations. Just as the use of expression (I.3,1) for the probability of realization of a state was based on the choice of equal a priori probabilities for equal volumes of phase space, so also the use of expressions (IV.1,21)—(IV.1,23) is based on the assumption of equality of a priori probabilities for all nondegenerate quantum states, as can be seen from the derivation of these expressions. This assumption will be discussed in section IV.3; there too the question will be considered of the extent to which the relations (IV.1,31) can be justified on the basis of fundamental principles*).

IV.2. The \(H\)-theorem in the quantum theory of ensembles**)

The consideration in the preceding section corresponded to the consideration of section I.3 in the classical case. We shall now be dealing with the quantum-mechanical analogue of part III. First of all, in the present section the fine-structure density will be considered. Then the coarse-structure density will be introduced and the quantity \(\Sigma\), related to \(\Sigma\) of section III.1, will be considered, with attention being paid to certain circumstances in which small differences from the classical case appear. After this, the influence of measurements will be briefly considered***). The question of representing ensembles and of a priori probabilities is deferred to section IV.3.

It is known that the so-called density matrix \(\boldsymbol{\rho}\)****), introduced by Neumann (\(^{27\mathrm{N}1}\); see also \(^{27\mathrm{N}2,\,29\mathrm{D},\,30\mathrm{D}1,\,30\mathrm{D}2,\,31\mathrm{D},\,32\mathrm{N}?,\,33\mathrm{P},\,35\mathrm{D},\,36\mathrm{D},\,37\mathrm{K},\,40\mathrm{H},\,54\mathrm{T}}\) ESM, ch. VII), plays in quantum statistics the role of the ensemble density \(\rho\) of classical statistics. This density matrix also plays an important role in other quantum-mechanical considerations; however, this is not the place to discuss the use of the density matrix for any other, nonstatistical purposes.

The fine-structure density matrix \(\boldsymbol{\rho}\) is introduced as follows. Let \(\psi^{k}\) be the normalized wave function describing the \(k\)-th system of the ensemble, and let \(\varphi_n\) be a complete set of orthonormal functions in the Hilbert space of wave functions*****).

*) Recently Stueckelberg and collaborators \(^{25\mathrm{S}2,\,54\mathrm{I}}\) have justified the decrease of \(H\) by means of unitarity of the so-called \(S\)-matrix.

**) Compare \(^{38\mathrm{T}}\), ch. XII, part A; in our treatment this work is used to a considerable extent.

***) For a detailed consideration of this question we refer to Neumann’s book \(^{32\mathrm{N}3}\) and to the recent work of Groenewold \(^{46\mathrm{G}2}\).

****) All matrices are denoted in boldface.

*****) For simplicity we shall suppose that the \(\varphi_n\) form a discrete sequence, so that in (IV.2,1) there will be a sum. It is not difficult to extend the formulas to the case when the sum must be replaced by an integral (a Stieltjes integral, if necessary); we leave this to the reader.

One can then expand \(\psi^k\) in the following way:

\[ \psi^k=\sum_n a_n^k \varphi_n, \tag{IV.2,1} \]

where

\[ a_n^k=\int \varphi_n^* \psi^k\,d\tau; \tag{IV.2,2} \]

here the asterisk denotes complex conjugation, and \(d\tau\) is the volume element in coordinate space*).

To describe the \(k\)-th system we may use \(a_n^k\) instead of \(\psi^k\). These two representations are equivalent; if \(\psi^k\) satisfied the Schrödinger equation

\[ \mathbf H \psi^k=i\hbar \dot{\psi}^k, \tag{IV.2,3} \]

then \(a_n^k\) satisfy the transformed Schrödinger equation

\[ i\hbar \dot a_n^k=\sum_m H_{nm}a_m^k; \tag{IV.2,4} \]

here \(H_{nm}\) is the matrix element of the Hamiltonian operator \(\mathbf H\)

\[ H_{nm}=\int \varphi_n^*\mathbf H\varphi_m\,d\tau. \tag{IV.2,5} \]

The physical meaning of the quantities \(a_n^k\) is that they are probability amplitudes; \(|a_n^k|^2\) is the probability that the \(k\)-th system is described by the function \(\varphi_n\). From the normalization of \(\psi^k\) and from the fact that \(\varphi_n\) form a complete orthonormal set, it follows that

\[ \sum_n |a_n^k|^2=1. \tag{IV.2,6} \]

We can now introduce the density operator \(\rho\), or the density matrix, by defining it by the corresponding matrix elements

\[ \rho_{mn}=\frac{1}{N}\sum_{k=1}^{N} a_m^k a_n^{k*}, \tag{IV.2,7} \]

where \(N\) is the number of systems in the ensemble.

The mean value \(\langle G\rangle\) of the physical quantity \(G\), which now corresponds to the operator \(\mathbf G\), is given by the expression

\[ \langle G\rangle=N^{-1}\sum_k \int \psi^{k*}\mathbf G\psi^k\,d\tau, \tag{IV.2,8} \]

*) In our formulas we confine ourselves to the case in which the wave functions are scalars. It is not difficult to extend the formulas to the case in which they are spinors or tensors.1

the averaging is carried out twice: first over the wave function and then over the ensemble.

Introducing the quantities \(a_n^k\), we obtain *)

\[ \langle G\rangle = N^{-1}\sum_k\sum_{m,n} a_m^k a_n^{k*}G_{nm} = \sum_{m,n}\rho_{mn}G_{nm} = \operatorname{Tr}(\rho G). \tag{IV.2,9} \]

As an example, from relation (IV.2,9) follows the normalization of \(\rho\). Setting \(G=1\), we have

\[ 1=\operatorname{Tr}\rho; \tag{IV.2,10} \]

the same relation follows directly from (IV.2,6).

The probability exponent \(\eta\) is again connected with \(\rho\) by the relation

\[ \rho=\exp\eta, \tag{IV.2,11} \]

where the right-hand side is a shorthand notation for the infinite exponential series

\[ \exp\eta=\sum (n!)^{-1}\eta^n. \tag{IV.2,12} \]

Here one may point out that the introduction of large ensembles into quantum statistics causes no additional difficulties. We may assume that secondary quantization has been carried out **) and that thus all wave functions are expressed in the form of products of the Jordan–Klein and Jordan–Wigner matrices. The formalism used thereby remains unchanged.

We now introduce the quantity \(\sigma\) by the relation

\[ \sigma=\langle\eta\rangle, \tag{IV.2,13} \]

or, using (IV.2,11) and (IV.2,9),

\[ \sigma=\operatorname{Tr}(\rho\ln\rho). \tag{IV.2,14} \]

As in Section III.1, we can show that \(\sigma\) possesses a number of minimal properties (cf. \({}^{36}\)). For this purpose we first choose, as a complete orthonormal system, the complete orthonormal system \(\varphi_n\) of eigenfunctions of the Hamiltonian operator and the particle-number operator ***). The state corresponding to \(\nu\) particles

*) \(\operatorname{Tr}A\) is an abbreviated notation for the trace, or spur, of the operator \(A\), i.e. the sum of its diagonal elements: \(\operatorname{Tr}A=\sum_n A_{nn}\).

**) The method of secondary quantization (\({}^{28J}\); see also \({}^{27J2,34F}\)) is set forth, for example, in Cramer’s monograph \({}^{38K}\), Section 72.

***) Again we restrict ourselves to the case when particles of one kind are present. The present treatment is easily extended to the case of particles of different kinds.

with the energy of the system \(E_k(\nu)\)*), we shall denote by the double index \(k,\nu\); a matrix element of the transition between two states, say \(k,\nu_1\) and \(l,\nu_2\), will be denoted by \((k,\nu_1;l,\nu_2)\), for example

\[ H(k,\nu_1;l,\nu_2). \]

The quantity \(\sigma\) has the following properties:

a) If the ensemble is such that all systems consist of the same number of particles \(N\) and the energies of all systems lie inside the prescribed interval \(E, E+\delta E\), then \(\sigma\) will be minimal if \(\rho\) is defined by the relations

\[ \left. \begin{aligned} \rho(k,\nu_1;l,\nu_2)&=c\hat{\delta}_{kl}\hat{\delta}_{\nu_1 N}\hat{\delta}_{\nu_2 N} \quad \text{for } E \leq E_k(N) \leq E+\delta E;\\ \rho(k,\nu_1;l,\nu_2)&=0 \quad \text{in all other cases;} \end{aligned} \right\} \tag{IV.2,15} \]

here \(\hat{\delta}_{kl}\) is the Kronecker symbol, and \(c\) is a constant.

b) If the ensemble is such that all systems consist of the same number of particles \(N\), but only the mean energy is specified, i.e. \(\rho\) satisfies the condition

\[ \operatorname{Tr}(\rho H)=E, \tag{IV.2,16} \]

then \(\sigma\) will be minimal if \(\rho\) is defined by the relation

\[ \rho(k,\nu_1;l,\nu_2)=\hat{\delta}_{kl}\hat{\delta}_{\nu_1 N}\hat{\delta}_{\nu_2 N}\exp\{\beta[\psi-E_k(N)]\}. \tag{IV.2,17} \]

c) If only the mean number of particles and the mean energy of the systems of the ensemble are specified, i.e. \(\rho\) satisfies condition (IV.2,16) and the condition

\[ \operatorname{Tr}(\rho \nu)=N, \tag{IV.2,18} \]

where \(\nu\) is the particle-number operator, then \(\sigma\) will be minimal if \(\rho\) is defined by the relation**)

\[ \rho(k,\nu_1;l,\nu_2)=\hat{\delta}_{kl}\hat{\delta}_{\nu_1\nu_2} \exp[-q+\nu_1\mu-\beta E_k(\nu_1)]. \tag{IV.2,19} \]

Let us prove the last of these assertions; the proof of the other two is completely analogous. As in Section III.1, the densities defined by expressions (IV.2,15), (IV.2,17), and (IV.2,19) correspond, respectively, to the ensemble with an energy shell, the macrocanonical ensemble, and the canonical grand ensemble, while the quantities \(\beta\), \(\psi\), and \(\mu\) have the same meaning as before.

*) For simplicity we shall assume that all energy levels are nondegenerate.

**) In this case, in fact, there is no need to make a special choice of \(\varphi_n\). In the general case (IV.2,19) will have the form

\[ \rho=\exp[-q+\mu\nu-\beta H]. \]

Proof of (b) is carried out in the following way (cf. $^{36D}$, Sec. 8):

Consider two density matrices $\rho_1$ and $\rho_2$, where $\rho_1$ is given by expression (IV.2,19), and $\rho_2$ by the relation

\[ \rho_2=\rho_1 \exp(\Delta\eta). \tag{IV.2,20} \]

The quantity $\sigma_2-\sigma_1$ has the form

\[ \sigma_2-\sigma_1=\operatorname{Tr}(\eta_2\exp\eta_2-\eta_1\exp\eta_1). \tag{IV.2,21} \]

The right-hand side of expression (IV.2,21) can be rewritten in the form

\[ \operatorname{Tr}[(\eta_2-\eta_1)\exp\eta_2]+\operatorname{Tr}[\eta_1(\rho_2-\rho_1)]= \]

\[ =\operatorname{Tr}[(\eta_2-\eta_1)\exp\eta_2] +\operatorname{Tr}[-q+\psi-\beta H)(\rho_2-\rho_1)]= \]

\[ \operatorname{Tr}\{(\eta_2-\eta_1)-\exp\eta_2\}; \]

here it has been used that $\rho_1$ and $\rho_2$ satisfy the conditions (IV.2,16), (IV.2,18) and the normalization condition (IV.2,10), (IV.2,11); moreover, the relations (IV.2,19) and (IV.2,11) have been used. We have, consequently,

\[ \sigma_2-\sigma_1=\operatorname{Tr}[(\eta_2-\eta_1)\exp\eta_2]. \tag{IV.2,22} \]

The right-hand side of expression (IV.2,22) is positive, except in the case $\eta_2=\eta_1$. This is seen from the following. First note that for any Hermitian operator $A$ the inequality holds (cf. $^{38P}$)

\[ [\exp A]_{kk}\geq \exp A_{kk}, \tag{IV.2,23} \]

with equality occurring in the case where the operator $A$ is diagonal.

Relation (IV.2,23) is proved by considering a unitary transformation $U$ which brings $A$ to diagonal form, with diagonal elements (eigenvalues) $A_k$. Then we have

\[ A_{kk}=\sum_k |U_{kl}|^2 A_l,\qquad \sum_k |U_{kl}|^2=1 \tag{IV.2,24} \]

and

\[ [f(A)]_{kk}=\sum_k |U_{kl}|^2 f(A_l)= \]

\[ =\sum |U_{kl}|^2\left\{f(A_{kk})+(A_l-A_{kk})f'(A_{kk})+\right. \]

\[ \left.+\frac{1}{2}(A_l-A_{kk})^2 f''\bigl(A_{kk}+\theta_{kl}[A_l-A_{kk}]\bigr)\right\}= \]

\[ =f(A_{kk})+\frac{1}{2}\sum |U_{kl}|^2(A_l-A_{kk})^2 f'', \tag{IV.2,25} \]

\((0 \leq \theta_{kl} \leq 1)\), and the relations (IV.2,24) have been used. Assuming \(f(A)\) to be an exponential function, we obtain (IV.2,23).

Let us now choose a representation in which \(\eta_2\) has diagonal form. Since from (IV.2,10) it follows that

\[ \operatorname{Tr}(\exp \eta_1)=\operatorname{Tr}(\exp \eta_2)\;(=1), \tag{IV.2,26} \]

we obtain

\[ \begin{aligned} \operatorname{Tr}\bigl[(\eta_2-\eta_1)\exp \eta_2\bigr] &=\operatorname{Tr}\bigl[(\eta_2-\eta_1)\exp \eta_2-\exp \eta_2+\exp \eta_1\bigr] \\ &=\sum \bigl[(\eta_{2i}-\eta_{1ii})\exp \eta_{2i}-\exp \eta_{2i}+(\exp \eta_1)_{ii}\bigr] \\ &\geq \sum \bigl[(\eta_{2i}-\eta_{1ii})\exp \eta_{2i}-\exp \eta_{2i}+\exp \eta_{1ii}\bigr] \\ &=\sum (\exp \eta_{1ii})\bigl[(\eta_{2i}-\eta_{1ii}-1)\exp(\eta_{2i}-\eta_{1ii})+1\bigr]\geq 0, \end{aligned} \]

where (IV.2,23) has been used, and then the properties of the function \(y\), defined according to (III.1,13), have been used.

Again \(\sigma\) will not change with time. This follows from the fact that from one instant \((t')\) we can pass to another instant \((t'')\) by means of a unitary transformation. This means that if \(\rho'\) and \(\rho''\) are the density matrices at the instants \(t'\) and \(t''\), then they are related by

\[ \rho''=U^{+}\rho' U, \tag{IV.2,27} \]

where \(U^{+}\) is the Hermitian conjugate of the matrix \(U\), and

\[ U^{+}U=UU^{+}=1. \tag{IV.2,28} \]

Any mean value \(\langle G\rangle\) is invariant with respect to a unitary transformation. This follows from the fact that

\[ \operatorname{Tr}(AB)=\operatorname{Tr}(BA) \]

and, consequently,

\[ \operatorname{Tr}(\rho''G'')=\operatorname{Tr}(U^{+}\rho'UU^{+}G'U) =\operatorname{Tr}(U^{+}\rho'G'U) =\operatorname{Tr}(G'UU^{+}\rho') =\operatorname{Tr}(G'\rho') =\operatorname{Tr}(\rho'G'). \tag{IV.2,29} \]

Since \(\sigma\) is the mean value of \(\eta\), then, as in the classical case, we have \(\sigma'=\sigma''\).

Let us now introduce, as in Section III.1, a coarse-structure density, or in the present case a coarse-structure density matrix \(P\). We choose as a complete orthonormal system of functions the eigenfunctions of the Hamiltonian of the system \(H\). We then subdivide the stationary states into groups, as before; however, now we shall assume the grouping to be carried out in accordance with the inaccuracy of our observations, i.e. we shall assume that, by means of the observational methods available to us, we can

establish differences between different groups, but not within them*). Let \(S_i\) be the number of levels in the \(i\)-th group. We define the coarse-structure density matrix \(\mathbf P\) by its matrix elements in the particular representation which we have just considered; we have

\[ P_{kl}=\delta_{kl}\sum_j \rho_{jj}/S_i, \tag{IV.2,30} \]

where the energy level \(E_k\) belongs to the \(i\)-th group and the summation is over all states of the \(i\)-th group. Since \(\rho_{jj}\) is the probability of finding the system of the ensemble in the state characterized by the function \(\varphi_j\), it is clear that \(S_i P_{kk}\) is the probability of finding the system of the ensemble in a state belonging to the \(i\)-th group. Expression (IV.2,30) defines the coarse-structure density in the particular representation chosen by us; the corresponding matrix elements in any other representation are obtained with the aid of the usual transformation rules**).

From (IV.2,30) and (IV.2,10) it follows that \(\mathbf P\) is also normalized,

\[ \operatorname{Tr}\mathbf P=1. \tag{IV.2,31} \]

We now define the quantity \(\Sigma\) by the expression

\[ \Sigma=\operatorname{Tr}(\mathbf P\ln \mathbf P)=\sum_k P_{kk}\ln P_{kk}, \tag{IV.2,32} \]

if we adhere to the representation chosen by us. Using (IV.2,30), we may write

\[ \Sigma=\sum \rho_{kk}\ln P_{kk}=\operatorname{Tr}(\boldsymbol{\rho}\ln \mathbf P)=\langle \ln \mathbf P\rangle . \tag{IV.2,33} \]

There are several different ways of studying the behavior of \(\Sigma\) in time. Probably the crudest is the method used by Born and Green \(^{48\mathrm{b}1,48\mathrm{b}2,49\mathrm{b}1,49\mathrm{b}2,49\mathrm{b}3}\), who, in our opinion, did not make the necessary distinction between fine-structure and coarse-structure densities and, moreover, did not take into account the fact that coarse-structure densities are introduced as a necessary step in preserving correspondence between the formalism used and experimental possibilities. In their works they define as entropy a quantity which is not proportional to \(-\sigma\) nor to \(-\Sigma\), but proportional to a certain combination of them, which we shall call \(-\sigma\Sigma\). The quantity \(\sigma\Sigma\) can be obtained

* There is a difference between the division into groups here and in Sec. IV.1. Earlier we grouped the energy levels of the particles constituting the system; now the levels of the entire system are grouped.

** A discussion of the choice of the coarse-structure density is also given in a recent paper by van Kampen \(^{54\mathrm{K}}\).

from \(\Sigma\), setting all \(S_i\) equal to unity, or from \(\sigma\), assuming \(\rho\) always diagonal,

\[ \sigma\Sigma=-\sum \rho_{kk}\ln \rho_{kk}. \tag{IV.2,34} \]

Below we shall see that \(\sigma\Sigma\) and \(\Sigma\) sometimes coincide, namely in the case of an ensemble representing a system on which a measurement has just been performed.

The second step in their reasoning consists in the observation that one is usually interested in systems that are in contact with their environment, so that the Hamiltonian of the system consists of two parts

\[ \mathbf H=\mathbf H_0+\mathbf V, \tag{IV.2,35} \]

where \(\mathbf H_0\) is the Hamiltonian of the system when interaction with the external environment is neglected, and \(\mathbf V\) is the interaction-energy operator. From the quantum-mechanical theory of perturbations it is well known that transitions between eigenstates of \(\mathbf H_0\) can occur under the influence of \(\mathbf V\). Let us now choose, as \(\varphi_n\), the eigenfunctions of \(\mathbf H_0\) and of any other operators commuting with \(\mathbf H_0\). For reasons that will presently become clear [cf. expressions (IV.2,37) and (IV.2,41)], we are especially interested in the case when the energy eigenvalues are degenerate. To indicate degeneracy, we attach to \(\varphi_n\) two indices: \(\varphi_{k,l}\), with the first index indicating the energy value \(E_k\), and the second index distinguishing the states belonging to one and the same \(E_k\).

Since we are dealing only with the diagonal elements of \(\rho\), which, as we have just seen, are the probabilities of finding the system of the ensemble in a particular state, we may use the quantum-mechanical formulas for the dependence of these probabilities on time; we obtain\(^*\)

\[ \rho(kl,kl;t)=\rho(kl,kl;t_0)+ \]

\[ +\sum_{k'l'} J(kl,k'l')\{\rho(k'l',k'l';t_0)-\rho(kl,kl;t_0)\}, \tag{IV.2,36} \]

where

\[ J(kl,k'l')=\delta_{kk'}(4\pi^2/h)\,\lvert V(kl,k'l')\rvert^2(t-t_0). \tag{IV.2,37} \]

In expression (IV.2,37), the quantity \(\delta_{kk'}\) indicates that \(J\) is equal to zero if energy is not conserved, i.e. transitions occur only between states with the same energy; \(h\) is Planck’s constant, and \(V(kl,k'l')\) is the matrix element of \(\mathbf V\), satisfying the relation

\[ V(kl,k'l')=V^*(k'l',kl) \tag{IV.2,38} \]

in view of the Hermiticity of the matrix \(\mathbf V\). By virtue of (IV.2,38) we have

\[ J(kl,k'l')=J(k'l',kl). \tag{IV.2,39} \]

\(^*\) Cf. Appendix V.

Substituting (IV.2,36) into (IV.2,34), we obtain*) to first order in \(J\)

\[ \sigma\Sigma(t)=\sigma\Sigma(t_0)-\frac{1}{2}\sum_{kl,k'l'} J(kl,k'l')Q(kl,k'l'), \tag{IV.2,40} \]

where

\[ Q(kl,k'l')=[\rho(kl,kl;t_0)-\rho(k'l',k'l';t_0)]\times \]

\[ \times \ln[\rho(kl,kl;t_0)/\rho(k'l',k'l';t_0)] . \tag{IV.2,41} \]

Since \(Q\) is nonnegative, and \(J\) is always a positive and, moreover, monotonically increasing function of \(t\), for \(\sigma\Sigma\) we have a monotonically decreasing function.

Against the argument just given one might object that expression (IV.2,37) is valid only for small \((t-t_0)\); however, as was shown by Green, the treatment can be modified in such a way that it is valid for arbitrary values of \((t-t_0)\). The most substantial objection is directed against the use of the function \(\sigma\Sigma\) as a measure of entropy. As Pauli indicated\(^{49p}\), if such a choice is nevertheless made, then all the rest will follow from it, as is clear from Klein’s work\(^{31K}\), presented in Appendix VI (see also below in the present section).

A second method for studying the behavior of \(\Sigma\) was proposed by Pauli\(^{28p}\). Let \(P_i\) be the probability of finding a system of the ensemble in the \(i\)-th group. We then have the following relation**):

\[ P_i=S_iP_{kk}, \tag{IV.2,42} \]

where \(P_{kk}\) is one of the \(S_i\) equal diagonal elements belonging to the \(i\)-th group. From (IV.2,32) we have, expressing \(\Sigma\) in terms of \(P_i\),

\[ \Sigma=\sum_i \sum (P_i/S_i)\ln(P_i/S_i)=\sum_i P_i\ln(P_i/S_i), \tag{IV.2,43} \]

where in the second term the summation is over all groups, and the second summation is over all \(S_i\) members of the group.

The derivative of \(\Sigma\) with respect to time is given by the expression

\[ d\Sigma/dt=\sum_i(\ln P_i-\ln S_i)(dP_i/dt), \tag{IV.2,44} \]

where it has been used that \(\sum P_i=1\), or \(\sum dP_i/dt=0\), which follows from (IV.2,31) and (IV.2,42).

*) Relation (IV.2,39) is used here 1) in order to show that the sum \(\sum J(kl,k'l')[\rho(k'l',k'l';t)-\rho(kl,kl;t)]\) is equal to zero and 2) in order to obtain the sum on the right-hand side of (IV.2,40) (cf. the derivation of equation (IV.1,14), the appearance of the factor \(1/2\) being obvious).

**) Cf. the discussion of relation (IV.2,30).

Let \(N_{ij}\) be the mean probability of transition from group \(S_i\) to group \(S_j\). In Appendix V it will be proved that \(N_{ij}\) is given by the relation

\[ N_{ij}=A_{ij}S_jP_i, \tag{IV.2,45} \]

and we shall assume that

\[ A_{ij}=A_{ji}. \tag{IV.2,46} \]

Then we have

\[ dP_i/dt=\sum_j (N_{ji}-N_{ij})=\sum_j A_{ij}[S_iP_j-S_jP_i], \tag{IV.2,47} \]

and, in the same way as expression (IV.2,40) was derived, we obtain

\[ d\Sigma/dt=\frac{1}{2}\sum_{i,j} A_{ij}[P_jS_i-P_iS_j]\ln(P_iS_j/P_jS_i). \tag{IV.2,48} \]

Again, from the form of the right-hand side of expression (IV.2,48) it follows that \(d\Sigma/dt\) cannot be positive.

One may note the fact that \(d\Sigma/dt\) will be equal to zero only in the case

\[ P_jS_i=P_iS_j \quad\text{or}\quad P_j/S_j=P_i/S_i, \tag{IV.2,49} \]

or, using (IV.2,45) and (IV.2,46),

\[ N_{ij}=N_{ji}; \tag{IV.2,50} \]

we see that in equilibrium there are as many transitions from group \(S_i\) to group \(S_j\) as conversely. This is a special case of the so-called principle of detailed equilibrium, discussed in greater detail in Appendix VII.

Before returning to the discussion of the second method of investigating the change of \(\Sigma\) with time, let us consider a third method, which is completely analogous to the analysis of the change of \(\Sigma\) with time in Section III.\(^*\)

Suppose that an observation has been made on our system at time \(t'\), giving us the coarse-structure density \(\mathbf{P}\). The fine-structure density \(\rho\) at time \(t'\) is given by the equality (cf. Section IV.3)

\[ \rho'=\mathbf{P}', \tag{IV.2,51} \]

whereas for \(\Sigma'\) we have

\[ \Sigma'=\operatorname{Tr}(\mathbf{P}'\ln\mathbf{P}')=\operatorname{Tr}(\rho'\ln\rho'). \tag{IV.2,52} \]

If the state at time \(t'\) does not correspond to equilibrium, so that \(\rho'\) (or \(\mathbf{P}'\)) does not satisfy the relations (IV.2,15), (IV.2,17), or (IV.2,19), then equality (IV.2,51) in the subsequent—

____________________

\(^*\) I express my gratitude to Prof. O. Klein for discussion of a number of questions connected with this method of consideration.

moment \(t''\), generally speaking, will not be satisfied, and we shall have

\[ \rho'' \ne P'' \tag{IV.2,53} \]

and, as a result,

\[ \Sigma'' < \Sigma' . \tag{IV.2,54} \]

Let us consider in somewhat more detail how relation (IV.2,54) follows from (IV.2,53)*. Take the expression

\[ \Sigma' - \Sigma'' = \operatorname{Tr}(P' \ln P') - \operatorname{Tr}(P'' \ln P'') = \operatorname{Tr}(\rho' \ln \rho') - \operatorname{Tr}(\rho'' \ln P'') = \sum_k \rho'_{kk}\ln \rho'_{kk} - \sum_n \rho''_{nn}\ln P''_{nn}, \tag{IV.2,55} \]

where we have used the fact that the coarse-structure density is always a diagonal matrix [see (IV.2,30)]. Now add to the right-hand side of (IV.2,55) the expression

\[ \sum_n \rho''_{nn}\ln \rho''_{nn} - \sum_k \rho'_{kk}\ln \rho'_{kk}, \tag{IV.2,56} \]

which, by Klein’s lemma, is nonpositive and is equal to zero only if \(\rho''\) is a diagonal matrix, as proved in Appendix VI. Next add to the right-hand side of (IV.2,55) the expression \(\operatorname{Tr} P'' - \operatorname{Tr}\rho''\), equal to zero, since both \(\rho''\) and \(P''\) are normalized. The final result is as follows:

\[ \Sigma' - \Sigma'' = \sum_n \left( \rho''_{nn}\ln \rho''_{nn} - \rho''_{nn}\ln P''_{nn} - \rho''_{nn} + P''_{nn} \right) \ge 0; \tag{IV.2,57} \]

the last inequality follows from the properties of the function \(y\), defined by expression (III.1,13), for, making the substitution \(x=\ln \alpha-\ln \beta\) and multiplying by \(\beta\), we obtain for \(\beta>0\)

\[ \begin{aligned} \alpha \ln \alpha - \alpha \ln \beta - \alpha + \beta &> 0, \qquad \alpha \ne \beta,\\ &= 0, \qquad \alpha=\beta . \end{aligned} \tag{IV.2,58} \]

One may expect that \(\Sigma\) will continue to decrease; this is based on considerations analogous to those set forth in Sec. III.1. This decrease will continue until a stationary state is reached, in which \(P\) is given by one of the three expressions (IV.2,15), (IV.2,17), or (IV.2,19), with \(\rho=P\).

We must now discuss the various ways of considering the dependence of \(\Sigma\) on time. As mentioned above, we are not convinced

* The treatment on p. 373 of ESM is too brief and rather unsuccessful.

in the soundness of the method of approach used by Born and Green, and we shall not discuss it here in greater detail. The second method of approach was used by Pauli. This method is in fact closely connected with the simple approach set forth in Sec. IV.1. This can be seen from the following. Let us write the expression (IV.2,43) for \(\Sigma\) in the form

\[ \Sigma=\sum_i P_i\ln P_i-\sum_i P_i\ln S_i . \tag{IV.2,59} \]

Now \(P_i\) is the probability of finding the system in one of the states corresponding to the \(i\)-th group, and \(S_i\) is the number of states in this group. The number of states in the group, however, is just equal to the function \(W(N_i)\) considered in Sec. IV.1; this is fairly easy to verify. Using the relation between \(H\) and \(\ln W\), leading to (IV.1,28), we can rewrite expression (IV.2,59) in the form

\[ \Sigma=\sum_i P_i H_i+\sum_i P_i\ln P_i=\langle H\rangle+\sum_i P_i\ln P_i, \tag{IV.2,60} \]

where \(H_i\) is the value of \(H\) when the system is in a state belonging to the \(i\)-th group\(^*\), and \(\langle H\rangle\) is the mean of \(H\) taken over the ensemble.

Using expression (IV.2,60), we see, first, that if we are dealing with an ensemble representing a system which, by measurement, has been found to be in some one definite state, so that one of the \(P_i\) is equal to unity and the others are equal to zero, then this expression reduces to

\[ \Sigma=H . \tag{IV.2,61} \]

Secondly, taking the derivative with respect to time of (IV.2,60), we obtain

\[ d\Sigma/dt=\langle dH/dt\rangle=\sum_i (dP_i/dt)\ln P_i, \tag{IV.2,62} \]

whence it is seen that the decrease of \(\Sigma\) is due to two circumstances: 1) the general decrease of \(H\) for any system of the ensemble and 2) the tendency toward a decrease of the quantity \(\sum_i P_i\ln P_i\) under a more homogeneous distribution of the systems of the ensemble among the various groups of states.

Let us now consider expression (IV.2,57). We see that \(\Sigma\) decreases, first, as a consequence of the fact that \(\rho\) and \(P\) no longer coincide; this cause of the decrease of \(\Sigma\) also occurred in the classical case. Secondly, \(\Sigma\) decreases because the expression \(\sum p_{hk}\ln p_{hk}\) decreases (Klein’s lemma). For the second cause there is no classical analogue, and it must be briefly discussed.

\(^*\) Let us note that one can derive an equation analogous to (IV.2,60) in the classical case. The discussion in this case will be analogous to the one carried out above.

This reason was called by Tolman[^38T] a quantum-mechanical change of the fine-structure probability. In this case there is no classical analogue, since the quantity \(\sum P_{kk}\ln P_{kk}\) is not the trace of a quantum-mechanical operator. The more states are occupied, the smaller the value of this sum. This decrease is connected with the irreversible perturbation that takes place when an observation of the system is performed2. Thus this effect will be revealed more distinctly in the case when the observations are such that they approach the limits established by the Heisenberg relation; and, conversely, it will be inessential in those cases when the perturbation of the system due to the observation may be neglected. In the latter case the decrease of \(\Sigma\) will be due to the increasing difference between \(\rho\) and \(P\).

It is interesting to note that some authors[^36D][^37S] consider only this influence of the measurement on the system, without paying due attention to the role of the fine-structure and coarse-structure densities and to the fact that the approach to equilibrium must not depend on whether or not we actually observe the system3.

In the next section it will be considered how ensembles representing ensembles in quantum statistics can be constructed; however, some aspects of the question of the consequences of observations on a physical system we wish to consider briefly already now, in particular the meaning of a pure case.

In classical statistics ensembles were introduced because our information about the physical systems under consideration is practically always very far from the maximum possible information. In quantum statistics, however, the situation is complicated by the statistical aspects inherent in quantum mechanics itself. In the ideal (and never attainable) case in classical statistics one may know the values of all constants of the motion, so that the problem of the behavior of the system reduces to a problem of classical mechanics. In the ideal (and never attainable) case in quantum statistics we may know that some one system is in a state proper to a certain quantum-mechanical operator, for which, for simplicity, we shall choose the energy operator4, so that we are dealing with a system in a stationary state. We have a pure case, and all the systems in the representing ensemble will possess the same wave function, namely the function corresponding to the eigenstate under consideration. Of the two averaging processes present in (IV.2,8), there remains

only one, namely quantum-mechanical averaging. We are still obliged to resort to a statistical consideration, although in this case our knowledge is maximal.

In the general case, however, the situation is not so favorable, and the construction of the representing ensemble becomes considerably more complicated. In this case both averages in (IV.2,8) must be taken into account. Here we are dealing with a mixed case.

In the pure case, if the state of the system is characterized by one of the functions forming a complete orthonormal set \(\varphi_n\), say \(\varphi_1\), then all \(\psi^k\) are equal to \(\varphi_1\), and we have

\[ a_n^k=\delta_{n1} \tag{IV.2,63} \]

and, consequently,

\[ \rho_{mn}=\delta_{mn}\delta_{m1}. \tag{IV.2,64} \]

In this case for \(\sigma\) we have

\[ \sigma=\operatorname{Tr}(\rho\ln\rho)=0. \tag{IV.2,65} \]

In the mixed case there is more than one nonzero quantity \(\rho_{mn}\), and therefore \(\sigma\) is less than zero). This, apparently, is the reason why Elsasser called \(\sigma\) an “indicator of mixing”*). The decrease of \(\sigma\) in passing from the pure case to the mixed case once again shows the connection between entropy, which is proportional to \(-\sigma\), and lack of information, since the decrease in the completeness of information causes a decrease of \(\sigma\) (cf. also the discussion in \(39^{K1}, 39^{K2}\)).

In the next section we shall see that the representing ensemble of a subsystem that is part of a large isolated system is the canonical grand ensemble. This means that, whereas the large system can be represented by a pure case, the subsystems correspond definitely to mixtures. This fact became the source of a lengthy discussion initiated by the paper of Einstein, Podolsky, and Rosen \(35^{E}\), in which the paradox bearing their name was formulated. We have no space to discuss this paradox, and we shall refer to the literature \(35^{B}, 35^{S1}, 35^{S2}, 35^{S3}, 35^{S4}, 36^{M}, 36^{F}, 36^{S}, 51^{B3}\); (cf. also \(27^{N1}\)).

\[ \text{*) Equality (IV.2,65) is a necessary and sufficient condition for the pure case }{}^{33P} \text{ and is equivalent to the usual condition } \rho^2=\rho. \]

\[ \begin{aligned} \text{**) In his work Elsasser considered how measurements can be used to construct representing ensembles. }\\ \text{Instead of making assumptions about a priori probabilities, Elsasser chose his ensembles so}\\ \text{that they gave a minimum of }\sigma\text{. In this, it appears, he obtained the same results as}\\ \text{we do by the more customary method, since the normal method leads to minimal values of }\sigma. \end{aligned} \]

IV.3. Representing ensembles

In this question the situation is again very similar to that in classical statistics. We have at our disposal only incomplete data, on the basis of which it is necessary to predict the future behavior of our system; hence we are obliged to resort to representing ensembles. At the same time the question of a priori probabilities again arises, and since in reality we need the quantities \(a_n^k\), the question of a priori phases also arises, as will be seen below. We use the assumption of equality of the a priori probabilities for all nondegenerate states*) and the assumption of the randomness of the a priori phases; below it will be clarified what we understand by the expression “random.” The following discussion will be analogous to the discussion in Sec. III.2. Let us suppose again that our system is part of a larger isolated system; instead of expressions (III.2,1)—(III.2,4) we now have

\[ \operatorname{Tr}\rho = 1, \tag{IV.3,1} \]

\[ \operatorname{Tr}\rho' = 1, \tag{IV.3,2} \]

\[ \operatorname{Tr}(H\rho) + \operatorname{Tr}(H'\rho') = \mathrm{const}, \tag{IV.3,3} \]

\[ \operatorname{Tr}(\nu\rho) + \operatorname{Tr}(\nu'\rho') = \mathrm{const}, \tag{IV.3,4} \]

where the unprimed quantities refer to the subsystem, and the primed quantities to the rest of the system. If equilibrium has been established, then \(\Sigma\) must be minimal, or

\[ \operatorname{Tr}\bigl[(\rho\rho')\ln(\rho\rho')\bigr] = \text{minimum}. \tag{IV.3,5} \]

Using the method of undetermined multipliers, we find (for more detail see, for example, \(^{40\mathrm{H}}\)**) the equilibrium expression

\[ \rho = \exp(-q + \nu\mu - \beta H), \tag{IV.3,6} \]

in agreement with what was obtained in the preceding section. As in Sec. III.2, expression (IV.3,6) in reality holds only for the coarse-structure density \(P\), for relation (IV.3,5) is satisfied only for \(\Sigma\), but not for \(\sigma\).

The situation here is also completely analogous to the classical case if only nonequilibrium situations are considered; therefore we refer to the discussion given in Sec. III.2, which is directly applicable in the present case as well.

) If a state is expressed \(g\) times, then it is counted as \(g\) nondegenerate states.
*) In this case both the characteristic values and the characteristic functions \(\rho\) are varied.

Let us make one more remark concerning the construction of the representing ensemble after an observation has been made on the system. We can distinguish only states belonging to different groups, as has already been said earlier, and therefore in a state determine only the coarse-grained density, or \(P_{kk}\). In accordance with the assumption of equality of the a priori probabilities for all nondegenerate states, it became necessary to set all diagonal elements of \(\rho\) corresponding to one group equal to one another and to equate them to the corresponding \(P_{kk}\). As for the nondiagonal elements of \(\rho\), since they are mean values of products of two \(a_n^k\), we have

\[ \rho_{mn}=\frac{1}{N}\sum a_m^k a_n^{k*} =\frac{1}{N}\sum r_m^k r_n^k \exp\{i(\Phi_m^k-\Phi_n^k)\}\quad (m\ne n), \tag{IV.3,7} \]

where \(r_m^k\) is the modulus of the quantity \(a_m^k\) and \(\Phi_m^k\) is its phase.

Using the assumption of the randomness of the a priori phases, one may expect that the mean of \(\exp[i(\Phi_m^k-\Phi_n^k)]\) is equal to zero, and thus, for the fine-grained density of our ensemble, obtain

\[ \rho_{kl}=P_{kk}\delta_{kl}, \tag{IV.3,8} \]

from which it is indeed seen that the quantities \(\sigma\), \(\Sigma\), and \(\sigma\Sigma\) are all equal to one another, as was noted in the discussion of expression (IV.2,34).

Let us now turn to the consideration of our choice of a priori probabilities and a priori phases. When the analogous question was considered in classical statistics, it was shown first of all that the a priori weights must be invariant with respect to the motion of the representative point along its orbit. It was also shown then that in the case of quasiergodic systems the a priori weights will be functions of the energy alone and, consequently, may be taken equal for equal volumes in phase space. The situation in quantum statistics is in some respects simpler, and in others more complicated, in comparison with classical statistics. We shall show, first, that an ensemble constructed without taking into account any conditions corresponding to the assumptions of equality of the a priori probabilities and of randomness of the a priori phases for different nondegenerate quantum states will be a stationary ensemble and, moreover, will correspond to the same ensemble if we pass to another representation. Secondly, the significance of the so-called adiabatic invariance of the weights of different quantum states will be briefly discussed in this connection. Further, we shall briefly consider the connection between our choice of a priori probabilities and the condition of nondegeneracy of energy states, which coincides with the condition of ergodicity (see Chapter V). Finally, it will be shown how, in the classical limit, the condition of equality of the a priori probabilities and of randomness

of the a priori phases of various nondegenerate states leads to the condition of equality of the a priori probabilities for equal volumes of phase space*).

Let us consider the so-called homogeneous ensemble, whose density matrix is given by the relation

\[ \rho_{kl}=\rho_0\delta_{kl}\quad \text{or}\quad \rho=\rho_0\mathbf{1}, \tag{IV.3,9} \]

where \(\mathbf{1}\) is the unit matrix. Relation (IV.3,9) is satisfied for \(\rho\) independently of the choice of a complete orthonormal system, which is easily shown by performing a unitary transformation \(U\)

\[ \rho'=U^+\rho U=U^+\rho_0\mathbf{1}U=\rho_0 U^+\mathbf{1}U=\rho_0 U^+U=\rho_0\mathbf{1}. \tag{IV.3,10} \]

One may consider that the homogeneous ensemble is obtained by applying the prescriptions concerning the equality of a priori probabilities and the randomness of a priori phases for different states, since from (IV.2,7) we have

\[ \rho_{mn}=(a_m a_n^*)_{\mathrm{av}}, \tag{IV.3,11} \]

where the averaging is carried out over all systems of the ensemble. Setting

\[ a_m=r_m\exp(i\alpha_m), \tag{IV.3,12} \]

where \(r_m\) is the modulus and \(\alpha_m\) the phase of the quantity \(a_m\), we obtain from (IV.3,11)

\[ \rho_{mn}=\{r_m r_n\exp[i(\alpha_m-\alpha_n)]\}_{\mathrm{av}}. \tag{IV.3,13} \]

Equality of the a priori probabilities means that

\[ r_m=\sqrt{\rho_0} \tag{IV.3,14} \]

independently of \(m\), so that (IV.3,13) reduces to the expression

\[ \rho_{mn}=\rho_0\{\exp[i(\alpha_m-\alpha_n)]\}_{\mathrm{av}}, \tag{IV.3,15} \]

which, in view of the randomness of the distribution of phases, reduces to (IV.3,9).

The homogeneous ensemble is also a stationary ensemble. This follows directly from (IV.2,27) [cf. relations (IV.3,10)].

We see here that the condition of equality of a priori probabilities and of randomness of a priori phases for different states is invariant both with respect to time and with respect to transformations from one representation to another. The assumption of randomness of a priori phases is necessary in order to establish this invariance with certainty, for the density matrix given by the expression

\[ \rho_{kl}=\rho_0\delta_{kl}+A_{kl}(1-\delta_{kl}), \tag{IV.3,16} \]

* Some other aspects of the transition from quantum to classical statistics are contained in Uhlenbeck’s article \(^{35U}\).

would satisfy the requirement of equality of the a priori probabilities, but would not be a stationary density, or a density invariant with respect to a unitary transformation.

It may be noted here that Dirac \(^{23D}\) (see also \(^{27N1}\)) singled out the homogeneous ensemble, since it remains invariant under an arbitrary small perturbation of the Hamiltonian and, thus, gives an a priori distribution in quantum statistics. He also drew attention to the fact that such an ensemble leads—as we have just seen—to equal a priori probabilities for all nondegenerate states.

The equality of the a priori weights for nondegenerate states follows from the so-called adiabatic principle, which is usually associated with the name of Ehrenfest. We cannot enter into a detailed discussion of this very important question here, and must refer to the literature*) (\(^{06E1, 11E1, 14E, 16E, 17B1, 17B2, 17B3, 17B4, 17E1, 18B, 23B2, 27U}\) and especially \(^{23E1}\), where further references may be found). This principle says that if the external parameters change adiabatically, then the relative weights of the various states cannot change. It follows directly from this that, in the case of systems possessing many periods (multiply periodic systems), each nondegenerate quantum state has one and the same a priori weight.

We have now arrived at the important question of whether the choice of a priori probabilities is completely determined if we are dealing with ergodic systems, as was the case in the classical situation. In Section III.3 we saw that in classical statistics this was so because the constants of motion distinct from the energy assume their possible values practically equally often in different parts of phase space, or, as is sometimes said, they are practically constant on the energy surface. This latter property is also valid for ergodic systems in quantum mechanics, as is not difficult to see. If all energy states are nondegenerate, the eigenfunctions of the energy form a complete orthonormal system. Since a constant of the motion corresponds to an operator commuting with the energy operator, it follows from quantum theory (see, for example, \(^{38K}\), Sec. 38) that the eigenfunctions of the energy are also eigenfunctions of all operators corresponding to integrals of the motion. Hence it follows further that in any energy state, which by virtue of nondegeneracy corresponds to one and only one eigenfunction, the integrals of the motion have quite definite constant values. Unfortunately, however, from this

*) In some of Ehrenfest’s papers the significance of this principle for the choice of a priori weights in classical statistics is also discussed.

does not necessarily imply our choice of a priori probabilities and a priori phases. As regards a priori probabilities, we hardly need the ergodicity of the systems under consideration, for, as we have just seen, the adiabatic principle admits only those a priori weights which we have constantly used. The choice of a priori phases, however, is more difficult to justify from first principles—and this is precisely what we are now trying to do.

Let us admit for a moment the plausibility of the existence of ergodic systems. From quantum theory it is known that degenerate levels are more the rule than the exception when we are dealing with such simple systems as atoms and molecules, and it seems unlikely that the situation would change in passing to the more complex systems considered in statistical mechanics. In this connection it is useful to recall that it has not even been possible to prove that the ground state of an arbitrary system is nondegenerate—a property which plays an important role in considering the third law of thermodynamics, the so-called Nernst heat theorem*). Consequently, we must also take into account the possibility that physical systems are not ergodic. Let us briefly consider here how far one can nevertheless justify the ergodic theorem of Part V, i.e. the equality of time and ensemble averages. This equality is based on relation (V.26), proved in Appendix V. From consideration of this expression it is clear that relation (V.26) can also be proved in the case where we can perform an averaging over phases on the right-hand side of expression (App. IV.7)**). One would like to think that this averaging is fully in accord with the basic ideas of quantum mechanics. It is well known that the (initial) phases of wave functions play no role in any physical quantities; hence it is tempting to suppose that every wave function should be regarded as an average (or, better, a mixture) of a large number of wave functions differing only in their phases and otherwise coinciding. This would not alter any results obtained from ordinary wave mechanics, and would ensure the validity of the ergodic theorem, and would also justify the method by which we constructed the representative ensembles. Thus the assumption of the randomness of the a priori phases is transferred from statistical mechanics to quantum mechanics itself.

In conclusion of the present section, let us consider in what way our classical a priori probabilities are a limiting case of the assumption concerning a priori weights and a priori phases made in quantum statistics. This transition is easy to justify in old quantum mechanics, at any rate when dealing with

*) For a discussion see 30S, 51S, 16S and ESM, Appendix III.
**) Cf. also E08.

systems possessing many periods (cf. 18В), since in the old quantum theory the energy levels were obtained precisely by subdividing phase space into cells of volume \(h^s\) (\(h\) is Planck’s constant, \(s\) is the number of degrees of freedom) and assigning one level to each cell.

In modern quantum mechanics the situation is more complicated. One often encounters the not entirely clear assertion that assigning to each state a volume \(h^s\) in phase space is connected with the Heisenberg relation \(^{27Н}\), which is then taken in the form

\[ \Delta p \cdot \Delta q \gtrless h, \tag{IV.3,17} \]

where \(p\) and \(q\) are a canonically conjugate momentum-coordinate pair, and \(\Delta p\) and \(\Delta q\) are the limits of accuracy with which they can be measured simultaneously. In these cases it is not explained why the volume should not be equal to \((h/4\pi)^s\), for the right-hand side of inequality (IV.3,17) in a more rigorous form should be \(h/4\pi\), and not \(h\) \(^{30Н1}\).

However, it is probably possible to justify the choice of volume in the following way (ESM, p. 60; cf. also \(^{35D}\), section 37, \(^{49M}\)). In classical mechanics the phase of a particle is represented by a point in phase space. This point may be regarded as a combination of a point in coordinate or \(q\)-space and a point in momentum or \(p\)-space. In quantum mechanics, however, it is necessary to consider probability densities. The two probability densities in \(p\)- and \(q\)-spaces are not independent. The probability density in \(q\)-space is determined by the wave function \(\psi\) and is equal to \(|\psi|^2\), the function \(\psi\) being normalized:

\[ \int |\psi|^2\,dq = 1; \tag{IV.3,18} \]

here \(dq\) denotes the volume element in \(q\)-space. The probability density in \(p\)-space is obtained, first of all, by transforming the wave function \(\psi(q)\) in the coordinate representation into the probability amplitude \(A(p)\) in momentum space. The probability density in \(p\)-space is then equal to \(|A|^2\). It follows from the theory of transformations that \(A\) and \(\psi\) are Fourier transforms of one another, if \(q\) and \(p/h\) are taken as variables*), and, consequently,

\[ \int |A|^2\,dp = h^s. \tag{IV.3,19} \]

We can now say that the total volume of phase space corresponding to one state of the system is obtained by multiplying the probability densities in \(p\)- and \(q\)-spaces and integrating over the entire phase space; in the final result we obtain the volume \(h^s\), as is seen from (IV.3,18) and (IV.3,19).

*) It is precisely at this point that Planck’s constant enters, either through the de Broglie relation or by means of the commutation relations, which also lead to the Heisenberg relation.

From the reasoning just given one may conclude that, in the limit \(h \to 0\), indeed, equal volumes of phase space containing an equal number of cells of volume \(h^s\) correspond to equal a priori probabilities, since we have assumed equality of a priori probabilities for all nondegenerate states. The necessity of the randomness of the a priori phases in this connection is not so obvious. The latter, however, follows from consideration of a homogeneous ensemble. It is desirable to obtain equal weights for nondegenerate states independently of the representation used, and this can be achieved only if we adopt the additional assumption of the randomness of the a priori phases, as was seen from the discussion in the present section.

V. THE ERGODIC THEOREM IN QUANTUM STATISTICS

We saw above that, alongside the approach to the problem under study by means of the \(H\)-theorem, the use of statistical methods can be justified with the aid of the ergodic (or quasi-ergodic) theorem and of the proof of the equivalence of averages with respect to time and with respect to the ensemble. Exactly the same choice of approach is also available in quantum statistics; in the present part the quantum-statistical ergodic theorem will be considered. By analogy with the old ergodic theorem (Section II.1), one may assume the existence of ergodic systems. We define these systems as those in which all states \(E_k\) lying within the prescribed energy interval

\[ E \leqslant E_k \leqslant E+\delta E, \tag{V,1} \]

can be reached starting from any other states within the same interval—even if this is possible only through other intermediate states*).

Let \(P_k\) again denote the probability of finding the system in the \(k\)-th state, and let \(w_{kl}\) be the probability of transition from the \(k\)-th to the \(l\)-th state. For the equation describing the change of \(P_k\) with time, we then obtain

\[ dP_k/dt=-\sum_l P_k w_{kl}+\sum_l P_l w_{lk} =\sum_l w_{kl}(P_l-P_k), \tag{V,2} \]

where the condition \(w_{kl}=w_{lk}\)**) has been used. The equilibrium solution of equation (V,2), reached exponentially, has the form

\[ P_k=\mathrm{const}=1/n, \tag{V,3} \]

*) See \(^{33}\) Section 22, \(^{40M}\) p. 56, \(^{46I}\) p. 152.
**) Van Kampen \(^{54K}\) showed that equation (V,2) can also be derived under the conditions \(w_{kl}\ne w_{lk}\).

if \(n\) is the number of states in the interval \((V,1)\). It follows from this that the mean time of residence in any of the \(n\) states is the same, and that in order to compute the time average of any quantity one may average it over the states of the interval \((V,1)\), or over the corresponding microscopic ensemble [cf. (V.2,15), where the density for such an ensemble is given]. Thus we have proved that for ergodic systems the two types of averages coincide. The question remains open whether ergodic systems exist; in the present case it is very probable that they do. In a closed system it is necessary to take account of radiation, which should make transitions from one state to another possible *).

In quantum mechanics, however, there is also another method of approach, analogous to the ergodic theorems of Birkhoff or Hopf. This method consists in a direct proof of the equivalence of time averages and ensemble averages. The original proof of Neumann \(^{29N}\) was subsequently simplified by Pauli and Fierz \(^{37P}\), and by Fierz (private communication) **).

As in the classical case, one can distinguish two different ergodic theorems according as “coarsening” of the system’s energy is or is not allowed. In the quantum-statistical case the second ergodic theorem is more important; however, first the first ergodic theorem will be briefly set forth, and we shall follow Rosenfeld’s presentation \(^{52R}\).

We shall first consider the pure case. Let \(\psi(t)\) be the wave function describing the system under consideration. Choose, as a complete orthonormal system of functions, the functions \(\varphi_n\), eigenfunctions of the energy operator \(\mathrm{H}\), and denote the corresponding eigenvalues by \(E_n\). Write

\[ \psi(0)=\sum_n a_n \varphi_n, \tag{V,4} \]

where

\[ a_n=r_n \exp(i\alpha_n), \tag{V,5} \]

with \(r_n\) the modulus and \(\alpha_n\) the phase of the quantity \(a_n\). For \(\psi(t)\) we obtain in the usual way

\[ \psi(t)=\sum r_n \exp(i\alpha_n)\exp(-iEt/h)\cdot\varphi_n . \tag{V,6} \]

*) In this connection one may mention the recent article of M. Klein \(^{52K}\), who made certain explicit assumptions concerning interactions that ensure the existence of transitions.

**) I am very grateful to Prof. M. Fierz for making available to me his latest investigations, discussed in 1953.

The density matrix, as a function of time, is given by the expression

\[ \rho_{mn}(t)=r_m r_n \exp [i(\alpha_m-\alpha_n)]\exp[-i(E_m-E_n)t/\hbar]. \tag{V,7} \]

The mean value of any physical quantity \(A\) is the quantum-mechanical mean given by the expression

\[ \langle A\rangle=\sum_{mn} r_m r_n \exp [i(\alpha_m-\alpha_n)-i(E_m-E_n)t/\hbar]A_{mn}, \tag{V,8} \]

where \(A_{mn}\) is given by the relation

\[ A_{mn}=\int \varphi_m^* A\varphi_n\,d\tau. \tag{V,9} \]

Now one may take the time average of \(\langle A\rangle\), and it is evident that if there is no degeneracy, i.e., if no two energy eigenvalues coincide, then

\[ \overline{\langle A\rangle}=\sum_n \rho_{nn}A_{nn}; \tag{V,10} \]

the latter does not depend on the initial phases \((\alpha_n)\).

We see that \(\overline{\langle A\rangle}\) will not depend on the initial conditions, in our case on the initial phases, only if there is no degeneracy. Thus the requirement of absence of degeneracy replaces the requirement of metric transitivity in the classical case.

Relation (V,10) gives \(\overline{\langle A\rangle}\) in the form of a statistical average; this can be seen by introducing the so-called projection operator \(\mathbf{R}(\varphi)\). This operator is defined by the relation

\[ \mathbf{R}(\varphi)\chi=\varphi \int \varphi^*\chi\,d\tau, \tag{V,11} \]

which must hold for any pair of functions \(\varphi\) and \(\chi\). The geometric meaning of the projection operator is easy to clarify by using precisely the terminology of Hilbert space. The integral \(\int \varphi^*\chi d\tau\) then corresponds to the scalar product of two vectors \(\chi\) and \(\varphi\) in Hilbert space, and the right-hand side of relation (V,11) therefore gives the “projection” of \(\chi\) onto \(\varphi\), if \(\varphi\) is normalized and consequently corresponds to a unit vector in Hilbert space*).

It is easy to see that in the case when the system is characterized by the wave function \(\psi\), the operator \(\mathbf{R}(\psi)\) is diagonal and its diagonal elements are equal to the diagonal elements of the density matrix,

*) The notation \(\mathbf{R}\) is used instead of \(\mathbf{P}\) in order to emphasize the connection with the density matrix.

so that one may rewrite (V,10) in the form

\[ \overline{\langle A\rangle}=\operatorname{Tr}(\mathbf{R}(\psi)A). \tag{V,12} \]

It also follows from (V,7) that \(R(\psi)\) can be obtained by averaging the density matrix over the various values of the initial phases, or

\[ \mathbf{R}(\psi)=\rho_{\mathrm{av}} . \tag{V,13} \]

The index “av” here denotes averaging over the values \(\alpha_n\), and expression (V,12) or (V,10) can be written in the form

\[ \overline{\langle A\rangle}=\operatorname{Tr}(\rho_{\mathrm{av}}A). \tag{V,14} \]

Speaking not quite rigorously, we may interpret this result in the sense that the change of the wave function with time corresponds to a uniform distribution over all possible values, in accordance with the classical picture of a trajectory completely covering the energy surface.

If we are dealing with an isolated system of constant energy, then \(\psi(0)\) coincides with one of the \(\varphi_n\); this is the pure case with the microcanonical density matrix. In the more general case considered here, averaging over time leads to the replacement of the pure case by a mixture obtained by averaging over the initial phases.

The first ergodic theorem is important chiefly in discussing the influence of observations on the state of the system and in discussing the transition from the pure to the mixed case as a result of observation. Some aspects of this problem were touched upon at the end of section (IV.2) and will not be discussed here in more detail.

The second ergodic theorem, like Hopf’s ergodic theorem, operates with energy layers. At the same time, it encounters the difficulty that, often in describing a system, a number of physical quantities are used that do not correspond to commuting operators, so that it is impossible to obtain a wave function that would be a proper function for all these operators. In view of this, it is necessary to introduce, instead of these microscopic operators, the so-called macroscopic operators, as was first done by Neumann*). Roughly speaking, a macroscopic operator may be obtained from the corresponding microscopic operator in the following way. Let \(A\) be a continuous variable that can take any value between \(-\infty\) and \(+\infty\), and let the macroscopic measurements be such that one can distinguish only the intervals

\[ k\leq A\leq k+1,\quad k=0,\ \pm1,\ \pm2,\ldots \tag{V,15} \]

*) Here it is necessary to mention the paper of van Kampen \(^{54\mathrm{K}}\), published after the writing of the present work; there the introduction of these macroscopic operators is discussed very thoroughly and rigorously.

Suppose, further, that \(f(x)\) is defined by the relation

\[ f(x)=k,\quad k\leq x<k+1,\quad k=0,\ \pm1,\ \pm2,\ldots \tag{V,16} \]

The macroscopic variable corresponding to \(A\) will then be \(f(A)\). As Neumann emphasizes*), macroscopic variables can always be measured simultaneously. It is useful to reproduce his own remark: “In a macroscopic simultaneous measurement of coordinate and momentum (or of another pair of quantities which cannot be measured simultaneously by virtue of the Heisenberg relation) we actually measure simultaneously and exactly two physical quantities; however, strictly speaking, these two quantities are not coordinate and momentum. In measuring, for example, the positions of the tips of two needles, or the positions of two images on a photographic plate, nothing prevents us from measuring them simultaneously and with arbitrary accuracy; however, the relation between them and the quantities of interest to us (\(q_k\) and \(p_k\)) is indeterminate, and the uncertainty of this relation is given by the Heisenberg relation.”

Let us first introduce energy layers, grouping the energy levels as was done earlier and assigning to each group of \(S_i\) levels one corresponding value of the energy \(E_i\). Within each layer let us introduce (phase) cells [cf. Section I.3], each containing, say, \(s_\nu\) states. In each cell the macroscopic variables of interest to us \(A, B,\ldots\) will have constant values. This means that we have introduced macroscopic operators of energy, \(A, B\), etc., in such a way that they all commute, and their eigenvalues are \(s_\nu\)-fold degenerate. If \(\omega_\tau\) are eigenfunctions of the macroenergy and of the macrooperators corresponding to \(A, B,\ldots\), then within a cell they correspond to identical eigenvalues. We then have

\[ \sum_\nu s_\nu=S_i, \tag{V,17} \]

where \(\nu\) runs over the values from \(1\) to \(N_i\); \(N_i\) is the number of cells in the layer.

Let \(P_\nu\) be the probability of finding the system in the \(\nu\)-th cell, and let \(p_\nu\) be the probability of finding the system in one of the states belonging to the \(\nu\)-th cell**). We then have [cf. relation (IV.2,42)]

\[ P_\nu=s_\nu p_\nu. \tag{V,18} \]

Define the quantity \(\Sigma(t)\) by the expression [cf. (IV.2,43)]

\[ \Sigma=\sum_\nu P_\nu \ln(P_\nu/s_\nu). \tag{V,19} \]

*) Cf. also Wightsecker \(^{49\mathrm{W}1}\).

**) \(p_\nu\) will be the coarse-grained density of the ensemble, in which all systems correspond to one and the same wave function, namely the wave function of the system under consideration.

It is necessary to bear in mind that the number of cells in the energy layer, i.e. the number of states that can be distinguished by an observer, will be of the order of magnitude, say, \(10^{20}\) (the very largest). On the other hand, the quantity \(-\Sigma\), which is a measure of the entropy of the system in units of \(k\) (Boltzmann’s constant), will also be of the order of \(10^{20}\), i.e. of the order of Avogadro’s number. This means that \(s_\nu\) will be monstrously large*)

\[ s_\nu \sim \exp(10^{20}), \quad \text{as also} \quad S_i \sim \exp(10^{20}). \tag{V,20} \]

As a result \(P_\nu\) will be of the order of magnitude \(10^{-20}\), while \(\ln s_\nu\) will be extremely large in comparison with \(-\ln P_\nu\), so that instead of (V,19) one may write

\[ \Sigma(t)=-\sum_\nu P_\nu \ln s_\nu . \tag{V,21} \]

Next, let us consider the quantity

\[ \langle \Sigma\rangle_{\mathrm{cp}}=-\sum_\nu (s_\nu/S_i)\ln s_\nu . \tag{V,22} \]

For \(\langle \Sigma\rangle_{\mathrm{cp}}\) the following inequalities can easily be proved:

\[ \ln S_i \geq \sum_\nu (s_\nu/S_i)\ln(s_\nu/S_i)+ \]

\[ +\sum_\nu (s_\nu/S_i)\ln S_i = -\langle \Sigma\rangle_{\mathrm{cp}}, \tag{V,23} \]

\[ -\langle \Sigma\rangle_{\mathrm{cp}} \geq \sum_\nu (1/N_i)\ln(S_i/N_i)=\ln S_i-\ln N_i . \tag{V,24} \]

The first inequality follows from the fact that \(0<s_\nu/S_i<1\), and the second from the fact that a sum of the type \(\sum a_\nu \ln a_\nu\), under the condition \(\sum a_\nu=1\), is minimal when all \(a_\nu\) are equal. According to (V,20), \(\ln S_i\) is of order \(10^{20}\), whereas \(\ln N_i\) is only of order 50, so that one may neglect \(\ln N_i\) in comparison with \(\ln S_i\) and write

\[ \langle \Sigma\rangle_{\mathrm{cp}}=-\ln S_i . \tag{V,25} \]

Since \(S_i\) is the number of states in the energy layer, \(-\langle \Sigma\rangle_{\mathrm{cp}}\) may be interpreted (apart from the factor \(k\)) as the entropy of the microcanonical ensemble (cf. the discussion preceding equation (IV.1,28)).

The ergodic theorem will be proved if we can prove that \(\langle \Sigma\rangle_{\mathrm{cp}}\) is the time average of \(\Sigma(t)\), given in (V,21),

*) I owe this observation to Prof. M. Fierz.

or, equivalently,

\[ \bar P_\nu=s_\nu/S_i; \tag{V,26} \]

the latter expression very much resembles relation (V,3).

It is possible that (V,26) is valid under averaging, the averaging being taken over all possible subdivisions of the layer into cells. Weyl \(^{25W}\) gave a method which makes it possible to assign definite weights to different subdivisions. In Appendix IV we shall discuss the proof of relation (V,26). The necessary condition for the proof is that there be no pronounced energy levels*).

We have now proved the equivalence of time averages and ensemble averages in the case of the function \(\Sigma\). However, if relation (V,26) holds, then one can also prove the equivalence of time and ensemble averages for other variables as well. This follows from Neumann’s consideration**), and also from an argument similar to that given in the Introduction; this argument is constructed as follows. From the fact that \(\bar\Sigma\) and \(\langle\Sigma\rangle_{cp}\) coincide, it follows that the system spends the greater part of its time in an equilibrium state and that, consequently, the time average of any quantity will correspond to the equilibrium value. It should also be noted that in proving expression (V,26) we have in fact proved more than just the equivalence of \(\bar\Sigma\) and \(\langle\Sigma\rangle_{cp}\): namely, we have made an assertion about the relative intervals of time spent in different cells; expression (V,26) corresponds to expression (II.1,5) of the classical case.

In conclusion, a few words may be said about a recent article by Klein \(^{52K}\). His point of view is similar to that of Born and Green in that only systems in completely definite states are considered. In this case it is necessary to introduce an interaction between the system and the external world. This interaction, provided that it satisfies certain requirements, will ensure the equivalence of time and ensemble averages. However, this method of treatment, in our opinion, does not give due weight to the fact that the observer is dealing with macroscopic variables. For a detailed acquaintance with Klein’s very interesting arguments we refer the reader to his work.

*) It should be noted here that if one uses equation (V,21) for \(\Sigma(t)\) (as Fierz did), then only the condition of nondegeneracy will be necessary, but not the very strong restriction of the absence of resonances. This is very satisfactory, since at a time when energy differences have physical meaning (they enter, for example, into expressions for relaxation times), second differences do not have such physical meaning. I am grateful to Prof. Fierz, who drew my attention to this.

**) See Appendix IV.

SUMMARY OF THE SITUATION IN QUANTUM THERMOSTATICS

Let us briefly summarize the results of the consideration of the two preceding parts (IV and V). There is no need to discuss the utilitarian approach again, since here there is no difference between the classical and quantum-mechanical cases.

In the case of the formalistic approach, the necessary conditions for the equality of time averages and ensemble averages are again investigated. In the quantum-mechanical case, the necessary condition is the nondegeneracy of all eigenstates of the energy.

In the case of the physical approach, we again dealt with representative ensembles and showed that the evolution of a nonequilibrium state will occur in such a way that the ensemble representing the system under consideration will, with the passage of time, approach the canonical ensemble. The problem here reduces to finding rules according to which representative ensembles could be constructed. These rules coincide with the assumptions of equal a priori probabilities for all nondegenerate states and of the randomness of the a priori phases for probability amplitudes. If such assumptions are made, then one may further use either the quantum-mechanical form of the statistical \(H\)-theorem, or use the \(H\)-theorem as applied to ensembles. The consideration here is very similar to that carried out in the classical case.

In the classical case, we have seen that, in the final analysis, the basic ideas of the physical and formalistic approaches are identical. If it is proved that physical systems are quasi-ergodic, then the ergodic theorem follows from this; one may, however, also justify the choice of a priori probabilities on the basis of fundamental principles. In the quantum-mechanical case the situation is not so simple. From fundamental principles one can justify the assumption of equality of the a priori weights, but there still remains the assumption of the randomness of the a priori phases. On the other hand, it is unlikely that physical systems are in fact ergodic in the quantum-mechanical sense, and in that case the quantum-mechanical ergodic theorem must also use an assumption concerning the distribution of phases. Since in quantum mechanics the (initial) phases have no physical meaning, one may say with some justification that in reality it is always necessary to average over these phases. In this case the quantum-mechanical ergodic theorem will always be valid, and the assumption of the randomness of the a priori phases is thereby also justified. As in the classical case, we see that the two different methods of approach are not so very different from one another as might have been thought at first glance.

APPENDICES

I. LORENTZ MODEL *)

In Section 1.1 a model of a physical system possessing the following properties was introduced. Particles of one kind are fixed in space and are randomly distributed with density \(n\) per unit volume. Particles of the second kind move through the lattice, undergoing elastic and isotropic scattering by particles of the first kind; interaction with particles of their own kind is absent. The density of particles of the second kind is equal to \(N\) per unit volume and their speed is equal to \(c\). The collision cross section is denoted by \(\sigma\).

Suppose that phase space, which in this case is the unit sphere, is subdivided into \(2m+1\) finite cells, each of size \(\delta \omega\),

\[ \delta \omega = 4\pi/(2m+1), \tag{App. I,1} \]

and renumber them by an index \(\upsilon\), running through the sequence \(-m, -m+1, \ldots, 0, \ldots, m-1, m\). Denote by \(f_\upsilon\) the number of particles of the second kind (we shall call them electrons by analogy with the Lorentz model of a metal) per unit volume which move in the direction characterized by the \(\upsilon\)-th element of solid angle. The quantities \(f_\upsilon\) satisfy the condition

\[ \sum_\upsilon f_\upsilon = N, \tag{App. I,2} \]

and their equilibrium values satisfy the condition

\[ f_\upsilon^e = N/(2m+1). \tag{App. I,3} \]

In order to characterize the state of the system, introduce the quantity \(\Delta\)

\[ \Delta = \sum_\upsilon (f_\upsilon - f_\upsilon^e)^2. \tag{App. I,4} \]

We shall now be interested in the probability \(w(\Delta)\, d\Delta\) of finding the value \(\Delta\) in the interval \(\Delta, \Delta + d\Delta\). To find the probability distribution of a quantity that is the sum of a series of functions whose probability distributions are known, we shall use Markov’s method (\(^{12\mathrm{M}}\); see also \(^{43\mathrm{C},\,46\mathrm{C}}\)). We give the result without proof (see \(^{43\mathrm{C}}\)). If

\[ \Phi = \sum_k \Phi_k(q_i), \tag{App. I,5} \]

where \(q_i\) are stochastic variables such that

\[ p(q_i)\,\prod dq_i \tag{App. I,6} \]

*) See \(^{54\mathrm{G}}\); cf. also \(^{55\mathrm{G},\,55\mathrm{H}2}\).

there is a probability of finding \(q_1\) inside the interval \(q_1, q_1+dq_1\), then the probability \(w(\Phi)d\Phi\) of finding \(\Phi\) between \(\Phi\) and \(\Phi+d\Phi\) is given by the expression

\[ w(\Phi)d\Phi=(d\Phi/2\pi)\int_0^\infty \exp(-i\rho\Phi)A(\rho)d\rho, \tag{App. I,7} \]

where \(A(\rho)\) is determined from

\[ A(\rho)=\int\cdots\int \exp(i\rho\Sigma\Phi_k)\,p(q_1)\prod dq_i. \tag{App. I,8} \]

In our case expression (App. I,4) is taken instead of expression (App. I,5), and the probability of finding a definite set of values \(f_\nu\) is a compound probability. Using (App. I,2), we find

\[ p(f_\nu)=[N!/\prod f_\nu!]\,(2m+1)^{-N}, \tag{App. I,9} \]

which, for large values of \(f_\nu\) close to \(f_\nu^e\), takes the form*)

\[ p(f_\nu)=[(2m+1)/2\pi N]^m(2m+1)^{1/2}\times \]

\[ \times \exp\{-(2m+1)\Sigma(f_\nu-f_\nu^e)^2/2N\}; \tag{App. I,10} \]

the latter expression is normalized.

For the function \(A(\rho)\) we obtain

\[ A(\rho)=\int\cdots\int p(f_\nu)\exp(-i\rho\Delta)\,df_{-m}\cdots df_m, \tag{App. I,11} \]

where the integral is \(2m\)-fold because, by virtue of (App. I,2), only \(2m\) of the \(f_\nu\) are independent. Since one of the \(f_\nu\) must be excluded (let it be \(f_0\), in order to preserve symmetry), the integration turns out to be rather cumbersome, but, fortunately, feasible, and gives

\[ A(\rho)=[1-2iN\rho/(2m+1)]^{-m}. \tag{App. I,12} \]

Substituting (App. I,12) into (App. I,7), we find, with the aid of contour integration,

\[ w(\Delta)d\Delta=[(2m+1)/2N]^m[\Delta^{m-1}/(m-1)!]\times \]

\[ \times \exp\{-(2m+1)\Delta/2N\}\,d\Delta, \tag{App. I,13} \]

which coincides with (I.3,13).

*) In deriving equation (App. I,10), Stirling’s formula is used,

\[ \ln x! = x\ln x - x + \frac{1}{2}\ln x + \frac{1}{2}\ln 2\pi \]

and the expansion in powers of \((f_\nu-f_\nu^e)/f_\nu^e\).

From (App. I,13) we obtain

\[ \Delta_{\mathrm{av}}=\int \Delta w(\Delta)\,d\Delta=2mN/(2m+1) \tag{App. I,14} \]

and

\[ (\Delta^2)_{\mathrm{av}}=\int \Delta^2 w(\Delta)\,d\Delta =2m(2m+2)N^2/(2m+1)^2, \tag{App. I,15} \]

whence the relations (I.3,16) follow, if in the final result 1 is neglected in comparison with \(2m\), since \(m\gg 1\).

Now we must calculate the probability \(w(\Delta,\Delta')\) that \(\Delta\) changes its value from \(\Delta\) to \(\Delta'\) during the time interval \(\tau\). The change in \(\Delta\) is related to changes in the values \(f_v\) by the relation

\[ \Delta'-\Delta =2\sum_v (f_v-f_v^e)(f'_v-f_v) =2\sum_v f_v(f'_v-f_v), \tag{App. I,16} \]

since from (App. I,2) it follows that \(\sum f'_v=\sum f_v\).

Let \(x_{vv'}\) be equal to the true number of electrons passing from cell \(v\) into cell \(v'\) during the time interval \(\tau\). If we could use the hypothesis on the number of collisions, we would obtain

\[ x_{vv'}=af_v, \tag{App. I,17} \]

where

\[ a=n\sigma c/(2m+1), \tag{App. I,18} \]

and we have also used the assumption of isotropy of scattering [cf. expression (I.1,15)].

We are interested, however, in fluctuations, and instead of expression (App. I,17) we must consider the distribution of the quantities \(x_{vv'}\) about the mean value given by expression (App. I,17). If the distribution of the fixed parts in space is random, we may assume that it is a Bernoulli distribution.

We choose the time interval \(\tau\) so that the probability that any electron undergo more than one collision during the time \(\tau\) is negligibly small, or

\[ \tau \ll 1/n\sigma c. \tag{App. I,19} \]

If \(N\) is chosen so large that \(aN\) is still large in comparison with unity, so that we have the inequalities

\[ N\gg aN\gg 1, \tag{App. I,20} \]

then for \(x_{vv'}\) one may use the Gaussian distribution

\[ p(x_{vv'})=(2\pi af_v)^{-1/2} \exp\left[-(x_{vv'}-af_v)^2/2af_v\right]. \tag{App. I,21} \]

The change in the quantities \(f_v\) during the time interval \(\tau\) is determined by the quantities \(x_{vv'}\), and we have

\[ f'_v-f_v=\sum_{v'}(x_{v'v}-x_{vv'}). \tag{App. I,22} \]

Combining (App. I,16)—(App. I,22), we see that we have returned to the problem leading to \(w(\Delta)\). The expression \(w(\Delta)w(\Delta,\Delta')\) now stands in place of \(\Phi\) in expression (App. I,5), and the \(\Phi_k\) are functions of \(f_v\) and \(x_{vv'}\). From (App. I,7) we now have

\[ w(\Delta)w(\Delta,\Delta') = \]

\[ = \frac{1}{4\pi^2}\iint \exp[-i\sigma\Delta - i\rho(\Delta' - \Delta)] A(\rho,\sigma)\,d\rho\,d\sigma, \tag{App. I,23} \]

where

\[ A(\rho,\sigma)=\int\ldots\int df_{-m}\ldots df_m\,dx_{-m,-m}\ldots dx_{mm}\times \]

\[ \times p(f_v)p(x_{-m,-m})\ldots p(x_{mm})\times \]

\[ \times \exp\left[i\sigma\sum(f_v-f_v^e)^2+2i\rho\sum(f_v^e-f_v)\right]. \tag{App. I,24} \]

Integration with respect to the \((2m+1)^2\) variables \(x_{vv'}\) is carried out without special difficulty after the expressions (App. I,21) and (App. I,22) have been used. Integration with respect to the \(2m\) variables \(f_v\) \((v\ne0)\) is more complicated. We introduce new variables \(\alpha_v\)

\[ \alpha_v=(f_v-f_v^e)/f_v^e \tag{App. I,25} \]

and neglect the cubic terms in the exponent in comparison with the quadratic terms. After this, a quadratic expression remains in the exponent, which can be integrated. The final result has the form

\[ A(\rho,\sigma)=[(2m+1)/2]^m[4Na\rho(2N\rho+2m+1)- \]

\[ -2Ni\sigma+2m+1]^{-m} \tag{App. I,26} \]

and from (App. I,23) and (App. I,13) we obtain the normalized expression

\[ w(\Delta,\Delta') = \]

\[ =(32\pi Na\Delta)^{-1/2}\exp\left[-(\Delta'-\Delta+2(2m+1)a\Delta)^2/32Na\Delta\right], \tag{App. I,27} \]

which coincides with expression (I.3,19).

Expression (I.3,21) follows from the relation

\[ \Delta'_{\mathrm{cp}}=\int \Delta' w(\Delta,\Delta')\,d\Delta' \tag{App. I,28} \]

when (App. I,27) is used.

In Part I it was mentioned that the following relation holds:

\[ w(\Delta)w(\Delta,\Delta')=w(\Delta')w(\Delta',\Delta). \tag{App. I,29} \]

It is necessary to draw attention to the fact that in proving this relation one neglects terms of higher order in \(a\), which is permissible in view of the inequality (App. I,19).

Using expressions (I.3, 27), (I.3,28), (I.3,31) and (App. I,27), we obtain expressions (I.3,19) and (I.3,32), again with accuracy up to terms of first order in \(a\). Similarly, expression (I.3,30) follows from (I.3,29) upon neglecting terms of order \(a^2\) and of higher order.

If the hypothesis on the number of collisions is valid, then for the rate of change of \(f_v\) one may use expressions (App. I,22) and (App. I,17), which give

\[ df_v/dt=(f'_v-f_v)/\tau=\sum_{v'}(x_{v'v}-x_{vv'})= \]

\[ =(a/\tau)\left[\sum_{v'} f_{v'}-(2m+1)f_v\right], \tag{App. I,30} \]

whence, after using the fact that \(\sum f_{v'}=\sum f^e_{v'}=(2m+1)f^e_v\), equation (I.3,34) follows.

II. PROOF OF THE EXISTENCE OF QUASIERGODIC SYSTEMS *)

Fermi’s proof of the quasiergodicity of canonical normal systems with more than two degrees of freedom consists of two parts. In the first part a proof is given that the energy surfaces \(\mathfrak{S}\) are the only families of surfaces in phase space possessing the property that any orbit issuing from any point of one of the surfaces of the family will always remain on that same surface. In the second part it is shown that if there are two arbitrary regions \(\sigma\) and \(\sigma^*\) on \(\mathfrak{S}\), then there will be orbits passing both through \(\sigma\) and through \(\sigma^*\).

The proof of the first part is rather complicated and proceeds as follows. A canonical normal system is a system in which one can introduce canonically conjugate coordinates \(x_i\) and \(y_i\) such that:

a) the energy does not depend on time;

b) the Hamiltonian \(\mathcal{H}\) of the system can be expanded in a power series in the parameter \(a\),

\[ \mathcal{H}=\mathcal{H}_0+a\mathcal{H}_1+a^2\mathcal{H}_2+\ldots; \tag{App. II,1} \]

c) the Hamiltonian, and consequently all the \(\mathcal{H}_i\), are periodic functions in all \(x_i\). One may, moreover, choose our \(x_i\) so that the period is the same for all \(x_i\) and equal, say, to \(2\pi\);

d) the first term in the expansion (App. II,1), \(\mathcal{H}_0\), does not depend on \(x_i\).

\[ \text{*) See }{}^{23}\mathrm{F}_1. \]

Suppose now that the equations

\[ \Phi(x, y; a)=0 \tag{Pr. II,2} \]

describe a family of surfaces \(\mathfrak{R}_a\), and that the functions \(\Phi\) are single-valued, analytic, and periodic in \(x_i\). We must show that each of the \(\mathfrak{R}_a\) coincides with the energy surface \(\mathfrak{S}\).

Expand \(\Phi\) in powers of \(a\)

\[ \Phi=\Phi_0+a\Phi_1+a^2\Phi_2+\cdots \tag{Pr. II,3} \]

If the \(\mathfrak{R}_a\) are prescribed, then nevertheless, within broad limits, the \(\Phi_i\) may be chosen arbitrarily; in particular, it can be shown*) that, for the given \(\mathfrak{R}_a\), one may choose \(\Phi_0\) independent of \(x_i\). Moreover, the \(\Phi_i\) may be chosen arbitrarily outside \(\mathfrak{R}_0\).

Since the \(\mathfrak{R}_a\) are such surfaces that an orbit issuing from any point of the surface remains on that very same surface, the equation

\[ \{\mathcal{H},\Phi\}=0 \tag{Pr. II,4} \]

must be a consequence of equation (Pr. II,2); here \(\{A,B\}\) denotes the Poisson bracket,

\[ \{A,B\}=\sum_i\bigl[(\partial A/\partial y_i)(\partial B/\partial x_i)-(\partial A/\partial x_i)(\partial B/\partial y_i)\bigr]. \tag{Pr. II,5} \]

From equation (Pr. II,4) and the series (Pr. II,1) and (Pr. II,3) it follows that the following equations must hold:

\[ \{\mathcal{H}_0,\Phi_0\}=0, \tag{Pr. II,6} \]

\[ \{\mathcal{H}_0,\Phi_1\}+\{\mathcal{H}_1,\Phi_0\}=0,\ldots \tag{Pr. II,7} \]

Equation (Pr. II,6) is trivial, since neither \(\mathcal{H}_0\) nor \(\Phi_0\) contains \(x_i\).

Since \(\Phi_1\) and \(\mathcal{H}_1\) are periodic functions in all \(x_i\), one may write

\[ \Phi_1=\sum A_m(y_i)\exp(\theta_m), \tag{Pr. II,8} \]

\[ \mathcal{H}_1=\sum B_m(y_i)\exp(\theta_m), \tag{Pr. II,9} \]

where

\[ \theta_m=\iota(m_1x_1+m_2x_2+\cdots+m_nx_n); \tag{Pr. II,10} \]

here, by means of \(m\), all \(n\) quantities \(m_i\) are denoted briefly, and the summations in (Pr. II,8) and (Pr. II,9) are taken over all \(n\) values \(m_i\).

*) See \({}^{23}\mathrm{F}1\), Section 3; the fact is used that the \(\mathfrak{R}_a\) are surfaces on which the orbit is situated.

From expressions (App. II,7) and (App. II,9) we obtain

\[ \sum \exp \theta_m\left[A_m\sum_i m_i\omega_i-B_m\sum_j m_j\chi_j\right]=0, \tag{App. II,11} \]

or

\[ A_m\sum_i m_i\omega_i=B_m\sum_j m_j\chi_j, \tag{App. II,12} \]

where

\[ \omega_i=\partial \mathscr{H}_0/\partial y_i,\quad \chi_j=\partial \Phi_0/\partial y_j. \tag{App. II,13} \]

If we exclude the possibility that all \(B_m\) may be equal to zero*), then it follows from (App. II,12) that at any point on \(\mathfrak{R}_0\) at which \(\sum m_i\omega_i\) is equal to zero, \(\sum m_j\chi_j\) must also be equal to zero.

Let us now take \(n=3\). The argument is also valid for \(n>3\), but not for \(n=2\) (cf. \(^{23\mathrm{F}1}\)). Generally speaking, of the three ratios \(\omega_2/\omega_1\), \(\omega_3/\omega_1\), \(\omega_3/\omega_2\), at most one ratio will be constant on \(R_0\). We may therefore suppose that the first two ratios vary on \(\mathfrak{R}_0\). In this case \(\mathfrak{R}_0\) is densely covered by points at which \(\omega_2/\omega_1\) takes rational values, so that one can find such \(m_1\) and \(m_2\) that

\[ m_1\omega_1+m_2\omega_2=0. \tag{App. II,14} \]

Further, at these points we have

\[ m_1\chi_1+m_2\chi_2=0, \tag{App. II,15} \]

so that

\[ \omega_2/\omega_1=\chi_2/\chi_1. \tag{App. II,16} \]

Since relation (App. II,16) is valid on a point sequence everywhere dense in \(\mathfrak{R}_0\), it holds identically on \(\mathfrak{R}_0\). The same is also true for the relation

\[ \omega_3/\omega_1=\chi_3/\chi_1. \tag{App. II,17} \]

From relations (App. II,13), (App. II,16), (App. II,17), and the fact that \(\mathscr{H}_0\) and \(\Phi_0\) do not depend on \(x_i\), it follows that on \(\mathfrak{R}_0\)

\[ \Phi_0=\mathscr{H}_0+c_0, \tag{App. II,18} \]

where \(c_0\) is a constant. Since \(\Phi_0\) can be chosen arbitrarily outside \(\mathfrak{R}_0\), we can choose it so that (App. II,18) is satisfied everywhere in phase space.

If our system is not degenerate, then the relation

\[ \sum m_i\omega_i=0 \tag{App. II,19} \]

*) Consideration of this case is given in \(^{92\mathrm{P}}\).

cannot hold identically on \(\mathfrak{R}_0\), and we are entitled to divide (Pr. II,12) by \(\sum m_i\omega_i\), which is equal to \(\sum m_j\chi_j\) by virtue of (Pr. II,18). As a result we have

\[ A_m=B_m, \]

except for the case when all \(m_i=0\). \(\tag{Pr. II,20}\)

We see, further, that on \(\mathfrak{R}_0\)

\[ \Phi_1(x_i,y_i)=\mathcal{H}_1(x_i,y_i)+f_1(y_i), \tag{Pr. II,21} \]

and again \(\Phi_1\) may be chosen outside \(\mathfrak{R}_0\) so that (Pr. II,21) is satisfied everywhere in phase space.

It is now easy to prove, by the method of complete induction, starting from equations (Pr. II,18) and (Pr. II,21), that

\[ \Phi_i=\mathcal{H}_i+c_i, \tag{Pr. II,22} \]

where \(c_i\) are constants. The proof of this is analogous to the proof of equations (Pr. II,18) and (Pr. II,21).

Equations (Pr. II,22) lead to the conclusion that (Pr. II,2) is equivalent to the equation

\[ \mathcal{H}(x_i,y_i;\alpha)=\sum c_n\alpha^n=c, \tag{Pr. II,23} \]

and, consequently, \(\mathfrak{R}_\alpha\) coincide with the energy surfaces.

The second part of Fermi’s proof is much less complicated. Let \(\sigma\) and \(\sigma^*\) be two arbitrary regions on \(\mathfrak{S}\). Let \(\sigma'\) be that part of \(\mathfrak{S}\) which is covered by the orbit issuing from somewhere in \(\mathfrak{S}\). If \(\sigma'\) covered all of \(\mathfrak{S}\), our proof would be completed. If \(\sigma'\) does not cover all of \(\mathfrak{S}\), then let \(\sigma''\) be that part of \(\mathfrak{S}\) which is not covered, and let \(\mathfrak{B}\) be the boundary surface between \(\sigma'\) and \(\sigma''\). One can verify that there cannot exist orbits having points both in \(\sigma'\) and in \(\sigma''\). Indeed, suppose \(P'\) in \(\sigma'\) is traversed at the moment \(t'\) by an orbit which at the moment \(t''\) traversed \(P''\) in \(\sigma''\). Since the solutions of the equations of motion are analytic functions, one can find such small regions \(\eta'\) in \(\sigma'\) around \(P'\) and \(\eta''\) in \(\sigma''\) around \(P''\) that an orbit passing through an arbitrary point \(Q'\) in \(\eta'\) at the moment \(t'\) would pass through \(\eta''\) at the moment \(t''\). Since, however, \(\eta'\) lies in \(\sigma'\), there will be found points \(Q'\) lying on the orbit issuing from \(\sigma\), and we have arrived at a contradiction.

Let now \(P\) be a point on \(\mathfrak{B}\), and let again \(P'\) and \(P''\) be points in \(\sigma'\) and \(\sigma''\), respectively. If we consider the orbits passing at the moment \(t\) through \(P, P', P''\), then these orbits at the moment \(t_1\) will pass through the points \(P_1, P'_1, P''_1\), at the moment \(t_2\) through \(P_2, P'_2, P''_2\), etc., where \(P'_1, P'_2,\ldots\) all lie in \(\sigma'\), and \(P''_1, P''_2,\ldots\) all lie in \(\sigma''\). One can choose \(P'\) and \(P''\) so close to \(P\) that \(P'_1\) and \(P''_1\) would be close to \(P_1\), the points \(P'_2\) and \(P''_2\) close to \(P_2\), etc.

It follows further from this that \(P, P_1, P_2,\ldots\) must all lie on \(\mathfrak B\), and that \(\mathfrak B\) thus contains an orbit. We see, however, that the only surface on which an orbit can be found is the energy surface \(\mathfrak S\). Consequently, \(\mathfrak B\) cannot exist, or else \(\sigma'\) must cover \(\mathfrak S\) completely. This means that \(\sigma^*\) will be a part of \(\sigma'\), and that there will be orbits issuing from \(\sigma\) and passing through \(\sigma^*\). Since both \(\sigma\) and \(\sigma^*\) can be chosen arbitrarily small, we have obtained a proof of the quasi-ergodicity of the systems under consideration.

III. PROOF OF THE CLASSICAL ERGODIC THEOREM*)

Birkhoff’s ergodic theorem shows the equivalence of the mean taken over the energy surface and the time mean taken practically over all orbits on the energy surface, under the condition that the energy surface is metrically indecomposable. Let \(f(P)\) be a phase function with sufficiently good behavior**), i.e., a function of the phase represented by the point \(P\) on the energy surface \(\mathfrak S\).

Consider the function \(\bar f(P,t_0,T)\), defined by the expression

\[ \bar f(P;t_0,T)=\frac{1}{T}\int_{t_0}^{t_0+T} f(P_t)\,dt, \tag{Pr. III,1} \]

where \(P_t\) is the point, at time \(t\), of the orbit passing at time \(t_0\) through \(P\). First of all we shall prove that, for almost all orbits, the following limit exists:

\[ \bar f(P;t_0,\infty)=\lim_{T\to\infty}\bar f(P;t_0,T). \tag{Pr. III,2} \]

It will then be shown that this limit does not depend on \(t_0\). Finally, we shall show that this limit is constant almost everywhere on \(\mathfrak S\).

To prove the first part, we subdivide the time scale into finite intervals and write

\[ T=n\tau, \tag{Pr. III,3} \]

\[ \bar f_n(P;t_0)=\bar f_0(P;t_0,n\tau). \tag{Pr. III,4} \]

If \(P\) were a phase point for which the limit (Pr.III,2) did not exist, then the lower bound \(L(P)\) of the quantities \(\bar f_n(P;t_0)\) and the upper bound \(U(P)\) of the quantities \(\bar f_n(P;t_0)\) would be different, and it would be possible—

*) The proof presented here was given by Kolmogorov (see \(^{25R,49K}\)).

**) Good behavior is understood here in the sense of summability on the energy surface, which is assumed to have finite volume.

but it would be possible to find two such quantities \(\alpha\) and \(\beta\) that

\[ L(P)<\alpha<\beta<U(P). \tag{Pr. III,5} \]

Moreover, if the sequence \(\mathfrak D\) of phase points for which the limit (Pr. III,2) does not exist were of positive measure, then one could find a subsequence \(\mathfrak D'\) of the sequence \(\mathfrak D\), also of positive measure, and such a sequence of values \(\alpha\) and \(\beta\) that the inequalities (Pr. III,5) would be satisfied for all points \(P\) in \(\mathfrak D'\). This, however, leads to a contradiction, as is seen from the following.

Let \(P_k\) be the phase point at the moment \(t_0+k\tau\), and let \(x_k(P)\) be defined by the expression

\[ x_k(P)=\frac{1}{\tau}\int_{t_k}^{t_{k+1}} f(P_t)\,dt. \tag{Pr. III,6} \]

By shifting the origin of time, we see that

\[ x_k(P)=x_0(P_k). \tag{Pr. III,7} \]

The time average \(\overline f_n(P;t_0)\) can be expressed in terms of \(x_k\) as follows:

\[ \overline f_n(P;t_0)=\frac{1}{n}\sum_{k=0}^{n-1} x_k(P). \tag{Pr. III,8} \]

Let us now integrate \(\overline f_n(P;t_0)\) over the sequence of points \(\mathfrak D^{(n)}\), which is a subsequence of \(\mathfrak D'\) such that for any point in \(\mathfrak D^{(n)}\) we have

\[ \overline f_n(P;t_0)>\beta . \tag{Pr. III,9} \]

The result of the integration has the form

\[ n\beta\,\mathfrak M(\mathfrak D_0^{(n)})<n\int \overline f_n(P;t_0)\,d\omega_0= \]

\[ =\sum_{k=0}^{n-1}\int x_k(P)\,d\omega_0 =\sum_{k=0}^{n-1}\int x_0(P_k)\,d\omega_k, \tag{Pr. III,10} \]

where \(\mathfrak M(\mathfrak D_0^{(n)})\) is the measure of the sequence \(\mathfrak D_0^{(n)}\), \(d\omega^{(n)}\) denotes integration over the sequence \(D_0^{(n)}\), and \(d\omega_k^{(n)}\) denotes integration over the sequence \(\mathfrak D_k^{(n)}\), obtained from \(D_0^{(n)}\) by the transformation \(P\to P_k\).

Suppose that the sequences \(\mathfrak{D}_{k}^{(n)}\) do not overlap, and let their sum be \(\mathfrak{D}^{(n)}\):

\[ \mathfrak{D}^{(n)}=\sum_k \mathfrak{D}_{k}^{(n)}; \tag{Pr. III,11} \]

from (Pr. III,10) we have

\[ \int x_0(P)\,d\omega>\beta\mathfrak{M}\bigl(\mathfrak{D}^{(n)}\bigr), \tag{Pr. III,12} \]

where \(\int d\omega^{(n)}\) denotes integration over \(\mathfrak{D}^{(n)}\), and use has been made of the fact that \(\mathfrak{M}(\mathfrak{D}_{k}^{(n)})=\mathfrak{M}(\mathfrak{D}_{0}^{(n)})\) for all \(k\), in view of Liouville’s theorem.

It can be shown that such sums \(\mathfrak{D}^{(n)}\) of nonoverlapping sequences can be found for each value of \(n\) in such a way that they exhaust \(\mathfrak{D}'\); from the inequality (Pr. III,12) it will then follow that

\[ \int x_0(P)\,d\omega' > \beta\mathfrak{M}(\mathfrak{D}'), \tag{Pr. III,13} \]

where \(\int d\omega'\) denotes integration over all of \(\mathfrak{D}'\). Similarly one can prove the inequality

\[ \int x_0(P)\,d\omega' < \alpha\mathfrak{M}(\mathfrak{D}'). \tag{Pr. III,14} \]

Combining the inequalities (Pr. III,13) and (Pr. III,14) with the assumption \(\mathfrak{M}(\mathfrak{D}')>0\), we arrive at a contradiction with our choice of \(\alpha<\beta\). We thus obtain that \(\mathfrak{D}\) has measure zero and that the limit (Pr. III,2) exists for practically all orbits.

The first part of the proof is completed by considering the quantity

\[ A=\left|\frac{1}{T}\int_{t_0}^{t_0+T} f(P_t)\,dt-\frac{1}{n\tau}\int_{t_0}^{t_0+n\tau} f(P_t)\,dt\right|, \tag{Pr. III,15} \]

where now (Pr. III,3) need not necessarily be satisfied, but \(n\) is the largest of the integers contained in \(T/\tau\). Split \(A\) into two parts

\[ A=A_1+A_2, \tag{Pr. III,16} \]

where

\[ A_1=\left|\left(\frac{1}{T}-\frac{1}{n\tau}\right)\int_{t_0}^{t_0+n\tau} f(P_t)\,dt\right|, \tag{Pr. III,17} \]

\[ A_2=\left|\frac{1}{T}\int_{t_0}^{t_0+T} f(P_t)\,dt-\frac{1}{T}\int_{t_0}^{t_0+n\tau} f(P_t)\,dt\right|. \tag{Pr. III,18} \]

Since \(\bar f_n(P;t_0)\) has a limiting value as \(n\to\infty\) (let us call it \(F\)), for \(A_1\) we have the limit

\[ A_1 \to \frac{n\tau - T}{T}\,F, \tag{Pr. III,19} \]

and we see that \(A_1 \to 0\) as \(T\to\infty\).

For \(A_2\) we have the inequality

\[ A_2=\frac{1}{T}\left|\int_{t_0+n\tau}^{t_0+T} f(P_t)\,dt\right| \leq \frac{\tau}{T}\left|x_n(P)\right| \tag{Pr. III,20} \]

and for any sufficiently good function \(f(P)\) the quantities \(x_n(P)\) will be bounded; thus \(A_2\to 0\) as \(T\to\infty\). This completes the first part of the argument.

The fact that \(\bar f(P;t_0,\infty)\) does not depend on \(t_0\) follows from the equalities

\[ \lim \frac{1}{T}\int_{t_0}^{t_0+T} f(P_t)\,dt = \lim \frac{1}{T}\int_{t_1}^{t_1+T} f(P_t)\,dt \]

\[ -\lim \frac{1}{T}\int_{t_0}^{t_1} f(P_t)\,dt - \lim \frac{1}{T}\int_{t_0+T}^{t_1+T} f(P_t)\,dt \tag{Pr. III,21} \]

and, for any sufficiently good function \(f(P)\), the last two limits will be equal to zero.

The final part of the proof of Birkhoff’s ergodic theorem follows from the fact that, if \(\bar f(P;t_0,\infty)\) were not constant practically everywhere on \(\mathfrak S\), then it would be possible to find such a value \(F\) of the quantity \(\bar f(P;t_0,\infty)\) that the conditions \(\bar f(P;t_0,\infty)<F\) and \(\bar f(P;t_0,\infty)>F\) would determine two sets of positive measure on \(\mathfrak S\), which would be invariant with respect to the transformations \(P\to P_t\), for \(\bar f(P;t_0,\infty)\) is invariant with respect to such transformations. This, however, would be in contradiction with the metric indecomposability of \(\mathfrak S\) assumed above.

IV. PROOF OF THE QUANTUM-MECHANICAL ERGODIC THEOREM\(^*\)

In part V it was established that the quantum-mechanical ergodic theorem will be proved if only it can be shown that the time average \(\bar P_\nu\) of the probability \(P_\nu\) of finding the system in the \(\nu\)-th cell prac-

\(^*\) See \(^{2*)N,37P,52R}\); I am grateful to Prof. M. Fierz for information on this question.

tically always equal to the ratio of the number of states \(s_\nu\) in the cell to the number of states \(S_i\) of the corresponding energy layer. In the discussion we used the fact that both \(s_\nu\) and \(S_i\) are quantities of the order of \(\exp(10^{20})\); this fact will also be used below. The expression “practically always” is understood in the sense that relation (V,26) holds for practically all subdivisions of phase space into phase cells; moreover, the weights of the different subdivisions will be determined below.

Let \(\psi\) be the wave function of our system, and let \(\varphi_\sigma\) be a complete orthonormal system corresponding to the Hamiltonian operator \(H\); and let \(\omega_\tau\) be a complete orthonormal system corresponding to the macroscopic operator considered in Part V. The sequence \(\omega_\tau\) will depend on the manner in which our phase cells are chosen. We may express \(\psi\) through \(\varphi_\sigma\) or through \(\omega_\tau\):

\[ \psi=\sum_\sigma r_\sigma \varphi_\sigma, \tag{Pr. IV,1} \]

\[ \psi=\sum_\tau t_\tau \omega_\tau . \tag{Pr. IV,2} \]

The quantities \(\varphi_\sigma\) and \(\omega_\tau\) are connected with one another by a unitary transformation

\[ \omega_\tau=\sum_\sigma U_{\tau\sigma}\varphi_\sigma,\qquad \varphi_\sigma=\sum_\tau U^*_{\tau\sigma}\omega_\tau, \tag{Pr. IV,3} \]

where the different methods of choosing the phase cells are reflected in different matrices \(U\).

From expressions (Pr. IV,1)—(Pr. IV,3) it follows that

\[ t_\tau=\sum_\sigma r_\sigma U_{\tau\sigma}, \tag{Pr. IV,4} \]

and since the \(\varphi_\sigma\) are eigenfunctions of the energy operator, we have

\[ r_\sigma=|r_\sigma|\exp(iE_\sigma t/\hbar). \tag{Pr. IV,5} \]

The quantity \(P_\nu\) is given by the expression

\[ P_\nu=\sum |t_\tau|^2, \tag{Pr. IV,6} \]

where the summation extends over the \(s_\nu\) levels of the \(\nu\)-th cell. From expressions (Pr. IV,4)—(Pr. IV,6) we obtain

\[ P_\nu=\sum_{\tau\sigma\rho}|r_\sigma|\,|r_\rho|\,U^*_{\tau\sigma}U_{\tau\rho} \exp[i(E_\sigma-E_\rho)t/\hbar], \tag{Pr. IV,7} \]

and for the time average

\[ \overline{P}_\nu=\sum_{\tau\sigma}|r_\sigma|^2|U_{\tau\sigma}|^2, \tag{Pr. IV,8} \]

where we have used the circumstance that no two energy levels coincide*).

*) It should be noted that the additional condition of the absence of resonances of degeneracy is not necessary.

Let us now consider

\[ \left|\bar P_\nu - s_\nu/S_i\right| = \left|\sum_\sigma |r_\sigma|^2(C_\sigma - s_\nu/S_i)\right|, \tag{App. IV,9} \]

where

\[ C_\sigma=\sum_\tau |U_{\tau\sigma}|^2, \tag{App. IV,10} \]

and we have used the fact that

\[ \sum_\sigma |r_\sigma|^2=1. \tag{App. IV,11} \]

From (App. IV,9) and (App. IV,11) it follows that

\[ \left|\bar P_\nu - s_\nu/S_i\right| < \max |C_\nu - s_\nu/S_i|. \tag{App. IV,12} \]

If we can show that, practically for all transformation matrices \(U\), the \(C_\sigma\) are practically equal to \(s_\nu/S_i\), then our theorem will be proved. We must, therefore, find the probability distribution of the quantity \(C_\sigma\), given by the expression (App. IV,10). The quantities \(U_{\sigma\tau}\) are the components of unitary unit vectors in Hilbert space, and the \(C_\sigma\) are the squares of the lengths of the projections of these vectors onto the subspace corresponding to the \(s_\nu\) states of the cell under consideration. If we now assume that the probability of finding subdivisions corresponding to unitary unit vectors \(U\) lying inside a given solid angle in Hilbert space is proportional to the measure of this solid angle (or to the area of the surface cut out on the unit sphere in Hilbert space), then the determination of the probability \(W(C)\,dC\) that \(C_\sigma\) lies between \(C\) and \(C+dC\) reduces to a problem of multidimensional geometry. Similar problems were considered by Neumann \(^{29N}\) and by Pauli and Fierz \(^{37P}\), and we shall give here only the final result for \(W(C)\):

\[ W(C)=k\cdot C^{s_\nu}(1-C)^{S_i-s_\nu}, \tag{App. IV,13} \]

where \(k\) is a normalization constant; the relation (V,20) has been used.

From (App. IV,13) it is clear, first, that the most probable value for \(C_\sigma\) is \(s_\nu/S_i\), and second, that the maximum is extremely sharp. Indeed, the function \(W(C)\) decreases by a factor of two at a distance \(S^{-1/2}\) from its maximum. It follows from this that for “practically all” subdivisions \(C_\sigma\) will have a value “practically” equal to \(s_\nu/S_i\), and that, consequently, relation (V,26) is satisfied “almost everywhere.”

V. TRANSITIONS PROPORTIONAL TO TIME *)

In this appendix we wish to outline the derivation of expressions (IV.2,45) and (IV.1,30), which played so important a role in our consideration of the quantum-mechanical \(H\)-theorem.

*) Cf. \(^{38T}\), Sections 99, 100.

Let us recall first of all that in the usual quantum-mechanical theory of perturbations, from the method of variation of constants (26D, 27D, 28P; 38K, Section 53) it follows that if \(a_k(0)\) are the amplitudes at the time \(t=0\), then at the subsequent time \(t\) the amplitudes are given by the formula

\[ a_n(t)=\sum_k V_{kn}\{\exp[i(E_n-E_k)t/\hbar]-1\}a_k(0)/(E_k-E_n), \tag{Pr. V,1} \]

where the summation is extended over all those states that were represented at the time \(t=0\), while the state \(n\) was not represented at the time \(t=0\); here \(V_{kn}\) is the matrix element of the perturbation operator \(\mathbf V\), by virtue of which transitions take place (cf. the discussion of Bohr’s and Dirac’s treatments in Section IV.2); \(E_k\) is the energy value of state \(k\).

Let us now consider the case when we want to know the number of transitions from one group \(S_i\) of energy levels to another group \(S_j\). Suppose first of all that at the time \(t=0\) we know from observation that the system is in one of the \(S_i\) states of the \(i\)-th group. In accordance with our basic assumption about the equality of a priori probabilities and the randomness of a priori phases, at the time \(t=0\) for the density matrix representing our system we have the expressions

\[ \left. \begin{aligned} \rho_{kl}&=\langle a_k a_l^*\rangle=\hat{\delta}_{kl}/S_i,\quad \text{if state } k\\ &\hspace{2.2cm}\text{belongs to the } i\text{-th group;}\\ \rho_{kl}&=0\quad \text{in other cases.} \end{aligned} \right\} \tag{Pr. V,2} \]

These expressions follow from the fact that if both \(k\) and \(l\) belong to the \(i\)-th group and are not equal to one another, then \(\langle a_k a_l^*\rangle=0\), if one averages over phases; the rest is obtained quite simply*).

The probability \(P'_j\) of finding the system at time \(t\) in a state of the \(j\)-th group follows from expression (Pr. V,1):

\[ P'_j=\sum_n a_n^*(t)a_n(t)=\sum_{k,l,n} V_{kn}^*V_{ln}\{\exp[-i(E_k-E_n)t/\hbar]-1\}\times \]

\[ \times \{\exp[i(E_l-E_n)t/\hbar]-1\}\times \]

\[ \times a_k^*(0)a_l(0)/(E_n-E_k)(E_n-E_l). \tag{Pr. V,3} \]

Averaging over the representing ensemble and using (Pr. V,2), we obtain for \(P_j\) the expression

\[ P_j(t)= \]

\[ =(4/S_i)\sum_{n,k}|V_{kn}|^2\left[\sin^2\frac{1}{2}\times(E_k-E_n)t/\hbar\right](E_k-E_n)^2. \tag{Pr. V,4} \]

*) It is taken into account that \(\rho\) is normalized.

In the expressions (App. V,3) and (App. V,4), the states \(k\) and \(l\) belong to the \(i\)-th group, and the states \(n\) to the \(j\)-th group. To compute the sum in (App. V,4), we replace it by a double integral and obtain (cf. \({}^{38\text{T}}\), section 99)

\[ P_j(t)=(2\pi/\hbar S_i)\sum_{k,n}|V_{kn}|^2\sigma_n(E)\sigma_k(E)\Delta E t, \tag{App. V,5} \]

where \(\sigma_n(\sigma_k)\) are the densities of energy levels in the \(j\)-th (\(i\)-th) group, and \(\Delta E\) is the energy interval corresponding to the \(j\)-th group.

Expression (App. V,5) can be rewritten in the form

\[ P_j=T_{ij}t/S_i, \tag{App. V,6} \]

whereas, if we were interested in the case when the initial observation showed that the system is in the \(j\)-th group and we wished to know the probability of finding the system in a state of the \(i\)-th group, we would obtain the result

\[ P_i=T_{ji}t/S_j, \tag{App. V,7} \]

and from (App. V,5) and the Hermitian character of \(\mathbf V\) it also follows that

\[ T_{ij}=T_{ji}. \tag{App. V,8} \]

If the observation showed only that at the moment \(t=0\) the different groups are occupied with probabilities \(P_i, P_j,\ldots\), then instead of (App. V,2) it is necessary to use the expression for \(\rho\)

\[ \rho_{kl}=\delta_{kl}P_i/S_i \]

(state \(k\) belongs to the \(i\)-th group),
\[ \tag{App. V,9} \]

and, by means of arguments analogous to those just given, we would find that the intensity of transition \(N_{ij}\) from the \(i\)-th group to the \(j\)-th group is given by the expression

\[ N_{ij}=T_{ij}P_i/S_i, \tag{App. V,10} \]

which reduces to (IV. 2,45) upon the substitution

\[ A_{ij}=T_{ij}/S_iS_j, \tag{App. V,11} \]

and from (App. V,8) and (App. V,11) expression (IV. 2,46) follows.

To derive expression (IV.1,30), let us recall that we are dealing with a system containing \(N\) practically independent particles, so that the Hamiltonian of the system can be written in the form

\[ \mathbf H=\mathbf H_0+\mathbf V, \tag{App. V,12} \]

where

\[ \mathbf H_0=\sum_i \mathbf H_i, \tag{App. V,13} \]

\(\mathbf H_i\) is the Hamiltonian of the \(i\)-th particle, \(\mathbf V\) is the interaction operator,

which is neglected in the first approximation, but which is necessary in order that transitions between different states of the system can occur*). The eigenfunctions of \(H_0\) have the form

\[ \Phi_k=\Pi_i\varphi_i,\qquad E_k=\sum_i\varepsilon_i, \tag{Pr. V,14} \]

where \(\varphi_i\) are the eigenfunctions of \(H_i\), \(E_k\) is the eigenvalue of the operator \(H_0\) corresponding to \(\Phi_k\), and \(\varepsilon_i\) are the eigenvalues of \(H_i\). The wave function \(P\Phi_k\), obtained from \(\Phi_k\) by permuting the arguments of the \(N\) particles, is also an eigenfunction of \(H_0\) belonging to \(E_k\).

If we are dealing with a system of Fermi–Dirac particles, then only those wave functions are admissible which are antisymmetric in the arguments of all \(N\) particles; consequently, in this case there is only one admissible combination \(P\Phi_k\), namely

\[ \Phi_{\mathrm{F.-D.}}=C\sum(-)^P P\Phi_k, \tag{Pr. V,15} \]

where \((-)^P\) is equal to \(+1\) or \(-1\) according as the permutation is even or odd.

In the case of systems of Bose–Einstein particles the single admissible wave function is completely symmetric, i.e.,

\[ \Phi_{\mathrm{B.-E.}}=C\sum P\Phi_k, \tag{Pr. V,16} \]

where in (Pr. V,15) and (Pr. V,16) the summation is over all \(N!\) permutations, while the factors \(C\) in both cases are normalization constants.

Below we shall confine ourselves to systems of Bose–Einstein particles. The case of Fermi–Dirac particles is simpler and is treated analogously. Let us assume first of all that \(V\) has the form

\[ V=\sum V_{\alpha\beta}, \tag{Pr. V,17} \]

where \(\alpha\) and \(\beta\) number the \(N\) particles, the summation extending over all pairs of the system \((\alpha<\beta)\). This means that only binary collisions are assumed to be essential, while triple and higher-order collisions may be neglected. Such an assumption appears natural**), since we

*) Cf. the analogous situation in the case of Boltzmann’s \(H\)-theorem, where the Maxwell–Boltzmann distribution, which actually presupposes independent particles, is established by means of collisions, i.e., by a path of interactions.

**) There are very few cases in which triple collisions have been taken into account. They enter into the calculation of the third virial coefficient \(^{39B,40B,41M}\), but up to the present time only the classical case has been studied; the difficulties expected in the quantum-mechanical case still appear to be scarcely surmountable.

we have already assumed that the system is so rarefied that, in the first approximation, \(\mathbf{V}\) may be neglected.

The wave function \(\Phi_k\) may contain \(n_1\) factors \(\varphi_1\), \(n_2\) factors \(\varphi_2,\ldots,n_i\) factors \(\varphi_i,\ldots\), where \(\varphi_i\) is a system of orthonormal eigenfunctions. To compute the factor \(C\) in (App. V, 16) it is necessary to take into account only those permutations which do not lead to a product of orthogonal \(\varphi_i\)’s with identical arguments*), so that for \(C\) one ultimately obtains the expression

\[ |C|^2=\left[N!n_1!n_2!n_3!\ldots n_i!\ldots\right]^{-1}. \tag{App. V, 18} \]

In the same way one can compute the matrix elements \(V_{kl}\) corresponding to the transition from the initial state to a state in which, instead of \(n_i\) quantities \(\varphi_i\), \(n_j\) quantities \(\varphi_j\), \(n_{i'}\) quantities \(\varphi_{i'}\), and \(n_{j'}\) quantities \(\varphi_{j'}\), there are \(n_i+1\) quantities \(\varphi_i\), \(n_j+1\) quantities \(\varphi_j\), \(n_{i'}-1\) quantities \(\varphi_{i'}\), and \(n_{j'}-1\) quantities \(\varphi_{j'}\). We introduce two integrals

\[ I_1=\int \varphi_i^*(\alpha)\varphi_j^*(\beta)V_{\alpha\beta}\varphi_{i'}(\alpha)\varphi_{j'}(\beta)\,d\tau_\alpha\,d\tau_\beta, \tag{App. V, 19} \]

\[ I_2=\int \varphi_i^*(\beta)\varphi_j^*(\alpha)V_{\alpha\beta}\varphi_{i'}(\alpha)\varphi_{j'}(\beta)\,d\tau_\alpha\,d\tau_\beta; \tag{App. V, 20} \]

we find

\[ V_{kl}=|I_1+I_2|^2(n_i+1)(n_j+1)n_{i'}n_{j'}. \tag{App. V, 21} \]

In the Fermi–Dirac case, instead of (App. V, 21) we would find the expression**)

\[ V_{kl}=|I_1-I_2|^2n_{i'}n_{j'}(n_i-1)(n_j-1). \tag{App. V, 22} \]

Up to now we have not introduced the groups of energy levels considered in Section IV.1. Let us now introduce them and consider the number of transitions, occurring per unit time, from \(Z_i\) and \(Z_j\) to \(Z_{i'}\) and \(Z_{j'}\), where, consequently, \(N_i,N_j,N_{i'},N_{j'}\) change to \(N_i-1,N_j-1,N_{i'}+1,N_{j'}+1\). Again we use the method of variation of constants, the consideration being very similar to that which led to expression (App. V, 6), so that it will only be outlined.

The probability \(P_f(t)\)***) of finding the system at time \(t\) in one of the states corresponding to \(N_i-1\), \(N_j-1\), \(N_{i'}+1\), \(N_{j'}+1\), if the system was initially in one of the states corre-

* ) We have only sketched the argument here and refer, for a more detailed derivation, to Tolman \({}^{38\mathrm{T}}\), Section 100.

** ) This formula can be found in the papers of Jordan \({}^{25\mathrm{J},\,27\mathrm{J}}\), Ornstein and Kramers \({}^{27\mathrm{O}}\), and Bose \({}^{28\mathrm{B}}\); the further development is due to Tolman.

*** ) The index \(f\) denotes the final state, while the indices \(o\) and \(o'\) denote the initial states.

corresponding \(N_i,\ N_j,\ N_{i'},\ N_{j'}\), is given by the expression

\[ P_{\mathrm f}=\sum_{\mathrm f}|a_{\mathrm f}(t)|^2 =\sum_{\mathrm f oo'}V^{*}_{\mathrm{fo}}V_{\mathrm{fo'}}\rho_{\mathrm{oo'}}(0) \{\exp[-i(E_{\mathrm f}-E_{\mathrm o})t/\hbar]-1\}\times \]

\[ \times\{\exp[-i(E_{\mathrm f}-E_{\mathrm{o'}})t/\hbar]-1\}/(E_{\mathrm{o'}}-E_{\mathrm f})(E_{\mathrm o}-E_{\mathrm f}), \tag{App. V,23} \]

where \(\rho_{\mathrm{oo'}}(0)\) is the density matrix at \(t=0\),

\[ \rho_{\mathrm{oo'}}(0)=a^{*}_{\mathrm{o'}}(0)a_{\mathrm o}(0), \tag{App. V,24} \]

and \(a_{\mathrm f}\) and \(a_{\mathrm o}\) are again probability amplitudes.

Let us suppose again that we made a measurement at the moment \(t=0\), which gave us information that at \(t=0\) one of the initial states had been realized, so that for the density matrix one may write

\[ \left. \begin{aligned} \rho_{\mathrm{oo'}}(0)&=(1/g_0)\delta_{\mathrm{o'o}},\qquad &&\text{if the state \(\mathrm o\) satisfies our requirements, and}\\ \rho_{\mathrm{oo'}}&=0 &&\text{otherwise;} \end{aligned} \right\} \tag{App. V,25} \]

here \(g_0\) is the number of states satisfying our requirements.

In (App. V,23) one may now substitute the expressions (App. V,21) and (App. V,22) for \(V_{kl}\) and write, for \(E_{\mathrm f}-E_{\mathrm o}\),

\[ E_{\mathrm f}-E_{\mathrm o}=\varepsilon_i+\varepsilon_j-\varepsilon_{i'}-\varepsilon_{j'}. \tag{App. V,26} \]

Writing, further,

\[ \left. \begin{aligned} n_i&=N_i/Z_i,\qquad &n_j&=N_j/Z_j,\qquad &n_{i'}&=N_{i'}/Z_{i'},\\ &&n_{j'}&=N_{j'}/Z_{j'}, \end{aligned} \right\} \tag{App. V,27} \]

i.e., replacing \(n_i,\ldots\) by their mean values in the corresponding group and integrating, we finally obtain for \(P_{\mathrm f}(t)\) the expression

\[ P_{\mathrm f}(t)=(2\pi/\hbar)[\,|I_1\pm I_2|^2/\Delta E\,]\,N_iN_j\times \]

\[ \times(Z_{i'}\pm N_{i'})(Z_{j'}\pm N_{j'})\,t, \tag{App. V,28} \]

where \(\Delta E\) is the width of the energy group. From (App. V,28) the expression (IV.1,30) is easily obtained, while the expression (IV.1,31) is now a consequence of the Hermitian character of the operator \(\mathbf V\).

Thus, we have derived the expressions (IV.1,30) and (IV.1,45), which played such an important role in the consideration of the quantum-mechanical \(H\)-theorem. We used: (1) the quantum-mechanical representation of time-dependent transitions and (2) the statistical assumption of the equality of a priori probabilities and the randomness of a priori phases. The second point was discussed in Section IV.3; as for the first point, we would like to conclude the present Appendix with a quotation from the classic paper of Pauli \(^{28}\): “Statisti-

ical laws for the frequencies of transitions between stationary states, having the same character as the laws of radioactive decay, by themselves guarantee exactly as much randomness as is necessary for the statistical interpretation of the second law of thermodynamics”*).

VI. KLEIN’S LEMMA**

In Sec. IV.2 the fact was used that the expression (VI.2,56) can never become positive. This property will be proved here. We have two density matrices \(\rho'\) and \(\rho''\), representing the system at the moments \(t'\) and \(t''\), with matrix elements \(\rho'_{kl}\) and \(\rho''_{kl}\), where

\[ \rho'_{kl}=\rho'_{kl}\delta_{kl}, \tag{Pr. VI,1} \]

however, \(\rho''\) is not necessarily a diagonal matrix. The diagonal elements \(\rho'_{kk}\) \((\rho''_{kk})\) are, as we saw in Sec. IV.2, the probabilities of finding the system at the moment \(t\) \((t'')\) in the \(k\)-th state. They, consequently, cannot be negative, and the expression

\[ Q_{kn}=\rho'_{kk}\left[\ln \rho'_{kk}-\ln \rho''_{nn}-1\right]+\rho''_{nn} \tag{Pr. VI,2} \]

likewise cannot be negative in view of relation (IV.2,58). The quantities \(\rho''_{kl}\) are determined, with the aid of (IV.2,7), through the probability amplitudes at the moment \(t''\), which in turn are obtained from the probability amplitudes at the moment \(t'\) by integration of the Schrödinger equation. It is known that the probability amplitudes \(a_n^k(t'')\) at the moment \(t''\) can be obtained from the amplitudes \(a_n^k(t')\) at the moment \(t'\) by a unitary transformation \(U\):

\[ a_n^k(t'')=\sum_l a_l^k(t')U_{nl}. \tag{Pr. VI,3} \]

Using (Pr. VI,3), we obtain from (IV.2,7)

\[ \rho''_{nn}=\sum_k |U_{nk}|^2\rho'_{kk}+\sum_{l\ne k}U^*_{nl}U_{nk}\rho'_{kl}, \tag{Pr. VI,4} \]

or, using (Pr. VI,1),

\[ \rho''_{nn}=\sum_k |U_{nk}|^2\rho'_{kk}. \tag{Pr. VI,5} \]

* After this work had been written, a paper by van Kampen \(^{54K}\) was published, in which the symmetry of the transition matrix was studied. Van Kampen drew special attention to the fact that if magnetic fields or Coriolis forces are present, this symmetry must be replaced by a somewhat weaker condition that does not follow from the Hermiticity of the Hamiltonian. In his paper van Kampen also considered the connection between Onsager’s relations \(^{31O1,31O2,45C}\) and the symmetry of the transition matrix.

** See \(^{31K,38T}\), Sections 101 and 106.

Multiplying (App. VI,2) by the positive quantity \(|U_{nk}|^2\) and summing over all values of \(n\) and \(k\), we obtain

\[ \sum_{k,n}|U_{nk}|^2\rho'_{kk}\ln\rho'_{kk} -\sum_{k,n}|U_{nk}|^2\rho'_{kk}\ln\rho''_{nn} -\sum_{k,n}|U_{nk}|^2\rho'_{kk} +\sum_{k,n}|U_{nk}|^2\rho''_{nn}\geq 0 . \tag{App. VI,6} \]

Using relation (App. VI, 5) and the property of unitarity

\[ \sum_n U^*_{nk}U_{nl}=\delta_{kl}, \tag{App. VI,7} \]

we reduce expression (App. VI, 6) to the form

\[ \sum_k \rho'_{kk}\ln\rho'_{kk} -\sum_n \rho''_{nn}\ln\rho''_{nn} -\sum_k \rho'_{kk} +\sum_n \rho''_{nn}\geq 0, \tag{App. VI,8} \]

or

\[ \sum_k \rho'_{kk}\ln\rho'_{kk} -\sum_n \rho''_{nn}\ln\rho''_{nn}\geq 0, \tag{App. VI,9} \]

since, by virtue of the normalization of \(\rho'\) and \(\rho''\) [expression (VI. 2,10)], we have

\[ \sum_k \rho'_{kk}=\sum_n \rho''_{nn}. \tag{App. VI,10} \]

Thus Klein’s lemma is proved. Pauli \({}^{49P}\) emphasized that the main content of this lemma consists in the fact that \(\operatorname{Tr}(\rho\ln\rho)\) increases when \(\rho\) is replaced by the diagonalized matrix in which the off-diagonal elements are set equal to zero. Since this has been established, the proof by Born and Green of the decrease of \(\sigma\Sigma\) becomes trivial and may be replaced by a reference to Klein’s lemma.

VII. THE PRINCIPLE OF DETAILED BALANCE *)

In Section IV. 2 it was mentioned that expression (IV. 2, 50) expresses the fact that in equilibrium there occur as many transitions from the \(i\)-th group of energy levels to the \(j\)-th group as in the reverse direction. This is an example of a principle, valid in an enormous number of cases, which asserts that in equilibrium the number of processes leading to the destruction of state \(A\) and to the formation of state \(B\) is equal to the number of processes forming \(A\) and destroying \(B\). Other examples of states in which the principle of detailed balance occurred were encountered in Section I.1 [expression (I.1,12)] in the case of a gas consisting of spherical molecules, in Section I.3 in the case of certain general models of a gas, and in

*) See \({}^{25F2,\,25T,\,38T}\) Section 50, ESM, p. 381.

, Section IV.1 [expressions (IV.1,4) and (IV.1,31) in combination with (IV.1,30) and (IV.1,27)].

From these examples one might have gained the impression that detailed balance always holds. However, in classical statistics Lorentz ^87L showed that it does not hold in the case of polyatomic molecules, which cannot be regarded as spheres (see also ^98B). He showed how the \(H\)-theorem can nevertheless be proved if, instead of collisions and their reverse collisions, one considers cycles of transitions *).

The same difficulty arises in quantum statistics ). Hamilton and Peng ^44H1 (see also ^44H2, ^52S2) showed that, for a system consisting of particles with spin and electromagnetic radiation, the principle of detailed balance is inapplicable *).

This principle is very important in applications, for example, in considering kinetic processes ^11K1, ^15M, ^23C2 (rate processes). If, for example, we wish to calculate the number of collisions in a gas that lead to excited states of a molecule—these states may be chemically active—then one can calculate the number of collisions leading to transitions to the normal state and use detailed balance to obtain the first quantity. Since the second number is easier to find than the first, the advantage is obvious.

In 1925, Lewis’s paper ^25L1 provoked a discussion ^25F2, ^25L2, ^25T concerning the principle of detailed balance, which Lewis called the principle of complete balance, or the law of reversibility in every detail ****). Fowler and Milne gave an impressive list of its applications. They noted that this principle was to a considerable extent based on Einstein’s classical work ^17E2 on transition probabilities. They also noted that the principle is in fact a generalization of Kirchhoff’s ideas ^60K and that it was first formulated, apparently, by Richardson ^14R1, ^24R.

In 1916 Langmuir ^16L, applying this principle to the problem of evaporation and condensation, said: “Since evaporation and condensation, generally speaking, are thermodynamically reversible phenomena, the mechanism of evaporation must be exactly the re—

*) For a discussion of this more general \(H\)-theorem we refer to Tolman’s monograph ^38T. On p. 119 of this monograph there is an example of a collision that does not possess a reverse collision.

**) The statement in ESM that in quantum mechanics the principle of detailed balance is always satisfied is incorrect.

***) See also the recent paper by van Kampen ^54K, which showed that the same applies to states in which magnetic fields are present. See also the paper by van Kleyn ^55K1, in which it is shown that detailed balance for nonequilibrium steady states, generally speaking, does not hold.

****) Tolman ^24T, ^25T sometimes uses the expression “principle of microscopic reversibility.”

equal mechanism of condensation*), even in the very smallest details.”

In 1921 Klein and Rosseland^21K used the principle of detailed equilibrium in their consideration of inelastic collisions between atoms and electrons. Franck and Cario^22F, 22C1, 22C2, 23C1, 23F4, 24F4 also applied it to collision processes. Fowler^24F1, 24F2, 24F3 applied it to the phenomena of capture and loss of electrons by $\alpha$-particles moving with high velocity. Becker^23B1, Kramers^23K and Milne^24M applied it to photoelectric processes of ionization and capture. Eddington^22E, 24E1 used it in investigations of absorption coefficients in stars. Pauli^23P and Einstein and Ehrenfest^23E2 used it to discuss the scattering of radiation by electrons. Dirac^24D used it in considering processes of multiple collisions. Lewis^25L2 applied it in discussing Planck’s radiation law. From this list it is clear that the principle of detailed equilibrium is undoubtedly one of the most important propositions that can be applied to a large number of different problems.

REFERENCES**)

38B. D. Bernoulli, Hydrodynamica (Dulsecker, Argentorati, 1738).
56K. A. Krönig, Ann. Physik 99, 315 (1856).
57C1. R. Clausius, Ann. Physik 100, 353 (1857).
57C2. R. Clausius, Phil. Mag. 14, 103 (1857).
58C. R. Clausius, Ann. Physik 105, 239 (1858).
60K. G. Kirchhoff, Ann. Physik 109, 148 (1860).
60M. J. C. Maxwell, Phil. Mag. 19, 19 (1860).
62C. R. Clausius, Ann. Physik 115, 2 (1862).
67M. J. C. Maxwell, Trans. Roy. Soc. (London) 157, 49 (1867).
68B. L. Boltzmann, Wien. Ber. 53, 517 (1868).
68M1. J. C. Maxwell, Phil. Mag. 35, 129 (1868).
68M2. J. C. Maxwell, Phil. Mag. 35, 185 (1868).
70C1. R. Clausius, Ann. Physik 141, 124 (1870).
70C2. R. Clausius, Phil. Mag. 40, 122 (1870).
71B1. L. Boltzmann, Wien. Ber. 63, 397 (1871).
71B2. L. Boltzmann, Wien. Ber. 63, 679 (1871).
72B. L. Boltzmann, Wien. Ber. 66, 275 (1872).
75B. L. Boltzmann, Wien. Ber. 72, 427 (1875).
76L. J. Loschmidt, Wien. Ber. 73, 139 (1876).
77B. L. Boltzmann, Wien. Ber. 76, 373 (1877).
77L. J. Loschmidt, Wien. Ber. 75, 67 (1877).
79M. J. C. Maxwell, Trans. Cambridge Phil. Soc. 12, 547 (1879).
87B. L. Boltzmann, J. Math. 100, 201 (1877).
87L. H. A. Lorentz, Wien. Ber. 95, 115 (1877).
90P. H. Poincaré, Acta Math. 13, 67 (1890).
91K. Lord Kelvin, Collected Works (Cambridge University Press, Cambridge, 1891), vol. IV.

) Langmuir discharge. It may be noted that Langmuir’s assertion is valid only at equilibrium.
*) The references are arranged in chronological order.

92P. H. Poincaré, Méthodes Nouvelles de la Mécanique Céleste (Gauthier-Villars, Paris, 1892), vol. I.
95B. L. Boltzmann, Nature 51, 413 (1895).
96B. L. Boltzmann, Vorlesungen über Gastheorie (Barth, Leipzig, 1896), vol. I.
96Z. E. Zermelo, Ann. Physik 57, 485 (1896).
97B. L. Boltzmann, Wien. Ber. 106, 12 (1897).
98B. L. Boltzmann, Vorlesungen über Gastheorie (Barth, Leipzig, 1898), vol. II.

00P1. M. Planck, Verhandl. Deut. physik. Ges. 2, 202 (1900).
00P2. M. Planck, Verhandl. Deut. physik. Ges. 2, 237 (1900).
02E. A. Einstein, Ann. Physik 9, 417 (1902).
02G. J. M. Gibbs, Elementary Principles in Statistical Mechanics (Yale University Press, New Haven, 1902).
03B. S. H. Burbury, Phil. Mag. 6, 251 (1903).
03E. A. Einstein, Ann. Physik 11, 170 (1903).
04B1. H. A. Bumstead, Phil. Mag. 7, 8 (1904).
04B2. S. H. Burbury, Phil. Mag. 8, 43 (1904).
06E1. P. Ehrenfest, Physik. Zeits. 7, 528 (1906).
06E2. T. and P. Ehrenfest, Wien. Ber. 115, 89 (1906).
06P. H. Poincaré, J. Phys. 5, 369 (1906).
07E. P. and T. Ehrenfest, Physik. Zeits. 8, 311 (1907).
07L. H. A. Lorentz, Abhandlungen über theoretische Physik (Teubner, Leipzig, 1907), p. 202.
08O. L. S. Ornstein, Dissertation, Leiden (1908).
09L. H. A. Lorentz, The Theory of Electrons (Teubner, Leipzig, 1909), note 29.
10E. A. Einstein, Ann. Physik 33, 1275 (1910).
11E1. P. Ehrenfest, Ann. Physik 36, 91 (1911).
11E2. P. and T. Ehrenfest, Encykl. math. Wiss. 4, No. 32 (1911).
11K1. P. Kohnstamm and F. E. C. Scheffer, Proc. Amsterdam Acad. Sci. 13, 789 (1911).
11K2. J. Kroon, Ann. Physik 34, 907 (1911).
12M. A. A. Markov, Wahrscheinlichkeitsrechnung (Teubner, Leipzig, 1912).
12S. M. von Smoluchowski, Physik. Zeits. 13, 1069 (1912).
13B1. N. Bohr, Phil. Mag. 26, 1 (1913).
13B2. N. Bohr, Nature 92, 231 (1913).
13P. M. Plancherel, Ann. Physik 42, 1061 (1913).
13R. A. Rosenthal, Ann. Physik 42, 796 (1913).
14E. P. Ehrenfest, Physik. Zeits. 15, 657 (1914).
14R1. O. W. Richardson, Phil. Mag. 27, 476 (1914).
14R2. A. Rosenthal, Ann. Physik 43, 894 (1914).
15M. R. Marcelin, Ann. Physik 3, 120 (1915).
16E. P. Ehrenfest, Ann. Physik 51, 327 (1916).
16L. J. Langmuir, J. Am. Chem. Soc. 38, 2221 (1916).
16S. O. Stern, Ann. Physik 49, 823 (1916).
17B1. J. M. Burgers, Ann. Physik 52, 195 (1917).
17B2. J. M. Burgers, Proc. Amsterdam Acad. Sci. 20, 149 (1917).
17B3. J. M. Burgers, Proc. Amsterdam Acad. Sci. 20, 158 (1917).
17B4. J. M. Burgers, Proc. Amsterdam Acad. Sci. 20, 163 (1917).
17E1. P. Ehrenfest, Proc. Amsterdam Acad. Sci. 19, 576 (1917).
17E2. A. Einstein, Physik. Zeits. 18, 121 (1917).
18B. J. M. Burgers, Dissertation, Leiden, 1918.
19S. A. Sommerfeld, Atombau und Spektrallinien (Vieweg, Brunswick, 1919).

20M1. R. von Mises, Physik. Zeits. 21, 225 (1920).
20M2. R. von Mises, Physik. Zeits. 21, 256 (1920).

21K. O. Klein and S. Rosseland, Zeits. f. Physik 4, 46 (1921).
22B. G. D. Birkhoff, Acta Math. 43, 113 (1922).
22C1. G. Cario, Zeits. f. Physik 10, 185 (1922).
22C2. G. Cario and J. Franck, Zeits. f. Physik 11, 161 (1922).
22D1. C. G. Darwin and R. H. Fowler, Phil. Mag. 44, 450 (1922).
22D2. C. G. Darwin and R. H. Fowler, Phil. Mag. 44, 823 (1922).
22D3. C. G. Darwin and R. H. Fowler, Proc. Cambridge Phil. Soc. 21, 262 (1922).
22E. A. S. Eddington, Monthly Notices Roy. Astron. Soc. 83, 32 (1922).
22F. J. Franck, Zeits. f. Physik 9, 259 (1922).
23B1. R. Becker, Zeits. f. Physik 18, 325 (1923).
23B2. N. Bohr, Zeits. f. Physik 13, 117 (1923).
23C1. G. Cario and J. Franck, Zeits. f. Physik 17, 202 (1923).
23C2. J. A. Christiansen and H. A. Kramers, Zeits. f. physik. Chem. 104, 451 (1923).
23D1. C. G. Darwin and R. H. Fowler, Proc. Cambridge Phil. Soc. 21, 391 (1923).
23D2. C. G. Darwin and R. H. Fowler, Proc. Cambridge Phil. Soc. 21, 730 (1923).
23E1. P. Ehrenfest, Naturwiss. 11, 543 (1923).
23E2. A. Einstein and P. Ehrenfest, Zeits. f. Physik 19, 301 (1923).
23F1. E. Fermi, Physik. Zeits. 24, 261 (1923).
23F2. R. H. Fowler, Phil. Mag. 45, 1 (1923).
23F3. R. H. Fowler, Phil. Mag. 45, 497 (1923).
23F4. J. Franck, Ergebn. exakt. Naturwiss. 12, 112 (1923).
23K. H. A. Kramers, Phil. Mag. 46, 836 (1923).
23P. W. Pauli, Zeits. f. Physik 18, 272 (1923).
24B. S. N. Bose, Zeits. f. Physik 26, 178 (1924).
24D. P. A. M. Dirac, Proc. Roy. Soc. (London) A106, 581 (1924).
24E1. A. S. Eddington, Monthly Notices Roy. Astron. Soc. 84, 104 (1924).
24E2. A. Einstein, Sitzber. preuss. Akad. Wiss., Physik.-math. Kl., p. 261 (1924).
24F1. R. H. Fowler, Phil. Mag. 47, 257 (1924).
24F2. R. H. Fowler, Phil. Mag. 47, 416 (1924).
24F3. R. H. Fowler, Proc. Cambridge Phil. Soc. 22, 253 (1924).
24F4. J. Franck, Naturwiss. 12, 1066 (1924).
24M. E. A. Milne, Phil. Mag. 47, 209 (1924).
24N. L. Nordheim, Zeits. f. Physik 27, 65 (1924).
24R. O. W. Richardson, Proc. Phys. Soc. (London) 36, 383 (1924).
24T. R. C. Tolman, Phys. Rev. 23, 693 (1924).
25B. M. Born and P. Jordan, Zeits. f. Physik 34, 858 (1925).
25E1. A. Einstein, Sitzber. preuss. Akad. Wiss., Physik.-math. Kl., p. 3 (1925).
25E2. A. Einstein, Sitzber. preuss. Akad. Wiss., Physik.-math. Kl., p. 18 (1925).
25F1. R. H. Fowler, Proc. Cambridge Phil. Soc. 22, 861 (1925).
25F2. R. H. Fowler and E. A. Milne, Proc. Natl. Acad. Sci. 11, 400 (1925).
25H. W. Heisenberg, Zeits. f. Physik 33, 879 (1925).
25J. P. Jordan, Zeits. f. Physik 33, 649 (1925).
25L1. G. N. Lewis, Proc. Natl. Acad. Sci. 11, 179 (1925).
25L2. G. N. Lewis, Proc. Natl. Acad. Sci. 11, 423 (1925).
25P. W. Pauli, Zeits. f. Physik 31, 776 (1925).
25T. R. C. Tolman, Proc. Natl. Acad. Sci. 11, 436 (1925).
25W. H. Weyl, Math. Zeits. 23, 271 (1925).
26D. P. A. M. Dirac, Proc. Roy. Soc. (London) A112, 661 (1926).

26F1. E. Fermi, Zeits. f. Physik 36, 902 (1926).
26F2. R. H. Fowler, Phil. Mag. 1, 845 (1926).
26F3. R. H. Fowler, Proc. Roy. Soc. (London) A113, 432 (1926).
26H. W. Heisenberg, Zeits. f. Physik 40, 501 (1926).
26J. P. Jordan, Zeits. f. Physik 40, 661 (1926).
26S1. E. Schrödinger, Ann. Physik 79, 361 (1926).
26S2. A. Smekal, Encykl. math. Wiss. 5, No. 28 (1926).
27D. P. A. M. Dirac, Proc. Roy. Soc. (London) A114, 243 (1927).
27E. P. Ehrenfest and G. E. Uhlenbeck, Zeits. f. Physik 41, 24 (1927).
27H1. W. Heisenberg, Zeits. f. Physik 43, 172 (1927).
27H2. P. Höllich, Zeits. f. Physik 41, 636 (1927).
27J1. P. Jordan, Zeits. f. Physik 41, 711 (1927).
27J2. P. Jordan and O. Klein, Zeits. f. Physik 45, 751 (1927).
27N1. J. von Neumann, Nachr. Akad. Wiss. Göttingen, Math.-physik. Kl., p. 245 (1927).
27N2. J. von Neumann, Nachr. Akad. Wiss. Göttingen, Math.-physik. Kl., p. 271 (1927).
27O. L. S. Ornstein and H. A. Kramers, Zeits. f. Physik 42, 481 (1927).
27P. W. Pauli, Zeits. f. Physik 41, 91 (1927).
27U. G. E. Uhlenbeck, Dissertation, Leiden (1927).
28B. W. Bothe, Zeits. f. Physik 46, 327 (1928).
28J. P. Jordan and E. Wigner, Zeits. f. Physik 47, 361 (1928).
28N. L. Nordheim, Proc. Roy. Soc. (London) A119, 689 (1928).
28P. W. Pauli, in Probleme der Modernen Physik, Sommerfeld Festschrift; P. Debye, ed. (Hirzel, Leipzig, 1928), p. 30.
29D. P. A. M. Dirac, Proc. Cambridge Phil. Soc. 25, 62 (1929).
29N. J. von Neumann, Zeits. f. Physik 57, 30 (1929).
29P. R. Peierls, Ann. Physik 3, 1055 (1929).
29S. L. Szillard, Zeits. f. Physik 53, 840 (1929).
30D1. P. A. M. Dirac, Proc. Cambridge Phil. Soc. 26, 351 (1930).
30D2. P. A. M. Dirac, Proc. Cambridge Phil. Soc. 26, 376 (1930).
30H1. W. Heisenberg, Die Physikalischen Prinzipien der Quantentheorie (Hirzel, Leipzig, 1930).
30H2. E. Hopf, Math. Ann. 103, 710 (1930).
30S. F. Simon, Ergeb. exakt. Naturwiss. 9, 222 (1930).
31B1. G. D. Birkhoff, Proc. Natl. Acad. Sci. 17, 650 (1931).
31B2. G. D. Birkhoff, Proc. Natl. Acad. Sci. 17, 656 (1931).
31D. P. A. M. Dirac, Proc. Cambridge Phil. Soc. 27, 240 (1931).
31K. O. Klein, Zeits. f. Physik 72, 767 (1931).
31O1. L. Onsager, Phys. Rev. 37, 405 (1931).
31O2. L. Onsager, Phys. Rev. 38, 2265 (1931).
32B. G. D. Birkhoff and B. O. Koopman, Proc. Natl. Acad. Sci. 18, 279 (1932).
32H1. E. Hopf, Proc. Natl. Acad. Sci. 18, 93 (1932).
32H2. E. Hopf, Proc. Natl. Acad. Sci. 18, 201 (1932).
32H3. E. Hopf, Proc. Natl. Acad. Sci. 18, 333 (1932).
32H4. E. Hopf, Sitzber. preuss. Akad. Wiss., Physik.-math. Kl., p. 182 (1932).
32N1. J. von Neumann, Proc. Natl. Acad. Sci. 18, 70 (1932).
32N2. J. von Neumann, Mathematische Grundlagen der Quanten mechanik (Verlag Julius Springer, Berlin, 1932).
33B1. N. Bohr, Nature, 131, 421 (1933).
33B2. N. Bohr, Nature, 131, 457 (1933).
33H. Handbuch der Physik 24, part 1 (Verlag Julius Springer, Berlin, 1933).

33J. P. Jordan, Statistische Mechanik auf Quanten theoretischer Grundlage (Vieweg, Brunswick, 1933).

33P. W. Pauli, Handbuch der Physik 24, vol. I, 149 (Verlag Julius Springer, Berlin, 1933).

34F. B. Fock, Zeits. f. Physik 75, 622 (1934).

34H. E. Hopf, J. Math. Phys. 13, 51 (1934).

35B. N. Bohr, Phys. Rev. 48, 696 (1935).

35D. P. A. M. Dirac, The Principles of Quantum Mechanics (Oxford University Press, New York, 1935).

35E. Einstein, Podolsky and Rosen, Phys. Rev. 47, 777 (1935).

35S1. E. Schrödinger, Naturwiss. 23, 807 (1935).

35S2. E. Schrödinger, Naturwiss. 23, 823 (1935).

35S3. E. Schrödinger, Naturwiss. 23, 844 (1935).

35S4. E. Schrödinger, Proc. Cambridge Phil. Soc. 31, 555 (1935).

35U. G. E. Uhlenbeck, J. Math. Phys. 14, 10 (1935).

36D. M. Delbrück and G. Moliere, Abhandl. preuss. Akad. Wiss., Physik-math. Kl. No. 1 (1936).

36E. P. S. Epstein, Collection Commentary on the Scientific Writings of J. Willard Gibbs (Yale University Press, New Haven, 1936).

36F. W. H. Furry, Phys. Rev. 49, 393 (1936).

36M. H. Margenau, Phys. Rev. 49, 240 (1936).

36S. E. Schrödinger, Proc. Cambridge Phil. Soc. 32, 446 (1936).

37E. W. M. Elsasser, Phys. Rev. 52, 987 (1937).

37H. E. Hopf, Ergodentheorie (Verlag Julius Springer, Berlin) (1937).

37K. E. C. Kemble, The Fundamental Principles of Quantum Mechanics (McGraw-Hill Book Company, Inc., New York, 1937).

37P. W. Pauli and M. Fierz, Zeits. f. Physik 106, 572 (1937).

37S. T. Saka j, Proc. Phys.-Math. Soc. Japan 19, 172 (1937).

38K. H. A. Kramers, Grundlagen der Quantentheorie (Akademische Verlag, Leipzig, 1938).

38P. R. Peierls, Phys. Rev. 54, 918 (1938).

38T. R. C. Tolman, The Principles of Statistical Mechanics (Oxford University Press, New York, 1938).

39B1. F. J. Belinfante, Physica 6, 849 (1939).

39B2. F. J. Belinfante, Physica 6, 870 (1939).

39B3. J. de Boer and A. Michels, Physica 6, 97 (1939).

39H. O. Halpern and F. W. Doermann, Phys. Rev. 55, 1077 (1939).

39K1. E. C. Kemble, Phys. Rev. 56, 1013 (1939).

39K2. E. C. Kemble, Phys. Rev. 56, 1146 (1939).

40B. J. de Boer, Dissertation, Amsterdam (1940).

40H. K. Husimi, Proc. Phys.-Math. Soc. Japan 22, 264 (1940).

40M. J. E. and M. G. Mayer, Statistical Mechanics (John Wiley and Sons, Inc., New York, 1940).

40P1. W. Pauli, Phys. Rev. 58, 716 (1940).

40P2. W. Pauli and F. J. Belinfante, Physica 7, 177 (1940).

40T. R. C. Tolman, Phys. Rev. 57, 1160 (1940).

41M. E. W. Montroll and J. R. Mayer, J. Chem. Phys. 9, 626 (1941).

41O. J. C. Oxtoby and S. M. Ulam, Ann. Math. 42, 874 (1941).

41W. A. Wintner, The Analytical Foundations of Celestial Mechanics (Princeton University Press, Princeton, 1941).

43C. S. Chandrasekhar, Revs. Modern Phys. 15, 1 (1943).

43P. M. Planck, Naturwiss. 31, 153 (1943).

44H1. J. Hamilton and H. W. Peng, Proc. Roy. Irish Acad. A49, 197 (1944).

44H2. W. Heitler, Quantum Theory of Radiation (Oxford University Press, New York, 1944).

44K. N. Krylov, Nature, 153 709 (1944).

45C. H. B. Casimir, Revs. Modern Phys. 17, 343 (1944).

46B. M. Born and H. S. Green, Proc. Roy. Soc. (London) A188, 10 (1946).

46C. H. Cramér, Mathematical Methods of Statistics (Princeton University Press, Princeton, 1946).

46G1. H. S. Green, Proc. Roy. Soc. (London) A189, 103 (1946).

46G2. H. J. Groenewold, Physica 12, 405 (1946).

46J. H. and B. S. Jeffreys, Method of Mathematical Physics (Cambridge University Press, Cambridge, 1946).

46K. J. G. Kirkwood, J. Chem. Phys. 14, 180 (1946).

47B1. M. Born and H. S. Green, Proc. Roy. Soc. (London) A190, 455 (1947).

47B2. M. Born and H. S. Green, Proc. Roy. Soc. (London) A191, 168 (1947).

47D. Б. Давыдов, J. Phys. U. S. S. R. 11, 33 (1947).

48B1. M. Born, Ann. Physik 3, 107 (1948).

48B2. M. Born and H. S. Green, Proc. Roy. Soc. (London) A192, 166 (1948).

48S. E. Schrödinger, Statistical Thermodynamics (Cambridge University Press, Cambridge, 1948).

48W. N. Wiener, Cybernetics (John Wiley and Sons, Inc., New York, 1948).

49B1. M. Born, Natural Philosophy of Cause and Chance (Oxford University Press, New York, 1949).

49B2. M. Born, Nuovo cimento 6, Suppl., 161 (1949).

49B3. M. Born and H. S. Green, A General Kinetic Theory of Liquids (Cambridge University Press, Cambridge, 1949).

49K1. A. I. Khinchin, Mathematical Foundations of Statistical Mechanics (Dover Publications, New York, 1949).

49K2. J. G. Kirkwood, Nuovo cimento 6, Suppl., 233 (1949).

49K3. H. A. Kramers, Nuovo cimento 6, Suppl., 158 (1949).

49M. J. E. Moyal, Proc. Cambridge Phil. Soc. 45, 99 (1949).

49P. W. Pauli, Nuovo cimento 6, Suppl., 166 (1949).

49R. G. S. Rushbrooke, Introduction to Statistical Mechanics (Oxford University Press, New York, 1949).

49S. A. J. F. Siegert, Phys. Rev. 76, 1708 (1949).

49S1. C. E. Shannon and W. Weaver, Mathematical Theory of Communication (University of Illinois Press, Urbana, Illinois, 1949).

49W1. C. F. von Weizsäcker, Zum Weltbild der Physik (Hirzel, Zürich, 1949), p. 80.

49W2. G. Wentzel, Quantum Theory of Fields (Interscience Publishers, Inc., New York, 1949).

50B. Bartlett, Nature 165, 727 (1950).

50H. P. Halmos, Measure Theory (D. Van Nostrand Company, Inc., New York, 1950).

51B1. J. A. Bearden and H. M. Watts, Phys. Rev. 81, 73 (1951).

51B2. R. Bellmann and T. Harris, Pacific J. Math. 1, 179 (1951).

51B3. D. Bohm, Quantum Theory (Prentice-Hall, Inc., New York, 1951).

51B4. L. Brillouin, J. Appl. Phys. 22, 334 (1951).

51D. J. W. M. DuMond and E. R. Cohen, Phys. Rev. 82, 555 (1951).

51S. F. E. Simon, Zeit. Naturforsch. 6a, 397 (1951).

52G1. H. Grad, J. Phys. Chem. 56, 1039 (1952).

52G2. H. Grad, Comm. Pure Appl. Math. 5, 455 (1952).

52H. T. E. Harris, Trans. Am. Math. Soc. 73, 471 (1952).

52K. M. J. Klein, Phys. Rev. 87, 111 (1952).

52R. L. Rosenfeld, Mimeographed Notes of Lectures given at the Summer School for Theoretical Physics at Les Houches, 1952.

52S1. M. Schönberg, Nuovo cimento 9, 1139 (1952).

52S2. E. C. Stueckelberg, Helv. Phys. Acta 25, 577 (1952).

53B. M. S. Bartlett, Proc. Cambridge Phil. Soc. 49, 263 (1953).

53D. J. L. Doob, Stochastic Processing (John Wiley and Sons, Inc., New York, 1953).

53H. D. ter Haar and C. D. Green, Proc. Phys. Soc. (London) A66, 153 (1953).

53S1. M. Schönberg, Nuovo cimento 10, 419 (1953).

53S2. M. Schönberg, Nuovo cimento 10, 697 (1953).

53T. J. S. Thompson, Phys. Rev. 91, 1263 (1953).

54G. C. D. Green, Dissertation, St. Andrews University, Scotland (1954).

54G2. A. Gamba, J. Appl. Phys. 25, 1549 (1954).

54H1. D. ter Haar, Elements of Statistical Mechanics (Rinehart & Company, Inc., New York, 1954).

54H2. D. ter Haar, Am. J. Phys. 22, 638 (1954).

54H3. Hirschfelder, Curtiss and Bird, Molecular Theory of Gases and Liquids (John Wiley and Sons, Inc., New York, 1954).

54I. Inagaki, Wanders and Piron, Helv. Phys. Acta 27, 71 (1954).

54K. N. G. van Kampen, Physica 20, 603 (1954).

54L. P. T. Landsberg, Phys. Rev. 96, 1420 (1954).

54L1. J. P. Lloyd and G. E. Pake, Phys. Rev. 94, 579 (1954).

54M1. D. K. C. MacDonald, J. Appl. Phys. 25, 619 (1954).

54S. L. Sartre, J. phys. radium 15 (April, 1954).

54T. T. Takabayasi, Progr. Theoret. Phys. 11, 341 (1954).

55G. C. D. Green and D. ter Haar, Physica 21, 63 (1955).

55G1. A. Gamba, Nuovo cimento 1, 358 (1955).

55H1. D. ter Haar, Am. J. Phys. (in press).

55H2. D. ter Haar and C. D. Green, Proc. Cambridge Phil. Soc. 51, 141 (1955).

55K. R. Kurth, Revs Modern Phys. (in press).

55K1. M. J. Klein, Phys. Rev. 97, 1446 (1955).

55M. E. W. Montroll and M. S. Green, Ann. Rev. Phys. Chem. 6 (1955).

  1. [[unclear: marginal/footnote marker “39B1”]] 

  2. Cf., for example, Bohr’s discussion.[^33B1][^33B2] 

  3. Cf. the criticism given by Pauli and Fierz.[^37P] We also mention the article by Davydov.[^47D] 

  4. It is easy to generalize the consideration to the case when the system is in a state proper to two or more commuting operators. 

Submission history

FOUNDATIONS OF STATISTICAL MECHANICS, Part II