Abstract
A review of the current state of the justification of quantum statistical mechanics has not yet existed, and the aim of the present work is to provide such a review. In a recently published monograph, the author has already touched upon some of the issues discussed below. It must be taken into account, however, that, on the one hand, a number of new results have been obtained since the publication of that monograph, and, on the other hand, that the emphasis in the present work is placed rather on presenting the existing state of the justification of statistical mechanics than on the historical development of these foundations, which was set out in greater detail earlier.
Full Text
FOUNDATIONS OF STATISTICAL MECHANICS *)
D. ter Haar
CONTENTS
Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 620
I. The \(H\)-theorem in classical statistical mechanics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 628
I.1. Kinetic aspect of the \(H\)-theorem; the hypothesis on the number of collisions . . . . . . . . . . . . . . 628
I.2. The reversibility paradox and the recurrence paradox . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 633
I.3. Statistical aspects of the \(H\)-theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 637
II. The classical ergodic theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 647
II.1. The ergodic and quasi-ergodic theorems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 647
II.2. Recent investigations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 651
III. The \(H\)-theorem in the classical theory of ensembles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 655
III.1. Fine-structure and coarse-structure densities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 655
III.2. Representative ensembles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 664
Summary of the situation in classical thermostatistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 669
IV. The \(H\)-theorem in quantum statistics. 1. The \(H\)-theorem in the elementary method of treatment. 2. The \(H\)-theorem in the quantum theory of ensembles. 3. Representative ensembles
V. The ergodic theorem in quantum statistics. Summary of the situation in quantum thermostatistics. Appendix I. Lorentz model. Appendix II. Proof of the existence of quasi-ergodic systems. Appendix III. Proof of the classical ergodic theorem. Appendix IV. Proof of the quantum-mechanical ergodic theorem. Appendix V. Transitions, proportional to the time. Appendix VI. Klein’s lemma. Appendix VII. Principle of detailed balance. References
) Foundations of Statistical Mechanics. Reviews of Modern Physics. 27*, 289, 1955. Translated by G. F. Zharkov.
INTRODUCTION
In recent years numerous attempts \(^{44K,\,46B,\,46G,\,46K,\,47B1,\,47B2,\,48B1,\,48B2,\,49B1,\,49B2,\,49B3,\,49K1,\,49K2,\,52K,\,52R,\,52S2}\) ) have been made to justify the use of statistical mechanics for the description of physical systems. As for classical, i.e. non-quantum, statistical mechanics, P. and T. Ehrenfest, in a well-known survey published in the Encyclopedia of Mathematical Sciences, analyzed its foundations and, as will be seen from what follows, their conclusions have, to a greater or lesser degree, remained valid up to the present time ). However, an analogous survey of the present state of the justification of quantum statistical mechanics has not existed up to now, and the purpose of the present paper is to give such a survey. In a recently published monograph \(^{54H}\) **) the author has already touched upon some of the questions discussed below. It should be borne in mind, however, that, on the one hand, since the publication of that monograph a number of new results have been obtained, while, on the other hand, in the present paper the emphasis is placed rather on an account of the existing state of the foundations of statistical mechanics than on the historical development of these foundations, which was set forth in greater detail earlier. Thus the present paper, to a considerable extent, supplements the consideration of the question that was carried out in ESM. This, of course, does not mean that here we shall completely ignore the historical aspects of the development of statistical mechanics.
It may be said to us that one can test the quality of a pie only by eating it, and that, since statistical mechanics proves able to predict sufficiently accurately and successfully the behavior of physical systems under equilibrium conditions and even, under certain circumstances, in nonequilibrium states ****), this fact should sufficiently justify the methods used. This ought to have been all the more true in our day, when the use of quantum mechanics has made it possible to explain such seemingly insurmountable difficulties as the paradox of the specific heat—
*) The references are placed at the end of the article in chronological order. Each reference number consists of the year of publication and the first letter of the author’s surname: for example, the reference to the paper by M. J. Klein, Phys. Rev. 87, 111 (1952) has the form \(^{52K}\).
**) R. Kurt \(^{55K}\) also considered the foundations of classical thermostatistics from a more axiomatic point of view. I must express my gratitude to him for the opportunity to become acquainted with his work before publication.
***) We shall use, as far as possible, the notation of this work, which we shall cite as ESM.
****) For an introduction to the statistical mechanics of nonequilibrium processes we refer the reader to the recently published review article by Montroll and Green \(^{55M}\) and to the bibliography given there.
capacity and the Gibbs paradox. However, it is logically quite natural that many physicists were not satisfied with the argument given and tried to show that the statistical formalism follows directly from classical or quantum mechanics. In the present review it will be set forth to what extent the attempts of these authors were successful. Even in those cases where success was limited, these attempts were extremely important and useful, since they revealed both the possibilities and the limitations of the statistical approach.
Before beginning the discussion of the various works devoted to the basic principles of statistical mechanics, it is necessary to note that we shall be concerned mainly with the following two questions:
(A) Why is it possible to describe the behavior of almost all physical systems by considering only equilibrium states?
(B) Why is it possible to describe the behavior of a real physical system by considering a large number of identical systems (in the theory of ensembles) and to identify the averaged behavior of this group of systems with the behavior of the physical system of interest to us?
From what follows it will become clear that these two questions are closely connected with one another. The first of them arises not only in statistical mechanics, but also in thermodynamics and kinetic theory, and the problem connected with it may be formulated in the form of the question:
(A1) In what way can an equilibrium state be defined?
The second question is of a purely statistical nature and in turn is connected with the question:
(B1) In what way can one construct such an ensemble as would “represent” the actual, given physical system?
This question of representative ensembles, which will play an important role in our discussion, was especially emphasized by Tolman\(^{38T,40T}\).
We have now reached the point at which it is necessary to introduce a definition of statistical mechanics. We shall do this by using the definition given by Kramers in his introductory address at the International Conference on Statistical Mechanics in Florence in 1949\(^{49K3}\): “For the physicist of our day, statistical mechanics is precisely that branch of physics which deals with the atomistic interpretation of the thermal properties of matter and radiation; for this reason it may also be called ‘thermostatistics.’” Our exposition will not explicitly deal with radiation; we shall confine ourselves to systems consisting of particles. However, in discussing transitions from one state to another, the existence of a radiation field will be tacitly assumed, which in many cases is the mechanism by means of which transitions can take place.
Statistical mechanics arose from the kinetic theory of gases, developed in the nineteenth century. The father of kinetic theory, apparently, should be called Krönig, although more than a hundred years before him Bernoulli had already connected the properties of a gas with the properties of individual particles. Krönig \(^{56K}\) drew attention to the fact that although the path of an individual atom in a gas is so irregular that it is utterly impossible to follow this atom, nevertheless, by using probability theory, such completely chaotic behavior can be reduced to the ordered behavior of the gas. The kinetic theory of gases was further developed in the works of Clausius \(^{57C1,\,57C2,\,58C,\,62C,\,70C1,\,70C2}\), Maxwell \(^{60M,\,67M,\,68M1,\,68M2}\), and Boltzmann \(^{68B,\,72B,\,96B,\,98B}\) ). In his “Theory of Gases” Boltzmann gave special consideration to question (A), and in 1872 proposed \(^{72B}\) his famous \(H\)-theorem in order to show that any nonequilibrium state will change in such a way that it will approach an equilibrium state. The answer to question (A1) in this case consists in the fact that the equilibrium state is the most probable—or, more precisely, the most probable state compatible with certain constraints. In its original unrestricted form, the \(H\)-theorem, if accepted, proved that any system would tend toward equilibrium, if there was no equilibrium at the initial moment *). This would mean that if one waited long enough, the system would necessarily find itself in a state of equilibrium and, moreover, this equilibrium state would then exist forever. It would follow from this that equilibrium is the rule, while nonequilibrium is the exception, and thus an answer to question (A) would have been given.
However, it was soon realized that the \(H\)-theorem in its original unrestricted form cannot be proved with absolute rigor, but is based on certain assumptions concerning the number of collisions that a given particle undergoes during some interval of time. Moreover, Loschmidt \(^{76L,\,77L}\) and Zermelo \(^{96Z}\) convincingly showed that the \(H\)-theorem in its original form cannot be correct. In his subsequent works Boltzmann, in view of this, carefully emphasized the statistical aspects of the \(H\)-theorem. This means that the \(H\)-theorem is a statement about the most probable behavior of the system. However, fluctuations about the equilibrium state are by no means forbidden. The statistical aspects of the \(H\)-theorem were especially emphasized by Ehrenfe—
*) For a more detailed acquaintance with the historical details one should consult the bibliographic references at the end of chapters I and II and at the end of the appendix I in ESM.
**) We exclude, of course, such hypothetical systems as completely ideal gases located between ideal walls. In such systems there would be no mechanism capable of producing a change in the distribution function, and a nonequilibrium state would exist for an unlimited time.
several works \(^{07E, 11E2}\), both from a general point of view and in the consideration of simple models*).
In the first part of the present article the foundations of statistical mechanics will be discussed. In section I.1 the unrestricted \(H\)-theorem will be discussed. Loschmidt’s and Zermelo’s arguments against the original formulation of the \(H\)-theorem are given in section I.2, while in section I.3 the statistical aspects of the \(H\)-theorem are discussed.
Besides introducing the \(H\)-theorem for the proof of (A), Boltzmann also attempted to prove that the averaged behavior of a system coincides with its equilibrium behavior. The content of this assertion is that the time average, taken over an infinite interval of time, of a phase function, i.e. of a function depending on the values of the coordinates and velocities which completely determine the physical state or phase of the system under consideration, must be equal to the value of the phase function in equilibrium. It is easy to convince oneself that this second way of approaching the answer to question (A) is equivalent to the first way. First, if the \(H\)-theorem is correct, then it follows from it that any nonequilibrium state will pass into an equilibrium state, which will then exist forever. Taking the time average of a phase function over an infinite interval of time, we thus obtain the value of this phase function in equilibrium. Second, if the time average coincides with the value in equilibrium, this means that the system must be in a state of equilibrium for the greater part of the time and, consequently, return to equilibrium from any nonequilibrium state). To prove that the averaged behavior and the behavior in equilibrium coincide, it is necessary to calculate the time average; serious difficulties are encountered here. Boltzmann tried to overcome these difficulties by assuming that the majority of physical systems are ergodic). An ergodic system is a system such that the representative point of it in \(\Gamma\)-space***) passes through every point of the energy—
*) The statistical nature of the \(H\)-theorem was insufficiently taken into account in a recent work by Sartre \(^{54S}\).
**) It is easy to see that this equivalence remains valid if the \(H\)-theorem in its statistical form is replaced by its unrestricted form.
) “Ergodic” is from the Greek words ergon (work; used here in the sense of energy) and hodos* (path): the representative point in \(\Gamma\)-space passes through all points of the energy surface. The term “ergodic” was introduced by Boltzmann in 1887 \(^{87B}\). Maxwell called the assumption of ergodicity the assumption of the continuity of the path.
****) A system with \(S\) degrees of freedom can be described by \(S\) (generalized) coordinates and \(S\) (generalized) momenta. The values of these \(2S\) quantities at a given instant of time determine a point (the so-called representative point) in a \(2S\)-dimensional space, called \(\Gamma\)-space (\(\Gamma\) from the word gas), or phase space. The \(2S-1\)-dimensional hypersurface in \(\Gamma\)-space, specified by the equation \(\varepsilon=\mathrm{const}\), where \(\varepsilon\) is the energy of the system, is called the energy surface.
cal surface corresponding to the energy of the system. The calculation of the time average in this case coincides with the calculation of the average over the energy surface, i.e. with the average over the aggregate of systems possessing one and the same energy. Such a specific aggregate of systems, or ensemble, if we use the term introduced by Gibbs, is the so-called microcanonical ensemble. Since, according to our assumption, the representative point of any system of the ensemble will pass through every point of the energy surface, the orbits described by the representative points of the systems of the ensemble will all coincide with one another, with the sole difference among the systems being that the exact instant of passage through a given point of the orbit will be different for different systems*). It follows from this that the time average, which is an average taken over one single orbit on the energy surface, through which different points pass at different instants of time, will give the same result as the average taken over the microcanonical ensemble and representing an average taken over the very same orbit, but where now all the representative points are considered at one and the same instant of time.
In considering averages over the energy surface, or over the microcanonical ensemble, Boltzmann departed from kinetic theory and truly entered the realm of statistical mechanics—if we use here the definition of statistical mechanics given by Gibbs in the preface to his famous monograph 02G**).
*) We use here the fact that the direction of the orbit at any point of $\Gamma$-space is uniquely determined by Hamilton’s equations of motion. It also follows from this that the orbit can never return to that point of $\Gamma$-space through which it has already passed, unless we are dealing with a strictly periodic orbit.
**) At this point we could advise everyone studying statistical mechanics to read and reread this preface attentively, since in it, more clearly than anywhere else, the basic ideas of rational thermodynamics, as Gibbs called it, are formulated. It is precisely in this preface that the term statistical mechanics itself was also introduced. In view of the difficulties then facing kinetic theory at the time of the appearance of Gibbs’s monograph (the paradox with specific heats, the difficulty in understanding Planck’s law of radiation, etc.), Gibbs attempted to create a rational system based on a small number of axioms and not necessarily connected with the phenomena of nature. Let us quote Gibbs here:
“It is well known that although theory assigns to each molecule of a gas six degrees of freedom, our experiments on heat capacity lead us to take into account no more than five degrees. Of course, he who bases his work on hypotheses concerning the structure of matter stands on an unreliable foundation.
The difficulties of this kind kept the author from attempting to explain the mysteries of nature and forced him to content himself with the more modest task of deriving certain more evident propositions relating to the statistical branch of mechanics. At the same time there can no longer be any error here from the point of view of the agreement of hypotheses with the facts of nature, for in this respect nothing …”
It must be noted, however, that the theory of ensembles enters, if one may put it this way, through the back door, since it is introduced merely as a mathematical device for computing the behavior of an isolated system.
At the beginning of this century, in connection with the development of the foundations of the theory, it was realized that ergodic systems are never realized; nevertheless, hope was expressed that most physical systems are quasi-ergodic. A quasi-ergodic system is a system that has an orbit in \(\Gamma\)-space that covers the energy surface everywhere densely, although in actuality it does not pass through every point of this surface. The hope was also expressed that quasi-ergodicity would prove sufficient to ensure equality of time averages and averages taken over an ensemble chosen in the corresponding way. Such, for example, was the point of view of Ehrenfest \(^{11\mathrm{E}2}\).
In 1913 Rosenthal \(^{13\mathrm{R}}\) and Plancherel \(^{13\mathrm{P}}\), independently of one another, proved the impossibility of the existence of ergodic systems, and in 1923 Fermi \(^{23\mathrm{F}1}\) showed that a certain class of systems is quasi-ergodic. He did not, however, prove the equality of time averages and ensemble averages). This equality was proved by Rosenthal \(^{14\mathrm{R}2}\), but his proof proved to be insufficiently rigorous). Recently the problem of the equivalence of time averages and ensemble averages has been studied mainly by mathematicians, and it is usually called the ergodic theorem*. In 1931 and 1932 Birkhoff \(^{31\mathrm{B}1,31\mathrm{B}2,32\mathrm{B}}\) and von Neumann \(^{32\mathrm{N}1,32\mathrm{N}2}\) established that, when certain plausible mathematical conditions are satisfied, these two types of averages give the same result; and then Oxtoby and Ulam \(^{41\mathrm{O}}\) recently showed that a very broad class of systems satisfies these mathematical conditions. The ergodic theorem will be discussed in Section II. In Section II.1 the works published before 1930 will be discussed, while later studies are considered in Section II.2.
Many physicists still regard the ergodic theorem as the most satisfactory, if not the only, justification for the use of statistical methods. These people believe that mechanics can deal only with isolated systems and that the statistical approach is a necessary mathematical device that is of no special physical interest, and is not assumed. The only error into which one can fall is insufficient agreement between premises and conclusions, and this is something that, with some caution, one may hope to avoid.
*) Below, however, we shall see that this equality in fact holds for all quasi-ergodic systems.
**) In this proof the passage to the limit was carried out twice, as was pointed out by Rosenthal himself in his reply to Epstein (see \(^{36\mathrm{E}}\), p. 478).
The author of the present article thinks that such a point of view does not meet the physical grounds for resorting to the statistical approach.*) These grounds consist in the fact that the enormous number of particles forming the majority of physical systems makes it impossible to determine the values of all the integrals of motion. Generally speaking, we can determine only the energy, the total momentum, and, perhaps, one or two more integrals. After this we must use the very meager information obtained about the system in order to predict its behavior in the future. Since our knowledge is wholly insufficient for predicting the future of the system with complete certainty, we must turn to statistical methods. It is precisely here that the so-called representative ensembles.**) enter into consideration.
Instead of considering one system, we shall consider a large number of systems possessing, to a degree of accuracy, the same values of those quantities that are known to us, and differing widely in all other respects. In order to construct such a representative ensemble satisfactorily, it is necessary to know what weight is to be assigned to the various systems of the ensemble, or, in other words, it is necessary to make assumptions concerning the a priori probabilities.
If we have agreed on the necessity of representative ensembles, then question (B) is resolved in ensemble theory by means of the proof that the overwhelming majority of the systems of the ensembles under consideration behave practically in the same way as one might have expected of an equilibrium system and, moreover, lead to the very same values of the phase functions.***) However, question (B1) still remains open. It always turns out that the canonical ensemble is a representative ensemble of a system in thermal equilibrium. To prove this, a generalized \(H\)-theorem is introduced and it is shown that if a representative ensemble at some instant of time is not a canonical ensemble, then—
*) It is also necessary to add that completely isolated systems are not of interest to physicists, since experiments cannot be carried out with them. Of course, if we are dealing with classical systems, one may always consider idealized experimental systems and thus include isolated systems in the discussion as well. Moreover, Pauli \(^{49P}\) pointed out that also in the case of quantum-mechanical systems one may neglect the influence of the act of observation on the system insofar as we are dealing with statistical mechanics, although for quantum mechanics itself this is of primary importance.
**) See sections 24, 86, 112, 117, and 120 of Tolman’s monograph \(^{38T}\) and part B of ESM, where the question of representative ensembles is considered in detail; see also sections III.2 and IV.3 of the present article.
***) A proof of this assertion may be found in any textbook on statistical mechanics that treats the theory of ensembles \(^{02G,\,38T,\,54H1}\).
the subsequent instants of time will be such that we shall be forced to use the representing ensemble, which has a much closer resemblance to the canonical one. This question of approximation to equilibrium, described by ensembles, and the question of representing ensembles will be discussed in Part III, insofar as this concerns the classical theory of ensembles.
The transition from classical to quantum statistics introduces no fundamental changes. In fact, the two cases may be considered in parallel, i.e. the classical and the quantum-mechanical cases, and their deep analogy will become apparent. There are, however, certain differences between them, and for this reason we prefer to discuss each of these cases separately. In the quantum-mechanical case we can again approach the problem either by means of the theory of ensembles or by means of the ergodic theorem. These two alternative methods of approach are discussed in Parts IV and V, respectively.
In conclusion of the introductory remarks I should like to express several considerations concerning certain general aspects of the second law of thermodynamics and its place in statistical mechanics.*) The fact that in all practical applications of thermodynamics the second law is valid, i.e. entropy increases, is due to the combined action of two effects. First of all, if at some instant of time the entropy is less than its equilibrium value, then the probability of its increasing is overwhelmingly large in comparison with the probability that it will decrease. Secondly, our observations always take place in such a way that, beginning with a given situation, we observe the further development; however, it is impossible for us to begin with a given situation and observe the preceding instants of time. The fact that we can do one thing and cannot do the other is connected with another fact, namely that we have a memory of the past and can therefore know what occurred at preceding instants of time, but not what will occur later. In this sense the irreversibility of the second law of thermodynamics is psychologically quite well founded.
Another remark concerns the fact that, so far as can be judged at the present time, all observations of phenomena in the universe are compatible with the idea of the development of the universe as a whole from some possible, though thermodynamically or statistically improbable, state removed into the distant past. This may be connected with cosmological ideas about a more or less singular state that existed, let us say, three or five billion years ago. On the other hand, one may adduce considerations, as was done, for example, by Boltzmann \(^{95B}\), that the idea of devel—
*) I must express my gratitude to Prof. R. Peierls for his valuable critical remarks on this question.
of the development of our world from a less probable to a more probable state is not necessarily in contradiction with the notion of the universe, as a whole, being in thermodynamic equilibrium. Let us quote Boltzmann: “We assume that the entire universe is, and has always been, in thermal equilibrium. The probability that one (and only one) part of the universe is in some state will be the smaller the farther this state is removed from the state of thermal equilibrium; however, this probability will be the greater the larger the universe itself is. If we suppose that the universe is sufficiently large, then the probability of finding a relatively small part of it in any given state (remote, however, from the state of equilibrium) may be arbitrarily large. We can also obtain a large probability for the fact that, although the entire universe is in thermal equilibrium, our world is so far removed from thermal equilibrium that it is impossible even to imagine the monstrously small probability of such a state. However, on the other hand, can we imagine how small a part of the whole universe this world occupies? If, then, the universe is sufficiently large, the probability that so small a part of it as our world is in its present state turns out to be by no means small.
If this supposition were correct, everything would gradually return to thermal equilibrium; however, since the entire universe is so large, it is possible that someday in the future some other world will deviate from equilibrium just as much as now takes place in our world. Then the above-mentioned \(H\)-curve*) would depict what is happening in the universe. The maxima of the curve would correspond to worlds in which visible motions and life would exist.”
In giving this quotation, we do not subscribe to Boltzmann’s point of view, but merely mention it. In discussing this point of view on its merits, it is necessary to investigate the difficult problem of how far it is possible to speak of thermal equilibrium in a nonclosed system such as our universe.
I. THE \(H\)-THEOREM IN CLASSICAL STATISTICAL MECHANICS
1.1. Kinetic aspect of the \(H\)-theorem; the hypothesis of the number of collisions
In 1872 Boltzmann \(^{72B}\) proposed his famous \(H\)-theorem in order to prove that any nonequilibrium distribution will tend to pass into the equilibrium one. We shall briefly set out its
*) In essence, this is a curve giving entropy as a function of time; cf. Section 1.3.
reasoning. He considered the case of an isolated system consisting of a gas (in particular, a monatomic gas in the absence of external forces), although subsequently he generalized his consideration to the case of polyatomic gases acted upon by external forces \(^{75}\).
We shall be interested in calculating the rate of change of the distribution function \(f\), which in our simplest case is a function only of the three (Cartesian) components \(u, v\), and \(w\) of the atom’s velocity \(c\) and of the time \(t\). We must, therefore, on the one hand take into account the rate of change of the velocities of atoms lying in the interval \(c\) and \(c+dc\), and, on the other hand, take into account the rate at which, in collisions, atoms are formed with velocities within the prescribed interval. In the formula
\[ \frac{\partial f}{\partial t}\,du\,dv\,dw\,dt=-A+B \tag{I.1,1} \]
the quantity \(f\,du\,dv\,dw\) represents the number of atoms per unit volume having velocities with components lying in the intervals \((u,u+du)\), \((v,v+dv)\), and \((w,w+dw)\); \(A\) is the number of atoms per unit volume with velocities lying within the indicated region which, during the time interval \(dt\), change their velocities; \(B\) is the number of atoms per unit volume which during the time \(dt\) change their velocities in such a way that after collision their velocities fall into the region indicated above. It can be shown (see, for example, ESM, Sec. I.4) that \(A\) is given by the expression
\[ A=f(u,v,w)\,du\,dv\,dw\,dt\int f(u_1,v_1,w_1)\,du_1\,dv_1\,dw_1\int a\,d\omega . \tag{I.1,2} \]
In expression (I.1,2), \(d\omega\) represents an element of solid angle enclosing the so-called line of centers \(\omega\), which is a vector directed from one of the colliding atoms to the other at the moment of collision; \(a\) is a quantity depending only on the absolute value of the relative velocity \(c_{\mathrm{rel}}\) and on the angle \(\theta\) between the vector of the relative velocity and \(\omega\); the integration is carried out over all values of \(u_1, v_1\), and \(w_1\), and over all possible directions of the line of centers. Expression (I.1,2) is obtained by multiplying the number of atoms with velocities lying within the selected interval, i.e. \(f(u,v,w)\,du\,dv\,dw\), by the number of atoms colliding with one selected atom during the time interval \(dt\). Consider atoms having velocities between \(c_1\) (components \(u_1, v_1, w_1\)) and \(c_1+dc_1\). If we assume that there is no correlation between the velocities and positions of different atoms, then the number of collisions between these atoms and any one selected atom, in the case when the line of centers lies within the solid angle \(d\omega\), will be equal to \(a f(u_1,v_1,w_1)\,du_1\,dv_1\,dw_1\,dt\,d\omega\), where the quantity \(a\) is a normalizing constant depending on \(c_{\mathrm{rel}}\) and \(\theta\). Expression (I.1,2) is then obtained after in-
integration over all possible values \(u_1, v_1, w_1\) and over all possible directions \(\omega\).
Similarly, for \(B\) we have the expression
\[ B=dt\int' du'\,dv'\,dw'\, f(u',v',w')\,du'_1\,dv'_1\,dw'_1\, f(u'_1,v'_1,w'_1)\int' a'\,d\omega . \tag{I.1,3} \]
The primes on the integral signs indicate that the integration is now carried out only over such values \(c'\) and \(c'_1\) and over such directions \(\omega'\) that the velocities after the collision, completely determined by the values \(c'\), \(c'_1\), and \(\omega'\), would satisfy the condition that one of them falls in the interval from \(c\) to \(c+dc\). Expressing the primed quantities in terms of the unprimed ones, instead of (I.1,3) one may write (cf., for example, ESM, Section I.4)
\[ B=du\,dv\,dw\,dt\int du_1\,dv_1\,dw_1\, f(u',v',w')\,f(u'_1,v'_1,w'_1)\int a\,d\omega, \tag{I.1,4} \]
where \(u'\), \(v'\), \(w'\), \(u'_1\), \(v'_1\), and \(w'_1\) are functions of \(u\), \(v\), \(w\), \(u_1\), \(v_1\), and \(w_1\), owing to the presence of the equations of conservation of momentum and energy. Combining (I.1,1), (I.1,2), and (I.1,4), we obtain, for the rate of change of the distribution function due to collisions,
\[ \frac{\partial f}{\partial t} = -\int (ff_1-f'f'_1)\,a\,du_1\,dv_1\,dw_1\,d\omega, \tag{I.1,5} \]
where, for brevity, the following notation has been used:
\[ \begin{aligned} f&=f(u,v,w), & f'&=f(u',v',w'),\\ f_1&=f(u_1,v_1,w_1), & f'_1&=f(u'_1,v'_1,w'_1). \end{aligned} \tag{I.1,6} \]
Introducing the quantity \(H\) by means of the equality *)
\[ H=\int f\ln f\,du\,dv\,dw, \tag{I.1,7} \]
we obtain, for the rate of change of \(H\) from (I.1,5),
\[ \frac{dH}{dt} = \int (f'f'_1-ff_1)(\ln f+1)\,a\,du\,dv\,dw\,du_1\,dv_1\,dw_1\,d\omega. \tag{I.1,8} \]
Because of the equivalence of the quantities \(c\), \(c_1\), \(c'\), \(c'_1\) in (I.1,8), this equation can be rewritten in the form (see, for example, ESM, Section I.5)
\[ \frac{dH}{dt} = \frac{1}{4}\int (f'f'_1-ff_1)\ln\frac{ff_1}{f'f'_1}\, a\,du\,dv\,dw\,du_1\,dv_1\,dw_1\,d\omega. \tag{I.1,9} \]
*) In his first paper devoted to the \(H\)-theorem, Boltzmann used the notation \(E\) (entropy) instead of \(H\).
Since
\[ (p-q)\ln\left(\frac{p}{q}\right)>0,\quad \text{if } p\ne q;\quad \text{and }=0,\quad \text{if } p=q \tag{I.1,10} \]
and since \(a\) is a positive quantity, we have from (I.1,9)
\[ \frac{dH}{dt}\leq 0, \tag{I.1,11} \]
where equality in (I.1,11) occurs only in the case when, for any two velocities \(\mathbf{c}\) and \(\mathbf{c}_{1}\), we have \(ff_{1}=f'f'_{1}\).
Thus, from (I.1,11) it follows that, provided our basic assumptions are valid, the quantity \(H\) will decrease monotonically until it reaches the minimum value attained as soon as the distribution function satisfies the condition
\[ ff_{1}=f'f'_{1} \tag{I.1,12} \]
for any pair of velocities \(\mathbf{c}\) and \(\mathbf{c}_{1}\). The minimum, or equilibrium, value of \(H\) is in most cases reached rapidly if \(H\) differs appreciably from this value. The relaxation time in the case, for example, of a gas at room temperature and a pressure of one atmosphere is of the order of \(10^{-7}\) sec. (ESM, p. 390).
The formula for the equilibrium distribution function can be derived either by using (I.1,12), or by taking as the equilibrium condition the requirement that at equilibrium \(H\) have a minimum—in both cases under the additional conditions of constancy of the total number of particles, constancy of the total energy, and constancy of the total momentum. As a result one obtains the generalized Maxwell distribution.
Unfortunately, it is difficult to consider the case of a real gas in all details, and although a derivation of the \(H\)-theorem is nevertheless possible, it is not an easy matter to subject to critical analysis the various assumptions made in obtaining (I.1,11), and the consequences of abandoning the additional conditions. It is possible, however, to introduce simplified models \(^{11\mathrm{E}2,\,53\mathrm{H},\,54\mathrm{G},\,55\mathrm{G},\,55\mathrm{H}2}\), and we shall consider in particular a model which in essence coincides with that used by Lorentz \(^{09\mathrm{L}}\) in his works on the electron theory of metals. In this model there are two types of particles. Particles of the first type (lattice sites, or metal ions) are fixed in space, and we shall assume that they are distributed randomly in space with density \(n\) per unit volume. Particles of the second type (electrons) move through the lattice, and their density is equal to \(N\) per unit volume. We shall neglect electron–electron collisions and further assume that collisions of electrons with the lattice are elastic and isotropic. This means that (a) in such collisions the electron does not change the absolute magnitude of its velocity, and that (b) if \(p(\theta,\varphi;\theta',\varphi')\) is the probability
that an electron moving in a direction determined by the polar angles \(\theta\) and \(\varphi\) changes its motion to a direction determined by the angles \(\theta'\) and \(\varphi'\), then the value \(p(\theta,\varphi;\theta',\varphi')\) does not depend on \(\theta'\) and \(\varphi'\). Since the absolute value of the velocity does not change in a collision, we can simplify our model by assuming that all electrons move with the same speed \(c\). Finally, let us denote by \(\sigma\) the cross section for the collision of an electron with the lattice.
To consider approach to equilibrium, introduce the distribution function. Let \(f(\theta,\varphi)\,d\omega/4\pi\) \((d\omega=\sin\theta\,d\theta\,d\varphi\) is an element of solid angle) be equal to the number of electrons per unit volume whose velocities lie within the solid angle \(d\omega\). The normalization condition is then obtained in the form
\[ \int f(\theta,\varphi)\,d\omega = 4\pi N, \tag{I.1,13} \]
where the integration is extended over the unit sphere.
According to our assumption concerning the model, for the number of electrons per unit volume \(N_{\omega,\omega'}\,dt\,d\omega\,d\omega'\) changing the directions of their velocities from the element of solid angle \(d\omega\) to the solid angle \(d\omega'\) during the time \(dt\), we have
\[ N_{\omega,\omega'}\,dt\,d\omega\,d\omega' = f(\theta,\varphi)\,nc\sigma\,\frac{d\omega}{4\pi}\,\frac{d\omega'}{4\pi}\,dt, \tag{I.1,14} \]
where we have used the assumption of isotropy of scattering.
Using (I.1,13) and (I.1,14), one can easily compute the rate of change of \(f(\theta,\varphi)\) and obtain
\[ \frac{df(\theta,\varphi)}{dt} = \alpha \left[ \frac{1}{4\pi}\int f(\theta',\varphi')\,d\omega' - f(\theta,\varphi) \right] = \alpha[N-f], \tag{I.1,15} \]
where
\[ \alpha = nc\sigma . \tag{I.1,16} \]
From (I.1,15) we find
\[ f(\theta,\varphi) = N-Ae^{-\alpha t}, \qquad A=A(\theta,\varphi)=N-f_{t=0}(\theta,\varphi), \tag{I.1,17} \]
i.e. exponential approach to the equilibrium value, if at \(t=0\) equilibrium was absent.
Again, as in the case of (I.1,11), we find a monotone approach to the equilibrium state. However, as before, we made a fundamental assumption in writing (I.1,14), and again this assumption concerns the number of collisions; such an assumption is often called the hypothesis on the number of collisions (Stosszahlansatz). Since collisions play a decisive role in this consideration, we shall call such a method of considering the problem kinetic. In the next section, however, it will be shown that the unrestricted hypothesis on the number of collisions leads to contradictions and that, in view of this, one must base one’s consideration rather on the so-called statistical method, which duly takes possible fluctuations into account.
I.2. The paradox of reversibility and the paradox of recurrence
Let us formulate the \(H\)-theorem in a somewhat different form. We have studied systems composed of a large number of identical particles, and considered the distribution function corresponding to such a system and the rate of its change. Another way of studying the behavior of a system consists in describing its state, or phase, at each instant of time by the representative point in \(\Gamma\)-space (see ESM p. 100, where the definition of \(\Gamma\)-space is given). To each state there corresponds, on the one hand, a certain distribution and, on the other hand, a certain point in \(\Gamma\)-space. Thus there is a one-to-one correspondence between points in \(\Gamma\)-space, on the one hand, and values of the distribution function, on the other*). Let us now consider the sequence of times \(t_1, t_2, \ldots, t_n, \ldots\). The phase of our system at the instant \(t_n\) will correspond to the point \(P_n\) in \(\Gamma\)-space, and the orbit described by the representative point of the system will pass through the sequence of points
\[ \ldots,\ P_1,\ P_2,\ldots,\ P_{n-1},\ P_n,\ P_{n+1},\ldots \tag{I.2,1} \]
Since each point corresponds to a distribution function, we can calculate the value \(H\) corresponding to each of the points. Let \(H_n\) be the value of \(H\) in the state corresponding to \(P_n\). If the \(H\)-theorem is adopted in its unrestricted form, it follows from this that the following inequalities must hold:
\[ \ldots > H_1 \geq H_2 \geq \ldots \geq H_{n-1} \geq H_n \geq \ldots, \tag{I.2,2} \]
where the equality sign occurs only in the case when equilibrium has been reached.
In the general case \(H\) is given, instead of (I.1,7), by the expression
\[ H=\int f\ln f\,d\omega, \tag{I.2,3} \]
where the integration is extended over the whole \(\mu\)-space (ESM, p. 30); here \(f\) is a function of the \(s\) generalized coordinates \(q_k\) and \(s\) generalized momenta \(p_k\) (\(s\) is the number of degrees of freedom of one particle); \(d\omega\) is the volume element of \(\mu\)-space:
\[ d\omega=\prod_{k=1}^{s} dp_k dq_k. \tag{I.2,4} \]
* This is true, of course, only insofar as we confine ourselves to the consideration of relatively nonessential specific phases. If generic phases are used, then to each distribution there corresponds an aggregate of points in \(\Gamma\)-space [cf. section (I.3)].
Let us now consider two states, corresponding to the depicting points \(P_i\) and \(P'_i\), which differ only in that all \(p_k\) have the same absolute values but opposite signs, while the \(q_k\) are the same in both states. Since the Hamiltonian \(\mathscr{H}(p_k,q_k)\) of the system is a homogeneous quadratic polynomial in the \(p_k\), it is invariant under the transformation \(q_k \to q_k,\ p_k \to p'_k\), which transforms \(P_i\) into \(P'_i\). From the canonical equations of motion
\[ \dot q_k=\frac{\partial \mathscr{H}}{\partial p_k},\qquad \dot p_k=-\frac{\partial \mathscr{H}}{\partial q_k} \tag{I.2,5} \]
it then follows that the transformation from \(P_i\) to \(P'_i\) corresponds to reversing the time axis. If we now consider a system passing through the sequence (I.2,1), then after the transformation we obtain a system passing through the sequence
\[ \ldots,\ P'_{n+1},\ P'_n,\ P'_{n-1},\ldots,\ P'_2,\ P'_1,\ldots \tag{I.2,6} \]
From the definition of \(H\) (I.2,3) it follows that
\[ H_i=H'_i. \tag{I.2,7} \]
This can be verified by carrying out the transformation under the integral sign and then passing to the new variables
\[ H'=\int f'(p',q')\ln f'(p',q')\,d\omega' =\int f(-p,q)\ln f(-p,q)\,d\omega = \int f(p'',q'')\ln f(p'',q'')\,d\omega''=H, \tag{I.2,8} \]
where \(p''_k=-p_k,\ q''_k=q_k\).
From (I.2,7) and (I.2,2) it follows that the sequence (I.2,6) corresponds to a sequence of values of \(H\) satisfying the inequalities
\[ \ldots \leq H'_{n+1}\leq H'_n\leq H'_{n-1}\leq \ldots \leq H'_2\leq H'_1\leq \ldots \tag{I.2,9} \]
We see, therefore, that to every system for which \(H\) decreases monotonically one can associate a system for which \(H\) increases monotonically. This fact constitutes the so-called Loschmidt reversibility paradox\(^{76L,77L}\) (Reversibility paradox).
The difficulty contained in it is easily discovered by considering the simple model introduced at the end of Section I.1. Reversing the direction of time, we obtain a monotonic departure from the equilibrium distribution, as is clear from (I.1,17).
Another difficulty was pointed out by Zermelo\(^{96Z}\), who used Poincaré’s theorem\(^{90P}\) (see also \(^{97B}\)). Poincaré showed that if a system enclosed in a closed volume passes through the sequence (I.2,1), say, from \(t_1\) to \(t_{n+1}\), then after the expiration of a finite-
will again pass through the same sequence with any prescribed accuracy*). More precisely, for any finite element of length \(\Delta s\) in \(\Gamma\)-space one can find such an interval of time \(T\) that the sequence of states at the moments \(t_1+T, t_2+T,\ldots,t_n+T,t_{n+1}+T\) will be given by the points
\[ \ldots, P_1'', P_2'',\ldots, P_n'', P_{n+1}'',\ldots, \tag{I.2,10} \]
where
\[ |P_i-P_i''|<\Delta s. \tag{I.2,11} \]
The interval of time may be extraordinarily large. For the case of \(10^{18}\) atoms in \(10^{-6}\ \mathrm{m}^3\), moving with an average velocity of \(5\cdot 10^6\ \mathrm{m/sec}^{-1}\), according to Boltzmann’s estimate \(^{96B}\) (see also \(^{43C}\)) more than \(10^{10^{19}}\) years will be required to reproduce the positions with an accuracy of \(10\ \text{\AA}\) and the velocities with an accuracy of \(10^4\ \mathrm{m\cdot sec}^{-1}\).
If \(\Delta s\) is sufficiently small, then for the values \(H\) corresponding to the sequence (I.2,10) we have
\[ H_i'' \cong H_i, \tag{I.2,12} \]
and, consequently, we have the following sequence in the times \(t_1,t_2,\ldots,t_n,\ldots,t_1+T,\ldots,t_n+T\):
\[ \ldots H_1, H_2,\ldots, H_n,\ldots, H_1'', H_2'',\ldots, H_n'',\ldots \tag{I.2,13} \]
Using (I.2,2) and (I.2,12), we are convinced that, going from \(t_n\) to \(t_1+T\), we obtain an increase of \(H\). This paradox is the so-called recurrence paradox. In the following section we shall show in what way the statistical consideration clarifies the situation that has arisen. However, already now we should like to emphasize that, to the extent that the reversibility paradox and the recurrence paradox refute the \(H\)-theorem in its unrestricted form (i.e., the assertion that \(H\) can never decrease), they show that the hypothesis concerning the number of collisions cannot be valid under all circumstances. Let us explain this by two simple examples.
The first example gives us consideration of the simplest model discussed in the preceding section. Suppose that the system changes with time; let \(\Delta t\) be an arbitrary interval during this time, and let \(f(\theta,\varphi)\) be the distribution function at the beginning of this inter—
*) For the proof of Poincaré’s theorem we refer to the literature \(^{96Z,\,43C}\), and also to ESM, p. 341; the last proof is not quite rigorous. The proof follows, in essence, from the fact that a finite region in \(\Gamma\)-space passes through all attainable parts of phase space (which, in the case of a system enclosed in a finite volume, is finite) during a finite interval of time. On the connection of Poincaré’s theorem with Birkhoff’s ergodic theorem, see Winter’s monograph \(^{41W}\), pp. 90–91.
was. Let us again denote by \(N_{\omega,\omega'}\Delta t\,d\omega\,d\omega'\) the number of collisions during this interval that lead to a change of velocity, from lying within the solid angle \(d\omega\), to the solid angle \(d\omega'\). Let us now consider the corresponding change of the system, reversing the direction of time, and introduce the corresponding interval. We shall denote the distribution function at the beginning of the time interval in the second case by \(f'(\theta,\varphi)\). If \(N'_{\omega',\omega}\Delta t\,d\omega\,d\omega'\) is the number of collisions leading to a change of velocity from \(\omega'\) to \(\omega\) during the interval \(\Delta t\) of the reverse change, then we have
\[ N'_{\omega',\omega}=N_{\omega,\omega'} \tag{I.2,14} \]
for any pair of directions \(\omega\) and \(\omega'\). It follows then from (I.1,14) that, if the hypothesis concerning the number of collisions were valid for the forward and reverse changes of the system, then one would have
\[ f(\theta,\varphi)=f'(\theta',\varphi'), \tag{I.2,15} \]
independently of the values of \(\theta,\varphi,\theta'\), and \(\varphi'\), i.e.
\[ f(\theta,\varphi)=\mathrm{const}. \tag{I.2,16} \]
We see, therefore, that the hypothesis concerning the number of collisions cannot be valid for at least one of the two directions of change, unless we are dealing with a state of equilibrium.
The second example is as follows.*) Consider a stream of point particles moving in one and the same direction, say, in the positive direction of the \(x\)-axis, and scattered by a system of impermeable spheres randomly distributed in a plane perpendicular to the \(x\)-axis. As a result of scattering, after their collision with the plane in which the spheres are located, the point particles form a system of particles with isotropically distributed velocities. Let us now consider the development of the system obtained from the preceding one by reversing the direction of all velocities of all particles. It is clear that after a certain time all the particles will be moving in the negative direction of the \(x\)-axis, i.e. the system will pass from a disordered state to a highly ordered one. The reason for this is easy to understand by considering the collisions of the particles of the second system with the spheres. Although this system appears to be distributed completely randomly, all collisions with the spheres occur in those parts of the hemispheres that face the negative direction of the \(x\)-axis, and it is clear that the collisions do not occur in a random manner. We see, therefore, that for this second system the hypothesis concerning the number of collisions definitely cannot be fulfilled.
*) I am very grateful to Dr. H. M. James, who pointed out this example.
I.3. Statistical Aspects of the \(H\)-Theorem
It has already been mentioned that the nature of the objects studied by statistical mechanics is such that we necessarily arrive at a statistical treatment. The first to use a purely statistical approach was Boltzmann\(^{77\mathrm{B}}\), and it was then that kinetic theory was transformed into statistical mechanics, although more than twenty years still passed before Gibbs definitively formulated its foundations.
In his statistical treatment Boltzmann sought above all to answer question (A1), in particular for the case of an isolated system, and for this purpose proposed two possible definitions of the equilibrium state:
(a) The equilibrium state is the most probable state for a given energy.
(b) The equilibrium state is the mean state in which the system finds itself over an infinitely long interval of time.
In this section we shall first consider the approximation to the most probable state. In order to prove that the most probable state corresponds to the equilibrium state characterized (in the case of systems of independent particles, such as those considered by Boltzmann) by the Maxwell–Boltzmann distribution, we proceed as follows.
Take an isolated system of \(N\) independent particles. To each particle there corresponds a representative point \(Q^{(k)}\) in \(\mu\)-space, where \(k\) numbers the particles of the system \((k=1,2,\ldots,N)\). The phase of the whole system, and hence also the representative point in \(\Gamma\)-space, is then determined by the \(N\) points \(Q^{(1)},\ldots,Q^{(N)}\). We now divide \(\mu\)-space into very small, but finite, cells of equal volume \(\omega\), which we number consecutively \(\omega_1,\omega_2,\ldots,\omega_i,\ldots\). This subdivision constitutes a fundamental step in the reasoning. The quantity \(\omega\) is such that, on the one hand, the size of the cells is small in comparison with the smallest macroscopically measurable dimensions, but, on the other hand, the number of representative points \(Q^{(k)}\) contained in any one of them is large*).
*) At first sight this subdivision appears artificial in classical statistics, whereas it proves natural for quantum statistics. Uhlenbeck\(^{27\mathrm{U}}\), in discussing this question, remarks: “It looks as though Boltzmann had foreseen the appearance of discrete quantum states in \(\mu\)-space.” However, it has a deeper foundation, and in a broad sense one may say that even in classical statistics this subdivision is natural, since it corresponds to the limited nature of our macroscopic measurements. In this connection one may compare the discussion of the question of the coarse-grained density of an ensemble in Section III.1 and at the end of the present section.
Let \(N_i\) be the number of representative points in the \(i\)-th cell. The state \(Z\)* is then completely described by specifying \(N_i\).
The description of the state by means of the numbers \(N_i\), instead of specifying the actual positions \(Q^{(k)}\), corresponds to our experimental possibilities, since the size of the cells is chosen so that we are not able in experiment to distinguish different points within one and the same cell.
The relation between the representative point of the system in \(\Gamma\)-space and the state \(Z\) is described by means of \(N_i\) as follows:
(a) To each point in \(\Gamma\)-space there corresponds one state \(Z\).
(b) For each \(Z\) there is a volume in \(\Gamma\)-space such that every point in this volume corresponds to that same \(Z\). We shall call such a region a \(Z\)-star, and its volume is given by the expression
\[ W(Z)=\frac{N!}{\prod_i N_i!}\,\omega^N . \tag{I.3,1} \]
Expression (I.3,1) can be obtained if one takes into account that, under the action of the following two operations, the set of numbers \(N_i\) will not change and, consequently, \(Z\) remains the same:
(a) Each of the \(N\) points \(Q^{(k)}\) moves within its own cell in such a way that the volume \(\Omega=\omega^N\) is filled in \(\Gamma\)-space.
(b) Any of the permutations of the \(N\) points, with the condition that only points in different cells are permuted.
Let us now introduce the quantity \(H(Z)\) by means of the expression
\[ H(Z)=-\ln W(Z). \tag{I.3,2} \]
Substituting expression (I.3,1) for \(W(Z)\) into (I.3,2), using Stirling’s formula for factorials,
\[ \ln x! = x\ln x - x \tag{I.3,3} \]
and neglecting additive constants, we obtain
\[ H(Z)=\sum_i N_i \ln N_i . \tag{I.3,4} \]
Comparing (I.2,3) and (I.3,4), we find that \(H(Z)\) coincides with Boltzmann’s \(H\), provided that instead of integration one takes the sum over the cells introduced by us in \(\mu\)-space.
It is easy to show (see ESM, Section 1.7) that if the equilibrium distribution \(N_i^e\) is defined as such a distribution for which—
* We use Ehrenfest’s notation \(^{11E2}\): \(Z\) from the word Zustandverteilung.
that \(W(Z)\) is maximal, then, in the case of constancy of the total energy \(E\) and of the total number of particles \(N\), i.e.
\[ N=\sum N_i, \tag{I.3,5} \]
\[ E=\sum N_i E_i, \tag{I.3,6} \]
where \(E_i\) is the energy corresponding to the \(i\)-th cell, we obtain that \(N_i^e\) satisfy the relation
\[ N_i^e=\exp(\mu-\beta E_i), \tag{I.3,7} \]
which is the so-called Maxwell–Boltzmann distribution.
In identifying the equilibrium distribution with the distribution for which \(W(Z)\) is maximal, we have defined the probability of a state as the corresponding volume in \(\Gamma\)-space. To prove that the equilibrium distribution is the distribution for which \(W(Z)\) is maximal or \(H(Z)\) minimal, we must introduce the \(H\)-theorem (i.e. show that \(\dfrac{dH}{dt}\) is always negative, except in the case \(H=H_{\min}\)), or something of this kind. However, as soon as we begin to discuss the behavior of \(H(Z)\) as a function of time, we become convinced that any function depending on the coordinates and momenta of the particles of the system through \(N_i\) will be a discontinuous function; for as soon as one of the representative points in \(\mu\)-space passes from one cell into another, two corresponding \(N_i\) change by one. Thus, for the time dependence \(H(Z)\) we have a step function), and its time derivative will have only three possible values: \(-\infty\), \(0\), and \(+\infty\)*).
From the step function we then obtain the so-called \(H\)-curve, by selecting a discrete sequence of points separated by constant time intervals. The length of the time interval \(\tau\) is chosen so that it is small in comparison with the time intervals measured experimentally, but sufficiently large that during the time \(\tau\) a large number of collisions occurs***).
Now we are in a position to use the properties of the \(H\)-curve in order to state certain assertions****). In doing so, we shall
\[
\text{*) It is necessary to note that this fact does not depend on the character}
\]
of the intermolecular forces and is a consequence solely of the introduction of finite cells \(\omega\).
\[
\text{**) More strictly: the time derivative is everywhere zero, except for points}
\]
of discontinuity of \(H\), where it is undefined.
\[ \text{***) Compare with the choice of the sizes of the cells in \(\mu\)-space.} \]
\[
\text{****) The numbering of the assertions and the content of some of them differ}
\]
from those in Ehrenfest \(^{11E2}\) or in ESM.
to use the simple model introduced in Section I.1.
It can be shown \(^{54G,55G}\) that this model is representative of a large class of models distinguished by the very same behavior*).
I. One can find a function \(\Delta\) which, on the one hand, measures the deviation from equilibrium and, on the other hand, is directly connected with the entropy of the system, or \(H\)**).
II. If one follows the history of the system in time, then the equilibrium distribution, for which \(\Delta\) is equal to its minimum value, will be realized much more often than any other distribution. Moreover, the chance of observing another distribution is so insignificant that the equilibrium distribution coincides with the mean distribution.
III. If the system is in a state corresponding to \(\Delta > \Delta_{\min}\), then the development of the system will proceed in such a way that \(\Delta\), with greatest probability, will decrease.
IV. The preceding assertion will be true independently of whether we follow the \(\Delta\)-curve—which is obtained from the step function \(\Delta(t)\) in the same way as the \(H\)-curve is obtained from \(H(t)\)—from \(t=-\infty\) to \(t=+\infty\), or else from \(t=+\infty\) to \(t=-\infty\).
V. The \(\Delta\)-curve will practically always be located near \(\Delta_{\min}\).
VI. The interval of time between repetitions of values of \(\Delta\) different from \(\Delta_{\min}\) increases sharply with the growth of \(\Delta\).
VII. If \(D_n\) is the mean of all values of \(\Delta\) at the moments \(t_0+n\tau\), beginning with \(\Delta_0\) at \(t_0\)***), then the \(D\)-curve, i.e. the sequence of values \(D_1, D_2,\ldots, D_n,\ldots\), will decrease monotonically from the value \(\Delta_0\) and asymptotically approach \(\Delta_{\min}\).
VIII. The overwhelming majority of \(\Delta\)-curves will follow the \(D\)-curve for an appreciable interval of time, but practically not one of them will follow the \(D\)-curve all the time.
We shall call the sequence of values \(\Delta'_1, \Delta'_2,\ldots,\Delta'_n,\ldots\), which follow from \(\Delta_0\) after application of the unrestricted \(H\)-theorem, based on the hypothesis concerning the number of collisions, the collision-number curve****).
*) See also \(^{51B2}\) and \(^{52H}\).
**) We have not proved that, apart from additive and multiplicative constants, \(H\), given by expression (I.2.3), is the entropy of the system. For such a proof the reader may consult the literature (see, for example, ESM, pp. 22 and 43).
***) In considering assertions VII and VIII (and partly assertion III), one should bear in mind that \(H\) and, consequently, \(\Delta\) do not determine the state completely. For a given value of \(\Delta\) there are several possible states \(Z\).
****) Ehrenfest considered \(H\) instead of \(\Delta\); therefore he called this curve the curve of the \(H\)-theorem (see also \(^{54G,55G}\)).
IX. The curve of collision numbers and the \(D\)-curve coincide. Let us consider the Lorentz model. The function \(\Delta\) mentioned in assertion I will be determined by the relation
\[ \Delta=\sum_{v=-m}^{+m}(f_v-f_v^e)^2, \tag{I.3,8} \]
where \(f_v^e\) are defined by means of (I.3,11). In (I.3,8) we have introduced a subdivision of \(\mu\)-space (which in this case reduces to the unit sphere) into finite cells. Each cell represents an element of solid angle of size \(\delta\omega\), where
\[ \delta\omega=\frac{4\pi}{2m+1}. \tag{I.3,9} \]
The \(2m+1\) expressions for \(f_v\), which are functions of time and together determine the state of the system, are specified in such a way that \(f_v\) is the number of electrons moving in the direction characterized by the \(v\)-th element of solid angle. Only \(2m\) of the quantities \(f_v\) are independent, since they satisfy the relation
\[ \sum_{v=-m}^{+m} f_v(t)=N. \tag{I.3,10} \]
The equilibrium values \(f_v^e\) of the quantities \(f_v\) are equal to one another and are given by the expression
\[ f_v^e=\frac{N}{2m+1}. \tag{I.3,11} \]
Since it can be shown that the \(f_v\) are always close to \(f_v^e\) [cf. (I.3,16)] and that, in view of this, one may assume \(f_v-f_v^e \ll f_v^e\), to accuracy up to quantities of second order in \((f_v-f_v^e)/f_v^e\) we have
\[ H=\sum_v f_v\ln f_v=\sum_v f_v^e\ln f_v^e+\frac{\Delta}{2N}=H_{\mathrm{eq}}+\frac{\Delta}{2N}. \tag{I.3,12} \]
It follows from (I.3,8) and (I.3,12), first, that there will be no difference whether we use \(H\) or \(\Delta\) to describe deviations from equilibrium, and, second, that \(\Delta\) is a nonnegative function of \(f_v\), vanishing only when all \(f_v\) are equal to their equilibrium values.
As for assertion II, we can first of all calculate the relative frequency of occurrence of a state of the system corresponding to a nonequilibrium value of \(\Delta\). In this connection, for
of the normalized probability \(w(\Delta)\,d\Delta\) that \(\Delta\) takes a value from the interval \((\Delta,\Delta+d\Delta)\), we find the expression*)
\[ w(\Delta)d\Delta= \left[\frac{2m+1}{2N}\right]^m \left[\frac{\Delta^{m-1}}{(m-1)!}\right] \exp\left\{-\frac{(2m+1)\Delta}{2N}\right\}d\Delta, \tag{I.3.13} \]
The function \(w(\Delta)\) plays the role of the function \(W(Z)\) introduced at the beginning of the section, and we see that \(w(\Delta)\) indeed attains its maximum value at \(\Delta=0\), i.e., for the equilibrium state.
Let us now make the assumption that the time \(t(Z)\) of residence in the state described by \(Z\) is proportional to \(W(Z)\)**). In this case the time average \(\overline{G}\) of the phase function \(G\) will be given by the expression
\[ \overline{G}= \frac{\displaystyle\sum_Z G(Z)W(Z)} {\displaystyle\sum_Z W(Z)}. \tag{I.3.14} \]
The assumption just made by us about the proportionality of \(t(Z)\) and \(W(Z)\) is necessary for the continuation of our reasoning. This assumption is also closely connected with the assumptions made in theories studying the outcome of events in various kinds of lotteries. In the latter case, since each, for example, of the throws of a die is assumed to occur independently of all those that took place earlier, the proportionality of the probability of a particular result to the number of times that such a result is realized in an infinite sequence of throws has fundamental significance and in fact is used to define probability. However, when we are dealing with processes that in fact necessarily follow from the preceding history, the proportionality of the time of residence in a certain state to the probability of this state must be assumed. Moreover, this circumstance cannot be used to define probability, since the processes under consideration are not even Markovian. The difficulty consists in the fact that \(t(Z)\) is connected with the long-term behavior of the system, and since the orbit is completely determined if one of its points in \(\Gamma\)-space is known exactly, then the more
*) For a detailed discussion and derivation of the formulas of the remaining part of this section, see Appendix I.
**) There is no need for us to emphasize the circumstance that this assumption in fact consists of two. First it is assumed that the region \(W_E(Z)\), cut out from the energy surface on which the representative point of our system is located, will be proportional to \(W(Z)\). Then the assumption is made that \(t(Z)\) is proportional to \(W_E(Z)\). It may be mentioned that Einstein \(^{10E}\) made the same (combined) assumption as we do.
the longer we move near the orbit, the greater the number of data necessary for the complete determination of the orbit will be at our disposal and, consequently, we shall no longer be dealing with a purely statistical case. In this connection one may point to Tolman’s consideration of an analogous question \((38^{T},\) p. 148; see also \(50^{B})\). Thus the assumption made by us is closely connected with the Markovianization of the processes corresponding to the development of the system (cf. \(49^{S}, 54^{G}, 55^{G}, 55^{H2}\)).
It is also necessary to pay attention to the fact that the proportionality of \(t(Z)\) and \(W(Z)\) has not been proved by the ergodic or quasi-ergodic theorem, as one might have hoped, since this theorem proves only the proportionality of \(t(Z)\) and \(W_E(Z)\), but not the proportionality of \(W_E(Z)\) and \(W(Z)\). Thus this remains one of the basic assumptions in the approach to the \(H\)-theorem in the form in which it is being discussed by us at the present moment. Moreover, we apparently have not come any closer to proving this proposition, despite the enormous amount of work done both with simplified models and on the basis of the ergodic theorem.
Continuing the consideration of assertions II—IX, we shall restrict ourselves mainly to those phase functions which depend on \(p\) and \(q\) only through \(\Delta\). For such functions one can rewrite (I.3,14) in the form
\[ \overline{G}=\int w(\Delta)G(\Delta)\,d\Delta, \tag{I.3,15} \]
since \(w(\Delta)\) is normalized. We can compute, for example, the mean value of \(\Delta\) and the dispersion about this mean value
\[ \Delta_{\mathrm{cp}}=N,\qquad \bigl\langle(\Delta-\Delta_{\mathrm{cp}})^2\bigr\rangle_{\mathrm{cp}}=\frac{N^2}{m^2}. \tag{I.3,16} \]
In order to consider the change of \(\Delta\) with time, introduce a time interval \(\tau\) and compute the probability \(w(\Delta,\Delta')\) that \(\Delta\) changes its value from \(\Delta\) to \(\Delta'\) during this time interval \(\tau\). Having found these transition probabilities, we obtain from the relation
\[ \Delta'_{\mathrm{cp}}=\int w(\Delta,\Delta')\,\Delta'\,d\Delta' \tag{I.3,17} \]
the mean value of \(\Delta'\) at the end of the interval \(\tau\), whereas at the beginning of the interval the value was \(\Delta\); here \(w(\Delta,\Delta')\) must be normalized with respect to \(\Delta'\). For the mean rate of change of \(\Delta\) (and hence of \(H\) as well,—cf. equation (I.3,12)) we have
\[ \left(\frac{d\Delta}{dt}\right)_{\mathrm{cp}} = \frac{\Delta'_{\mathrm{cp}}-\Delta}{\tau}. \tag{I.3,18} \]
In the special case of the Lorentz model, for \(w(\Delta,\Delta')\) we find
expression (see Appendix I)
\[ w(\Delta,\Delta')=(16\pi NA\Delta)^{-1/2} \exp\left\{-\frac{[\Delta'-\Delta+A\Delta(2m+1)]^2}{16NA\Delta}\right\}, \tag{I.3,19} \]
where \(A\) is given by the expression
\[ A=2\alpha\tau, \tag{I.3,20} \]
and \(\alpha\) is determined by (I.1,16). From (I.3,17)—(I.3,19) we obtain
\[ \Delta'_{\mathrm{cp}}=(1-A)\Delta, \tag{I.3,21} \]
\[ \left(\frac{d\Delta}{dt}\right)_{\mathrm{cp}}=-\frac{A\Delta}{\tau}. \tag{I.3,22} \]
Assertion III follows directly from (I.3,22). Here there is the same exponential approximation to equilibrium as was found in Section I.1 [equation (I.1,17)]. We shall soon be convinced of the close connection existing between (I.1,15) and (I.3,22).
To prove assertion IV it is necessary to compute the mean value \(\Delta''_{\mathrm{st}}\) of the function \(\Delta\), which after a time interval \(\tau\) assumes the value \(\Delta\). This quantity is given by the expression
\[ \Delta''_{\mathrm{st}}= \frac{\int \Delta'' w(\Delta'')w(\Delta'',\Delta)\,d\Delta''} {\int w(\Delta'')w(\Delta'',\Delta)\,d\Delta''}, \tag{I.3,23} \]
where the denominator is not equal to unity, since the quantity \(w(\Delta'')w(\Delta'',\Delta)\) will be normalized only after integration over both variables \(\Delta\) and \(\Delta''\). From (I.3,23), (I.3,13), and (I.3,19) we find
\[ \Delta''_{\mathrm{st}}=(1-A)\Delta, \tag{I.3,24} \]
which proves assertion IV*).
Assertion V follows from the expression found for \(w(\Delta)\) and from the assumption that the probability of finding some value from the sequence \(\Delta\) is proportional to \(w(\Delta)\).
To prove assertion VI we use Chandrasekhar’s formula\(^{43\mathrm{C}}\) for \(\Theta(\Delta)\)—the recurrence time of a state characterized by the value \(\Delta\)**):
\[ \Theta(\Delta)= \frac{\tau[1-w(\Delta)]} {w(\Delta)-w(\Delta)w(\Delta,\Delta')}, \tag{I.3,25} \]
\[ \underline{\hspace{2.5cm}} \]
*) It is necessary to note here that in an earlier work\(^{53\mathrm{H}}\) we used, in order to prove the equivalence of the sequences with \(t=-\infty \to t=+\infty\) and sequences with \(t=+\infty \to t=-\infty\), the relation \(w(\Delta)w(\Delta,\Delta')=w(\Delta')w(\Delta',\Delta)\), which also holds in the present model.
**) The objections raised by Bartlett\(^{50\mathrm{B},\,53\mathrm{B}}\) against Chandrasekhar’s formula do not apply if one considers sequences that are continuously Markovized (see \(^{50\mathrm{B}}\)).
or
\[ \Theta(\Delta)\approx \frac{\tau}{w(\Delta)}, \tag{I.3,26} \]
whence assertion VI follows (see (I.3,13)).
To compute \(D_n\), i.e. the mean value of the quantities \(\Delta\) at the instant \(t_0+n\tau\), starting from \(\Delta_0\) at \(t_0\), it is first of all necessary to compute the normalized probability \(w(\Delta_n;\Delta_0,\Delta_1,\Delta_2,\ldots,\Delta_{n-1})\) of a sequence of values \(\Delta_0\) at \(t_0\), \(\Delta_1\) at \(t_0+\tau\), \(\Delta_2\) at \(t_0+2\tau,\ldots,\Delta_{n-1}\) at \(t_0+(n-1)\tau\), and \(\Delta_n\) at \(t_0+n\tau\). This quantity is given by the expression (cf. \({}^{55}G\))\(^*\)
\[ w(\Delta_n;\Delta_0,\Delta_1,\ldots,\Delta_{n-1}) = w(\Delta_0,\Delta_1)\,w(\Delta_1,\Delta_2)\cdots w(\Delta_{n-1},\Delta_n). \tag{I.3,27} \]
For \(D_n\) we then obtain
\[ D_n=\int \Delta_n w(\Delta_n;\Delta_0,\Delta_1,\Delta_2,\ldots,\Delta_{n-1})\,d\Delta_1\,d\Delta_2\cdots d\Delta_n, \tag{I.3,28} \]
whence it follows that
\[ D_n=(1-A)^n\Delta_0, \tag{I.3,29} \]
and, putting \(t=n\tau\), we have
\[ D_n\approx \Delta_0 e^{-At/\tau}=\Delta_0 e^{-2\alpha t}. \tag{I.3,30} \]
Assertion VII follows directly from (I.3,29) or (I.3,30).
One can also compute the dispersion \(\sigma_n\) of the quantities \(\Delta_n\), for which we have
\[ \sigma_n=\int (D_n-\Delta_n)^2 w(\Delta_n;\Delta_0,\Delta_1,\ldots,\Delta_{n-1})\,d\Delta_1\cdots d\Delta_n. \tag{I.3,31} \]
or
\[ \sigma_n=0. \tag{I.3,32} \]
From (I.3,32) and from the fact that nonequilibrium values of \(\Delta\) recur with frequency \(\Theta(\Delta)^{-1}\), assertion VIII follows.
Finally, let us consider the curve of the numbers of collisions. We note that from (I.3,08) it follows that
\[ \frac{d\Delta}{dt}=\sum_v 2(f_v-f_v^e)\frac{df_v}{dt}. \tag{I.3,33} \]
Under the condition that the hypothesis on the number of collisions is valid, the rate of change of \(f_v\) with time is given by the equation
\[ \frac{df_v}{dt}=\frac{A}{\tau}(f_v^e-f_v) \tag{I.3,34} \]
—cf. equation (I.1,15); the proof of equation (I.3,34) is given in Appendix I.
\[ \text{*) It may be noted that in writing (I.3,27) the assumption of continuous Markovization has been used.} \]
Using equations (I.3.33) and (I.3.34), we obtain for the collision-number curve the differential equation
\[ \frac{d\Delta}{dt}=-\frac{A\Delta}{\tau}. \tag{I.3.35} \]
From comparison of (I.3.30) and (I.3.35) follows the proof of assertion IX for the model under consideration.
Let us summarize the consideration of the existing state of the \(H\)-theorem in classical statistics. We note first of all that, although since Ehrenfest’s very detailed consideration of this problem\(^{11E2}\) the situation has improved thanks to the study of simplified models, much still remains to be done. As far as this concerns the problem of approach to equilibrium and fluctuations about the equilibrium state, the study of simplified models has shown that the situation is in many respects as expected (assertions II—VI), provided the processes can be treated as purely Markovian. We discussed the behavior of the system mainly from the point of view of fluctuations; however, we could do the same for the return to equilibrium. Siegert\(^{49S}\), for example, showed that, starting from any \(\Delta\) at \(t=t_0\), the probability of finding the value \(\Delta\) at \(t=\infty\) is always given by the quantity \(w(\Delta)\). We tacitly assumed that we are always dealing with a system which has essentially “forgotten” its initial state. This can be done when the relaxation time of the system is sufficiently small. The relaxation time \(t_{\mathrm{rel}}\) is essentially determined by the expression \(\tau/A\), which gives the value \(t_{\mathrm{rel}}\sim 10^{-8}\ \mathrm{sec}\) under the assumption \(n=10^{30}\ \mathrm{m}^{-3}\), \(c=10^{4}\ \mathrm{m}\cdot\mathrm{sec}^{-1}\), \(\sigma=10^{-20}\ \mathrm{m}^{2}\), \(m=10^{6}\) (cf. ESM, p. 390). Such rapid relaxation, together with large time intervals between repetitions of a noticeable deviation from equilibrium\(^*\), are the reasons why there are very few exceptions to the second law of thermodynamics, or, expressed somewhat differently, why fluctuations very rarely play an essential role in the behavior of macroscopic physical systems. According to Smoluchowski\(^{12S}\), a system will behave irreversibly if its initial state is characterized by a mean recurrence time that is large in comparison with the time interval accessible to experiment. In this connection one may see confirmation of Tolman’s supposition (\(^{38T}\), p. 179) that the relaxation time will be very small in comparison with the period of a Poincaré cycle (Section I.2), provided that the size of the elementary cells is not too large. Since \(A\) is inversely proportional to the number of cells \(2m+1\), in fact—
\(^*\) For a mean deviation from \(f_v^{e}\) of order \(10^{-8} f_v^{e}\), we obtain for \(\Theta(\Delta)\) a value of order \(10^{10^{9}}\) years!
It is therefore clear that \(t_{\mathrm{rel}}\) increases with \(m\), and we have just convinced ourselves that \(t_{\mathrm{rel}}\) assumes an acceptable value for an acceptable value of \(m\).
The chief remaining problem is the justification of the Markovianization of processes [Ehrenfest’s assertion VII \(^{11E2}\) or assertion IX ESM]. This problem appears very difficult, but with the use of the powerful computing machines now available it is possible to compute the \(H\)-curve of a given system for a set of initial conditions, using the exact equations of motion.
In conclusion we should like to point out that, although the ergodic—or, better to say, quasi-ergodic—theorem has in a certain sense resolved the question of the equivalence of time averages and ensemble averages, nevertheless, owing to the incompleteness and inaccuracy of the experimental data available, we can determine the state only with limited precision. Even if we are dealing with an isolated system, so that its representative point remains on the energy surface, we still do not know on precisely which surface; and the only true representative ensemble will be one that consists of all systems whose representative points at the instant \(t_0\) fill a \(Z\)-star. Consequently, the \(H\)-theorem still remains connected with the problem that we are now discussing.
II. THE CLASSICAL ERGODIC THEOREM
II.1. Ergodic and quasi-ergodic theorems
As was already mentioned in the introduction, many authors prefer to approach the ergodic theorem by means of the \(H\)-theorem. At the end of Section I.3 the objections to such a point of view were set forth, and we shall return to this question in Part III. However, the ergodic theorem plays so important a role in the development of statistical mechanics that it seems unquestionably useful to discuss this problem in the present part as well.
At the beginning of Section I.3 it was mentioned that Boltzmann gave two different definitions of the equilibrium state, namely as the most probable state or as the mean state. The first possibility was considered by us above; now we shall discuss the second possibility. Boltzmann \(^{71B1,\,71B2}\) (see also \(^{68B}\)) and Maxwell \(^{79M}\) believed that, more or less consistently, the following theorem could be proved: the average behavior of a system of independent particles corresponds to the Maxwell–Boltzmann distribution. The average is taken in the sense of averaging over time for an infinite period.
Boltzmann believed that there exist so-called ergodic systems, and in this he saw a basis for hopes concerning the validity of the theorem just stated. An ergodic system—
is such a system that its representative point in \(\Gamma\)-space will pass through every point of the corresponding energy surface. If the existence of ergodic systems is admitted, it is easy to see the advantages of considering time averages. First of all, one must bear in mind that in any physical experiments we always measure time averages. Moreover, since relaxation times are usually small in comparison with the period during which the experimentally determined quantities are measured, it is quite immaterial whether we take the average over a finite time interval or over an infinite period. Further, if it were possible to prove that real physical systems are ergodic, then for any phase function \(\varphi(p,q)\) the following equalities would hold:
\[ \langle \varphi\rangle_{\mathrm{cp}}=\overline{\langle \varphi\rangle}_{\mathrm{cp}}=\langle \overline{\varphi}\rangle_{\mathrm{cp}}=\overline{\varphi}, \tag{II.1,1} \]
where the bar denotes the time average, and the brackets \(\langle\rangle_{\mathrm{cp}}\) denote the average taken over the microcanonical ensemble. The second and third terms in (II.1,1) give, respectively, the time average of the ensemble average and the ensemble average of the time average. We shall now outline a proof of the equalities (II.1,1). We note here only that from these equalities follows the possibility of computing ensemble averages instead of time averages; and since the former are obtained much more simply, this gives us an undoubted advantage.
The time average \(\overline{\varphi}\) of a phase function is defined by the equality
\[ \overline{\varphi}=\lim_{T\to\infty}\frac{1}{2T}\int_{t-T}^{t+T}\varphi(p,q)\,dt, \tag{II.1,2} \]
where \(p\) and \(q\) are the values of the coordinates of the representative point of the system in \(\Gamma\)-space at the time \(t\). In order to verify whether or not the average behavior of some system corresponds to the Maxwell—Boltzmann distribution, we choose a sequence of phase functions which together determine the distribution, and compare their time averages with the values corresponding to the Maxwell—Boltzmann distribution.
The ensemble average \(\langle\varphi\rangle_{\mathrm{cp}}\) is taken over the microcanonical ensemble. This average is given by the expression
\[ \langle \varphi\rangle_{\mathrm{cp}}= \frac{\int \varphi(p;q)\,\sigma(p,q)\,d\mathfrak{S}} {\int \sigma(p,q)\,d\mathfrak{S}}, \tag{II.1,3} \]
where both integrals extend over the entire energy surface \(\mathfrak{S}\); \(\sigma(p,q)\) is the surface density corresponding
...to the microcanonical ensemble. This surface density is given by the expression*)
\[ \sigma(p,q)= \left\{ \sum_{k=1}^{S} \left[ \left(\frac{\partial \mathscr{H}}{\partial p_k}\right)^2+ \left(\frac{\partial \mathscr{H}}{\partial q_k}\right)^2 \right] \right\}^{-1/2}, \tag{II.1.4} \]
where \(S\) is the number of degrees of freedom of the system and, consequently, the dimension of \(\Gamma\)-space; \(\mathscr{H}\) is the Hamiltonian of the system.
From Liouville’s theorem (ESM, p. 102) it follows (see, for example, ESM, p. 355) that the only stationary ensembles on the energy surface will be those for which the surface density along an orbit satisfies (II.1.4).
Let us now consider the equalities (II.1.1). The first equality follows immediately from the fact that the microcanonical ensemble is a stationary ensemble. The second equality follows from the possibility of interchanging the order of averaging. There remains the last equality, which is valid for ergodic systems, since \(\overline{\varphi}\) is the same for all systems of the ensemble. This is seen from the following. Since for an ergodic system the orbit of the representative point passes through each of the points of \(\mathfrak{S}\), and since from the canonical equations of motion (I.2.5) it follows that at each point the direction of the orbit is uniquely determined, there is therefore only one orbit on \(\mathfrak{S}\). The only difference for different systems of the ensemble will be that the exact moment of time at which the orbit passes through a point will be different for them.**) Since all systems follow one and the same orbit, it follows from (II.1.2) that \(\overline{\varphi}\) is the same for all systems and thus we obtain the last equality in (II.1.1).
For an ergodic system one may go even one step further and prove that the following equality holds\(^{71\mathrm{B}}\):
\[ \lim_{T\to\infty}\frac{dt}{T} = \frac{\sigma\, d\mathfrak{S}}{\int \sigma\, d\mathfrak{S}}, \tag{II.1.5} \]
where \(dt\) is the time during which the representative point was inside the surface element \(d\mathfrak{S}\) over the period \(T\). From (II.1.5), in principle, one can calculate the frequency of occurrence of nonequilibrium states.
The proof is now reduced to establishing the ergodicity of mechanical systems. However, it is easy to see that, apparently, none of the systems is ergodic (\(38^{\mathrm{T}}\), p. 67). Recall that a system for which equations (I.2.5) hold possesses \(2S\) integrals of motion. One of these integrals is the energy; another fixes the time along the orbit. Choosing
*) The expression inside the braces in (II.1.4) is the absolute value of the “velocity” of the representative point in \(\Gamma\)-space.
**) We exclude strictly periodic orbits.
point \(P\), on the energy surface, we fix the values of the remaining \(2S-2\) integrals of motion. Taking other values for these constants, but the same value of the energy, we obtain another point \(Q\) on the same energy surface; however, there is no orbit connecting the points \(P\) and \(Q\). To complete the proof it was therefore necessary to show that \(2S-2\) constants of motion are not constants almost everywhere on the energy surface, which seems implausible.
By using the expression “almost everywhere,” we have entered the domain of measure theory ) and it is precisely with the aid of this theory that one can prove the nonexistence of ergodic systems\(^{13R,13P}\) (see also \(^{52R}\)). The essence of this proof is that the points of an orbit form a sequence of measure zero on the energy surface, whose measure is different from zero ). This can be seen in the following way\(^{52R}\). Consider a point \(P\) of the orbit of the representative point and take a small region \(\mathfrak A\) of finite measure around \(P\). Since it is clear that the orbit cannot remain inside \(\mathfrak A\) for an unlimited time if it is to cover all of \(\mathfrak S\), the only part of the orbit within \(\mathfrak A\) will be a collection of separate segments **), corresponding to a sequence of finite time intervals. Since the sequence of such time intervals is countable, the segments of the orbit inside \(\mathfrak A\) can also be renumbered and, consequently, have measure zero.
Although the impossibility of the existence of ergodic systems was shown by Rosenthal and Plancherel only in 1913, P. and T. Ehrenfest (\(^{11E2}\)—especially note 89a,—see also \(^{91K}\), p. 484) were evidently quite convinced of the implausibility of ergodic systems and of the fact that the points of an orbit and all points of the energy surface have different measures. To avoid this difficulty it was assumed that mechanical systems may be quasi-ergodic. The representative point of a quasi-ergodic system ****) will come arbitrarily close to any—
*) It may be noted here that many mathematicians working in statistics (see, for example, \(^{53D}\)) regard statistics merely as a branch of measure theory. For discussion of measure theory we refer to the recently published book by Halmos \(^{50H}\) (see also \(^{53D}\)).
**) The measure of a point sequence \(\mathfrak m\) on \(\mathfrak S\) is defined as follows: let \(f(P)\) be a function equal to 1 if \(P\) belongs to \(\mathfrak m\), and equal to 0 otherwise. The Lebesgue integral \(\int f\,dS\), extended over all of \(\mathfrak S\), is then called the Lebesgue measure of \(\mathfrak m\) on \(\mathfrak S\) and is denoted by \(\mathfrak M\mathfrak m\).
***) They must be distinguished, since for a nonperiodic orbit no point on \(\mathfrak S\) is ever traversed twice.
****) It may be noted here that, in view of the impossibility of ergodic systems in the original sense of the word, modern authors often call quasi-ergodic systems ergodic.
...any point of the energy surface, without, however, passing through every point.
The use of the concept of quasi-ergodic systems did not, however, solve the problem of proving the equalities (II.1.1), although Rosenthal \(^{14\mathrm{R}2}\) did give a not quite rigorous proof of the relation (II.1.5), based on the quasi-ergodicity of mechanical systems, and also showed the equivalence of time averages and ensemble averages. Fermi \(^{23\mathrm{F}1}\) proved that a certain class of systems, the so-called canonical normal systems*) are quasi-ergodic. His proof is briefly set forth in Appendix II.
For further progress along this path it was necessary to find an independent proof of the equivalence of time averages and ensemble averages, which became the aim of more recent investigations devoted to the ergodic theorem, which are described in the following section. The proof of the equivalence of time averages and ensemble averages is now referred to as the ergodic theorem, although from the historical point of view it would apparently be more correct to use the term quasi-ergodic theorem.
II. 2. Recent investigations**)
Modern ergodic theory begins with Birkhoff’s article \(^{31\mathrm{B}2}\), in which he considered the following time averages:
\[ \bar f(P; t_0, T)=\frac{1}{T}\int_{t_0}^{t_0+T} f(P_t)\,dt, \tag{II.2.1} \]
where \(P_t\) is the point of the orbit passing through \(P\) at the moment \(t_0\), at time \(t\). He showed, first of all, that for practically all points \(P\) of the energy surface \(\mathfrak{S}\) there exists the limit of the quantity \(\bar f(P; t_0, T)\) as \(T\to\infty\). Moreover, Birkhoff showed that \(\bar f(P; t_0, T)\) in the limit \(T\to\infty\) is independent of \(t_0\) for practically all points \(P\). A proof of this property of time averages on the energy surface is given in Appendix III. “Practically all” is understood here, as was proposed by Birkhoff in 1922 \(^{22\mathrm{B}}\) (see also \(^{26\mathrm{S}2}\)), in the sense that the exceptional orbits form on \(\mathfrak{S}\) a set of measure zero.
Further, Birkhoff showed that \(\bar f(P;\infty)\), i.e., the time average taken over an infinite interval, is constant almost everywhere on \(\mathfrak{S}\), provided that the group of transformations \(P\to P_t\) is
*) It is interesting to note that Einstein \(^{02\mathrm{E},\,03\mathrm{E}}\), independently of Gibbs, who proposed the theory of ensembles, himself restricted himself to such systems; see also \(^{11\mathrm{K}2}\).
**) See also \(^{20\mathrm{M}1,\,20\mathrm{M}2,\,27\mathrm{H}2}\).
metrically transitive, or, in other words, under the condition that \(\mathfrak{G}\) is metrically indecomposable). The constancy of \(\overline f(P;\infty)\) follows from the fact that, if it were not satisfied, one could find such a value \(F\) of the quantity \(\overline f(P;\infty)\) that the conditions \(\overline f(P;\infty)<F\) and \(\overline f(P;\infty)\geq F\) would define on \(S\) two sequences of positive measure, both invariant with respect to the transformations \(P\to P_t\)*).
Birkhoff went even further and proved a relation equivalent to equation (II.1,5), namely, that for practically all points \(P\) on \(\mathfrak{G}\) we have
\[ \lim_{T\to\infty}\frac{dt(P;\mathfrak{m})}{T}=\frac{\mathfrak{M}\mathfrak{m}}{\mathfrak{M}\mathfrak{G}}, \tag{II.2,2} \]
where \(dt(P;\mathfrak{m})\) is the time which the representative point of the orbit passing through \(P\) spends in the region \(\mathfrak{m}\) during the interval \(T\). Let us emphasize once again that (II.2,2) is valid only in the case when the group of transformations \(P\to P_t\) is metrically transitive***).
Birkhoff’s proof is somewhat stronger than that given by Neumann \(^{32N1,\,32N2}\), who showed that the convergence \(\overline f(P;t_0,T)\) holds in the mean, whereas Birkhoff showed that the convergence holds practically for all \(P\).
The problem connected with the ergodic theorem has now been reduced to the problem of proving that energy surfaces are, as a rule, metrically indecomposable. In 1941, Oxtoby and Ulam \(^{41O}\) proved this for a fairly general class of surfaces that are polyhedra of dimension equal to three or more. Thus they showed, in their own words,
*) A group of transformations \(P\to P_t\) is called metrically transitive, and \(\mathfrak{G}\) is called metrically indecomposable, if \(\mathfrak{G}\) cannot be divided into two parts \(\mathfrak{G}_1\) and \(\mathfrak{G}_2\), both of positive measure and both invariant with respect to every transformation of the group.
**) We do not wish here to enter into a discussion of the complications caused by the presence of other integrals of the motion, such as momentum, angular momentum, etc. They lead, in the sense expressed by Rosenfeld, to an inessential decomposition of \(\mathfrak{G}\), whereas we are interested only in the essential indecomposability of \(\mathfrak{G}\). In our exposition we shall assume that no other homogeneous integrals of the motion, apart from energy, exist; moreover, we shall consider metric indecomposability in the physical sense. For a discussion of the complications thereby avoided, we refer to the literature (see, for example, \(^{52R}\)). It is also necessary to point to the work of Grad \(^{52G1,\,52G2}\), who examined in detail the statistical mechanics of systems possessing integrals different from energy.
***) From (II.2,2) one sees the relation between Birkhoff’s ergodic theorem and Poincaré’s recurrence theorem (cf. \(^{41W}\), pp. 90–91). Equation (II.2,2) expresses the fact that the representative point will be in any region of positive measure for a finite interval of time.
“that the ergodic hypothesis in its modern form of metric transitivity is, at least, free from any objections based on topological considerations.” According to Gamow \({}^{49K1}\), p. 54; the discharge belongs to Gamow) this instills confidence “that in a certain sense almost every continuous transformation is metrically transitive.” Gamow, unfortunately, gives no arguments in support of his assertion. However, the energy surfaces considered by Oxtoby and Ulam do not fully correspond to physical systems, and in view of this one may say that, although Birkhoff’s ergodic theorem has brought the problem very close to its solution, nevertheless for the time being we are still dealing with a hypothesis. Birkhoff and Koopman \({}^{32B}\) say: “The quasi-ergodic hypothesis has been replaced by its modern version: the hypothesis of metric transitivity.” It seems, however, to have been overlooked that the metric indecomposability of the energy surface has been proved for a certain class of physical systems, some of which exist. To make this clear, let us note, first, that for any quasi-ergodic system the energy surface is metrically indecomposable; the latter justifies Ehrenfest’s assumption that for quasi-ergodic systems relation (II.1,5) holds. Indeed, quasi-ergodicity assumes that each orbit on the energy surface will pass through any chosen region \(\mathfrak A\) of positive measure. Since an orbit is formed by the set of transformations \(P \to P_t\), where \(t\) varies from \(-\infty\) to \(+\infty\), it follows that the condition of metric indecomposability is satisfied for the energy surface of any quasi-ergodic system. Secondly, let us note that Fermi \({}^{23F1}\) proved the quasi-ergodicity of canonical normal systems and that Poincaré \({}^{92P}\) showed that, in the simplified three-body problem, we are dealing precisely with such a system*). Of course, in this case the number of degrees of freedom is very small, and it would be of definite interest to show that other physical systems also belong to the same class.
Canonical normal systems are characterized by the following properties. After excluding homogeneous integrals of motion, such as the components of the total momentum, it is possible to introduce a sequence of canonically conjugate variables \(x_i\) and \(y_i\), such that: 1) the energy \(E\) does not depend on time, 2) in the system there exists a parameter \(\alpha\) such that the Hamiltonian \(\mathscr H\) can be represented in the form of a power series in \(\alpha\)
\[ \mathscr H=\mathscr H_0+\alpha \mathscr H_1+\alpha^2\mathscr H_2+\cdots, \tag{II.2,3} \]
*) The parameter \(\alpha\) occurring in (II.2,3) is, in this case, the ratio \(m_2/m_1\), where the three masses \(m_1, m_2, m_3\) satisfy the inequalities \(m_3 \ll m_2 \ll m_1\).
3) the first term \(\mathscr H_0\) in (II.2,3) does not depend on \(x_i\), but the remaining terms may depend both on \(x_i\) and on \(y_i\); 4) all \(\mathscr H_i\) are periodic quantities with common period in \(x_i\).
Before briefly setting forth Hopf’s ergodic theorem, we should like to draw attention to the following two circumstances. First, Einstein and other authors \(^{02E,03E,06P,11K2}\) explicitly indicated that they believed all physical systems to belong to the class of canonical normal systems. Secondly, we wish to point out that in the case of systems satisfying the condition of metric indecomposability of their energy surfaces, the various integrals of motion are almost*) constant on the entire energy surface and that, therefore, the Boltzmann–Maxwell assumption of ergodicity, which, as we have seen, requires exact constancy of the integrals of motion almost everywhere on the energy surface, has not been so completely eliminated from modern considerations as one might have thought.
Birkhoff’s ergodic theorem is restricted to the consideration of orbits lying on a single energy surface. However, as we emphasized above, a consideration closer to physical reality would be one that takes into account the totality of energy surfaces corresponding to energies lying within a finite interval, for example:
\[ E_0-\delta < E < E_0+\delta \tag{II.2,4} \]
(cf. ESM, p. 99), where \(\delta \ll E_0\), so that one could still speak of a system with more or less definite energy. For such an energy layer Hopf \(^{37H}\) (see also \(^{30H2,32H1,32H2,32H3,32H4,34H,52R}\)) gave a slightly modified ergodic theorem, which again proves equality of the time and ensemble averages. A necessary condition for the validity of Hopf’s ergodic theorem is that not only must each energy surface in the layer be metrically indecomposable, but so must also be almost every energy surface of any arbitrary space obtained by the introduction of a phase space whose coordinates consist of the aggregate of pairs of coordinates of the \(\Gamma\)-space*). If this condition is fulfilled, any distribution in the energy layer will in the end become more or less homogeneous (see also Part III). Hopf’s ergodic theorem does not introduce anything essentially new. The most important thing, not to mention its
*) It is undesirable for us to make this assertion with greater rigor, since this would lead to lengthy arguments.
**) These references contain an extensive bibliography relating to the mathematical aspects of the quasi-ergodic theorem.
***) The phase space is, so to speak, doubled. This question is discussed in more detail in Hopf’s works (see also \(^{52R}\)).
of value from the mathematical point of view lies in the fact that the quantum-mechanical analogue of the ergodic theorem is more closely connected with Hopf’s theorem than with Birkhoff’s theorem, as will be seen in Part V.
III. The \(H\)-Theorem in the Classical Theory of Ensembles
III.1. Fine-Structure and Coarse-Structure Densities
Up to now we have considered only the behavior of isolated systems. As has already been mentioned several times, such an approach, in our opinion, does not do justice to a purely statistical treatment. It is an approach characteristic rather of the kinetic theory of gases, and in view of this it is not surprising that both the \(H\)-theorem in its original form—which we discussed in Part I—and the classical ergodic theorem (Part II) trace their origin to the works of Boltzmann, the culmination of which was his “Theory of Gases” \(^{96\mathrm{B},98\mathrm{B}}\). When Gibbs developed his statistical mechanics and introduced the theory of ensembles, a new element was brought into consideration. Gibbs tried to show (\(^{02\mathrm{G}}\), especially Chapter XII; see also \(^{07\mathrm{L}}\)) that an ensemble of systems will evolve in such a way that it begins to approach the micro- or macrocanonical ensemble*). For this it is necessary to introduce a somewhat modified \(H\)-theorem, and this will be discussed in the present section. Our treatment will be based on the so-called grand ensembles (grand ensemble), partly because, in our opinion, the consideration of such ensembles is a logical consequence of the statistical approach. Such ensembles were first introduced by Gibbs in the last chapter of his monograph (\(^{02\mathrm{G}}\), Ch. XV). Having recognized the importance of the concept of representative ensembles, we automatically arrive at grand ensembles, as will be seen from the next section (see also ESM, p. 235 and \(^{55\mathrm{H}}\)).
To simplify the discussion we shall have in mind only systems consisting of particles of one kind. Let \(\nu\) denote the number of particles in the system. A system consisting of \(\nu\) particles, each of which has \(s\) degrees of freedom, can be described by \(\nu s\) generalized coordinates \((q)\) and \(\nu s\) generalized momenta \((p)\). Its phase space, or \(\Gamma\)-space, therefore has dimension \(2s\nu\). In considering a grand ensemble we deal with a system having a variable number of degrees of freedom. Let \(D(\nu; p, q)\,d\Omega_\nu\) denote the number of systems in the ensemble with \(\nu\) particles and with representative points lying within the volume \(d\Omega_\nu\), in
* We often use the term macrocanonical for ensembles which Gibbs called canonical. The reasons for this departure from the classics are given elsewhere \(^{54\mathrm{H}2}\).
in the corresponding phase space (the representative point is called a point of the phase space of the system, with coordinates equal to the values of the generalized coordinates and momenta of the particles of the system). Here we have
\[ d\Omega_\nu=\prod_{k=1}^{s\nu} dp_k dq_k . \tag{III.1,1} \]
The dependence of \(D\) on \(p_k\) and \(q_k\) is denoted by the symbols \(p\) and \(q\). Let \(N_{\mathrm{ens}}\) be the number of systems of the ensemble, so that
\[ N_{\mathrm{ens}}=\sum_{\nu=1}^{\infty}\int D(\nu;p,q)\,d\Omega_\nu, \tag{III.1,2} \]
where the integration is in each case extended over the whole phase space. We now introduce the density of the ensemble \(\rho(\nu;p,q)\)*:
\[ \rho=\frac{D}{N_{\mathrm{ens}}}; \tag{III.1,3} \]
from (III.1,2) and (III.1,3) it follows that \(\rho\) satisfies the normalization condition
\[ \sum_\nu \int \rho(\nu;p,q)\,d\Omega_\nu=1. \tag{III.1,4} \]
The probability index \(\eta\) is defined by the relation
\[ \eta=\ln\rho. \tag{III.1,5} \]
The mean value \(\langle G\rangle\) of any phase function \(G(\nu;p,q)\) is given by the expression
\[ \langle G\rangle=\sum_\nu \int \rho G\,d\Omega_\nu . \tag{III.1,6} \]
Up to now we have not yet taken into account the complication introduced by the fact that in a system containing only particles of one kind, we are dealing with \(\nu\) identical quantities, in view of which it is necessary to decide whether states differing only by a permutation of some particles are to be regarded as different, or as identical. If they are regarded as different, then we are dealing, so to speak, with specific phases (specific phases); if they are regarded as identical, then we are dealing with generic phases (generic phases). It is easy to see that each generic phase contains \(\nu!\) specific phases. Although even before the appearance of quantum mechanics there were weighty grounds for preferring generic phases to specific phases (see, for example, 02G or ESM, p. 142), the decisive argument in favor of generic phases consists in the fact that only they lead to formulas,
\[ \text{*) Gibbs called } D \text{ the (phase) density, and } \rho \text{ the probability coefficient.} \]
which are limiting cases of the quantum-mechanical formulas. Below we shall deal only with generic phases and, moreover, shall assume that integrations such as those in (III.1,2), (III.1,4), or (III.1,6) extend only over all distinct generic phases, i.e. of all \(\nu!\) distinct species phases only one will be taken into account.
Let us now consider the properties of the quantity \(\sigma\), defined by the relation
\[ \sigma=\sum_{\nu}\int \rho \ln \rho \, d\Omega_{\nu}. \tag{III.1,7} \]
First of all it is clear that \(\sigma\) is the mean value of the logarithm of the probability \(\eta\). We shall show below that \(\sigma\) has the following properties:
a) If our ensemble is such that all systems consist of the same number of particles \(N\), and the energies \(\varepsilon\) of all systems lie in the interval \(E \leq \varepsilon \leq E+\delta E\), then \(\sigma\) will be minimal when \(\rho\) is given by the equalities:
\[ \begin{aligned} \rho &= \mathrm{const.}\,\delta(\nu-N) \quad &&\text{for } E \leq \varepsilon \leq E+\delta E;\\ \rho &= 0 &&\text{for } E>\varepsilon \text{ or } E+\delta E<\varepsilon, \end{aligned} \tag{III.1,8} \]
where \(\delta(x)\) is the Dirac \(\delta\)-function (\(^{35}\), Sections 20, 21).
b) If our ensemble is such that all systems consist of the same number of particles, but only the mean energy is specified, i.e. \(\rho\) is subject to the condition\(^*\)
\[ \sum_{\nu}\int \varepsilon \rho \, d\Omega_{\nu}=E, \tag{III.1,9} \]
then \(\sigma\) will be minimal when \(\rho\) satisfies the relation
\[ \rho=\delta(\nu-N)e^{\beta(\psi-\varepsilon)}, \tag{III.1,10} \]
where \(\beta\) and \(\psi\) are constants that can be found from (III.1,4) and (III.1,9).
c) If only the mean number of particles and the mean energy of the systems of the ensemble are specified, i.e. \(\rho\) is subject to condition (III.1,9) and the condition
\[ \sum_{\nu}\int \rho \, d\Omega_{\nu}=N, \tag{III.1,11} \]
then \(\sigma\) will be minimal when \(\rho\) satisfies the relation
\[ \rho=\exp[-q+\nu\mu-\beta\varepsilon]. \tag{III.1,12} \]
Again \(\beta\), \(\mu\), and \(q\) are constants, i.e. they do not depend on \(p_k\), \(q_k\), or \(\nu\), and they can be found from (III.1,4), (III.1,9), and (III.1,11).
\[ \text{*} \]
\(^*\) The summation over \(\nu\) is carried out trivially in this case, since only one term of the sum contributes.
Before proving these properties of \(\sigma\), let us recall once again that the values of the density given in these three cases refer, respectively, to the ensemble with an energy shell\(^*\), the macrocanonical ensemble, and the grand canonical ensemble. It follows from the usual consideration (see, for example, ESM, Secs. 5.3 and 6.1) that the quantity \(1/\beta\), which coincides with the modulus of the ensemble, is equal to \(kT\) (\(k\) is Boltzmann’s constant, \(T\) is the absolute temperature), that \(\psi\) is the free energy of the system represented by the macrocanonical ensemble, that \(\mu\) is equal to \(\beta\) times the partial free energy or the partial thermodynamic potential, that \(q\) coincides with Kramers’ \(q\)-potential, which for a homogeneous system is equal to \(\beta pV\) (\(p\) is the pressure, \(V\) the volume), and that in both cases (a) and (c) \(\langle \eta \rangle\) is equal to \(-S/k\) (\(S\) is the entropy)\(^**\).
The proof of the minimality properties of \(\sigma\) is based mainly on the fact that the function \(y\), given by the expression
\[ y = xe^x - e^x + 1, \tag{III.1,13} \]
is positive for \(x \ne 0\) and equal to zero for \(x = 0\). This property of \(y\) follows most simply from the facts that 1) \(y(0)=0\) and 2) \(\dfrac{dy}{dx}(=xe^x)\) has the same sign as \(x\).
Let us now compare, for the three cases (a), (b), and (c), two densities \(\rho_1\) and \(\rho_2\), where the density \(\rho_1\) in each of the cases corresponds to the minimal value of \(\sigma\), i.e. \(\rho_1\) is given respectively by relations (III.1,8), (III.1,10), and (III.1,12), while \(\rho_2\) is given by the expression
\[ \rho_2 = \rho_1 e^{\Delta \eta}, \tag{III.1,14} \]
where \(\Delta \eta\) may be an arbitrary function of \(p_k\) and \(q_k\), and in case (c) also of \(\nu\). Both quantities \(\rho_1\) and \(\rho_2\) satisfy relation (III.1,4)
\[ \sum \int \rho_1 d\Omega = \sum \int \rho_2 d\Omega = 1, \tag{III.1,15} \]
and in cases (b) and (c) they both satisfy (III.1,9):
\[ \sum \int \varepsilon \rho_1 d\Omega = \sum \int \varepsilon \rho_2 d\Omega = E. \tag{III.1,16} \]
In case (c) they satisfy (III.1,11)
\[ \sum_{\nu} \int \rho_1 d\Omega = \sum_{\nu} \int \rho_2 d\Omega = N. \tag{III.1,17} \]
\(^*\) Such an ensemble is sometimes called microcanonical, although the microcanonical ensemble is in fact only a particular case of an ensemble of energy-shell type.
\(^**\) This is another example of the advantage of generic densities over specific densities, since this relation between \(S\) and \(\rho\) does not hold for a specific density.
Finally, we have
\[ \sigma_2-\sigma_1=\sum \int (\rho_2\ln\rho_2-\rho_1\ln\rho_1)\,d\Omega . \tag{III.1,18} \]
From (III.1,18), for all three cases one can obtain the relation*)
\[ \sigma_2-\sigma_1=\sum \int \rho_1\left[\Delta\eta e^{\Delta\eta}-e^{\Delta\eta}+1\right]d\Omega\geqslant 0, \tag{III.1,19} \]
the last inequality being a consequence of the function defined in (III.1,13).
Thus, we have proved that \(\sigma\) is minimal in the case of an energy layer, a microcanonical ensemble, or a large canonical ensemble, under the fulfillment of certain conditions (the physical meaning of these conditions will be discussed in the next section). Comparing expression (III.1,7) for \(\sigma\) with expression (I.2,3) for \(H\), one might think that we are in a position to prove that \(\dfrac{d\sigma}{dt}\) is negative and thereby to establish a tendency of evolution of canonical ensembles. However, it is easy to show that \(\dfrac{d\sigma}{dt}=0\). This was clarified by Weber soon after the publication of Gibbs’s monograph. To verify this, let us write
\[ \sigma(t'')=\sum \int \rho''\ln\rho''\,d\Omega'' =\sum \int \rho'\ln\rho'\,J\,d\Omega' = \]
\[ =\sum \int \rho'\ln\rho'\,d\Omega'=\sigma(t'); \tag{III.1,20} \]
here \(\rho''\), \(\rho'\) are abbreviated notations for the quantities \(\rho(p'',q'';t'')\), \(\rho(p',q';t')\); the relation between \(p'',q''\) and \(p',q'\) is that the representative point \(p',q'\) at time \(t'\), with the passage of time (at the instant \(t''\))**) passes into the point \(p'',q''\); finally, \(J\) is the Jacobian of the transformation from \(p'',q''\) to \(p',q'\), which is equal to unity by Liouville’s theorem (see, for example, ESM, p. 102).
*) To obtain relation (III.1,19) in case (a), add to the right-hand side of (III.1,18) the expression
\(\sum \int (1-\ln\rho_1)(\rho_1-\rho_2)\,d\Omega\), which is equal to zero by virtue of (III.1,15) and the fact that \(\rho_1=\mathrm{const}\). In case (b), add to the right-hand side of (III.1,18) the expression
\(\sum \int [\beta(\psi-\varepsilon)+1](\rho_1-\rho_2)\,d\Omega\), which is equal to zero by virtue of (III.1,15) and (III.1,16). Finally, in case (c), add the expression
\(\sum \int (iq-\psi+\beta\varepsilon-1)(\rho_2-\rho_1)\,d\Omega\), equal to zero by virtue of (III.1,15)—(III.1,17).
**) Since we are dealing with systems consisting of a single component, the representative point of the system will remain in the same \(\Gamma\)-space and will not leave it. This circumstance will not occur if we consider systems in which chemical reactions may take place. In that case, the situation is greatly complicated; in their consideration one may use the method of second quantization (Schönberg 52S1, 53S1, 53S2).
Although we have convinced ourselves that \(\sigma\) is constant, nevertheless in a certain sense we have an approximation to the stationary case. Let us explain this by means of the example indicated by Gibbs \(^{02G}\) (see also \(^{03B,\,04B1,\,04B2,\,06E2,\,11E2}\)). Consider some vessel with a liquid, for example water, to which some insoluble dye has been added, consisting, for instance, of colloidal particles. It is well known from experience that if in the initial state the coloring substance was distributed nonuniformly throughout the volume, then practically any stirring will lead to a state in which the dye is distributed uniformly, insofar as this can be determined visually. This means that the stirring will lead to an “equilibrium” state. However, if one looks more carefully, we shall still find that in microscopic volumes part of the space is occupied by water and part by colloidal particles. Thus, although the coarse distribution is uniform, the fine-structure distribution still proves to be nonuniform.
From this example it follows that in a number of cases it may be useful, alongside the fine-grained density \(\rho\) (fine-grained density), to introduce the coarse-grained density \(P\) (coarse-grained density), defined as follows. For each \(\nu\) we divide the corresponding \(\Gamma\)-space into finite, but small cells*) \(\Omega_i^{(\nu)}\) with volume \(W(\Omega_i^{(\nu)})\), and let \(P_i^{(\nu)}\) be the mean value of \(\rho\) in \(\Omega_i^{(\nu)}\),
\[ P_i^{(\nu)}=\frac{\int \rho\,d\Omega_\nu}{W(\Omega_i^{(\nu)})}, \tag{III.1,21} \]
where the integration is extended over the cell \(\Omega_i^{(\nu)}\)**).
We now introduce the coarse-grained density \(P(\nu; p,q)\), taking it to be constant in each cell \(\Omega_i^{(\nu)}\) and equal to \(P_i^{(\nu)}\). From (III.1,4) we obtain, as is easy to verify,
\[ \sum_\nu \sum_i P_i^{(\nu)} W(\Omega_i^{(\nu)})=1, \tag{III.1,22} \]
or
\[ \sum \int P\,d\Omega=1; \tag{III.1,23} \]
we see that \(P\) is normalized.
*) Cf. the introduction of cells in Section I.3.
**) We can somewhat generalize the definition of the coarse-grained density by also subdividing the possible values of \(\nu\) into intervals containing some number of integers each. However, in order not to complicate the formulas, we shall not do this, since in the case of systems containing only one component this generalization is unnecessary.
We next introduce, instead of $\sigma$, the function $\Sigma$,
$$ \Sigma=\sum_{\nu}\sum_i W\left(\Omega_i^{(\nu)}\right)P_i^{(\nu)}\ln P_i^{(\nu)}, \tag{III.1,24} $$
or, expressing it in terms of the coarse-structure density $P$,
$$ \Sigma=\sum\int P\ln P\,d\Omega. \tag{III.1,25} $$
If we take into account that $\ln P$ is constant within any cell $\Omega_i^{(\nu)}$ and that the integration of $P$ over $\Omega_i^{(\nu)}$ coincides with the integration of $\rho$ over $\Omega_i^{(\nu)}$, we obtain
$$ \Sigma=\langle\ln P\rangle . \tag{III.1,26} $$
The properties of $\sigma$ indicated in cases (a), (b), and (c) are also valid for $\Sigma$, if $\rho$ is replaced everywhere by $P$. However, now $\Sigma$ is no longer a constant; indeed, it can be shown that $\Sigma$ decreases. Thus one can find a tendency toward the establishment of canonical ensembles—in any case insofar as this concerns the coarse-structure density. Here it may be noted that, although in Gibbs’ monograph$^{02G}$ a distinction is introduced between fine-structure and coarse-structure densities, in the exposition itself these two densities are not distinguished quite clearly. The same largely applies to the exposition of Berberian$^{03B,04B2}$ and Bumstead$^{04B1}$; we also note the works of Ehrenfest$^{06E2,11E2}$, Poincaré$^{06P}$, Lorentz$^{07L}$, and Kroo$^{11K2}$.
Let us now consider the change of $\Sigma$ with time. For this it is necessary to use some of the considerations presented in the next section, where representing ensembles are discussed. Suppose that we have made certain observations on a physical system at the time $t'$. Since such observations never give the maximum possible information, we shall construct an ensemble whose average properties at the time $t'$ correspond to the observed properties of the system under consideration at that same instant. Owing to experimental limitations it is possible, at best, to determine $\rho(\nu;p,q)$, which changes from one cell to another, if only the sizes of the cells are chosen in accordance with the experimental limitations (this will be assumed). Thus suppose that the fine-structure density in each cell is constant, and at $t'$ we have:
$$ P'=\rho' \tag{III.1,27} $$
and for $\Sigma$
$$ \Sigma'=\sum\int P'\ln P'\,d\Omega=\sum\int \rho'\ln\rho'\,d\Omega. \tag{III.1,28} $$
If the state at the time $t'$ already corresponds to equilibrium, then $\Sigma$ will take its minimum value and any further changes in the
cannot be expected, since all three ensembles corresponding to the minimum of \(\Sigma\) in cases (a), (b), and (c) will be stationary (see, for example, ESM, pp. 105 and 137). Let us therefore assume that \(\rho'\) does not correspond to a stationary ensemble. After some time \(*\), at the moment \(t''\), equality (III.1,27) will no longer hold, and we obtain
\[ \rho'' \ne P'' . \tag{III.1,29} \]
The latter follows from the fact that, although \(\rho\) remains constant under deformations in the phase space of a cell that preserves the unchanged volume \(W(\Omega_i^{(\nu)})\), the character of these deformations will change so that, at subsequent instants of time, each of the cells will be filled with points which at the instant \(t'\) belonged to a set of different cells.
Write relation (III.1,29) in the form
\[ \rho'' = P'' e^\Delta , \tag{III.1,30} \]
where \(\Delta\) is a function of \(\nu, p\), and \(q\). For \(\Sigma\) we shall have
\[ \Sigma'' = \sum \int P'' \ln P'' d\Omega ; \tag{III.1,31} \]
for the change of \(\Sigma\) we obtain \(**\)
\[ \Sigma' - \Sigma'' = \sum \int P'' \left[\Delta e^\Delta - e^\Delta + 1\right] d\Omega > 0 . \tag{III.1,32} \]
Hence it is clear that \(\Sigma''\) is less than \(\Sigma'\), owing to the fact that \(\rho''\) and \(P''\) are now not equal everywhere. Comparing the result obtained with the case of a dye in a liquid, one may expect that with the passage of time \(\rho\) and \(P\) will differ more and more, and that \(\Sigma\) will decrease until the equilibrium distribution \(P\) is established, i.e.
case (a)
\[ \left. \begin{aligned} P &= \delta(\nu-N)\cdot \mathrm{const}, && E < \varepsilon < E+\delta E;\\ P &= 0, && \varepsilon < E \ \text{or}\ E+\delta E < \varepsilon; \end{aligned} \right\} \tag{III.1,33} \]
\(*\) Strictly speaking, it is necessary to say “at another time,” without prejudging whether \(t''\) is later or earlier than \(t'\). However, since we have constructed our ensemble to represent the system observed at \(t'\), we are interested only in predictions that can be checked experimentally, i.e. at later times \(t''\). In this connection it is interesting to read the remarks of Berberian and Ehrenfest on this question \({}^{03\mathrm{B},\,04\mathrm{B}2,\,06\mathrm{E}2}\) (see also \({}^{02\mathrm{G},\,04\mathrm{B}1}\)).
\(**\) Relation (III.1,32) is obtained by subtracting expression (III.1,31) from (III.1,28), replacing in the integrand of (III.1,28) \(\rho'\) by \(\rho''\), which can be done since \(\dfrac{d\rho}{dt}=0\) [cf. (III.1,20)], and adding the expression \(\sum \int (P''-\rho'')\,d\Omega\), equal to zero by virtue of (III.1,4) and (III.1,23).
case (b)
\[ P=\partial(\nu-N)e^{\beta(\psi-\varepsilon)}, \tag{III.1,34} \]
case (c)
\[ P=\exp[-q+\psi-\beta\varepsilon]. \tag{III.1,35} \]
As soon as the distributions (III.1,33)—(III.1,35) are established, \(\Sigma\) will attain its minimal value and thereafter, consequently, will no longer decrease. In this case, having made the observation, we shall come to the conclusion that the state of the system under consideration could best be described with the aid of microcanonical or canonical grand ensembles—which are stationary. From this one can derive an additional justification for using these ensembles to describe systems in thermodynamic equilibrium*).
The time required for the establishment of an equilibrium distribution is, generally speaking, of the order of the relaxation time, and not of the order of the recurrence times of the Poincaré cycle (Section I.2), as P. and T. Ehrenfest proposed to assume \(^{11E2}\). Apparently, they forgot to take into account the indistinguishability of the particles making up the system. This means that the Poincaré period should be divided by \(N!\), after which intervals of the order of the relaxation time are obtained.
In concluding this section it is necessary to note two further circumstances. First, we have proved that if at the moment \(t'\) we constructed the representative ensemble, then the quantity \(\Sigma\) will decrease from its value at the moment \(t'\) to a smaller value at the moment \(t''\); however, we have not proved, but only illustrated with the example of a dye in water, that \(\Sigma\) will continue to decrease. The situation here is very reminiscent of that with which we dealt in Section I.3, where the statistical aspects of the \(H\)-theorem were discussed. A more detailed investigation of the behavior of \(\Sigma\) as a function of time, analogous to what was done in Section I.3, would deserve attention. It is interesting to note that Gibbs \(^{02G}\) himself did not assert that “\(\Sigma\)” would decrease monotonically, but asserted only that
\[ \lim_{t\to\infty}\Sigma(t)\leq \Sigma(t'); \tag{III.1,36} \]
he himself did not prove the latter relation, but for a special case it was proved by Poincaré \(^{05P}\) and Kroon \(^{11K2}\). In this connection one may point to Tolman (\(^{38T}\), especially Sections 49 and 51), and also to the material contained in the next section.
The second circumstance to which we would like to draw attention is the following: the decrease of \(\Sigma\) follows from the fact that we pass from a state in which \(\rho\) and \(P\) are equal to a state in which \(\rho\) is no longer everywhere equal to \(P\). Since \(P\) is obtained by experimental observations, one may say that
*) The justifications that are usually given may be found, for example, in Gibbs’s monograph \(^{02G}\) (especially Chapters IV, X, and XV) or ESM, Sections 5.3, 5.7, and 6.1 (cf. also ESM, Sections 5.6 and 6.3).
at the moment \(t'\) we “know” \(\rho\). However, at the moment \(t''\) we no longer possess such a volume of concrete information. Since \(\Sigma\) may be regarded as “coarse-structural” \(\sigma\), and since \(\sigma\) is related to entropy by the relation
\[ \sigma=\langle\eta\rangle=-\frac{S}{k}, \tag{III.1,37} \]
it is once again clear from this how an increase in the degree of absence of information corresponds to an increase of entropy*).
III.2. Representing ensembles
In the preceding section the behavior of ensembles has already been considered and, thus, we have begun to discuss the question posed in the Introduction (B), i.e. the question: why is it possible to describe the behavior of a system by considering the average behavior in an ensemble? This, apparently, is the most important question in statistical mechanics. However, for rather unconvincing reasons it has not been considered as thoroughly as it deserves. This happened in part because Gibbs himself introduced the notion of ensembles rather for using them in statistical considerations than for illustrating the behavior of physical systems, although he himself did not strictly adhere to his original intentions as they are set forth in the introduction to his monograph. In fact, there are certain indications that Gibbs had in mind the possibility that the average behavior of a system in a micro- or macrocanonical ensemble can represent the actual behavior of a physical system in equilibrium (see, for example, Bumstead’s remark \(^{04B1}\)). In the same year as Gibbs, and independently of him, Einstein \(^{02E, 03E}\) proposed the theory of ensembles; he definitely had in mind the possibility of equivalence between the average behavior in an ensemble and the actual behavior of a physical system. However, P. and T. Ehrenfest \(^{11E2}\), Ornstein \(^{08O}\), and Uhlenbeck \(^{27U}\) regarded the theory of ensembles chiefly as a mathematical device for facilitating the calculation of mean values of phase functions. Uhlenbeck, for example, definitely placed ensemble theory in the same category as the Darwin–Fowler method (\(^{22D1, 22D2, 22D3, 22F2, 23F3, 23D1, 23D2, 25F1, 26F2, 26F3}\); a brief review is given in ESM, Appendix IV, and in \(^{48S}\), Chapter VI). Only after Tolman introduced the concept of representing ensembles did ensemble theory acquire a firm physical foundation; however
*) We shall not trace in greater detail the correspondence between entropy and the absence of detailed information, or the volume of ignorance, but refer to the literature: see, for example, ESM, p. 160, and the papers by Szilard \(^{29S}\) and Brillouin \(^{51B4}\) (see also \(^{48W, 49S1, 54M1, 54G1, 55G1}\)).
his point of view was not always taken into account to the proper extent. In the present section Tolman’s method will be set forth in fairly full detail as applied to classical statistical mechanics; the corresponding material relating to the quantum-mechanical case is given in Section IV.3.
It has already been mentioned more than once above that, in view of the severe limitation of our capabilities, we cannot obtain sufficiently adequate information, for example, about the initial state of the physical systems we study. If one recalls that one mole of gas contains \(6\cdot 10^{23}\) atoms, this means that, instead of knowing \(36\cdot 10^{23}\) values of coordinates and momenta, necessary for a complete description of the state (assuming each atom to be a point particle with three degrees of freedom), we can obtain only a few relations between them by measuring, for example, the pressure, momentum, and angular momentum of the system. Moreover, we are quite unable to know the full number of particles in the system; let us only recall that Avogadro’s number is known with an accuracy of only a few ten-thousandths\(^{51\mathrm{B1},51\mathrm{B2}}\). Having obtained such meager information, we want to make predictions about the behavior of the system in the future on the basis of this information. It is clear that the only possible way to do this is to use probability theory and statistical methods. If we could obtain the values of all coordinates and momenta at one instant of time—and if we could solve the equations of motion—the behavior of the system would be completely determined, since then we would be dealing with a problem of classical mechanics and with a sufficient number of boundary conditions to determine the orbit of the representative point of the system in phase space. However, in view of the impossibility of obtaining these values, we confine ourselves to assertions concerning the most probable behavior of the system. To determine this most probable behavior, we compare different systems possessing identical values of those quantities that are measured, but otherwise capable of differing within wide limits. This means that we construct an ensemble representing our system. The question remains as to how such a representative ensemble can be constructed; this question is connected with the problem of a priori probabilities. In constructing the classical ensemble, equality of a priori probabilities was assumed for equal volumes in \(\Gamma\)-space). In considering the so-called small ensembles (petit ensemble) this leads to quite definite prescriptions. However, in the case of constructing representative large ensembles, it is necessary to remember that we are dealing with \(\Gamma\)-spaces of different numbers of dimensions, and that spaces with a larger number of dimensions have larger a priori*
*) This assumption was tacitly adopted when we wrote down expression (I.3,1) for the probability of the state \(Z\).
weight. Now we shall try to justify, or at least to show the plausibility of, our choice of a priori probabilities.
In order to reveal the type of ensemble that best represents the system under consideration, we must, first of all, bear in mind that in practice, under almost any circumstances, we measure only the properties of small subsystems. When measuring the pressure of a gas, only that part of the gas which is near the wall takes part in the formation of the force acting on the wall and measured by our instruments. When determining the density of a system by an optical method, the beam of light passes only through part of the system (for example, the case of critical opalescence). When measuring the electrical conductivity of a semiconductor, the electrode is applied to a part of the sample, which, in fact, is often used to detect differences between states at different places in the sample, and so on. It follows from this that we are actually interested only in predictions concerning the behavior of subsystems. For brevity we shall assume that the complete system (which may be either the entire physical system enclosed in our apparatus, or even this system plus a part of the external surroundings) is isolated, so that it cannot exchange either energy or particles with the external surroundings. Let \(\rho\) be the density of the ensemble describing some subsystem, and \(\rho'\) the density of the ensemble describing the “rest of the system,” i.e. the complete system with the subsystem removed. Let, further, \(\varepsilon\) and \(\nu\) denote the energy and the number of particles in the subsystem, and \(\varepsilon'\) and \(\nu'\) the corresponding quantities in the rest of the system. The densities \(\rho\) and \(\rho'\) will then satisfy the relations:
\[ \sum \int \rho\, d\Omega = 1, \tag{III.2,1} \]
\[ \sum \int \rho'\, d\Omega = 1, \tag{III.2,2} \]
\[ \sum \int (\varepsilon+\varepsilon')\rho\rho'\, d\Omega\, d\Omega' = \text{const}, \tag{III.2,3} \]
\[ \sum (\nu+\nu') \int \rho\rho'\, d\Omega\, d\Omega' = \text{const}, \tag{III.2,4} \]
where relations (III.2,3) and (III.2,4) express the fact that in an isolated system the total energy and the total number of particles are constant. Suppose, further, that a state of equilibrium has been established, i.e. \(\Sigma\) is minimal for the complete system, or
\[ \sum \int (\rho\rho')\ln(\rho\rho')\, d\Omega\, d\Omega' = \text{minimum}. \tag{III.2,5} \]
In writing (III.2,3) and (III.2,5) use was made of the fact that the ensemble of the complete system is a product-
subsystem and the rest of the system. Using (III.2,1)—(III.2,4), one can rewrite (III.2,3) and (III.2,4) in the form
\[ \sum \int \varepsilon \rho\, d\Omega+\sum \int \varepsilon' \rho'\, d\Omega'=\text{const}, \tag{III.2,6} \]
\[ \sum^\nu \int \rho\, d\Omega+\sum^{\nu'} \int \rho'\, d\Omega'=\text{const}. \tag{III.2,7} \]
The equilibrium density now follows from the condition of minimality (III.2,5), which must be satisfied simultaneously with (III.2,1), (III.2,2), (III.2,6), and (III.2,7). This leads, by the usual method (the method of Lagrange multipliers; see, for example, ESM, p. 20 or 444), to the following expression for the equilibrium value of \(\rho\)
\[ \rho=\exp[-q+\nu\mu-\beta\varepsilon]. \tag{III.2,8} \]
It may be noted here that, since (III.2,5) is in fact valid only for the coarse-grained density, the expression (III.2,8) actually refers to \(P\).
Further, one may point out that our assumption concerning a priori probabilities was used in writing (III.2,1)—(III.2,4) without introducing, alongside \(\rho\), a density function that would give the a priori weights of different volumes in phase space.
Above we have shown, in a way somewhat different from that used in the preceding section, that the grand canonical ensemble will represent any subsystem in equilibrium. We must, however, devote some attention to the problem of nonequilibrium systems; in this matter we shall follow Tolman’s exposition (\(^{38T}\), especially Sections 79b, 51c, and 112). At the end of this section the choice of equal a priori probabilities for equal volumes in phase space will be discussed.
Suppose that our system is in a nonequilibrium state, as we have ascertained by making the corresponding observations of the system at the moment \(t'\). To predict its further behavior, let us construct an ensemble representing the system, using, on the one hand, information obtained from experiment and, on the other hand, the assumption of equality of a priori probabilities (or equal weights) for equal volumes in phase space. As was indicated in the preceding section, we can at most determine the coarse-grained density \(P'\) and, using the assumption of equal weights for equal volumes, choose the fine-grained density \(\rho'\) in accordance with (III.2,27). At a subsequent time \(t''\), the fine-grained density \(\rho''\), following from \(\rho'\), will no longer be constant in each of the cells into which we have divided phase space. As was indicated earlier, this follows from the fact that \(P''\) and \(\rho''\) are now no longer equal. From (III.1,32) we see that
\(\Sigma' - \Sigma''\) is the greater, the more \(P''\) and \(p''\) differ from each other and, consequently, we have the right to expect that \(\Sigma\) will decrease continuously until the distributions (III.1,33)—(III.1,35) are established.
In order to verify the approach to equilibrium, we can carry out the corresponding observations on the system. From the preceding discussion one might suppose that these observations will lead to coarse-grained densities such that \(\Sigma\) will be a monotonically decreasing function of time. However, in making such an assumption, we first of all tacitly suppose that the mixing of the various cells of phase space occurs as completely as the mixing of a dye in water. Further, in doing so we forget to take into account the fact that each time we make an observation the amount of information available to us increases, and, consequently, our choice of the representative ensemble at the moment \(t'''\), for example, will no longer be as simple as at the moment \(t'\), since it is necessary to take into account the information obtained both at the moment \(t'\) and at the moment \(t''\). With regard to mixing, it is clear that we return to the question touched upon in Hopf’s ergodic theorem. Indeed, starting from identical basic assumptions, we find that the treatment by means of the ergodic theorem and the treatment by means of the theory of ensembles do not differ from one another as greatly as might have seemed at first glance.
Let us now consider the problem of a priori probabilities. For simplicity, let us restrict ourselves to small ensembles. Consider two regions \(\mathfrak{A}_0\) and \(\mathfrak{B}_0\) in \(\Gamma\)-space. Let these two regions at the moment \(t_0\) be filled with representative points of the ensemble, and let \(p(\mathfrak{A}_0)\) and \(p(\mathfrak{B}_0)\) be the a priori probabilities for the two regions. At the moments \(t_1, t_2, \ldots\), the representative points that at \(t_0\) filled the regions \(\mathfrak{A}_0\) and \(\mathfrak{B}_0\) will fill the regions \(\mathfrak{A}_1\) and \(\mathfrak{B}_1\), \(\mathfrak{A}_2\) and \(\mathfrak{B}_2\), … with a priori probabilities \(p(\mathfrak{A}_1)\) and \(p(\mathfrak{B}_1)\), \(p(\mathfrak{A}_2)\) and \(p(\mathfrak{B}_2)\), … . It is obvious that only such a choice of a priori probabilities is admissible for which
\[ \frac{p(\mathfrak{A}_0)}{p(\mathfrak{B}_0)} = \frac{p(\mathfrak{A}_1)}{p(\mathfrak{B}_1)} = \frac{p(\mathfrak{A}_2)}{p(\mathfrak{B}_2)} = \cdots, \tag{III.2,9} \]
and these relations are satisfied if we take
\[ p(\mathfrak{A}) = W(\mathfrak{A}), \qquad p(\mathfrak{B}) = W(\mathfrak{B}), \tag{III.2,10} \]
where \(W(\mathfrak{A})\) or \(W(\mathfrak{B})\) are the volumes of the regions \(\mathfrak{A}\) or \(\mathfrak{B}\) in \(\Gamma\)-space. Let us note here that the relations (III.2,10) will satisfy (III.2,9) only in the case where we are dealing with a phase space consisting of canonically conjugate \(p\) and \(q\); this is one of the reasons for choosing canonical coordinates to describe the phase of a system.
Instead of (III.2,10), one may use the more general form
\[ p(\mathfrak{A}) = \int F(\varepsilon,\varphi_i)\,d\Omega, \tag{III.2,11} \]
where the integration is extended over the region \(\mathfrak{A}\) and where \(F(\varepsilon,\varphi_i)\) is a once-and-for-all chosen function of the energy \(\varepsilon\) and of the other integrals of motion \(\varphi_i\).
The reason why \(F(\varepsilon,\varphi_i)\) is taken to be equal to unity (with (III.2,10) being chosen instead of the more general relation (III.2,11)) was originally historical and was based on the assumption of the existence of ergodic systems. Since we are now convinced that mechanical systems are quasi-ergodic, it follows that all the different possible values of \(\varphi_i\) occur practically equally often in any part of phase space. Consequently, we may at once restrict ourselves, in Ehrenfest’s expression, to the ergodic choice
\[ p(\mathfrak{A})=\int F(\varepsilon)\,d\Omega . \tag{III.2,12} \]
However, since the densities \(\rho\) representing ensembles of equilibrium systems are functions only of the energy, the choice (III.2,12) instead of (III.2,10) would correspond to the density
\[ \rho'(\varepsilon)=\frac{\rho(\varepsilon)}{F(\varepsilon)} \tag{III.2,13} \]
instead of \(\rho(\varepsilon)\), given by the expressions (III.1,8) or (III.1,10). The quantity \(F(\varepsilon)\) would appear in all integrals, but the final results would be the very same.
It may be mentioned here that Ehrenfest’s remark that the ergodic choice is often inadvisable no longer concerns the essence of the matter, for it was based on the difficulties encountered by thermostatics before the appearance of quantum mechanics *).
SUMMARY OF THE SITUATION IN CLASSICAL THERMOSTATICS
Before beginning the discussion of the role and status of the \(H\)-theorem and of the ergodic theorem in quantum statistics, we wish briefly to summarize the content of Parts I–III. There are, in essence, three different approaches to the problem of justifying the method of statistical mechanics: 1) the utilitarian (applied) approach, 2) the formalistic approach, and 3) the physical approach **), although the last two may in some respects overlap.
The utilitarian approach justifies the statistical method by the results obtained and does not concern itself with any further foundation. In this approach, calculations carried out by any method are welcomed, whether it be the Darwin–Fowler method, the use of the microcano-
*) One may also refer to Ehrenfest’s paper \(^{14\mathrm{E}}\) in the discussion of the possible choice of the function \(F(\varepsilon,\varphi_i)\).
**) The reader, I hope, will forgive me for such a classification, which clearly testifies to my own prejudices and tastes.
ical ensembles, the use of canonical ensembles or of large ensembles—provided only that they lead to useful results. The fact that for physical systems with many degrees of freedom all these calculations lead to the same results is used in the sense that the choice of the final method is merely a matter of convenience*).
With the formalistic approach, arguments are usually adduced to the effect that a mechanical system must be isolated and, consequently, must possess a definite value of the energy, as well as a number of particles in it, since only such systems are considered in classical mechanics. The only stationary ensemble admitted for consideration is thus the microcanonical small ensemble, although, as a mathematical device for simplifying calculations, one may use the macrocanonical, i.e., large, ensemble. Moreover, the microcanonical ensemble was admitted only after it had been established that the time average of any phase function, calculated for a single system, apparently coincides with the average taken over the microcanonical ensemble. To prove this it is necessary to introduce the ergodic, or, better said, quasi-ergodic theorem. In this connection adherents of the formalistic school of thought place the ergodic theorem at the center of their consideration.
With the physical approach, the fact is emphasized that we are able to obtain only a limited amount of information about the system of interest to us, and, consequently, we must work with representative ensembles. In this case it is necessary to prove that the development of the nonequilibrium state will proceed in such a way that the ensemble representing the observed system will approach ever more closely the canonical ensemble.
Attempts to justify the assumptions underlying the last two methods of approach fall into three groups. One may, for example, try to use the \(H\)-theorem directly, but, as was seen in Part I, difficulties arise in doing so. The direct \(H\)-theorem would provide another alternative to the ergodic theorem; however, the \(H\)-theorem in its statistical form does not lead to this, unless we wish to introduce various kinds of assumptions concerning transition probabilities which we are not in a position to confirm more rigorously. Nevertheless, the statistical form of the \(H\)-theorem is useful in that it indicates the relative frequency with which fluctuations occur, and the rate at which a nonequilibrium state can return to equilibrium.
*) It is interesting to note that even Tolman used this utilitarian point of view when he justified the assumption of equality of volumes in \(\Gamma\)-space.
Thus, we arrive at the ergodic theorem, and, as was seen in Part II, the problem at present reduces to proving the metric indecomposability of the corresponding phase spaces. If it could be shown that physical systems are, generally speaking, quasi-ergodic, then the ergodic theorem would be proved; this, however, has not been done. It is interesting to note that both Einstein^[02E] and Poincaré^[06P] openly stated that, in their opinion, energy is the only single-valued analytic integral of a physical system that does not depend on time. If one introduces a somewhat greater degree of plausibility by allowing the energy to vary in a small but finite region, then one may resort to Hopf’s ergodic theorem, in which a mixing of different energy surfaces takes place. Let us note once again that the metric indecomposability of the phase spaces under consideration has still not been proved.
If one prefers the approach by means of representative ensembles, one can use the \(H\)-theorem in the form that corresponds to ensemble theory, i.e., identifying \(H\) with the mean value of the natural logarithm of the coarse-grained density. The approach to equilibrium then follows quite simply; however, the argument is based on a special choice of a priori probabilities. As we saw in Section III.2, this choice can be justified only if the systems under consideration are quasi-ergodic.
Thus, we see that in all methods there is no proof of the quasi-ergodicity of physical systems. If such a proof were given, one could use either the formalistic approach, since in that case the (quasi-)ergodic theorem is proved, or the physical approach, since the choice of a priori probabilities would be justified. In conclusion, one may say that although at first glance there seem to be major differences between the formalistic and physical approaches, these differences in fact turn out not to be so significant, and both approaches prove vulnerable at one and the same point.
(To be continued in the next issue)