Abstract
The first of three lectures delivered at the Franklin Institute in Philadelphia.
Full Text
METHODS OF STATISTICAL MECHANICS1
Richard Tolman, Pasadena.
1. Introduction
In the three lectures which I propose to deliver before you, I shall describe certain methods and results obtained with the aid of the comparatively young science—so-called statistical mechanics—a difficult science and, if you will, a tangled one, but nevertheless one that at the present time constitutes an essential branch of mathematical physics.
If we are dealing with some complex system, as is for the most part the case in the study of chemical or biological processes, then it is very difficult, if not impossible, to observe experimentally or to predict theoretically exactly, in all details, the behavior of all the elements of any given system. Even in the case in which a chemist could specify, with complete exactness, the position, orientation, and velocity of all the molecules and atoms of which a reacting mixture consisted at some definite instant, and could give a completely precise account of the laws governing their individual behavior, the complexity of the calculations would prevent him from predicting, even for a fraction of the following second, the positions and chemical combinations formed from these elements. If, however, we speak of a biologist dealing with incomparably more complex systems, consisting not of
of an isolated reacting mixture, but from a combination of reacting zones that complement one another and constitute a single whole in the biological process under study, with the interplay of the physicochemical processes underlying it, then, of course, the problem becomes wholly insoluble.
Nevertheless, despite the impossibility of exact observation and prediction, both the chemist and the biologist have not been frightened off by the difficulty of the problem. Even if they are unable to predict exactly the behavior of each individual element1, they are able to observe and predict, with remarkable confidence, the behavior of their systems taken as a whole (macrobehavior). Every time a chemist touches with a lit match a soap bubble filled with a mixture of hydrogen and oxygen, he obtains an explosion. Every time a bacteriologist heats his glass plates in an autoclave, he knows that the bacteria in it will be killed; and I think that whenever any war occurs, any sociologist will be able to predict without error that rich people will profit greatly from it, while the working people will become impoverished. All the aforementioned tangled phenomena do not lie beyond the domain of scientific law and prediction.
To investigate the laws by means of which one can describe the macrobehavior of systems consisting of an extraordinarily large number of molecules is the chief aim of statistical mechanics. The essence of the methods applied by this science is that we deliberately renounce any attempt at an exact prediction of the “life history” of each individual molecule of the given system, in order thereby to determine all the more accurately, by statistical methods, what the average behavior of a very large number of systems of similar structure will be. The introduction of the methods of mathematical statistics into the field of physical chemistry is therefore explained by the same reasons that led statistics to find application in the biological sciences. More and more often the chemist is compelled to make use of
methods of statistical mechanics, he must henceforth develop this science theoretically, since he always has to deal with so complex a conglomerate of atoms that other methods may prove too difficult to apply.
2. Comparison of the methods of ordinary and statistical mechanics.
We shall obtain the clearest idea of the methods of statistical mechanics if we compare its methods with those of ordinary mechanics, and then with the methods of thermodynamics. For a concrete illustration let us consider a gas consisting of a large number of similar molecules, moving in all directions between the walls of the vessel containing them. For simplicity let us suppose that the molecules are very small elastic spheres of invariable volume.
Applying the methods of ordinary mechanics to a gas of this kind, we would have to assume as known, for some given instant, the positions (coordinates) and velocities of all the molecules. We could then predict those collisions which would take place both between the molecules themselves and between the molecules and the walls of the vessel in which they are contained; and, applying the laws of conservation of energy and of quantity of motion in calculating these impacts, we could predict the results caused by them. Thus, theoretically, we would be able with the mind’s eye to follow the subsequent behavior of the gas and to predict its state for every instant of future time. In reality, however, we would become lost in our calculations even for the nearest interval of one thousandth of a second.
The difficulties arising in such an application of the methods of ordinary mechanics compelled the introduction into the old kinetic theory of gases of various simplifying assumptions of an extremely artificial character, for which we now feel no necessity. Thus, for example, it was assumed that all molecules possess equal velocities, while their directions are in no way ...
are not connected with one another, and that thereby they neglected the influence of collisions between the molecules themselves on the change in velocity. With the aid of simplifications of this kind, the methods of ordinary mechanics could lead to a large number of results which, though only approximate, were extremely valuable for science.
In applying the methods of statistical mechanics to the investigation of the behavior of a gas in the sense in which we postulated it above, we do not attempt to study the behavior of a single specimen of gas. Instead, we reason about an infinitely large quantity, or a multitude (ensemble), of specimens of gas, the molecules of which occupy all possible positions and possess all possible velocities in all conceivable directions. Even if we are not able to predict the exact behavior of any one specimen of gas, nevertheless—surprising as this may at first appear—we can, by applying the methods of mechanics, obtain information about the statistical behavior of the entire multitude (ensemble), and thus draw conclusions about the possible behavior of a single specimen. Thus, for example, among all the specimens of the ensemble only an infinitely small fraction will have molecular velocities whose distribution differs appreciably from the given Maxwell distribution law, and from this we conclude that each given specimen of gas will probably obey the Maxwellian distribution of velocities; and if, owing to some conditions, the distribution of velocities departs from the Maxwellian one, then the gas, left to itself, will tend toward it, i.e. the Maxwellian distribution corresponds to a stable, stationary state.
At first glance it may seem surprising that the indicated method is so powerful. For a single specimen of gas, with its very large number of molecules, proved to be so complicated a system that we were quite unable to predict its future behavior; yet, having passed to the consideration of an incomparably more complex system, consisting of an infinite multitude of such
of specimens, we seem to obtain such great possibilities for predicting its behavior. However, this apparent paradox arises from the fact that in the first of the considered cases of a single specimen we tried to predict the actual behavior of the given individual specimen, whereas with the aid of our new method we content ourselves with determining the statistical behavior of an ensemble, and consequently with predicting the probable behavior of an individual specimen.
Statistical methods occupy the same place in solving problems confronting the biological sciences. Thus, for example, a given person between the ages of 34 and 35 is so complex a system of atoms and molecules that it is quite impossible to predict whether the given subject will be alive or will die in the following year. Nevertheless, in a given country we have perfectly correct statistics for the number of people per thousand dying during one year between the ages of 34 and 35; and, on the basis of these data, we can predict with sufficient satisfaction the probability that this person will live during the year. In this biological case the prediction was based on statistical laws that are the result of observation, whereas in the physico-chemical problems of interest to our audience the statistical laws are obtained as a consequence of deduction from the laws of mechanics. Nevertheless, the reason that impels us to appeal to statistical methods is identical in both cases: namely, it is the complexity of the systems, which makes it difficult to apply more direct methods and procedures.
Of course, we must not think that the methods of statistical mechanics can be applied usefully only to systems that are too complex for the application of the methods of ordinary mechanics. Even in the case of a system with a small number of degrees of freedom, obeying simple equations of motion, we can mentally imagine a multitude of such systems and discuss the statistical behavior of such an ensemble, as was in particular indicated
Prof. E. Wilson.^1 The application of this method may be useful in the case where the initial values of the coordinates and velocities of the system at some given instant are unknown to us or are of no interest. But for the most part the chief driving force in the present development of statistical mechanics is the possibility of applying it to physico-chemical systems with a large number of degrees of freedom, corresponding to a large number of molecules; and the guiding practical aim must be the attainment of the possibility of elucidating the laws governing the behavior of such complex systems by a method which in many respects is better than any other.
3. Comparison of the methods of statistical mechanics and thermodynamics.
In the theoretical treatment of processes occurring in complex physico-chemical systems, in the preceding period of the development of science the methods of thermodynamics were applied more often than statistical ones. Therefore a comparison of these two methods and some assessment of their relative advantages are not without interest.
If, first of all, we compare these two sciences from the point of view of the factual indubitability of their foundations and the rigor of their development, then we must, of course, say that all the advantages are on the side of thermodynamics.
The whole sum of information provided by classical thermodynamics is based entirely on two postulates, called the first and second principles; in other words, we must regard thermodynamics as a logical system of theorems derived in one way or another from these postulates, in the process of applying them to various types of physico-chemical systems. Moreover, the certainty of these two fundamental laws, limited, to be sure, to the macroscopic behavior of physico-chemical systems, has the right to be considered entirely unquestionable, since never before
^1 E. B. Wilson. Annals of Math. 10, 129, 149, 1909.
...so far no deviations from these laws have been observed in macroscopic systems, and the entire past history of the successive successes of thermodynamics supplies us with a sufficient quantity of evidence for the correctness of its starting points. We must regard the theorems of thermodynamics as belonging to the most reliable possessions of science. On the other hand, if we turn to statistical mechanics, we shall find a much less definite foundation and, in part, a less rigorous development of the methods provided by this science.
The first assumption underlying this science is the hypothesis of the atomistic and molecular structure of matter. This general hypothesis at the present time need no longer withstand the attacks of the old school of physical chemists, guided by Ostwald’s views, which consisted in the belief that the scientist should deal exclusively with phenomena accessible to observation. However, although we feel the justice of accepting, as one of our points of departure, the general theory of atomistic structure, in the development of the ideas of statistical mechanics we are compelled to resort to special hypotheses that refine our conception of the nature of atoms and molecules, which at best may be regarded as necessary simplifying assumptions.
Proceeding from the hypothesis of the atomistic structure of matter, the method of statistical mechanics consists in applying the laws of dynamics to determine the behavior of a multitude of systems, where each system, in turn, is composed of atoms and molecules. In the time of Gibbs and Boltzmann, there was little doubt as to the correctness of such a mode of action. The laws of dynamics, expressed, for example, in the form of Hamilton’s principle, are, of course, a sound and well-founded generalization of observations and of experiments actually performed on moving bodies of macroscopic dimensions, and the danger of applying these laws to the study of molecules was not yet clearly recognized at that time. Now, however, the development of quantum theory indicates that the most satisfactory theory of atoms...
and molecules requires a complete revision of the laws of dynamics, as well as, perhaps, of our notions concerning the continuity of space and time. At present we shall try to avoid this difficulty and shall first give an elementary exposition of so-called classical statistical mechanics, based exclusively on the laws of dynamics. We may then regard the results obtained in this way as a limiting case that is valid when the high frequencies of periodic motion are not included in the given problem, and afterward we shall introduce, sometimes perhaps even arbitrarily, those changes that are most suitable for the theory of a large group of quantum phenomena.
Next we must mention one more shortcoming of the method of statistical mechanics from the standpoint of the requirements of theoretical rigor. At the very outset we are compelled to introduce the famous ergodic hypothesis, or the principle of the continuity of path. The introduction of this hypothesis is, of course, justified by the correctness of the results obtained with its aid, and the arguments demonstrating the usefulness of introducing it in a suitably modified form will, I hope, appear convincing. Nevertheless, the above-mentioned hypothesis still presents a tangled problem for further study.
Let us summarize our enumeration of the deficiencies in the rigor of the method of statistical mechanics. These are: first, the premise of the atomistic hypothesis itself—not so much the assumption of the atomistic structure of matter in general as the specific assumptions about atomic properties that are made from time to time in considering particular problems; second, the application of the laws of classical dynamics to describe the behavior of atomistic systems, these laws being modified in a special way so as to obtain the necessary agreement with quantum phenomena; and, third, the introduction of the ergodic hypothesis, since with the aid of this hypothesis a whole series of questions of statistical mechanics, as applied to physicochemical systems, is solved very simply.
It must be noted that this lack of rigor is not inherent in the science of statistical mechanics per se, but appears only when we come to actual applications in solving problems of physics and chemistry. Nevertheless, it is interesting to compare the methods of thermodynamics and statistical mechanics as applied to real physico-chemical systems, and, of course, since our aim is rigor, there can be no shadow of doubt that thermodynamics at present stands higher. However, if we turn our attention to the results that can be obtained by the one method and by the other, then the scales tip entirely in favor of the younger science.
First of all, as was shown by the works of Boltzmann and Gibbs, both laws of thermodynamics may themselves be regarded as consequences of the principles of statistical mechanics and interpreted as an inevitable result of atomistics. Therefore all possible conclusions obtained with the aid of thermodynamics can, at least indirectly, also be obtained with the aid of statistical mechanics, although their direct derivation often has advantages, since the theorems receive a more intimate atomistic interpretation.
Secondly, the properties of matter expressed in the equations of thermodynamics are simply empirical parameters, whose values are determined by actual measurements carried out on macroscopic systems. Meanwhile, in the treatment of statistical mechanics the physical meaning of these quantities is often revealed by means of their atomistic interpretation. Thus, for example, in the thermodynamic derivation, the specific heats of gases and solids, heats of reaction, and equilibrium constants enter into the equations and are brought into relation with one another. But their absolute values remain unknown. In the treatment of these problems by statistical mechanics, the numerical values of specific heats are obtained from consideration of the number of degrees of freedom of the system under consideration and the degree of their excitation; the values of heats of reaction have gradually been
are brought into connection with the energy levels of atoms and molecules, as determined by spectroscopy; finally, the absolute values of the equilibrium constants are obtained from consideration of the relative probability of the various arrangements of the atoms that form the molecules1.
Third, even in the domain of equilibria, whose theory is the special province of thermodynamics, thermodynamics must accept the laws appropriate to the given case, but has no possibility of justifying them. Thus, in particular, the gas laws and equations of state for matter in any form are for thermodynamics only empirical facts, whereas for statistical mechanics they have either already been derived or, under known circumstances, are derivable by means of methods whose further development we can now foresee.
Fourth, in the case of a system not in equilibrium, thermodynamics can only tell us that changes are by nature permitted which entail an increase of entropy. Thermodynamics is not in a position to inform us about the rate at which these permitted changes of state will proceed; hence it is clear that in reality thermodynamics gives us no possibility either to foresee or, consequently, to predict what will actually occur in a system not in equilibrium. Such problems as the rates of diffusion or evaporation, the coefficients of thermal and electrical conductivity, the coefficients of transfer of momenta (forces), and the coefficients of thermal and photochemical reactions are wholly inaccessible to solution by means of thermodynamic reasoning. They are, however, amenable to treatment by the methods of statistical mechanics. Finally, when objects of an exclusively atomistic character arise before us for investigation, it is clear that thermodynamics is utterly powerless to help us in such cases.
METHODS OF STATISTICAL MECHANICS
Brownian motion of particles, density fluctuations in a liquid, differences in the velocities of electrons flying out of an incandescent filament, and the distribution of the energy of thermal radiation as a function of various frequencies—all these problems posed to us by nature can be investigated by means of the methods of statistical mechanics; but what can non-atomistic thermodynamics undertake for the cognition of the laws underlying these most urgent problems of modern natural science?
This somewhat protracted exposition by no means exhausts the extraordinarily varied content of statistical mechanics. Nevertheless, I shall attempt to give some idea of the power of this method and of the sphere of its application. The future of theoretical chemistry depends on the further expansion of the domain of its application, and there is no doubt that the interaction of the two sciences will have the most beneficial influence on their further development.
We are now prepared to follow the development of statistical mechanics. In the short time at my disposal, I hope that I shall be able to sketch only the outline of the basic methods of statistical mechanics, then present a number of arbitrarily selected results and consider some of the most important applications.
4. Ensemble (set) and phase.
As has already been clarified, we shall reason not about the behavior of an isolated system, but about the behavior of a set or ensemble of systems containing an enormous number of identical systems. Knowing the behavior of this set of samples, we can then draw conclusions concerning the probable behavior of each system.
The various systems making up the ensemble have the same structure, i.e. they consist of the same number of atoms of one or of different kinds, enclosed in identical vessels, but they differ in the positions and velocities of their constituent elements. According to the terminology introduced by
Gibbs, we must say that they differ from one another in phase.
If each system in the ensemble possesses \(m\) generalized coordinates \(Q_1, Q_2 \ldots Q_m\) and \(m\) corresponding momenta \(P_1, P_2 \ldots P_m\), then the instantaneous phase of the system may be determined by means of the instantaneous values of these \(m\) coordinates and momenta. Thus, for example, if we have a gas consisting of \(N\) point particles flying inside a vessel, we may take as the \(m\) coordinates the \(3N\) Cartesian coordinates \(x, y, z\) of each particle, and as the \(m\) momenta the quantities corresponding to the \(3N\) components of the linear momenta \(m\dot{x}, m\dot{y}\), and \(m\dot{z}\) for each particle. The grounds for choosing generalized coordinates and momenta, rather than ordinary coordinates and velocities, will to some extent be clarified in the subsequent exposition.
5. Phase spaces.
It is extremely convenient to follow the behavior of such a system, i.e. the change of its phase with time, if one represents the phase as the position of a point (a phase point) in a \(2m\)-dimensional space (phase space), corresponding to the \(2m\) coordinates and momenta whose values are given to us. Over the course of some time the phase point will describe a certain trajectory in this \(2m\)-dimensional space, and one may regard the phase points of all the systems of the ensemble as describing continuous lines in this space. A supposition of this kind has the advantage that at our disposal we obtain a ready-made geometrical language for discussing the question of the behavior of the given ensemble.
If our system consists of molecules, then it often proves convenient to represent the values of the coordinates and momenta of an individual molecule in a phase space with a smaller number of dimensions than is necessary for the whole system. Thus the phase of an individual molecule, having \(2n\) coordinates and momenta \(q_1, q_2 \ldots q_n, p_1 \ldots p_n\), may be repres—
represented by the position of a point in a space of \(2n\) dimensions.
In order to avoid the possibility of confusing these two different “phase spaces,” it is sometimes convenient to use Ehrenfest’s nomenclature and to call the phase space employed for representing the entire system or “gas” \(\gamma\)-space, and the phase space for a single molecule \(\mu\)-space. For example, if our system is composed of \(N\) monatomic molecules, then the phase space, or \(\mu\)-space, for each individual molecule will have six dimensions, corresponding to the three coordinates and the three momenta that determine the state of each molecule, while the \(\gamma\)-space will have \(2n = 6N\) dimensions. Thus, if I have a system of axes at my disposal, I may think of the \(x\)-coordinate of a given molecule as the abscissa, and the corresponding momentum \(m\dot{x}\) as the ordinate. I may then draw an axis perpendicular to the plane of the drawing for plotting the values of the \(y\)-coordinate, and, in the fourth dimension, an axis perpendicular to the first three for the mental plotting of the values of \(m\dot{y}\), and in that case two more axes in the fifth and sixth dimensions for \(z\) and \(m\dot{z}\). The fact that I cannot construct a genuine six-dimensional space in order to form the totality of the values of my point should not hinder the use of the convenient geometrical language represented by this mental method.
Having finished with the first molecule, I can now proceed further and construct, on the next six axes, the state of each of the remaining \(N - 1\) molecules, and thus obtain a \(6N\)-dimensional phase space, or \(\gamma\)-space, for the whole gas. The position of a point in this latter \(\gamma\)-space will therefore depict the instantaneous positions and momenta for all the molecules of the gas; and since the molecules fly in all directions and collide both with one another and with the walls, the point representing the state of the gas will move around in this \(\gamma\)-phase space, describing a trajectory determined by the laws of mechanics.
Richard Tolman
6. Distribution of the Ensemble by Phases.
If we now return to the consideration of our ensemble of systems, taken as a whole, it is evident that each system of the ensemble gives us a point in \(\gamma\)-space. We may at the outset distribute these points in any desired manner and then follow their motion, since they describe continuous lines in this space. We shall define the density of distribution \(\rho\) for some position in phase space as the number of points in a unit \(2m\)-dimensional volume representing the positions of our systems.
We make use of various methods of initially distributing the points in space, depending on different purposes. The three kinds of distribution most often used are: the canonical ensemble, introduced by Gibbs and also used by other investigators,
\[ \rho = N_i e^{\frac{\psi - E}{\theta}}, \tag{1} \]
where \(N\) is the total number of systems entering into the ensemble, \(E\) is the energy of one system, and \(\psi\) and \(\theta\) are constants having the dimension of energy1; the microcanonical ensemble, used by Gibbs, Boltzmann, and other scholars:
\[ \begin{aligned} \rho &= \text{const} \quad &&\text{(when the values of the energy lie between } E \text{ and } E+dE),\\ \rho &= 0 &&\text{(for other values of the energy)} \end{aligned} \tag{2} \]
and the surface ensemble, used by Ehrenfest, where the phase points are distributed on the surface in the space of constant energy with surface density:
\[ \sigma = \frac{\text{const}} {\sqrt{\left(\frac{\partial E}{\partial p_1}\right)^2 + \cdots + \left(\frac{\partial E}{\partial p_m}\right)^2}} . \tag{3} \]
All these three kinds of distribution may be used for obtaining conclusions and inferences concerning the probable behavior of a single system.
At first glance it may seem that the surface ensemble is the most suitable method for our work, since in this case we assume that all systems of the ensemble possess the same energy as the system that interests us.
However, from the standpoint of mathematical simplicity, distribution over a surface is not so easily considered as other kinds of distributions, and we shall not use it here. From various points of view, the canonical distribution of phases \(\Gamma\) and \(b\) is mathematically the simplest and theoretically the clearest. In turning to this distribution in order to predict the behavior of some single system, we must proceed from the proposition that the majority of systems of the ensemble will evidently be in such a state that their store of energy, without appreciable error, is equal to the mean value of the energy for all systems entering the ensemble under consideration, since the fluctuations of energy will be of the same order as the possible oscillations that we might expect for a system placed in a thermostat.
It follows from this that any system from the given ensemble may be regarded as a specimen of a single system left to its own fate. The canonical distribution of phases therefore has very great significance, for by means of it one can clarify the relation between statistical mechanics and thermodynamics—a topic which, unfortunately, lies outside the scope of our conversations.
The microcanonical distribution of phases is most often adhered to in investigations, and it satisfies our purposes quite tolerably. Taking as the limit of the fluctuations of energy \(dE\), and assuming that only the phase points satisfying this series are distributed in \(\gamma\)-space, and decreasing \(dE\) at will, we can, in the limiting case, endow all systems with the same amount of energy without appreciable error, and thus we again meet the preceding conclusion, which stated that in such a case we have the right to regard the system,
snatched at random from this ensemble as a sample of an individual system, supplied on its own with a store of energy whose numerical value satisfies the above-mentioned small sphere of oscillations.
7. Change of the density of distribution as a function of time. Liouville’s theorem.
Up to now we have spoken only about how theory permits us, at the initial instant, to place the points characterizing the systems that form an ensemble, in order to obtain one of the three distributions known by the names canonical, microcanonical, and surface; but we have still said nothing about the behavior of these points subsequently. It is clear that we must at least know what the distribution density \(\rho\) will be, in order to make use of the concept of an ensemble in drawing conclusions. The investigation of this important problem is possible because the motion of points in phase space is, of course, directly connected with the motion of the individual system of the ensemble, and the sequence of this motion is governed by the laws of mechanics. This problem of the change in the density of distribution may be considered for any arbitrary initial distribution of points, and its solution is known under the name of Liouville’s theorem. For our purpose we shall need to consider only one special simple case of Liouville’s theorem, namely, we shall reason about an ensemble if the initial distribution of points is uniformly dense throughout the whole phase space.
In order to clarify the content of this theorem, let us mentally imagine a certain cubic element in the phase space under consideration,
\(dQ_1 dQ_2 \ldots dQ_m dP_1 \ldots dP_m\),
and let us turn our attention to a pair of opposite surfaces perpendicular to the axis \(Q_1\). The area of each of these surfaces, situated at a distance \(dQ_1\) from one another, will obviously be equal to the expression \(dQ_2 \ldots dP_m\). Consequently, the number of points passing each second through the first of
of the above-mentioned surfaces, the one situated at \(Q_1\), obviously, can be expressed in the following way:
\[ \rho \frac{dQ_1}{dt} dQ_2 \ldots dP_m . \]
And for the number of points leaving the opposite surface at \(Q_1+dQ_1\), we shall have:
\[ \rho \left(\frac{dQ_1}{dt}+\frac{\partial}{\partial Q_1}\cdot \frac{dQ_1}{dt}\, dQ_1\right)dQ_2\ldots dP_m . \]
Subtracting the second expression from the first, we obtain, as the increase in the number of points in one second within the cubic element,
\[ -\rho \left(\frac{\partial}{\partial Q_1}\cdot \frac{dQ_1}{dt}\right)(dQ_1 dQ_2\ldots dP_m). \]
We now have the right, proceeding from similar considerations, to obtain an analogous result for another pair of parallel surfaces, and, adding the separate results, we obtain for the complete change in the number of points in one second
\[ \frac{dN}{dt} = -\rho \left( \frac{\partial}{\partial Q_1}\cdot \frac{dQ_1}{dt} +\ldots+ \frac{\partial}{\partial Q_m}\cdot \frac{dQ_m}{dt} + \frac{\partial}{\partial P_1}\cdot \frac{dP_1}{dt} +\ldots+ \frac{\partial}{\partial P_m}\cdot \frac{dP_m}{dt} \right) (dQ_1\ldots dP_m) \]
or, dividing by the volume of the element \(dQ_1\ldots dP_m\), we may write for the magnitude of the change of density within the element under consideration
\[ \frac{d\rho}{dt} = -\rho \left( \frac{\partial}{\partial Q_1}\cdot \frac{dQ_1}{dt} +\ldots+ \frac{\partial}{\partial Q_m}\cdot \frac{dQ_m}{dt} + \frac{\partial}{\partial P_1}\cdot \frac{dP_1}{dt} +\ldots+ \frac{\partial}{\partial P_m}\cdot \frac{dP_m}{dt} \right) \tag{4} \]
But the values of \(\frac{dQ_1}{dt}\), \(\frac{dP_1}{dt}\), etc., are known to us from the equations of motion, which for any dynamical system may be written in Hamiltonian form
\[ \frac{dQ_i}{dt}=\frac{\partial H}{\partial P_i}; \qquad \frac{dP_i}{dt}=-\frac{\partial H}{\partial Q_i}; \tag{5} \]
where \(H\) is the energy of the system, expressed as a function of the coordinates and momenta. Hence we have:
\[ \frac{\partial}{\partial Q_1}\cdot \frac{dQ_1}{dt} = \frac{\partial}{\partial Q_1}\cdot \frac{\partial H}{\partial P_1} \tag{6} \]
and
\[ \frac{\partial}{\partial P_1}\cdot \frac{dP_1}{dt} = -\frac{\partial}{\partial P_1}\cdot \frac{\partial H}{\partial Q_1}. \]
And since the result does not depend on the order of differentiation, we see that in equation (4) the terms in brackets cancel pairwise and we obtain:
\[ \frac{d\rho}{dt}=0 \tag{7} \]
This result is sometimes also called Liouville’s theorem or the principle of conservation of density in phase.
In other words, if we begin with our symbolic points uniformly distributed in phase space, they will permanently remain in this uniform distribution. Moreover, since the given system can change its stock of energy only as a result of its own motion, it is evident that the density of the distribution remains unchanged with the passage of time also in the case when the phase points are distributed in some functional dependence on the energy, but uniformly in each region between the limits \(E,\ E+dE\). Thus we may write
\[ \frac{d\rho}{dt}=0 \]
whenever the distribution is only a function of the energy of the system.
Such ensembles, in which the density of the distribution does not change with time, are designated as being in statistical equilibrium. It is necessary to draw attention to the circumstance that the canonical and micro-
canonical distributions mentioned above evidently give ensembles that are in statistical equilibrium, and it can also be shown that the surface distribution will likewise be in statistical equilibrium.
The possibility of obtaining the result indicated above depends on our choice of generalized coordinates and moments as the \(2m\) axes for our phase space. If we had thought to place our points in a phase space whose axes represented coordinates and velocities, then, generally speaking, we would not have been able to complete our investigation with such a simple result, since we would not have had the right to write the simplifications that follow from the introduction of equations (5), which depend on the use of the equations of motion in Hamiltonian form, including generalized moments.
This circumstance is the reason that determines the importance of the equations of motion in canonical Hamiltonian form for statistical mechanics, and explains why, in works on statistical mechanics, moments are used in practice in preference to velocities. We owe to Boltzmann the elucidation of the full importance of being able to get by with expressions for moments when discussing questions of statistical mechanics.
8. Applications to molecular systems.
The simple result mentioned above gives us a supply of almost all the prerequisites necessary for the use of the methods of statistical mechanics, and we are now prepared to discuss the behavior of actually existing systems of molecules.
Suppose that we consider a system of molecules enclosed in an appropriate vessel; we shall denote the store of energy of this system by the letter \(E\). In general, this energy and the components of the linear and angular moments of the system taken as a whole, which we shall regard as equal to zero, will be the only dynamical quantities about which we can speak definitely, since the motions
individual particles are too complex for us to be able to follow them.^1 Consequently, since we do not know precisely the molecular configuration and velocities of our system, we must consider all possible phases in which the given system, possessing a known store of energy \(E\), may reside; and in order to do this, we are entitled to appeal for help to a microcanonical ensemble of systems similar to the system that interests us, with their representative points uniformly distributed everywhere within the shell bounded by surfaces determined, in the \(2m\)-dimensional space, by the values of the energy \(E\) and \(E+dE\).
9. Introduction of the ergodic hypothesis (Ergodenhypothese).
Up to this point we have introduced no hypothesis whatever, except that the system moves according to the laws of dynamics in canonical form. In order to proceed further, we must introduce an assumption consisting in the assertion that such a microcanonical ensemble of systems, possessing energies whose magnitudes lie between \(E\) and \(E+dE\), gives a true representation of the various phases through which, in the course of time, an individual system with a store of energy lying between \(E\) and \(E+dE\) may pass. In order to formulate this hypothesis in a more special form, we must assume that results obtained from a multitude of systems, chosen at random from the ensemble, will be practically the same as in the case when we consider the given system at moments chosen in a completely arbitrary manner.
We obtain a certain empirical justification for accepting this kind of hypothesis from the circumstance,
^1 The fact that these motions lie outside the domain of our methods of observation is of secondary importance. The statistical method would likewise be necessary in the investigation of systems of the same complexity even if the particles composing these systems could be readily observed.
that many conclusions developed on this basis have indeed found their confirmation in experience.
The theoretical justification of this hypothesis is a more difficult and less definite matter, despite its plausibility. First of all it should be noted that the state of uniform density under the microcanonical distribution of phases, ensured by equation (7), means that the points representing the states of the systems will not have any tendency to cluster, to any degree, in special regions of phase space. If such a clustering were to occur, then the point representing one of the given systems could, in any case, at some moment find itself with a greater degree of probability in such a region of concentration than somewhere else in phase space. In the absence of this kind of concentration we may state definitely that a given system, taken at a randomly chosen moment, has equal chances of being in any one of the various microscopic states through which it actually passes in the course of its existence.
This entirely legitimate conclusion is, of course, not identical with the assertion that the system has equal chances of being in every microscopic state lying in the phase space determined by the limits \(E\) and \(E + dE\), since the given system, even when its complete history is considered, may be incapable of extending its trajectory through the entire region comprised between the energies \(E\) and \(E + dE\). To accept this hypothesis of “ergodic systems,” as it was named by Boltzmann, or the “principle of continuity of the trajectory,” as Maxwell designated it, means explicitly to accept that a system consisting of molecules must in fact pass through all microscopic states compatible with the energy contained in it before it fully completes its cycle of motions. If this hypothesis is correct, then the results of a random sampling of specimens from the microcanonical ensemble must be absolutely identical with the results obtained by considering the given system at moments
time, taken quite arbitrarily, and from this follows the justification of our further arguments, to which we now turn.
Be that as it may, it must be admitted that the “ergodic” hypothesis, in its least complicated form, is scarcely plausible. True, closed orbits in astronomy have accustomed us to the possibility of an infinitely large number of closed orbits for a given dynamical system, all of them possessing the same store of energy. However, we cannot point to a single example of a dynamical system with more than one degree of freedom in which, in its motion, the system would pass through all possible phases corresponding to the energy contained in it. It seems evident that we must reject the ergodic hypothesis in its elementary form as insufficiently plausible, and investigate other similar assumptions by which we might be guided in order to obtain equivalent results. We may consider three additional, comparatively simple possibilities:
a) One may suppose that, for a molecular system containing a given quantity of energy, there is a known number of possible closed paths, but that one of them is infinitely longer than all the others taken together, so that we may neglect the possibility of the system’s remaining on any trajectory except the longest one, and the microscopic states corresponding to this trajectory will include practically all microscopic states in the region between \(E\) and \(E + dE\).
A possibility of this kind, it seems, was first advanced by Jeans, but we must abandon this assumption, since in essence it is no more plausible than assigning to the system only a single closed trajectory.
b) Another possibility consists in the fact that, for a given store of energy contained in a system of molecules, the motion may take place along one of numerous closed trajectories, but the system cannot by itself
to pass from one of these closed trajectories to another. Nevertheless, it may be transferred at random from one trajectory to another under the influence of external causes, in such a way that the same probability is obtained for the existence of any microscopic state belonging to any system of the corresponding microcanonical ensemble. This point of view was also expressed by Jeans, who, having developed it in greater detail, showed the probability that one of the molecules may be the external perturbing cause acting on the system, which in that case must be regarded as consisting of all the other existing molecules. Indeed, an assumption of this kind leads to the desired results.
c) The third possibility, which seems to me the most satisfactory and the most conducive to achieving the aim we have set ourselves, consists in accepting the fact that in reality there exists more than one closed trajectory for the motion of a system of molecules containing a given amount of energy—there may even be a large number of such trajectories. Along each of the closed trajectories we have a sequence of microscopic states through which the system must pass before it returns to the initial point; this process is completely inevitable, since we must take regions \(dQ_1 \ldots dP_m\), which determine distinct states, though of very small size, but nevertheless finite, which excludes the possibility of an infinitely large number of possible states.
Let us now turn to the discussion of the motion of this system, left to its own fate. The point symbolizing it will move along one of the possible trajectories corresponding to the energy contained in the given system. We cannot indicate exactly along which trajectory; but, as we pointed out above, we have the right to say that at any moment, taken quite arbitrarily, it has equal chances of being in any
of microscopic states located along this trajectory.
Reasoning in this way, we have not yet adopted any hypothesis. We shall now take into account the possible character of the various closed trajectories. It is obvious that, for a real existing system consisting of a large number of molecules, the greater part of these closed trajectories will be extremely long. Suppose, for example, that we are dealing with a gas consisting of a large number of molecules flying back and forth between the walls of the vessel enclosing them. Even if the molecules collided exclusively with these walls and did not collide with one another at all, it is quite obvious that in this case an exceedingly long time would elapse before we could again find each individual molecule in its initial position; and if we also take into account the collisions that take place among the molecules themselves, this period will be lengthened to an enormous degree. It follows from this that the vast majority of trajectories will consist of an extremely large number of microscopic states, of those situated in the whole layer between the energies \(E\) and \(E + dE\).
Thus, in the end, we find a set of states and, consequently, have the right to suppose that the microscopic states located along the special closed trajectory along which the system of interest to us moves constitute a complete collection of all the various kinds of states present in the whole region of phase space between the limits \(E\) and \(E + dE\). If we agree with this premise, then any significant collection of samples chosen at random from the microcanonical ensemble between \(E\) and \(E + dE\) will have practically the same properties as the collection obtained by us by observing the given system at arbitrarily chosen instants; the use of the microcanonical ensemble as giving a perfectly accurate representation of the sequence of states passed through in time by a single system is therefore justified.
Thus, for example, in the case of a gas it proves possible to show that the vast majority of all states of the corresponding microcanonical ensemble will possess a distribution of molecular velocities not noticeably differing from Maxwell’s law. Hence we are confirmed in the conviction that this will probably be valid for an extremely large number of states lying along the individual trajectory on which the given sample of gas is found.
We may therefore accept that the microcanonical ensemble makes it possible to obtain a correct representation of the successive states of an actually existing molecular system.
I hope that the description given here of the point of departure in investigations of statistical mechanics does not appear to you too confused or to contain excessive technical details. The use of a space of more than three dimensions—indeed, a space of \(6N\) dimensions, where \(N\) is a number of the order of trillions—for constructing the position of a point seems, at first glance, exceedingly alarming. Nevertheless this concept is introduced solely in order to make it possible to use the language of geometry, familiar to everyone, which makes it possible to set out analytical conclusions in shorthand form; and, as Gibbs indicated from the very beginning, these results may be used even though they are expressed in a language foreign to analysis in the strict sense of the word. I am confident that, in any case, you have grasped the main essence of this method. A single system, consisting of numerous atoms and molecules, is so complex that we cannot hope to be able to understand or describe its behavior in all details.
It follows from this that we must change the method of our reasoning and consider a multitude of systems constructed in complete similarity with the system that interests us, but whose molecules are arranged in all possible ways, of course in agreement with the energy of the system. In studying
of the statistical behavior of such an ensemble, we can make predictions concerning the probable behavior of an individual system. You can see directly that such a device is a powerful method for investigating this very difficult area of science. The future development, above all, of theoretical chemistry, and also of a large part of theoretical physics, will, I am convinced, depend most closely on the successful use of the methods of statistical mechanics.