Full Text
PROBABILISTIC JUSTIFICATION OF THE ERGODIC HYPOTHESIS
B. M. Gessen, Moscow.
§ 1. Statement of the Problem. The Ergodic Hypothesis in Statistical Mechanics.
In studying an ensemble we may pose two kinds of problems: one may study the structure of the ensemble, the distribution of its elements at a given moment of time. If not one ensemble is given, but a finite or infinite number of ensembles, then one may set oneself the aim of establishing regularities in the ensemble of the given ensembles. The statistical character of such an investigation will consist in the fact that we shall determine how often a given element, or a given set of elements, occurs within the limits of the given ensemble, or we shall determine how often an ensemble of a definite structure occurs within the limits of the given collective of ensembles.
The characteristic feature of this kind of problem is that we investigate the ensemble not in time, but in space.
The second kind of problem consists in the fact that we investigate the given ensemble in time and set ourselves the task of establishing how the states of the ensemble are distributed in time.
Does the structure of the ensemble (or of the ensembles) at a given moment of time reflect the structure of the time series for the given ensemble? What relation, and under what conditions,
conditions exist between the spatial and temporal structure of the ensemble?
We determine, for example, the frequency with which 5 Brownian particles are encountered in the field of view of a microscope at a given instant of time.
But we may ask how often 5 particles will appear at a given point, if the ensemble of Brownian particles is left to itself over the course of a definite interval of time.
The essential difference between the first problem and the second is that, in order to determine the frequency of a spatial distribution, it is sufficient to observe the Brownian particles at a single instant, each time mixing the emulsion; whereas to determine the frequency of distribution in time, one must leave the emulsion to itself for a definite interval of time and observe how the initial configuration of the particles changes during that interval.
One of the fundamental techniques of statistical mechanics consists in our striving to replace the study of the temporal behavior of an ensemble by the study of its spatial structure. In order for such a method of study to be admissible, we must show that such a replacement is indeed possible and lawful. If this is proved, then we can reduce the study of the regularity of the behavior of an ensemble in time to finding the probability of a definite structure, i.e. to a combinatorial problem.
In the classical kinetic theory of gases of Boltzmann–Maxwell, we had two kinds of hypotheses: 1) the hypothesis of the atomistic structure of the gas and 2) the theoretical-probabilistic hypothesis, consisting in the fact that a certain regularity was ascribed to the complex motions of molecules in the form of a statement about the relative frequency of various configurations and motions of molecules.
The criticism of the theorem \(H\) was directed chiefly against that assumption of a theoretical-probabilistic character which is usually designated as the assumption concerning the number of collisions—“Stosszahlansatz.”
At its basis lies the following proposition on equal probability of distribution: in considering the collisions of two groups of molecules, it is assumed that, for any unit volume of space swept out in their motion by the molecules of the second group, there corresponds the same number of molecules as for any unit volume of the remaining space.
Thus, by means of the supposition of the equipossibility of elementary processes, an assertion is derived concerning the relative frequency of non-equipossible states (events). The attempt to do without these basic suppositions about the probability of elementary events, in particular to construct kinetic theory without the Stosszahlansatz, led Boltzmann to the creation of statistical mechanics. “The first investigations in the field of statistical mechanics were devoted to a somewhat restricted field of inquiry, since they concerned particles of one and the same system. Later the investigation was extended to the phases (or states with respect to configurations and velocities) through which the system passes in the course of time” (Gibbs).
The principal task in statistical mechanics is to find the time average (Zeitmittel) of a certain phase function. To find this average, the spatial average (Scharmittel, Zahlmittel) is computed, and then equality between these two averages is established.
Instead of considering a successive series of states of a system in time, in which each subsequent state is a consequence of the preceding one, we consider the ensemble of systems that is obtained from the given system if all possible values of the parameters, differing from one another by arbitrarily small amounts, are considered. We thus obtain an ensemble of systems not as a temporal succession of states of the given system, but as a certain ensemble of systems in space.
In contrast to an ensemble of states in time, connected with one another, we obtain a series of independent systems.
...from one another. Suppose that our model of a gas consists of \(N\) identical polyatomic molecules, each of which has \(r\) degrees of freedom. Then the configuration of the system—a phase in the Gibbsian sense—at a given instant of time will be determined by specifying \(2rN\) parameters: coordinates and momenta. Suppose that the given system possesses energy \(E\). If we imagine a series of systems possessing one and the same energy, but independent of one another and differing at the given instant of time only in the values of the parameters \((p,q)\), then we obtain the so-called virtual ensemble. Physically the problem is usually posed in such a way that we seek not the relative frequency within the virtual ensemble, but the average duration of the system’s stay in a definite state.
The essence of the ergodic hypothesis in the general sense consists in the fact that the average frequency within the virtual ensemble is set equal to the average duration of the system’s stay in the given state (for example, in a stationary state).
What conditions must a system satisfy in order that a numerical average (spatial average) may be equated to a time average?
Let us first reveal what assumptions must be made concerning the system in order to be able to assert the equality of these averages.
Let us denote the time average of some phase function \(u(p\cdot q)\) by \(\overline{u(p\cdot q)}^{\,t}\), and the numerical average by \(\overline{u(p\cdot q)}\). It is necessary to show under what conditions
\[ \overline{u(p\cdot q)}^{\,t}=\overline{u(p\cdot q)}. \tag{1} \]
Since we assume that the ensemble of systems is stationary, the phase density \(\rho\) does not change with time, and consequently the time average of the numerical average is equal to the numerical average,
\[ \overline{\overline{u(p\cdot q)}}^{\,t}=\overline{u(p\cdot q)}. \tag{2} \]
Further, since the formation of the time average is independent of the numerical average, we may interchange the order in which the averages are formed, i.e.
\[ {}^{t}u(p\cdot q)=\overline{u(p\cdot q)}^{\,t}. \tag{3} \]
Combining this with the preceding equation (2), we have
\[ \overline{u(p\cdot q)} = \overline{u(p\cdot q)}^{\,t} = u(p\cdot q)^{t}. \tag{4} \]
Consequently, for a stationary ensemble the numerical average is equal to the time average of the numerical average and to the numerical average of the time average.
What does \(u(p\cdot q)^{t}\) mean? We first take the time average of one system along its phase trajectory, then the time average of another system along its trajectory, and so on. Then we form the numerical average of all these averages. In the general case \(\overline{u(p\cdot q)}^{\,t}\) will change from trajectory to trajectory, since each trajectory is determined by its own constants of the time-independent integrals of the system of Hamilton equations.
If all trajectories are determined by the same constants, then the time average will be one and the same for all systems and is equal to the numerical average of the time average
\[ \overline{u(p\cdot q)}^{\,t}=u(p\cdot q)^{t}. \tag{5} \]
Comparing (5) with (2), we obtain
\[ \overline{u(p\cdot q)} = \overline{u(p\cdot q)}^{\,t} = \overline{u(p\cdot q)^{t}} = u(p\cdot q)^{t}; \tag{6} \]
thus we obtain:
\[ \boxed{\overline{u(p\cdot q)}=u(p\cdot q)^{t}} \tag{7} \]
We have thus arrived at the conclusion that the numerical average and the time average are equal thanks to the assumption,
that $\overline{u(p\cdot q)}^{\,t}$ does not change from trajectory to trajectory or, in other words, all trajectories are determined by the same values of the constants $(\varphi_2 \ldots \varphi_{2rN})$, independent of time, of Hamilton’s equations. Since Hamilton’s equations give a unique solution for each point of space, only one trajectory passes through each point, and equality of the constants for all phase trajectories can occur only in the case if the trajectories of all systems form one trajectory.
Thus, through each point of the energy surface there must pass only one trajectory; on the other hand, all trajectories represent one single trajectory. This is equivalent to the assertion that the phase trajectory passes through all points of the energy surface,—in other words, the system passes through all states compatible with the given energy.
This property of a system Boltzmann called ergodic.1 Only by virtue of the assumption that all trajectories represent one trajectory have we been able to prove that
$$ \overline{u(p\cdot q)}=u(p\cdot q)^{t}. $$
We assume that our system consists of a sufficiently large number of material particles, whose motions obey Hamilton’s equations. Thus, if we suppose that each particle possesses $r$ degrees of freedom, and the given system consists of $N$ particles, then for a complete description of the motion of the system we shall have $2rN$ equations of the form:2
$$ \frac{dq_s^k}{dt}=\frac{\partial E}{\partial p_s^k}, \qquad \frac{dp_s^k}{dt}=-\frac{\partial E}{\partial q_s^k}. \tag{8} $$
The principal feature of systems characterized in this way is Liouville’s theorem on the conservation of phase volumes.
Liouville’s theorem is a consequence of the special form of phase space \((\Gamma)\) and of the form of Hamilton’s equations. Thus, at its basis lies only the assumption that the material particles constituting the system move according to the laws of mechanics.
Of particular importance for the study is the stationary distribution of systems in phase space, to which there corresponds a stationary state of the system in time.
A necessary and sufficient condition for stationarity will be
\[ \frac{\partial \rho}{dt}=0, \tag{9} \]
where \(\rho\) will be a function of \(p, q\), and \(t\). Since, on the basis of the form of Hamilton’s equations,
\[ \sum\left(\frac{\partial \dot q_s^{\,k}}{\partial q_s^{\,k}}+ \frac{\partial \dot p_s^{\,k}}{\partial p_s^{\,k}}\right)=0, \]
and
\[ \frac{d\rho}{dt} = \frac{\partial \rho}{\partial t} + \sum_{s,k} \left( \frac{\partial \rho}{\partial q_s^{\,k}}\dot q_s^{\,k} + \frac{\partial \rho}{\partial p_s^{\,k}}\dot p_s^{\,k} \right), \tag{10} \]
it follows from this that the necessary and sufficient condition for stationarity is
\[ \sum_{s,k} \left( \frac{\partial \rho}{\partial q_s^{\,k}}\dot q_s^{\,k} + \frac{\partial \rho}{\partial p_s^{\,k}}\dot p_s^{\,k} \right)=0. \tag{11} \]
The density \(\rho\) must be such a function of \(p\) and \(q\) that it does not change with time. Hamilton’s equations admit only \(2rN-1\) time-independent integrals. Thus the necessary and sufficient condition for stationarity must have the form
\[ \rho(p\cdot q)=F(E,\varphi_2,\ldots,\varphi_{2rN-1}) \tag{12} \]
for a spatial distribution, or
\[ \sigma(p\cdot q)=\frac{1}{\operatorname{grad} E}F(E,\varphi_2,\ldots,\varphi_{2rN-1}) \tag{13} \]
for a surface distribution.
The condition of stationarity will also be satisfied if we set
\[ \rho(p\cdot q)=F(E) \tag{14} \]
or
\[ \sigma(p\cdot q)=\frac{\mathrm{Const}}{\operatorname{grad} E}. \tag{15} \]
This condition will in general be sufficient, but not necessary. We shall show that for ergodic systems, i.e. for such systems in which the numerical average (the spatial average) is equal to the time average, this condition will be not only sufficient but also necessary.
On the basis of the ergodic hypothesis, all points of the surface lie on one and the same trajectory. On the other hand, since the ensemble is stationary, the density along any trajectory must be constant. In the general case of stationarity the density is determined as (12) or (13).
But since \(\sigma \operatorname{grad} E\) must be constant along a phase trajectory, the only form of stationary density for ergodic systems will be (14) for the volume density, and (15) for the surface density.
These special forms of density are called ergodic density distributions.
In statistical mechanics we operate almost exclusively with ergodic densities. This is explained by the fact that only this special kind of densities is compatible with the ergodic hypothesis.
Thus, at the basis of all the conclusions of statistical mechanics lies the ergodic hypothesis. This fundamental assumption about the special character of systems is equivalent to the assumption concerning the number of collisions (Stosszahlansatz) in classical kinetic theory.
Like classical kinetic theory, statistical mechanics cannot be constructed solely on the basis of the assumption that the material particles of which the system consists obey the laws of mechanics.
The ergodic hypothesis makes it possible to give an expression for the time during which the phase point representing the system is found in different regions of the energy surface over the unlimited interval of time for which the system is left to itself:
\[ \lim_{T}\frac{dt}{T}=\frac{\sigma dS}{\int \sigma dS} \tag{16} \]
[\(dS\) is an element of the energy surface, and \(\sigma\) is defined by (15)]. If one accepts the ergodic hypothesis, then this expression for the relation between the temporal behavior of a gas and its spatial distribution is a purely mechanical theorem, wholly independent of any theoretical-probabilistic considerations.
The relation given determines (in the event that the ergodic hypothesis is correct) not only statements about average behavior, but also the relative interval of time during which the gas remains in various states.
If, however, the ergodic hypothesis is not accepted, then there is no basis for asserting that equality (16) is valid. Thus all of statistical mechanics is called into question.
Up to the present time it has not only proved impossible to provide even a single example of an ergodic system, but in 1918 Rosenthal and Plancherel showed that the ergodic hypothesis contains a contradiction. Therefore the possibility of equating the time average with the spatial average must be proved from other considerations.
If ergodic systems do not exist, then the equality of averages cannot be proved as a general theorem for all systems and is not a simple consequence of the mechanical structure of the system.
This, however, does not mean that the equality of the time average and the spatial average cannot occur at all.
This equality must be justified in each case, and the justification of this equality will be of a theoretical-probabilistic character.
§ 2. A New Foundation of R. Mises’ Theory of Probability
The difficulties of physical statistics are connected mainly with the uncertainty inherent in the classical concept of probability. A large body of literature has been devoted to the analysis and criticism of the foundations of the classical theory of probability, but very little has been done toward a satisfactory foundation of probability theory. Quite recently R. Mises (R. v. Mises) proposed a new foundation of probability theory, constructed on a conception of probability entirely different from that of the classical theories.
Mises’ new conception of probability also leads to a new conception of physical statistics.
We cannot dwell here on the exposition of the basic propositions of Mises’ theory, since our task is to show the concrete application of the new foundation of physical statistics to the problem of the ergodic hypothesis. Therefore, referring the reader to Mises’ original works and to expositions of his ideas,^1 we shall briefly formulate those basic propositions which are guiding for the entire theory.
The classical concept of probability is based on the concept of equally possible cases. This concept of equally possible cases is the weakest point of the classical theory of probability, giving the concept of probability a subjective character.
In Mises’ theory the concept of probability is connected not with the indefinite concept of equipossibility, but with the strictly defined concept of a collective, which has an objective meaning.
By a collective Mises means an infinite aggregate of objects satisfying the following two requirements: 1. Within the given aggregate there exist limits for the frequencies of elements with specified attributes. 2. If, from the given aggregate, we form a new aggregate by selecting elements in any manner,
^1 See the literature, Nos. 7–10.
not connected with the attribute of the element being selected, then within the new aggregate obtained by means of this selection the same limits for the frequencies are preserved as in the original aggregate.
The limit of the frequency of occurrence of a given attribute within an aggregate satisfying these two requirements will be the probability of occurrence of the given attribute within the aggregate.
Thus the concept of probability is inseparably connected with the concept of a collective, and acquires definiteness only when we can indicate the collective to which the given concept of probability pertains.
The aggregate of probabilities, i.e. the limits of frequencies within a given collective, is called a distribution. A collective is specified and characterized by specifying its distribution. For example, the aggregate of throws of a fair die gives a collective with six attributes. This collective, obtained by throwing a fair die, is characterized by the fact that its distribution, consisting of the six probabilities of obtaining one or another number of points, is uniform, i.e. all the probabilities are equal to one another. If the die were unfair, then we would no longer have a uniform distribution, but some other one, characterizing the collective obtained by throwing an unfair die.
From a given collective, by means of four basic operations, new collectives may be obtained. The task of probability theory is to determine the distribution in a new collective, obtained by means of various operations from a given collective, in the case where the distribution in the original collective is known.
As for the values of the probabilities (the distribution) in the original collective, they must be given. The determination of the distribution of the original collective does not enter, and cannot enter, into the range of problems of probability theory, just as the determination of the initial velocities and positions of bodies does not enter into the problem of mechanics.
Thus the fundamental assumptions concerning the original
probabilities must be made before we proceed to the solution of the probability-theoretic problem. Therefore, in statistical mechanics as well it is impossible to dispense with certain assumptions about the initial probabilities. The correctness of the assumptions we have made about the initial probabilities in the original ensemble is confirmed by the agreement with experiment of the consequences obtained under these assumptions.
The essence of the classical ergodic hypothesis consisted in the fact that, by making a certain assumption of a mechanical character concerning the gas model, we thereby obtained the possibility of identifying the spatial and temporal aggregate of states.
The essence of the probability-theoretic justification consists in the fact that we regard the aggregate of spatial and temporal states as two different ensembles.
We then show that, under certain assumptions concerning the distributions in the original temporal and spatial ensemble, the probabilities of the appearance of a definite attribute within the new spatial and temporal ensembles obtained by means of simple operations from the original spatial and temporal ensemble will be equal to one another.
Thus the place of the assumption about the special character of the mechanical system is taken by a certain assumption concerning the distributions within two original ensembles—the spatial and the temporal.
It has already been noted above that, in accepting the ergodic hypothesis, we impart to our conclusions the character of a dynamical law, independent of any probability-theoretic considerations.
In the probability-theoretic justification of the ergodic hypothesis, our statement about the equality of probabilities within the spatial and temporal ensemble likewise has a probability-theoretic character, i.e. it is given in the form of a statistical regularity and in fact not for all cases without exception, but only for the overwhelming majority of cases.
This must be borne in mind, since probabilistic-theoretical premises are often introduced implicitly in the derivation. The result, however, is expressed in the form of a dynamical law, which, of course, leads to contradictions. The entire history of the criticism of the \(H\)-theorem may serve as an example of such supposed contradictions.
Thus, the equality of the numerical (spatial) mean and the time mean for certain gas models has a statistical character and can be rigorously proved if we make certain assumptions concerning distributions in the two initial ensembles—the spatial and the temporal.
In this consists the probabilistic-theoretical justification of the ergodic hypothesis.
§ 3. The ergodic hypothesis in Brownian motion.
The ergodic hypothesis was first formulated by Boltzmann and Maxwell.
Despite the fact that only by means of the ergodic hypothesis are derivations of the fundamental theorems possible—for example, the theorem on the uniform distribution of energy over degrees of freedom—it was often introduced tacitly, and even Boltzmann himself does not mention it in his work, although he does give the density distribution (14), which he calls ergodic.
He simply gives the ergodic distribution as “the simplest case,” without referring to the ergodic hypothesis, under which, as was shown above, this distribution is the only one.
Despite all the importance of the theory of Brownian motion, until the very last time it had not been pointed out that in Brownian motion as well we apply the ergodic hypothesis. P. Lévy first pointed out that the theory of Brownian motion is based on the ergodic hypothesis, and gave it a probabilistic-theoretical justification.
Smoluchowski and Svedberg showed that the probability that in a given part of the volume at a given moment
time there will be exactly \(x\) particles has the following form:
\[ W(x)=\frac{e^{-\nu}\nu^x}{x!}; \tag{17} \]
where \(\nu\) denotes the number of particles falling within the given part of the volume in the case where the particles are distributed uniformly throughout the entire volume.
Expression (17) for the probability \(W(x)\) was obtained on the basis of combinatorial considerations, i.e. we operate only with the spatial ensemble, and not with its behavior in time.
Smoluchowski\(^1\) sees direct experimental confirmation of this formula in Svedberg’s work on colloidal solutions and emulsions. As is known, the verification of formula (17) amounted to the fact that Svedberg observed a layer of emulsion under a microscope at various intervals of time (39 times per minute, Perrin every 30 sec.). Throughout the experiment the emulsion was not stirred. By means of a diaphragm, a small region was isolated from the field of view, and in it the number of particles was counted.
The frequencies computed by Svedberg on the basis of 518 observations agreed very well with the probabilities given by formula (17).
It is clear that confirmation of formula (17) by means of such an experiment can be seen only in the case where we accept the ergodic hypothesis for the motion of Brownian particles.
Indeed, the probability obtained from combinatorial considerations as applied to a spatial ensemble of particles is verified by observing this ensemble in time. We assume that the time average is equal to the spatial average, which holds only for ergodic systems.
If we abandon the ergodic hypothesis, then the verification of formula (17) in Svedberg’s experiments can
\(^1\) M. v. Smoluchowski. Abhandlungen über die Brownsche Bewegung etc. Ostwald’s Klassiker, No. 207, p. 42.
recognized only if the equality of probabilities in the spatial and temporal ensemble of observations has first been proved.
We shall now proceed to this proof.
We start from the following model of Brownian motion: in a bounded part of space there are \(n\) Brownian particles. We observe the positions of these particles in a plane lattice consisting of \(N\) cells, at certain equal intervals of time, which are multiples of some unit of time \(\tau\). The coordinates of the lattice points are multiples of some unit \(\alpha\).
We have to establish the relation between the probability that, in a single observation, among all \(n\) particles, in \(z\) points of the lattice there are \(x\) particles each, and the probability that, when observing over the course of some interval of time \(m\tau\) (\(m\) sufficiently large), any arrangement of particles will, \(z\) times, pass into such an arrangement in which a certain cell of the lattice will contain \(x\) particles.
In the first case we seek the combinatorial probability corresponding to the spatial average. To obtain it we observe the particles at one instant of time, determine in how many places \(x\) particles are concentrated, then mix the particles, observe again, and so on. From the resulting series of observations our spatial ensemble will be composed. It is clear that we perform the mixing of the particles in order to obtain precisely a spatial ensemble, since mixing gives us a series of states independent in time.
In the second case we observe the transition of one arrangement of particles into another in time; that is, starting from some arrangement of the particles, we leave them to themselves for an interval of time \(m\tau\), note the arrangements that they have assumed after the interval of time \(m\tau\) (in all, therefore, there will be \(m\) arrangements), mix them, and again observe over the course of an interval of time \(m\tau\), note the new arrangements, mix again, and observe over the course of time \(m\tau\), and so on.
The totality of these observations forms an ensemble in time, in which we establish the probability of the sought constellation of particles. An element of this ensemble will be \(m\) consecutive observations at equal intervals of time.
I. Spatial ensemble.
1. Initial ensemble.
Starting from the indicated model of Brownian motion, we compose the initial ensemble, an element of which will be the observation of a single Brownian particle at a definite moment of time. The attribute of an element in this ensemble will be the number of the cell in which our particle is located.
With respect to the distribution of this ensemble we make the following basic assumption:
We take the distribution to be uniform; in other words: the probability that any particle, at a given moment of time, is in any cell of the lattice is equal to \(\frac{1}{N}\) (\(N\) is the number of cells).
Since we are dealing with a spatial ensemble, each observation, lasting one instant of time, is preceded by a fundamental mixing of the particles.
This basic spatial ensemble serves us as the starting point for the formation of new ensembles necessary for solving the problem posed above. The assumption we have made concerning the probability (distribution) in the initial ensemble makes possible the calculation of probabilities in the derived ensembles.
2. First derived ensemble.
Taking \(n\) ensembles identical with the initial one, we form a derived ensemble in which an element will be the observation of a group of \(n\) Brownian particles. In this ensemble we carry out the operation of mixing, combining together those groups of elements in which there will be \(x\) particles in some cell of the lattice. Thus, the attribute of the new ensemble, composed by joining \(n\) basic
collectives, there will be a group of \(x\) particles having one and the same cell number.
What will be the probability \(w_n(x)\) that in the given cell, for example at the origin of coordinates, at the given instant of time we shall find \(x\) particles?
Since the distribution of the initial collective is known to us, it is easy to find the value of \(w_n(x)\):1
\[ w_n(x)=\binom{n}{x}(N-1)^{n-x}N^{-n} =\binom{n}{x}\left(\frac{1}{N}\right)\left(1-\frac{1}{N}\right)^{n-x} \]
If \(n\) and \(N\) are sufficiently large numbers, but \(\frac{n}{N}=\nu\) is finite, then formula (17) can be simplified
\[ w_n(x)=\frac{1}{x!}\left(\frac{n}{N}\right)\left(1-\frac{1}{N}\right)^{N\nu} \left[ \frac{1}{1-\frac{1}{N}}\, \frac{1-\frac{1}{n}}{1-\frac{1}{N}} \cdots \frac{1-\frac{x-1}{n}}{1-\frac{1}{N}} \right] \cong \frac{\nu^x e^\nu}{x!} \tag{18} \]
3. Second derivative collective. Our task consists in determining the probability of finding \(x\) particles in \(z\) points of the lattice. Denoting by \(y\) the relative number of points in which there are \(x\) particles each, we have \(z=Ny\). The sought probability \(f_n(x,y)\) will be some function of \(x\) and \(y\). In order to determine it, we subject the collective, in which an element consists of a single observation of \(n\) particles, to a new mixing, which consists in the fact that we join together those elements in which \(x\) particles, located in one and the same cell of the lattice (having one and the same number), occur \(z\) times.
Since the distribution of the initial collective is known to us, we could directly calculate the probability \(f_n(x,y)\). But this calculation is very complicated, whereas for our purpose it is sufficient to find the mean value \(a\) and the variance \(S^2\) of \(f_n(x,y)\). As is known from the elements of probability theory, the following relations will hold:
Justification of the Ergodic Hypothesis
\[ \sum_{(y)} f_n(x,y)=1 \qquad \sum_{(y)} y f_n(x,y)=a \tag{19} \]
\[ \sum_y (y-a)^2 f_n(x,y)=\sum_y y^2 f_n(x,y)-a^2=S^2 . \tag{20} \]
To find the mean value \(a\), we proceed as follows. From all \(n\) particles we choose \(x\) particles. This can be done in \(\binom{n}{x}\) ways. These selected particles we place in some one cell of the lattice—this can be done in \(N\) ways. The remaining \((n-x)\) particles we place in the remaining \(N-1\) cells, which can be done in \((N-1)^{n-x}\) ways.
Among all possible arrangements of the \(n\) particles obtained in this way, there will also be those in which \(x\) particles are contained in \(z\) cells. But all such arrangements, evidently, in our method of placement will be counted \(z\) times. We may therefore write
\[ \sum_y z f^n(x\cdot y)=\binom{n}{x} N (N-1)^{n-x} N^{-n} \tag{21} \]
\(N^{-n}\) is the probability of any arrangement of \(n\) particles. Since \(z=Ny\), substituting this value in (21) and canceling \(N\), we obtain:
\[ a=\sum_y y f_n(x\cdot y)=\binom{n}{x}(N-1)^{n-x}N^{-n} \tag{22} \]
But the expression we have obtained for the mean value is precisely equal to the probability \(w_n(x)\), found above, of finding \(x\) particles in a given cell; consequently
\[ \boxed{a=w_n(x)} \tag{23} \]
In order to find the variance \(S^2\), we proceed in a similar way. We choose from \(n\) particles \(2x\) in \(\binom{n}{2x}\) ways. We divide the \(2x\) particles into groups, each containing \(x\) particles, which can be done in \(\binom{2x}{x}\) ways. We place the two groups obtained in any two cells out of \(N\), which can be done in \(\frac12 N(N-1)\) ways, and finally we place the remaining \(n-2x\) particles in the remaining \(N-2\) cells, which can be done in \((N-2)^{n-2x}\) ways. The product
\[ \binom{n}{2x}\binom{2x}{x}\frac12 N(N-1)(N-2)^{n-2x} \]
will give all possible arrangements of particles in which \(x\) particles occur at least in two places of the lattice, and, moreover, each arrangement in which \(x\) particles occur at two points will be counted \(\frac{1}{2}z(z-1)\) times.
Thus we obtain:
\[ \frac{1}{2}\sum z(z-1) f_n(x,y) = \binom{n}{2x}\binom{2x}{x}\frac{1}{2}N(N-1)(N-2)^{n-2x}N^{-n} \tag{24} \]
Substituting \(z=Ny\) and transforming, we obtain:
\[ S^2=\frac{1}{N^2}\left[\sum_y 2 f_n(x,y)+\sum_y 2(2-1)f_n(x,y)\right]-a^2 \]
\[ S^2=\frac{a}{N}+ \left(1-\frac{1}{N}\right) \left(1-\frac{2}{N}\right)^n \frac{(N-2)^{-2x}n!}{(n-2x)!x!x!} \left[ 1- \frac{\left(1-\frac{1}{N}\right)^{2n-2x-1}} {\left(1-\frac{2}{N}\right)^{n-x}} \frac{n!(n-2x)!}{(n-x)!(n-x)!} \right]; \]
since the expression in square brackets is negative for sufficiently large \(N\), and \(S^2\) is always positive, it follows that
\[ S^2<\frac{a}{N}=\frac{w_n(x)}{N} \tag{25} \]
that is, as \(N\) increases, the variance approaches \(0\) without bound and, consequently, the values \(f_n(x,y)\) cluster near the mean value \(a\).
As a result we arrive at the following proposition. For sufficiently large \(N\) (the number of cells of the lattice), with sufficiently high probability one may expect that the relative number of cells \((y)\) in which exactly \(x\) particles are found is, in its value, close to the probability \(w_n(x)\) of finding \(x\) particles in a definite cell of the lattice.
II. Temporal ensemble.
1. Initial ensemble.
The element of the initial temporal ensemble will be the observation of one Brownian particle at the beginning and at the end of a time interval \(\tau\). The attribute of each element will be the magnitude of the displacement, expressed by three numbers giving the co-
ordinates of the new position at the end of the time interval \(\tau\) relative to the coordinates at the beginning of the time interval \(\tau\).
Thus an element of our initial ensemble is characterized by the set of numbers \((x,\lambda,\mu)\).
With respect to the distribution in our initial ensemble we make the following basic assumptions:
- The probability of a particle’s displacement by \(x,\lambda,\mu\) along the coordinate axes, which we shall denote by \(p_{x\lambda\mu}\), is symmetric with respect to the sign of \(x\lambda\mu\), i.e.
\[ \sum_x x p_{x\lambda\mu} = \sum_\lambda \lambda p_{x\lambda\mu} = \sum_\mu \mu p_{x\lambda\mu} =0 \tag{1} \]
- All three terms of the dispersion are equal and different from zero:
\[ \sum x^2 p_{x\lambda\mu} = \sum \lambda^2 p_{x\lambda\mu} = \sum \mu^2 p_{x\lambda\mu} = r^2 \ne 0 \tag{2} \]
- Of all possible displacements in the three coordinate directions, there are always possible at least two displacements differing by only one unit.
Moreover, obviously,
\[ \sum_{x\lambda\mu} p_{x\lambda\mu}=1 \tag{3} \]
Conditions (1) and (2) show that all three coordinate directions are equivalent, i.e. we consider the motion of particles in the absence of an external field of forces.
As regards boundary conditions, we either assume that the boundaries are so far removed that they exert no influence on the motion, or we may suppose that a particle, on reaching a boundary, is reflected from it, and in this way establish a definite correspondence between points lying outside our region and inside it.
Thus our basic assumption about the initial ensemble may be summarized as follows: the displacement of a particle during a time interval \(\tau\) by \(x,\lambda,\mu\) along the three coordinate axes has probability \(p_{x\lambda\mu}\), satisfying conditions 1–3.
2. First derived ensemble. Combine \(m\) initial ensembles into a new ensemble. An element of this ensemble will be the observation of one particle during \(m\) successive time intervals.
A characteristic will be a set of \(m\) displacements, characterized by \(3m\) numbers
\[ (x_1\lambda_1\mu_1),\ (x_2\lambda_2\mu_2),\ (x_3\lambda_3\mu_3)\ldots (x_m\lambda_m\mu_m). \]
In this collective we carry out mixing, combining into one group all those elements for which the sums of the displacements along the axes
\[ x_1+x_1+\ldots+x_m=k \qquad \lambda_1+\lambda_2+\ldots+\lambda_m=l \]
\[ \mu_1+\mu_2+\ldots+\mu_m=q \]
are equal to the same quantities \(k, l, q\).
Let us denote the probability of a particle’s displacement by \(kx, lx, qx\) as
\[ V_m(k,l,q). \]
It can be proved,\(^1\) under the assumptions made about the distribution \(p_{kl\mu}\) in the initial collective, that the probability \(V_m(k,l,q)\) in the first derivative collective will have the form:
\[ V_m(k,l,q)=\frac{1}{\sqrt{(2\pi r^2 m)^3}}\, e^{-\frac{k^2+l^2+q^2}{2r^2m}} \tag{26} \]
The verification of this formula must proceed as follows. We mix the particles. We observe a particle at the beginning and at the end of the time interval \(m\tau\), mix again and observe during the time \(m\tau\), mix again, and so on. The difference in observation, as compared with the spatial collective, consists in the fact that in it we observe the given cell at one moment and determine the number of particles in it. In the case of the temporal collective, however, we observe the interval of time \(m\tau\), during which the particle, left to itself, successively passes from one position to another.
- The second derivative collective. Let us renumber all the cells of our lattice from 1 to \(N\). Then each arrangement of \(n\) particles can be expressed by a set—
\(^1\) For the case of one coordinate, the form of the probability is well known, and the derivation may be found in any course on probability theory and, for example, in De Haas-Lorentz, Die Brownsche Bewegung etc., Chap. 2. For the case of three independent coordinates, the proof was given by Mises, Fundamentalsätze d. Wahrscheinlichkeitsrechnung, Mathemat. Zeitschr. 4, 24, 68, 1919.
ness, by \(N\) numbers \(n_1, n_2 \ldots n_N\), where \(n_i\) expresses the number of particles in the \(i\)-th cell. For brevity we shall denote the aggregate of the \(N\) numbers \(n_1 \ldots n_N\) by \(\bar n\).
An element of the derived ensemble will be the observation of \(n\) particles at the beginning and at the end of the time interval \(\tau\).
To each possible system of changes in the positions of the particles over the time interval \(\tau\) there will correspond a certain probability, expressed by the product \(\rho_{\lambda\mu}\) of probabilities.
A characteristic will be the new arrangement \(\bar n^{(i)}\), in comparison with the preceding arrangement \(\bar n^{(i-1)}\), which existed at the beginning of the observation interval \(\tau\). The probability that a given initial arrangement \(\bar n^0\) at the end of the time interval \(\tau\) passes into \(\bar n^{(1)}\) depends, obviously, not only on \(\bar n^{(1)}\), but also on \(\bar n^0\). Therefore we shall denote this probability by \(V(\bar n^{(1)}\bar n^0)\).
4. Third derived ensemble. Let us combine \(m\) similar ensembles. An element of the new ensemble will be the observation of \(n\) particles over the time interval \(m\tau\).
The probability that during an observation lasting \(m\tau\) we shall have a transition of definite arrangements \(\bar n^{(0)}, \bar n^{(1)}, \bar n^{(i)} \ldots \bar n^{(m)}\), successively one into another, will be expressed by the product of probabilities
\[ V(\bar n^{(1)},\bar n^{(0)}),\; V(\bar n^{(2)},\bar n^{(1)}) \ldots V(\bar n^{(m)},\bar n^{(m-1)}). \]
In this ensemble, whose element consists of the observation of \(n\) particles over the time interval \(m\tau\), by means of the operation of mixing we form a new ensemble, combining into one group all those elements in which there are \(x\) particles in one cell (by the end of the observation interval \(m\tau\)). In order to calculate the probability \(fmn^{(x)}\) that by the end of the time interval \(m\tau\) in a definite cell there will be exactly \(x\) particles, we must form the sum of the products
\[ V(\bar n^{(0)},\bar n^{(1)})\; V(\bar n^{(2)},\bar n^{(1)}) \ldots V(\bar n^{(m)},\bar n^{(m-1)}) \]
under the condition that in a definite cell, for example in the first, by the end of the observation, i.e. in the arrangement \(n^{(m)}\), there will be exactly \(x\) particles.^1
We can arrive at the same result in another way. The desired final arrangement can be obtained only in the case where, of the \(n_i\) particles which initially had coordinates \(\varkappa_i a,\lambda_i a,\mu_i a\), exactly \(z_i\) receive the displacement \((\varkappa_1-\lambda_i)a,(\lambda_1-\lambda_i)a,(\mu_1-\mu_i)a\), with
\[
z_1+z_2+\cdots+z_N=x.
\]
Let the cell in which \(x\) particles are to be found be at the origin of coordinates, i.e. \(\varkappa_1=\lambda_1=\mu=0\). For brevity denote the probability of the required displacement \(v_m(\varkappa_i,\lambda_i\mu_i)\) at the end of the time interval \(m\tau\) by \(v_i\). In this case the required probability \(f_{mn}(x)\) will be equal to the coefficient of \(t^x\) in the expansion
\[ \prod_{i=1}^{N}(1-v_i+t v_i)n_i \tag{27} \]
Indeed, this coefficient contains all products of the form \(v_i^{z_i}(1-v_i)(n_i-z_i)\) under the condition \(z_1+z_2+\cdots+z_N=x\). The numerical factor attached to this product corresponds to the number of ways in which one can choose \(z\) particles from \(n_i\) particles.
Let us compute this coefficient under the assumption that \(n\) is sufficiently large, and
\[ \nu=\sum_{i=1}^{N} n_i v_i, \tag{28} \]
is finite, so that the \(v_i\) may be regarded as fractions of sufficiently small magnitude.
We have^2
\[ f_{mn}(0)=\prod(1-v_i)n_i \]
^1 Since we characterize each \(i\)-th arrangement by the totality of numbers \((n_1^{(i)}, n_2^{(i)},\ldots,n_N^{(i)})\) (\(i\) is the number of the arrangement), this condition can be expressed thus: \(n_1^{(m)}=x\).
^2 \(f_{mn}(0)\) denotes the probability that after the time interval \(m\tau\) there will not be a single particle at the origin of coordinates; i.e. all displacements will be such that the particles formerly there will leave the origin, and none will come there. The probability of a displacement taking a particle away from the origin will obviously be \(1-v_i\); hence we obtain the expression written.
Whence
\[ \Pi (1-v_i)^{n_i}=e^{-\sum n_i \ln(1-v_i)} \sim e^{-\sum n_i v_i}\sim e^{-\nu}, \tag{29} \]
if we neglect the higher powers of \(v_i\).
Further, for finite \(x\) we have
\[ f_{mn}(x)=f_{mn}(0)\sum_{\alpha_1=1}^{n}\ldots\sum_{\alpha_x=1}^{N} \frac{v_{\alpha_1}v_{\alpha_2}\ldots v_{\alpha_x}} {(1-v_{\alpha_1})(1-v_{\alpha_2})\ldots(1-v_{\alpha_x})} \sim \]
\[ \sim \frac{1}{x!}e^{-\nu}\left[\sum \frac{v_i}{1-v_i}\right]^x \tag{30} \]
\[ \frac{1}{x!}e^{-\nu}\left[\sum \frac{v_i}{1-v_i}\right]^x \sim \frac{e^{-\nu}\nu^x}{x!}. \tag{31} \]
In order to establish the equality of \(f_{mn}(x)\) and \(w_n(x)\), it remains for us to show that, for sufficiently large \(m\), the expression for \(\nu\), defined in (28) as \(\sum n_i v_i\), passes into the expression
\[ \nu=\frac{n}{N}. \]
We denoted by \(v_i\) the probability of displacement by \((\kappa_i,\lambda_i,\mu_i)\) over \(m\tau\) intervals of time.
The ratio of the probabilities of two displacements \((\kappa,\lambda,\mu)\) and \((\kappa_1,\lambda_1,\mu_1)\), according to expression (26), for the probability \(v_i\) is equal to
\[ e^{-\frac{(\kappa-\kappa_1)^2+(\lambda-\lambda_1)^2+(\mu-\mu_1)^2}{2mr^2}}. \]
For sufficiently large \(m\), the value of this ratio is close to 1. In other words, for sufficiently large intervals of observation the probabilities \(v_i\) become more and more equal to one another and, consequently, each is exactly equal to \(\frac{1}{N}\). Therefore (28) may be rewritten as
\[ \nu=\frac{1}{N}(n_1+n_2\ldots n_N)=\frac{n}{N}. \]
And consequently,
\[ \boxed{w_n(x)=f_{mn}(x)} \tag{32} \]
V. M. GESSEN
It is necessary to note that, as we have shown, the value \(f_{mn}(x)\) does not depend on the initial arrangement.
Thus, we arrive at the following conclusion: the probability that, for sufficiently large \(n\) and \(m\), the arrangement of the particles over a time interval \(m\tau\) will pass into such an arrangement in which, at a definite place of the grating, there will be exactly \(x\) particles is equal to the probability that, observing the arrangement of the particles at one moment of time, we shall find \(x\) particles in the given cell.
The fundamental difference between \(f_{mn}(x)\) and \(W_n(x)\) consists in the fact that \(f_{mn}(x)\) represents the probability under prolonged observation, i.e. gives an idea of the behavior of the system in time, whereas \(W_n(x)\) represents the result of single observations and gives an idea of the spatial character of the system.
It remains for us to show that the probability, among \(m\) trial observations, of encountering \(my\) arrangements of \(n\) particles in which there will be \(x\) particles in a definite cell is equal to the probability of finding \(x\) particles in a definite cell under a single observation.
We have shown that, for a sufficiently large interval of time (\(m\) sufficiently large), the probability of a given arrangement does not depend on the initial arrangement. Therefore, in order to compute the desired probability, instead of the unit of time \(\tau\) we shall take some other unit of time of such magnitude that, with sufficient approximation, \(w_n(x)\) may be taken instead of \(f_{mn}(x)\).
In Svedberg’s experiments the observations were made at the beginning and at the end of approximately 2 seconds.
Then the probability that any arrangement of particles will pass into one in which there will be \(x\) particles in the given cell may be taken equal to \(w_n(x)\).
The probability that, in \(m\) successive observations, \(x\) particles will be encountered \(my\) times in the given cell will be equal to the coefficient of \(t^{my}\) in the expansion:
\[ [1-w_n(x)+t w_n(x)]^m \]
Thus, the required probability \(f_{mn}(y,x)\) is equal to
\[ f_{mn}(y,x)=\binom{m}{my} w_n^{my}(x)\,[1-w_n(x)]^{m-my}. \]
This is nothing other than the distribution of the well-known Bernoulli problem, and it has, as is known from probability theory, the following mean value \(a\) and variance \(=s^2\):
\[ a=w_n(x) \]
\[ s^2=\frac{2w_n(x)\,[1-w_n(x)]}{m}. \tag{33} \]
For sufficiently large \(m\) the variance approaches 0 and, consequently, we arrive at the following result:
For sufficiently large \(m\) and \(n\), with probability close to certainty, it may be assumed that among \(m\) consecutive observations (following one another at sufficiently large intervals of time) of arrangements of \(n\) particles, arrangements in which \(x\) particles will be found in a specified cell occur with a relative frequency equal to \(w_n(x)\).
It is thereby proved that the relative frequency within the time ensemble is, with probability close to certainty, equal to the relative frequency in the spatial ensemble, and the ergodic hypothesis in Brownian motion is justified. It should be emphasized that the equality proved has the character not of a dynamical, but of a statistical regularity and, consequently, holds not without exception, but in the overwhelming majority of cases.
§ 4. Probability-theoretic study of Brownian motion and study by means of differential equations.
The treatment of Brownian motion given above differs from the usual one in that our results have a probabilistic character, whereas the usual treatment of Brownian motion starts from the differential equation of motion of a particle and, consequently, its results have the character of a dynamical regularity, independent of any probabilistic considerations. It is necessary to—
show that our conclusions are not in contradiction with the usual theory of Brownian motion. This is not hard to do if we take into account that, as was already mentioned above, in the theory of Brownian motion one usually makes tacit use of the ergodic hypothesis.
Let us trace the principal stages in the derivation of the well-known formula for Brownian motion.
We write the equation of motion of a particle of mass \(m\) under the influence of a force \(S\) and a frictional force \(f\), proportional to the velocity of the particle:
\[ m\frac{d^{2}x}{dt^{2}}=-f\frac{dx}{dt}+S . \tag{84} \]
Multiplying both sides of equation (84) by \(x\), we form the mean over all particles. After simple transformations equation (34) is rewritten as follows:
\[ \frac{m}{2}\frac{d}{dt}\left[\frac{d}{dt}\overline{(x^{2})}\right] - m\overline{\dot{x}^{2}} = -\frac{f}{2}\frac{d}{dt}\overline{(x^{2})} + \overline{Xx}. \tag{85} \]
Next one usually assumes that, according to the law of equipartition of energy, \(m\overline{\dot{x}^{2}}=kT\), and that \(\overline{Xx}\) is equal to 0. Then (35) is rewritten as
\[ \frac{m}{2}\frac{d}{dt}\left(\frac{d\,\overline{(x^{2})}}{dt}\right) + \frac{f}{2}\frac{d}{dt}(x^{2}) = kT . \tag{86} \]
Integrating and assuming that \(Ce^{-\frac{ft}{m}}\), owing to the smallness of \(m\), may be neglected after a sufficiently long interval of time, we obtain, after a second integration over the limits from \(0\) to \(\tau\),
\[ \overline{\Delta x^{2}}=kT\frac{2}{f}t, \]
or
\[ \frac{\overline{\Delta x^{2}}}{t}=\frac{2kT}{f}. \tag{87} \]
\(\overline{\Delta x^{2}}\) denotes, of course, not the actual path of the particle, but the mean square of its displacement.
This result contains a definite statement about the motion of a Brownian particle and admits no devi-
tions on the possibility of any displacement of the particle. According to this result, for example, it is impossible for the particle, after a sufficiently long interval of time, to return to its initial position. Such an assertion is in contradiction with the expression given above for the probability of displacement of the particle (26), according to which any displacement has a definite probability.
This contradiction is the result of a hidden application of the ergodic hypothesis in passing from (35) to (36).
Indeed, in replacing the mean kinetic energy by the expression \(kT\), we rely on the law of equipartition, which can be derived only under the condition of the ergodic hypothesis. If, however, we abandon the ergodic hypothesis and the interpretation of the law of equipartition based upon it, and instead of an absolute character give our statements a theoretical-probabilistic character, then result (37) must be interpreted as follows: if one observes the particle sufficiently often over a prolonged interval of time \(t\), stirring the emulsion before each observation, then on average we shall obtain that mean displacement which is given in (37).
Thus from (34) there in fact follows not (37), but the following expression:
\[ \lim_{t=\infty}\frac{\Delta x^2 v(x)}{t}=\frac{2kT}{f}, \tag{38} \]
in which \(v(x)\) denotes the probability of a displacement by the amount \(\Delta x\), and the summation extends over all possible values of the quantity \(\Delta x\). This equation is in complete agreement with our equation (26).
Indeed, in the case under consideration of motion in one dimension, one must put \(x=ka\), \(t=m\tau\), and for \(v(x)\) we obtain
\[ v(x)=\frac{1}{\sqrt{2\pi r^2 m}}\,e^{-\frac{x^2}{2r^2a^2m}}; \]
whence, for sufficiently large \(m\),
\[ \sum x^2 v(x)=r^2a^2m=\frac{r^2a^2}{\tau}\,t. \]
Thus (88) will be satisfied if we put
\[ \frac{r^2 \alpha^2}{\tau}=\frac{2kT}{f}. \]
With the proper interpretation, the result of the mechanical interpretation of Brownian motion is not only not in contradiction with the result obtained from theoretical-probabilistic considerations, but also makes it possible to establish a connection between the physical quantities \(T\) and \(f\), and the quantities \(r^2\alpha,\ \tau\) introduced by us.
Conclusion.
Ehrenfest concludes his well-known report on the foundations of statistical mechanics with the indication that “at the present time every investigation of the structure of physical theory inevitably leads to the question of the nature of theoretical-probabilistic hypotheses.”1
Statistical mechanics, in contrast to the kinetic theory, rests not on assumptions of a theoretical-probabilistic character like Stosszahlansatz, but on the ergodic hypothesis, which is not a theoretical-probabilistic proposition.
The ergodic hypothesis not only contains an internal contradiction, but also leads to disagreement with experiment, for example, in the question of the uniform distribution of energy over degrees of freedom.
It is clear that the way out must be sought in replacing the ergodic hypothesis by theoretical-probabilistic assumptions. We thus return to the proposition that it is impossible to construct statistical mechanics on the basis solely of assumptions about the mechanical character of the system, without any assumptions of a theoretical-probabilistic character.
The essence of Stosszahlansatz consisted in the fact that, proceeding from assumptions about the equal probability of certain states, we
lead to the law of distribution of the relative frequencies of non-equipossible phenomena (the Maxwell–Boltzmann distribution).
It is precisely the statistical character of the Stosszahlansatz that makes it possible to give an interpretation of irreversible phenomena by means of a mechanical model.
In rejecting the ergodic hypothesis, we must again put in its place certain assumptions about initial probabilities within the original ensembles.
These assumptions do not follow from the mechanical character of the system and cannot be obtained from considerations of probability theory. They must be specified on the basis of considerations that are derivable neither from mechanical equations nor from operations of probability theory. Therefore statistical mechanics cannot be reduced either to pure mechanics or to pure statistics.
Literature
-
Ehrenfest, P., and T. Begriffliche Grundlagen der Statistischen Mechanik. Enzykl. d. Math.-Wiss. IV.
-
Lorentz, H. Theories Statistiques en Thermodynamic.
-
Wassmuth. Grundlagen d. Statistischen Mechanik.
-
Gibbs. Grundlagen d. Statistischen Mechanik.
-
Rayleigh. The Law of Partition of kinetic energy. Scientific papers, Vol. 4.
-
De Haas-Lorentz. Die Brownsche Bewegung etc.
-
Mises, R. Grundlagen der Wahrscheinlichkeitsrechnung. Mathematische Zeitschr. 5. S. 52–99, 1919.
-
Mises, R. Ausschaltung d. Ergoden Hypothese in der physikalischen Statistik. Phys. Zs. 225, 256, 1929.
-
Mises, R. Wahrscheinlichkeit, Statistik und Wahrheit. (A Russian translation is being printed: “Probability, Statistics, Truth.”)
-
Tëssen, V. The Statistical Method in Physics and the New Foundation of the Theory of Probability of R. Mises. Estestvoznanie i Marksizm No. 1, 1929.
-
Khinchin, A. R. Mises’s Doctrine of Probabilities and the Problems of Physical Statistics. Uspekhi fizicheskikh nauk No. 2, 1929.