Abstract
The article presented to readers was published after Smoluchowski’s death in 1917, in the anniversary issue of the journal Naturwissenschaften dedicated to the 60th anniversary of M. Planck’s birth.
Full Text
On the Concept of Randomness and the Origin of Probability Laws in Physics¹
M. Smoluchowski.
I.
The theory of probability, which at the beginning of its development found application with great success in biology and sociology—fields not readily accessible to mathematical treatment—has in recent times won for itself a new and extremely important sphere of application: physics. By the sphere of physics we now understand not the theory of errors of physical measurements, which since the time of Gauss has developed into an entire auxiliary discipline, but the fundamental constructions of this science—the whole system of theoretical physics. Probability theory was first introduced by Clausius and Maxwell in 1857–1860 as an auxiliary mathematical means for developing the kinetic theory of gases. After a brief period of stagnation, owing to the final victory of the atomistic view, probability theory acquired fundamental importance for physics and to this day remains the most important instrument of investigation in the field of the new theories of matter, the electron theory, radioactivity, and the theory of radiation. The spirit of probability theory is fully in keeping with the tendency, increasingly manifest in recent times, to reduce all the laws of physics,² following the model of the kinetic theory of gases, to the statistics of hidden elementary processes. The “simplicity” of these elementary
¹ The article here brought to the attention of readers was published after Smoluchowski’s death, in 1917, in a jubilee issue of the journal Naturwissenschaften dedicated to the 60th anniversary of M. Planck’s birth (Naturwissenschaften, 17, 1918). Although 10 years have passed since the writing of this article, the ideas set forth in it have not only not lost their freshness; precisely now, in connection with the posing of the problem of causality in the new quantum mechanics, Smoluchowski’s ideas acquire special interest. — The article was translated by B. M. Gessen. Ed.
² The statistical interpretation has so far not extended only to the equations of Lorentz’s electron theory, the law of conservation of energy, and the principle of relativity; but it is quite possible that in time here too exact laws may be replaced by statistical regularities.
processes must be regarded as a “secondary consequence” of the “law of large numbers.” Despite the gigantic expansion of the domain of application of probability theory, an exact analysis of the concepts underlying it has achieved only the most insignificant successes. To this day the assertion remains entirely valid that no mathematical discipline rests on so unclear and shaky a foundation as probability theory. Thus, for example, the fundamental questions concerning the subjectivity and objectivity of the concept of probability, and concerning the definition of the concept of chance, are resolved by different authors in diametrically opposite senses. There is a complete absence of a general and mathematically precise definition of the characteristic conditions for the criterion of the possibility of applying probability theory. In this respect, for the most part, one relies on intuition.
The present sketch is an attempt to investigate the fundamental concepts and those difficulties which arise in an especially acute form when probability theory is applied in physics. It must be acknowledged that it was precisely the unsatisfactory character of such investigations in recent works—works otherwise deserving of full attention—that gave the impetus to the present article. It goes without saying that the present article in no way lays claim to a complete and final resolution of the entire complex of philosophical questions connected with the concept of probability; however, it may perhaps prompt further investigations in a definite direction, in that it brings to the fore and properly illuminates the principal guiding idea concerning the objective aspect of the concept of probability, to which almost no attention has hitherto been paid.
II.
To the question of which events fall within the sphere of investigation of probability theory, the usual answer is that they are those events the occurrence of which depends on chance. The investigation of this latter concept is, consequently, in any case primary and fundamental, and we shall therefore first of all try to clarify what characterizes the essence of chance.
Connected with this question are two problems which have aroused much controversy, and whose difficulty is especially felt in the exact mathematical constructions of theoretical physics. They may be formulated as follows:
-
In what way is the result of chance amenable to calculation, and, consequently, in what way do chance causes have lawful actions as their consequence?
-
In what way can chance arise at all, if everything that occurs must be reduced solely to lawfully acting ...
laws of nature? In other words, in what way can law-governed causes produce random effects?
If one considers, in popular terms, chance as the negation of lawfulness, then the contradiction cited above is completely insoluble; however, such a conception of the notion of chance is wholly irreconcilable with the determinism that predominates in contemporary science. It is often desired to explain the contradiction by assuming that between a given cause and effect there exists a lawful causal connection; but because of the complexity of the phenomena we cannot determine the character of this connection, and therefore there arises an apparent absence of lawfulness. In this sense chance appears as an “unknown to us partial cause.” Meinong’s conception apparently comes close to this view1, according to which the random is the realization of something “not necessary.” In this case the necessity that is denied must be either internal or external (in relation to a definite complex of objective phenomena). But if, from the point of view of determinism, cause and effect are regarded as constantly linked by the internal necessity of elementary processes, then one can speak of the “not necessary” only in a relative sense—namely, insofar as the objective necessity is not yet known, i.e., insofar as part of the operative causes remains undetermined.
This usual conception, which reduces chance to our ignorance of the operative laws or causes, could at best give an answer to the second of the questions posed above. But the first question—why it is possible to calculate the action of unknown partial causes—still remains unresolved.
The numerous philosophical works concerned with the analysis of the concept of probability provide no explanation in this respect. In general, the philosopher is concerned with a different side of the question than the physicist. The philosopher directs his attention above all to the subjective, psychological side of the concept of probability, analyzes its epistemological significance, and investigates how probable statements, together with certain and false statements, enter into the system of formal logic. But he does not touch upon the question of the character of the objective phenomena that underlie the concept of probability.
In contrast to the philosopher, exact natural science is interested not in statements and not in subjectively justified or unjustified assumptions, but in objective, or “mathematical,” probability, i.e. in the relative frequency of the occurrence of definite random events. Exact natural science—as Meinong aptly observed—employs an undefined concept
probability in a very narrow sense, to which this author and other philosophers would more readily give the designation of degree of possibility. But only in this narrow sense is the concept of probability amenable to exact mathematical development. With this concept there occurs the same thing as with the terms: force, work, energy, heat, which the physicist understands in a sense entirely different from that in which they are used in everyday life.
It is perfectly clear that, insofar as the matter concerns application in theoretical physics, all theories of probability that regard randomness as an unknown partial cause must be acknowledged in advance as unsatisfactory. The physical probability of an event can depend only on the conditions influencing its occurrence, and not on the degree of our knowledge.
I am fully aware that what I have said above contradicts the usual widespread opinion, which considers partial ignorance of causes to be the most essential point; therefore, in support of my thought I shall note the following: the calculus of probabilities, as applied to the kinetic theory of gases, would be fully justified even in the event that the structure of molecules, their initial positions, etc., were known to us with absolute exactness, and we were able, with mathematical exactness, to trace their motion at every instant of time. The theory of probabilities would then remain at least just as rational a mathematical tool as abbreviated multiplication or the use of tables of logarithms (or of a slide rule), alongside ordinary exact multiplication.
How, however, do the defenders of the usual view of probability explain the possibility of calculating the effect of unknown partial causes? They usually refer to the “law of large numbers” as a principle which, although unprovable, is empirically quite justified. Timerding1, for example, says the following:
“… the continuous causality of all phenomena of nature may be preserved, but it is insufficient fully to explain the regularity of everything occurring in the world. It is necessary to introduce one more proposition, called the law of large numbers, whose consequence is that deviations from regularity, which are introduced by random events, disappear in the final result. Our reason resists accepting this principle only because for the most part it proves to be correct. It strives above all to find a basis for such a leveling out of random phenomena. However, such a basis is impossible to find.”
In truth this is a very unsatisfactory solution of the question, and we must try to find another way out of the dilemma. Poincaré observes that even in pure mathematics one can often speak of laws of probability; thus, for example, the frequency of the digits 1, 2, 3 in the last place of a column of numbers in a table of logarithms follows the usual law of probability of equipossible cases. Will the mathematician really be satisfied with admitting here the action of an incomprehensible, purely empirical law of large numbers?
III.
An indication of the solution of the question, as it seems to me, lies in the fact that the definitions of chance given above as an unknown partial cause1, and in general all similar definitions, are undoubtedly too broad.
When Leverrier noticed that the motion of Uranus did not quite agree with the computations made in advance, he did not say: “this is chance.” We have absolutely no knowledge of when a magnetic disturbance may occur, but we by no means consider its occurrence to be a matter of chance.
In all these examples there is absent the essential feature of what in everyday life and in science we designate as chance, namely that which may briefly be defined as follows: small causes—great effects.
The slightest difference in the initial push of the roulette wheel—and as a result the winning or loss of an enormous sum of money. Poincaré2, who especially emphasized this point, at the same time indicates two further features of chance: the complexity of many simultaneously acting causes, or else the interaction of two phenomena which usually belong to independent domains. But I think that all cases of this kind, after careful analysis, can be brought under the definition given above.
The feature indicated above appears especially vividly in all cases where the matter concerns a state of unstable equilibrium. Let us imagine a die of ideally exact cubic form, standing on one of its vertices; then the slightest deviation of the center of gravity from the vertical fully determines onto which of the three adjacent faces the cube will fall. We say that the number of points that comes up depends on chance. Expressed mathematically, the action \(y\) (the number that will be on top) depends on the cause \(x\) (the position of the center
gravity), so that the function \(y=f(x)\), for the value of \(x\) corresponding to the position of equilibrium, has a discontinuity. Let us note in passing that in this case the cause consists essentially of two variables: if we project the center of gravity \(O\) and the three edges, mutually intersecting at the lower corner \(E\), onto the horizontal plane, then, obviously, the distance \(r=OE\) in the projection thus obtained determines the velocity of fall of the die. The angle \(\theta\), determining the direction of the vector \(OE\) with respect to the three edges, determines the number that will turn up.
A randomness of this kind is not amenable to preliminary calculation and therefore cannot serve as a basis for applying the theory of probabilities. Indeed, so long as the definite causes are not known to us with sufficient accuracy (in the present case, the position of the center of gravity), we can say nothing in advance about their action. If, however, they are known, then we can predict their action with certainty, and then there remains no place for probability.
As an example of randomness not subject to calculation, one may cite an artilleryman who has at his disposal a gun acting with mathematical precision, but who must fire at a target whose distance from him is unknown. He lacks knowledge of one quantity on which the correct elevation of the gun depends, and if he hits the target, this will be a matter of blind chance. There can be no question either of any preliminary calculation or of any probability in our sense so long as the psychology of this artilleryman is unknown to us.
But as soon as it becomes known to us that this artilleryman uses definite methods of firing or definite mechanical auxiliary devices, of which we shall speak below (for example, rotation of the body of the gun about its axis), the problem becomes quite definite, and it is possible to indicate a definite “probability of hitting” (in accordance with the size of the target, its distance from the gun, etc.).
Thus randomness, or, one might say, “ordered” randomness, which makes possible the application of the theory of probabilities, differs from randomness in the broad sense by an essential feature, namely, by a known regularity of action under frequent repetition of the phenomenon, independently of the special character of the cause. For example, if in the above case of the die we make it fall from a height of one meter onto an absolutely smooth (not ideally elastic) board, then the phenomenon changes in an essential way. The die rebounds, falls, rises again, and repeats this motion several times, the height of rise becoming ever smaller, while the rotational motions, seemingly quite arbitrary, keep increasing, until at last it falls on one of its six faces. On ca-
which face will ultimately fall depends on its initial position. But the function \(y=f(r,\theta)\), expressing this dependence, must be such that, under a continuous change of the independent variables \(r\) and \(\theta\), which determine the initial position of the cube, the regions corresponding to all possible final positions of the cube are traversed extremely rapidly. This change of the final positions as a function of the change in the conditions determining the initial position occurs in such a way that already within an extremely small region of variation \(V\) of the original arrangements of the axes of the cube (with respect to the perpendicular to the plane) the range of variation of the numbers \(1\)—\(6\) turns out to be densely covered. The magnitude \(V\) could be designated by the term “region of equalization.” If, before throwing the die, we tried, by means of any auxiliary devices, to orient it in a definite way, then even with the greatest care errors in the setting would be inevitable. We shall denote the region of these inevitable errors as the region of deviation \(\Omega\); it may be assumed that the distribution function \(\varphi(r,\theta)\), which represents the relative frequency of these errors in an innumerable number of repeated throws of the die, will have a regular “analytic” character. If, therefore, the region \(V\), determined by the form of the function \(f(r,\theta)\), is small in comparison with the region of deviations, then it is easy to see that in the end all the figures from \(1\)—\(6\) must have the same probability, independently of the special character of the initial setting of the die and of the form of the function \(\varphi(r,\theta)\). Each individual event cannot be foreseen, but with continually continuing repetitions it is quite possible to foresee the general distribution of events. In such a case chance reigns in a regular manner.
Simpler than the case of the die is the example of roulette, on which Poincaré1 develops a similar argument, or else the example of a rotating disk divided into black and white sectors and serving as a target for a shooter. Whether he hits the black or white sectors depends on the moment at which the shot is fired from the (stationary) gun. But it is always possible to impart to the disk serving as the target such rapid rotation that the factor of the shooter’s accuracy will be excluded. Whatever moment he may choose to fire, from the moment of the decision to the moment of the shot there will always pass an indeterminate interval of time, though varying within definite limits, so that the probability that the shot will occur precisely at the moment \(t\) is determined by some function \(\varphi(t)\), which within the limits of deviation from \(t\) to \(t+\tau\) is different from zero. The form of this function is indifferent to us, but we assume that it has no discontinuities
of continuity and of an infinitely large number of maxima and minima.
If sufficiently many revolutions of the circle fall within the range of deviations of the time intervals, then the influence of the form of the distribution function \(\varphi(t)\) is eliminated; the probability of falling into a white or black sector then depends exclusively on the relative size of the area occupied by them. It is usually customary to speak simply of this probability, without paying attention to the function \(\varphi(t)\). But the assumption made above concerning it is tacitly accepted. This reasoning, based on the concept of probability, becomes meaningless if the gun is connected to the rotating circle by an electrical contact.
In the final analysis the whole argument rests on the fact that every (differentiable) function, in the domain of sufficiently small changes of the independent variable, changes approximately in proportion to these changes; this can be explained by a simple geometrical analogy. Let us rule a sheet of paper into narrow, equal-width, white and black strips. If we then draw by hand an arbitrary (but not too small and not too irregular) closed curve, then the “black” and “white” areas cut out by it will, with fairly high accuracy, be approximately equal, independently of the shape of the curve. The shape of the curve corresponds to what we called the individual range of deviation, while the manner of dividing the paper into differently colored strips is determined by the necessary form of the causal relation.
Thus we see how, as a result of the action of randomness, a definite law may be obtained independently of the special form of the unknown distribution function. In this way the first of the contradictions indicated in the second chapter finds its explanation. Of course, it must be admitted that our reasoning by no means exhausts the essence of randomness, since it is based on the assumption of a certain distribution function \(\varphi\) for the random deviations of the cause, and, moreover, we assume that this function possesses certain properties (regular variations). This circumstance finds expression in one statement, on the whole quite apt, by which mathematicians\(^{1}\) distinguish themselves from an answer to the question of the essence of randomness: the task of probability theory consists not in explaining the probability of an event, but in calculating its probabilities on the basis of other probabilities, namely on the basis of the known probability of a simpler phenomenon that is the cause of the more complex one.
\(^{1}\) See, for example, E. Borel, Le hasard. Paris, Alcan, 1914, p. 15 (Russian translation: Borel, Chance. Contemporary Problems of Natural Science, book 8. GIZ, 1923).
IV.
Let us sum up everything said above in a more general form. We call chance a special kind of causal connection. It is usually said that an event \(y\) depends on chance if it is a function of some variable cause (the magnitude of which is unknown or to which attention is deliberately not paid) or of a partial condition \(x\), and this dependence is such that the occurrence or non-occurrence of the event depends on a very small change in \(x\) (“small” in relation to the region of variation of \(x\)).
This usual formulation of the concept of chance, however, is quite insufficient to serve as the basis for an exact definition of the concept of probability. One can speak of the form of the mathematical law of probability \(W(y)\), relating to the quantity \(y\), only when the causal connection, expressed by the relation \(y=f(x)\), besides the properties mentioned above, also possesses the following special property: the distribution of \(y\), at least within certain limits, does not depend on the character of the distribution function \(\varphi(x)\), which determines the relative frequency of \(x\) (according to the supposition, \(\varphi(x)\) must vary “lawfully”).
These conditions can easily be formulated mathematically for the case of one variable, if we keep in mind the examples given above.
It is sufficient that the function \(y=f(x)\) have an “oscillatory” character of such a kind that:
-
For every value \(x_0\) in the region of variation \(\Omega\), it would be possible to indicate \(\Delta x\), so small in comparison with \(\Omega\), that the function \(y=f(x)=f(x_0+\varepsilon \Delta x)\) assumes all values, while \(\varepsilon\) assumes all values from 0 to 1.
-
All parts of the region \(\varepsilon\) corresponding to a definite region of values of \(y\), for all points \(x_0\) lying within \(\Omega\), are (approximately) equal in size.
To each \(x\) there corresponds some smallest region \(\Delta x\), to which the whole aggregate of changes of \(y\) corresponds, and the magnitude of the region \(\Delta x\) determines in a known way the structure of the causal relation \(f(x)\). The more “fine-grained” the structure of the causal connection, i.e. the smaller \(\Delta x\), the less significant are the requirements that we must impose on the “regularity of variation” of the primary distribution function \(\varphi(x)\) in order to obtain for the distribution \(W(y)\) a result independent of the character of \(\varphi(x)\).
It goes without saying that, conversely, each value of \(y\) may appear as the consequence of a whole aggregate of different values of \(x\), i.e. the inverse function is highly multivalued: the same effect may be caused by the most varied combina-
M. SMOLUCHOWSKI
relations of causes—this is likewise a very characteristic feature of those causal relations which give rise to the emergence of probability laws. Special cases of such a functional dependence are easy to indicate: for example, \(y=\sin\left(\dfrac{x}{a}\right)\). Suppose that \(a\) is extremely small in comparison with the range of variation of the “cause” \(x\); then \(\Delta x=\dfrac{2\pi}{a}\) will also be very small, and as a result for the “effect” \(y\) we obtain a frequency distribution independent of the probability of \(x\):
\[ W(y)\,dy=\frac{1}{\pi}\frac{1}{\sqrt{1-y^{2}}}\,dy. \]
Still simpler is the case considered by us above, with the rotating target. Here we take as \(x\) the time \(t\) at which the shot is fired; \(y\) here will be the angular distance \(\theta\) of the point on the circle at which the projectile strikes. Thus \(\theta=ct-2n\pi\), where the angular velocity \(C\) must be very large and \(n\) is chosen so that \(\theta\) lies between 0 and \(2\pi\). The interval \(\Delta x\) in this case too is, obviously, equal to \(\Delta x=\dfrac{2\pi}{c}\), and all angles \(\theta\) will be equally probable if this quantity is small in comparison with the range of variation of the cause.
There exists, moreover, a whole series of cases, not so easily accessible to mathematical analysis, in which by a purely physical arrangement one can attain, with any desired approximation, independence of the resulting probability law from the character and causes of the primary deviations. We shall examine in somewhat greater detail the following examples, as the most characteristic.
- Galton’s board. This apparatus consists of an inclined board with a large number of pins, which are arranged in regular horizontal rows, their arrangement being such that the pins of each row are placed opposite the gaps formed by two neighboring rows of pins. If, from a certain place on the board, balls of the corresponding size are made to roll down (their diameter must be somewhat smaller than the distance between two neighboring pins), then, owing to collisions with each pin, they will deviate from their path in a disorderly manner and in the end, after they have passed through all the rows of pins, will collect in a special receiver arranged at the lower edge of the board. The position which they occupy in this receiver can directly serve as a measure of the probability of the corresponding position of the balls.
It turns out that the positions of the balls in the receiver are distributed according to Gauss’s law of error distribution, \(y=Ae^{-\alpha x^{2}}\), in such a way that the greatest number collects at the place correspond-
THE CONCEPT OF RANDOMNESS AND THE LAWS OF PROBABILITY
the corresponding point from which the balls emerged; their number on both sides of this position decreases in accordance with the Gaussian curve. Mathematically this result is easily explained if we assume that each ball, after leaving the opening between two pins, can with equal probability pass to the right or to the left of the pin situated beneath it. If this phenomenon occurs completely at random, with equal probability for a deviation to the right or to the left, then the probability that, in passing through the \(m\)-th row of pins, the ball will have a deviation from the mean line equal to the distance between the pins increased \(n\) times is expressed by Bernoulli’s well-known formula:
\[ W(n)=\left(\frac{1}{2}\right)^m \frac{m!}{\left(\frac{m}{2}-n\right)!\left(\frac{m}{2}+n\right)!}. \]
For large values of the number \(m\), this formula is approximately equal to the expression given above. Thus a complex aggregate of phenomena has been reduced to simple elementary phenomena; but it still remains to explain why we may regard these latter as completely random, although in essence the initial position and the initial velocity of the ball unambiguously determine all its subsequent motions.
In order to exclude the action of uncontrollable secondary circumstances, we idealize our case by making the following assumptions: let us assume that the board is absolutely smooth, that the arrangement of the pins is perfectly regular, that the balls have a geometrically regular shape; let us further suppose that their diameter is almost exactly equal to the distance between the pins and that the impact of the ball against a pin is inelastic. Then it is quite clear that, after the ball has emerged between two pins, the remaining horizontal component of velocity wholly determines whether it will strike the next pin on the right or on the left side, i.e. whether it will pass on one side of it or the other. In turn, the horizontal component of the velocity is the result of the numerous reflections of the ball from a pair of pins and is determined unambiguously by the position of the line of centers relative to the corresponding row of pins at the first impact. The slightest deviation of the line of centers is enough to change the value of the horizontal component of velocity to the opposite one. With a further extremely small change of position, this component may again take the opposite value, and so on.
In the experiment described we recognize the characteristic features of “ordered” randomness:
I. Small causes—large effects.
II. The oscillatory character of the causal connection, which can be expressed, not quite precisely, in the words: “different causes—identical effects.”
III. An approximately uniform distribution of chances in elementary events. In the limit, when the diameter of the ball is exactly equal to the free distance between the pins, the function expressing the connection between the totality of initial conditions and the final position of the ball loses its analytic character. The chances for a positive and a negative deviation at each impact become equal, and we obtain the Gaussian distribution curve, quite independently of how small the fluctuations in the totality of initial conditions for the balls may be (provided that they are not exactly equal to zero). Thus we obtain a model, so to speak, of an ideally random phenomenon. The phenomenon described, let us note in passing, provides an excellent illustration of a whole class of physical phenomena which we usually designate as diffusion and thermal conductivity. Without going into details, let us note that the deviations to the side which the ball receives in passing through successive rows of pins correspond exactly to the deviations in so-called Brownian molecular motion. If in our experiment we limited the width of the Galton board by two partitions, and if from the right half of the upper row we dropped white balls and from the left half black ones, then after passing through all the rows of pins the balls would gradually mix in exactly the same way as two gases do in diffusion in the well-known experiment of Loschmidt. If the “bounded” Galton board were of sufficient length, then the result would have to be a uniform distribution.
- A mathematically more complicated, but physically simpler, example will be the following: let us imagine a vessel of any irregular shape with perfectly reflecting walls, into which, through a very small opening in the wall, we throw a mercury ball (best of all, a gas molecule). Let us try to calculate when the ball must again come out through this opening. Since the opening is sufficiently small in relation to the surface of the vessel’s walls, the ball, in general, owing to numerous reflections from the walls, must describe an extremely complicated zigzag path before it again reaches the opening. It is perfectly clear that the slightest change in direction upon entering the opening will cause very large changes in the form of the path, for the traversal of which the ball will have to spend more time, and this will cause considerable changes in the interval of time necessary for its exit from the opening. It is just as easy to see that by means of the very same
of different combinations of initial conditions, identical intervals of time can be obtained for the ball to emerge. For this it is only necessary, in reverse order, to trace the segments of the path at exit.
Thus there appears, as it were, the possibility of applying the theory of probability. It is true that in this case an exact mathematical analysis has not yet been carried out, but physical considerations from the field of the kinetic theory of gases, and also from the theory of radiation, in which this same problem appears in a somewhat different form, make very plausible the assumption that, under any distribution of initial directions, an equalization of probabilities occurs with time. This equalization occurs in such a way that every element of volume inside the vessel represents an equal probability for the location of the ball; furthermore, with equal probability it can move in any direction, and on average it collides equally often with any element of the surface of the vessel.
If the speed of the ball is denoted by \(c\), the volume of the vessel by \(V\), and the cross-section of the opening by \(\omega\), then, by analogy with the calculations of the kinetic theory of gases, it is easy to show that the probability of the ball’s exit from the vessel during an interval of time \(\tau\) is expressed by
\[ W=\frac{\omega c\tau}{4V}. \]
Thus the mean interval of time during which the ball remains in the vessel is equal to
\[ T=\frac{4V}{\omega c}. \]
The characteristic features of (ordered) randomness are revealed to an even greater degree when one deals with the motion of an aggregate of balls enclosed in a closed vessel. In this case, the mutual collisions of the balls have as their consequence the disorderly disturbance of the originally existing state of motion.
This will be a special case of the tendency toward molecular disorder indicated by Boltzmann, which constitutes a general property of molecular systems. The kinetic interpretation of the law of entropy is based on this tendency.
V.
The considerations by which, in Chapters III and IV, we attempted to characterize the essence of randomness and the regularity of its action seem to me not entirely satisfactory in two respects:
- We assumed that the cause \(x\) follows a definite law of probability \(\varphi(x)\); thus this concept was presupposed as
something primary. It was necessary to explain only the invariability of the law of probability for the resultant action.
2. We assumed known properties of the function \(\varphi(x)\), which we denoted as “regularities.”
These two observations compel us to draw attention to yet another shortcoming of our reasoning. Indeed, what does it mean when we say that the probability of the occurrence of \(x\) (the movement of the hand when spinning a roulette wheel, the position of a die at the beginning of its fall, the position of a ball on Galton’s board) is determined by a regular distribution function \(\varphi(x)\)?
If the matter concerns an \(x\) which we could not reduce to primary causes, then the law \(\varphi(x)\) was known only empirically. What is immediately given to us is a discrete aggregate of individual cases, and only by abstraction on the basis of an innumerable quantity of individual cases does it become possible to determine the function \(\varphi(x)\), with respect to which we suppose that it possesses property (2).
It would therefore be more rational to leave entirely aside the intermediate abstract concept of the distribution function \(\varphi(x)\) and to introduce directly into consideration a definite number of unit cases. Let us therefore try, in place of the formulations of Chapter IV, to put the following proposition: one may speak of mathematical probability in the case where the function \(y=f(x)\), representing the causal connection between the random1 cause \(x\) and the action \(y\), is such that to any distribution of the aggregate of values \(x\) there always corresponds approximately one and the same distribution of the corresponding values \(y\). The word “approximately” means that the invariability of the distribution of \(y\) can be expected only for an innumerably large number of unit cases.
These relations appear most clearly in the example of a rotating target. In general, the target proves to be approximately uniformly covered with traces of the bullets that hit it, if a sufficiently large number of shots has been fired at any intervals of time, and the distribution of the density of the traces of hits on the target will be the more uniform the greater the number of shots. Obviously, entirely exceptional deviations are also possible. If, for example, all the intervals of time were commensurable with the period of rotation of the target, then all the traces of hits would be concentrated in definite places, and the remaining places on the target would remain empty. This would be a decisive objection to the possibility of applying our proposition in the formulation given above. But here the consideration comes to our aid that such a distribution of time intervals pred-
constitutes only special and exceptional cases, whose frequency, in relation to all possible cases of the distribution of intervals of time, is vanishingly small. In set theory it is proved, as is known, that—speaking popularly—there exists an infinitely many times greater quantity of irrational numbers than of integers, and, consequently, those intervals of time which are commensurable with the period of revolution constitute an infinitely small part of all possible intervals of time. Therefore, if one chooses various intervals of time at random, it is infinitely improbable that we shall hit precisely upon such intervals of time as are commensurable with the period of revolution; as a result, in general we shall obtain a uniform covering of the target.
Analogous reasoning is applicable in other cases as well. If, for example, the vessel mentioned in Chapter IV had the form of a mathematically exact cube, then it is easy to see that a ball thrown inside it, however numerous its reflections from the walls, could move only in eight definite directions. It is enough, however, for the angle of inclination of the walls to deviate by an arbitrarily small amount from the exact position in order, after a sufficiently long interval of time, to destroy this definiteness in direction and to make all possible directions in space equally probable for the motion of the ball. Consequently, if we do not select a special, mathematically exactly constructed vessel, then for an aggregate of balls their reflection from the walls of the vessel (and also their mutual collisions) will cause a uniform distribution of the directions of motion in space.
In the finest details, similar relations can be traced in the example of two dimensions, in which one may avoid the discontinuity connected with reflection from the walls of the vessel. Let us imagine a point which, under the action of arbitrarily chosen but mutually independent elastic forces \(x\) and \(y\), performs a complex oscillatory motion: \(x = a \sin \alpha t\), \(y = b \sin \beta t\), as happens in acoustics in the representation of Lissajous figures. If we were able to tune the corresponding elastic systems (tuning forks) so that their frequencies were commensurable, then the point would periodically describe a closed curve, not passing through the other parts of the rectangle \(ab\). If we require the commensurability of the ratios to be mathematically exact, then such a ratio would represent a completely exceptional case, which one could not hope to realize with the means at our disposal, since it is infinitely more probable that we shall obtain an irrational ratio of frequencies. In general, therefore, we shall obtain a non-closed curve, which comes infinitely close to any of the points situated inside the rectangle \(ab\). It is easy to find that the relative frequency (equal to the relative interval of time) of passage
points through some place \(x, y\) of the surface element is expressed as
\[ W(x\cdot y)\,dxdy=\frac{1}{\pi^2}\frac{1}{\sqrt{a^2-x^2}\sqrt{b^2-y^2}}\,dxdy. \]
Moreover, this law for the probabilities, as we see, depends not at all on the assumptions concerning the frequencies of the oscillations (or the forces \(x\) and \(y\)).
Let us note also that, according to the equations of oscillation given above, to each point of the plane there corresponds a definite direction of motion and a definite velocity. If, instead of only one point starting from the zero position, we have a whole aggregate of points initially distributed arbitrarily on the plane and moving according to the indicated formulas, then, repeating the argument given above, we arrive at the conclusion that after a sufficiently long interval of time the traces of the initial positions of the points disappear, and as a result one obtains a distribution of points in accordance with the probability law given above and entirely independent of the manner of the initial arrangement of the points.
In a similar way it is easy to see that prolonged mixing in a vessel of two initially separated solutions of coloring substances has as its result the production of a homogeneous mixture; further, it is also clear that an aggregate of gas molecules which were distributed in any manner whatever in a closed space, in general with the passage of time becomes distributed in it as though their positions were completely random (with equal probability for all elements of volume) and entirely independent of the initial positions. This justifies the application of the usual methods of the kinetic theory of gases to the calculation of such quantities in which the average action of a large number of molecules appears.
In all similar phenomena special exceptional cases are theoretically possible, but, owing to their vanishingly small probability, they may in practice be disregarded.
If, in order to forestall reproaches of inaccuracy, we wish to refine the formulation of the definition of probability given above, then in it we must replace the word “always” by the expression “in general,”—i.e. with the exception of a vanishingly small, in percentage terms, number of exceptional cases.
It is possible that the following more precise form should be preferred: a probability law is possible for the action \(y\), depending on a not wholly determinate cause \(x\), in the case when the function \(y=f(x)\) representing the corresponding causal connection possesses the following properties: 1) small changes of \(x\), in general
cause large changes in \(y\); 2) the sets of such groupings of the values \(x\) to which, approximately, one and the same grouping of the values \(y\) corresponds, are immeasurably more numerous than the set of groupings of \(x\) to which a noticeably deviating distribution of the values \(y\) corresponds.
From the mathematical point of view this proposition ought to be formulated more rigorously, but the formulation we have given aims, in a simple and intelligible form, to emphasize the basic idea that interests us in the present exposition. We again draw attention to one circumstance which appears quite clearly both throughout the exposition and in almost all the examples we have given: complete randomness and, corresponding to it, the frequency interpretation of probability evidently constitute an ideal case, to which in reality we have a greater or lesser approximation.
In the practical application of probability theory one most often makes do with a very rough approximation.
VI.
Even more important than the question with which we were chiefly occupied in the preceding chapters, and which has rather a formal character, seems to me the question of the origin of randomness. We approached this question in Chapter V, when we spoke of the shortcomings of our definition of the essence of randomness. An answer to this question can in part be found in the examples cited and in the explanations given for them. The random variability of causes, on which our initial explanation of the law of large numbers was based, becomes self-evident if the matter concerns experiments performed by human hands. In this case, in the final analysis, randomness is reduced to psychophysiological primary causes. But is the application of the concept of probability excluded if we suppose that human actions, with their capricious psychology and physiology, are excluded, and that the circumstances determining a physical phenomenon are established with complete exactness? For the most part an affirmative answer is given to this question, whereas the examples given above show something quite different. If a single ball is released in a perfectly definite manner on a “bounded” Galton board, with a very large number of rows of pins, and then we compile statistics of the places in which the ball passes through each of the rows, then we shall find that all values of the abscissas occur, approximately, equally often. They are equally probable, and this assertion represents an objective, human-independent
fact. In the second example it is theoretically possible to calculate in advance where in the vessel a ball thrown into the vessel in a definite direction will be found; but without further explanations it is clear that, with the passage of time, all possible directions occur with equal frequency, and thus the ball will pass equally often through all parts of the vessel.
In a similar way, in the example of a complex oscillatory motion (Chapter V), we defined with perfect clarity probability as the relative frequency of the movable point’s presence (over a sufficiently long interval of time) at a definite place in the plane, although in all our arguments the initial conditions of the motion played no role whatever.
In an analogous manner, the concept of objective probability may be extended to all similar, not completely determined (“random,” in the sense given above) phenomena, which are characterized by the fact that one and the same character of elementary phenomena is repeated again and again with the passage of time.
As is known, statistical mechanics shows that such cases of motion are by no means rare; on the contrary, according to Poincaré’s theorem, the motions of all “finite” mechanical systems of conservative character belong here. They are all “quasi-periodic” (in special cases exactly periodic), i.e. any initial position is repeated with time to any degree of approximation.
If it is a question of the motion of molecular systems, then the frequency of repetition of identical cases increases extraordinarily owing to the fact that the chemical nature of identical molecules is entirely indifferent for physical phenomena. In order to make still clearer the laws of physical randomness and the concept of an objective probability entirely independent of human knowledge, let us consider, in conclusion, one more phenomenon, which may be regarded as the most perfect type of what we have called “random,” namely, the radioactive decay of an atom. As is known, with the passage of time radium atoms undergo transformation, each emitting an $\alpha$-particle and turning into atoms of emanation; at the same time, in the radium atoms there is not observed the slightest progressive evolution on the model of the aging of an organism. When an arbitrary atom that we are observing undergoes transformation is a matter of absolute chance.
We cannot influence the transformation by any means, and we cannot predict it in advance. The probability that the process of decay will occur precisely in the time interval $dt$ is equally great for young as for old atoms and, consequently, is expressed mathematically by the simple relation: $Wdt=\lambda dt$, where $\lambda$
denotes a constant, whose value we cannot change by any means available to us.
On the basis of what has been said above, one can at once give a model of chance that appears in this case. This will be the vessel we have often mentioned and a little ball thrown into it. We have already noted earlier that for the ball we shall always have an invariable value of the probability of its exit through the opening in the vessel, and it is necessary only to equate the value of this probability to the constant of radioactive decay
\[ \lambda = \frac{\omega C}{4V}. \]
If we had a large number of similar vessels of equal volumes, and if into each a ball were thrown in a different direction, then both phenomena—the exit of the ball from one of the vessels and the emission of an \(\alpha\)-particle by one of the atoms of radium \(^{1}\)—would proceed in exactly the same way.
It goes without saying that I by no means suppose that atoms of radium are in reality constructed like the vessel just mentioned. What is important for us is only the principled possibility of constructing a physical model of ordered chance. The possibility of such a construction proves, at any rate, that the apparent contradiction which we emphasized in the second of the questions posed in Chapter II does not in reality exist, and that chance, in the sense in which this word is used in physics, can always be caused by perfectly precisely determined lawful causes.
Accordingly, a similar kind of chance plays a decisive role in the world of molecules, and there are many phenomena belonging here, such as, for example, the Brownian motion of molecules, in which this can be traced with extraordinary clarity.
If we contrast such cases with phenomena caused by arbitrary intervention of an organism, then one might speak of “molecular” and “physiological” chance. And both these kinds often intertwine in more complex random phenomena.
If, for example, we stretch a wire more and more strongly, or else increase more and more the pressure inside a hollow sphere, then it is customary to say that the place where the rupture will occur and the shape of the fracture depend on chance. The true cause may be small irregularities of thickness and the like, which can indirectly be reduced to physiological chance that occurred in the selection of the corresponding things. But even if, thanks to machine devices and extremely great precision, these irregularities were made arbitrarily small, there would still remain random irregularities in the structure of the material, depend-
\(^{1}\) The number of atoms is taken to be equal to the number of vessels.
…from molecular randomness. However carefully the hollow sphere is cast, such irregularities must inevitably occur. Solidification is based on the formation of centers of crystallization in the supercooled casting. The number and disposition of these centers, apart from regular influences (the rate of cooling, etc.), are determined mainly by molecular randomness.
It is precisely this, therefore, that is responsible for the actually occurring microcrystalline structure of the casting, on which the strength properties depend. The fact that here aggregates of positions of molecules entail such noticeable consequences is based, in turn, on the fact that in the final analysis what is involved is a disturbance of positions of unstable equilibrium.
We shall not investigate in greater detail the questions whether all random phenomena reduce to the two types indicated above, and to what extent, in the end, physiological random events also have their roots in molecular random events. In general it must be repeated once more that our investigation in no way claims to be an exhaustive analysis of all the problems connected with the concept of probability. It seems to us that it will also be a very important result for the philosopher if—even if only in the very limited domain of mathematical physics—it can be shown that the concept of probability, in the ordinary sense of the regular value of the frequency of random phenomena, has a strictly objective meaning. Moreover, it should also be of great significance that one can precisely establish the concept and the origin of randomness, while remaining all the time strictly on the standpoint of determinism. Here the law of large numbers appears not as some mystical principle and not as a purely empirical proposition, but as a simple mathematical consequence of that special form in which, in such cases, causal dependence is represented.
It is perhaps not superfluous to note in conclusion that the calculation of probabilities, in the sense of the conception set forth here, is not a new principle of investigation independent of other methods of knowledge of nature. The calculation of probabilities is a simplified statistical schematization of certain functional relations very often encountered in nature, whose exact investigation, owing to their extraordinary complexity, is very difficult. In the development of modern physics, a characteristic feature of which is the decomposition of physical phenomena into “hidden” partial phenomena, randomness and probability play an important role as a vivid and clarifying auxiliary device. But, if necessary, one can perfectly well do without them, if these schematiza-
methods would be replaced by exact statistical computations1.
The theory we have sketched provides a key to understanding why the use of the concepts of randomness and probability, even when all the details of the functional dependence are unknown, usually yields sufficiently accurate results. It is quite clear that these methods provide an invaluable auxiliary means of investigation in those empirical sciences where exact mathematical study of elementary phenomena is excluded.
-
The essential difference between the kinetic theory of gases (Maxwell, Boltzmann, and others) and statistical mechanics (Gibbs) consists in the fact that the former is based on definite—though highly plausible, yet not rigorously proven—conceptions of randomness and probabilities, whereas the latter (at least in its program, if not entirely in its execution) avoids all such notions and is built upon exact statistical methods. ↩↩↩↩↩↩
-
H. Poincaré. Calcul de probabilité. Paris, 1912, Introduction. ↩