On Causal and Statistical Regularity in Physics[^1]
R. de Misès
Submitted 1930 | SovietRxiv: ru-193001.80580 | Translated from Russian

Abstract

Report at the Fifth Congress of Physicists and Mathematicians in Prague, September 16, 1929.

Full Text

On Causal and Statistical Regularity in Physics1

R. Mises, Berlin

[Problems of the philosophy and methodology of physics occupied a rather prominent place at the 5th Congress of German Physicists and Mathematicians in Prague.

The report by Ph. Frank, “What Do Contemporary Physical Theories Give to the Theory of Knowledge,” was devoted to questions of gnoseology. The report by Mises, the translation of which we are publishing in the present issue of our journal, is devoted to the question of physical regularity.

The development of quantum mechanics has posed in a new way the question of the interrelation between dynamical and statistical regularity. The problem of the statistical method, as a special way of expressing physical regularities, has come to the fore. The extraordinary development of statistical physics has also drawn attention to the problems of probability theory. The inadequacy of the classical foundation of probability theory, based on the subjective conception of chance, became increasingly clear.

In the article printed here, Mises attempts to approach the question of the relation between statistical and dynamical regularity from the standpoint of the conception of probability theory developed by him, based not on a subjective understanding of chance, but on the concept of a collective.

As for Mises’s gnoseological views, in the main he adheres to Ph. Frank, who takes Machist positions.

The unsatisfactory character of Mises’s gnoseological views could not but be reflected in his treatment of the question of lawfulness in physics.

On the cardinal question of causality in physics, Mises, although he does not take the extreme point of view defended by Heisenberg and Dirac, which denies the elementary conditionedness of physical phenomena, nevertheless does not give a sufficiently clear answer to this question, although he does attempt to show that the concept of statistical lawfulness does not exclude the concept of causality.

Also unacceptable is Mises’s reasoning on the relation between physical and philosophical concepts, which in essence repeats Frank’s thought about “school and positive philosophy.”

The healthy kernel of protest against philosophical dogmatics and natural-philosophical fantasies is completely devalued by crude pragmatism in the solution of this problem. If one takes Mises’s point of view and admits that the whole task of philosophy consists only in adapting itself to those solutions to questions that physics provides, then it is completely incomprehensible why a philosophical consideration of the question is needed at all and why philosophical problems are discussed at a congress of physicists. All this happens because Mises does not distinguish between philosophical dogmatics and true philosophy.

The unsatisfactory character and inconsistency of Mises’s philosophical conception lead him to a metaphysical interpretation of the concept of the limit and of the limiting case, which play a large role in the uncertainty relation.

Mises does not pose the question of the relation of an infinite approximation to a given limit and of the limit as something given at once in each individual act of observation. Therefore he does not find the correct solution to the problem of divisibility and indivisibility, which has great significance for the uncertainty relation.

Despite all these shortcomings, Mises’s article is of great interest as an attempt to pose the question of the relation between statistical and dynamical regularity from a new point of view, in contrast to theories that see the sole solution of the question in the expulsion of the concept of causality from physics.
B. G.

In the very most recent period theoretical physics has achieved success in two very different, almost opposite, directions.

On the one hand, Einstein’s general field theory strives to bring classical physics to completion in its strictest form, while at the same time advancing by one step the application of differential equations in the sense of their generality and economy of use. In the theory of the atom, however, there has become established a conception that emphasizes more decisively than ever the shakiness of the classical point of view and the inevitability of a statistical outlook based on the theory of probability—although this was preceded by a brief interval when, in Schrödinger’s wave equation, a return to determinism was seen. In the newest atomic physics, images of a completely new kind play a fundamental role, such as the “cloud of probabilities”; they arouse a distrustful interest among philosophers, who see that one of their most important principles, the law of causality, which had seemed to them the foundation of all knowledge, is in danger. Therefore it would be appropriate to approach the consideration of the question of the relation between the causal and the statistical explanation of nature from the standpoint of that discipline which for more than a century has been concerned with statistical and, consequently, not wholly deterministic phenomena—the theory of probability. I shall preface this with several remarks not determined directly by probability theory.

1. The Law of Causality, Determinism, Atomistics

A few words about the so-called law of causality. Considering the vague and changeable definitions which this principle has received at various periods from leading philosophers, one must come to the conclusion that there is neither danger, nor even possibility, of coming into conflict with it. If in the first edition of the Critique of Pure Reason it says: “Everything that happens (begins to exist) presupposes something after which it follows according to some rule,” or in the second edition this formulation is replaced by the following: “All changes take place according to the law of the connection of cause and effect,” then it proves not too difficult to clothe any regularities whatsoever, including statistical ones, in a form suitable to this “principle.”

Everything depends on what, in the given case, is designated as a change or as an event, and what is understood by cause and effect.

When Galileo discovered by observation the law of inertia, the conception of causality of that time was, of course, contradicted by the fact that prolonged motion can occur of itself, without a cause continuing to act. But after generations of physicists had recognized the law of inertia, with all its consequences, as a suitable basis for the systematic description of the phenomena of motion, the philosophical mode of thought also joined in with this. Now the law of inertia, in the view of all philosophers, not only is in agreement with the principle of causality, but, according to Schopenhauer and many others, represents an inevitable, necessary consequence of the latter: an acting cause is necessary only for a change in the state of velocity, and not of place.

If, from the observations of physicists, it should turn out that only the third derivatives of coordinates with respect to time are determined by independent circumstances, which would then be called forces, then without doubt they would consider them...

...would be a result, or even a necessary consequence, of the principle of causality: the proposition that a body, left to itself, moves along a parabola without the action of external causes.

For the propositions of statistics, too, one can find such forms as correspond to the law of cause and effect. If, in prolonged play with two dice, double six appears on average once in 36 throws, then this phenomenon has as its “cause” the fact that, on average, each side of both dice comes up with equal frequency; and the “cause” of any noticeable deviation from the frequency \(1/36\) would be that one or the other die is faulty. Or if balls falling on Galton’s board show a Gaussian curve, then the “cause” of this consists in the fact that an individual ball, thrown often enough, passes to the right and to the left of a nail with equal frequency. Of course, one may object to this that, for a single event in a game of dice or on Galton’s board, the cause is absent, or cannot be indicated here, or else that deviations in small series cannot be reduced to causes. But after all, in the principle of inertia we have ultimately abandoned the cause of change of place which the naïve conception demanded, and are content only with a cause to explain change of velocity. Thus the following view seems appropriate to me: as soon as physics (or natural science in general, based on continuing observations) fully assimilates the statistical mode of reasoning and inference and recognizes it as a necessary auxiliary means for itself, then after some time no one will any longer have the impression that here some requirement of logic has not been fulfilled or some philosophical need has not been satisfied. In a word: the principle of causality is changeable and submits to what physics requires.

Of course, this does not mean that there is no difference between the two modes of describing nature, i.e., between the physics of differential equations on the one hand and physical [[unclear: continuation cut off]] on the other. But it would simply be too little

there is no sense in supposing that the difference lies in the fact that in one method the principle of causality is fulfilled, while in the other it is violated. It is more correct to call, as is usual, the classical propositions deterministic, and the statistical theory indeterminism. However, considerable difficulties are encountered at once when one attempts to explain, in a general but sufficiently precise form, what exactly constitutes the essence and claims of determinism. What exactly does the famous Laplacean demon accomplish, that official for errands under determinism, the sovereign over all mathematical difficulties? If we confine ourselves at first to questions of the mechanics of a point, or better, of the mechanics of small solid bodies, then here he performs the following task. Given, for all points, the masses and the exact values of the initial position and initial velocity, as well as the forces acting on these points during a known interval of time; from this the final positions and final velocities of all bodies must be computed with any desired degree of accuracy. I do not wish now to enter into a discussion of the insidious meaning of the twice-used word “exact”—I shall return to this later—nor to speak of the difficulty of determining the initial state, for in my view the true limits of the incapacity of the Laplacean demon are set by the proper meaning of the forces that appear on the right-hand side of Newton’s equations.

Laplace, after all, was thinking above all only of the motion of celestial bodies, for which Newtonian mechanics was first and foremost created: here, or in the Galilean problem of a stone falling onto the surface of the earth, no difficulty arises with respect to the forces, which are exhausted by the simple formula of gravitation. But if one takes a stone smaller and smaller and reduces it to the dimensions even of a Brownian particle, then, in order even somewhat to reconcile its fall with Galileo’s laws, or in general to discern any regularity in its motion, we must already attain great perfection in producing the vacuum in which the motion takes place. What will happen if the perfection of the vacuum lags behind the reduction

of the dimensions of the falling body? Then they acquire significance as incidental circumstances: the actions of forces, unpredictably changing in direction and magnitude and originating from the surrounding particles of air; because of them, the motion of the body under investigation will prove to be entirely irregular. No one will say that this motion contradicts Newton’s principle. But in order to apply the Newtonian differential equation, in order to have it as a suitable means for describing the motion, one would have to know all the extremely tangled forces acting on the body; one would have to introduce them into the equation as functions of position, time, etc. Meanwhile, it is obvious that they are known only as an aggregate (in the statistical sense). One may also say that here it is easier to establish the very course of action than the separate instantaneous forces that determine it and are caused by the current of air; that, in a certain sense, the variety of force laws is greater, or in any case no less, than the variety of forms of motion. This is the decisive point: Newtonian mechanics is suitable for the causal explanation of nature only so long as comparatively simple laws of force have as their consequence more complex processes of motion—just as all motions of the heavenly bodies follow from the simple principle of gravitation. To “explain” precisely means: to reduce to something simpler. Without losing its validity, the mechanics of a point ceases to be an effective instrument of deterministic physics as soon as it becomes more difficult to establish the data of the problem than to determine what its solution, i.e. its result, should be. And the same thing that occurs in the fall of small particles occurs in many other cases, where the matter does not by any means always concern very small bodies under investigation. Let us recall the Galton board already mentioned. No one will wish to admit that the motion of the individual balls does not obey Newton’s law of motion. But what does this law say for the instant when a ball is situated directly above one of the nails and it is necessary to decide whether it will run farther to the right or to the left?—that the motion will follow according to

in the direction in which a negligible breath of air or a minimal elastic effect will tip the balance. By this, however, nothing definite has been said about the actual further motion, since we have no means whatever of foreseeing the direction of the next breath of air, and so on.

It may be objected that in the present case it is wrong to confine oneself to the mechanics of a point, and that Laplace’s demon also rules in the mechanics of continuous media, as well as in hydromechanics and in the mechanics of an elastic medium. Here, too, he can broaden the range of application of mechanics by adjoining the appropriate partial differential equations; therefore he will not need to apply complicated force laws. However, this is no salvation. For what corresponds, in the mechanics of continuous bodies, to the forces in point mechanics are boundary conditions. This mechanics yields much, so long as it is possible to derive complicated forms of motion from roughly known boundary conditions. But in this sense only certain special cases of flow, and only certain problems of the mechanics of rigid bodies, can be interpreted in a truly deterministic way. For such delicate details as have to be resolved even in the case of the Galton board, or in the fall of very small particles, the decisive factors are again boundary conditions, such as, for example, the very finest roughnesses, invisible to the eye, of the surfaces bounding the flow. Therefore here, too, there can be no question of reducing the more complex to the simpler. I have already once had occasion to set forth (in a report at a meeting in Jena eight years ago)1—and I do not wish to repeat myself here—that both hydromechanics and the mechanics of rigid bodies lead to relations knowable only statistically as soon as we go beyond the elementary formulation of the questions. I have long since come to the conviction that the decisive problems of microscopic hydrodynamics (such as the problem of turbulence) are solvable only in a statistical sense—perhaps even in connection with new conceptions—

1 Naturwiss., 1922, 25; ZS angew. Math. u. Mech., 1, 425, 1921.

... notions of parallelism between the processes of the continuum and the point. Exactly the same applies to electrical and magnetic phenomena, where similar difficulties are created, on the one hand, by boundary conditions, and on the other—by the so-called material constants in the field equation.

The result of these considerations may be expressed as follows: the deterministic propositions of classical physics can be preserved purely formally or, more precisely, in idea, throughout the entire domain of directly observable phenomena; but in many cases, which may be subdivided from different points of view, they as it were run idle; they lose the character of causal explanation and no longer contribute anything to the cognition, description, or prediction of a phenomenon. The philosophical evaluation of this circumstance will differ depending on what position is taken with regard to the fundamental concepts of theoretical physics. Whoever sees in ponderomotive forces, in densities, in dielectric constants, things existing independently of the task of describing nature, will consider determinism in principle unshakable and only practically excluded. For the one who conceives these images merely as an auxiliary means, introduced through differential equations and, together with them, serving only for better orientation in the world of phenomena, the limits of the applicability of determinism and the limits of determinism itself coincide.

Until now I have spoken exclusively of motions accessible to “direct observation,” of the motion of visible masses, which are the usual objects of classical mechanics. Of course, it had long since been noticed that, in the realm of phenomena occurring with “terrestrial” bodies, it is insufficient to confine oneself only to simple regularities, as in the case of celestial mechanics. And Laplace, when creating his “demon,” was already thinking that the latter, in solving his problem on earth, would encounter quite different difficulties than in the case of celestial bodies, and he gives hints as to the direction in which a way out should be sought. The problems immediately surrounding us are by no means...

are, in the sense of physics, “pure” problems; they must first be reduced to these latter. The means for this is atomistics, the assumption that there exist the smallest particles, whose behavior obeys simple laws; while the less visually evident macroscopic phenomena occur only as a result of the superposition and overlapping of such simplest processes. It is important to remember that the starting point of all rational atomistics, at least since the time of Daniel Bernoulli (1738), was the impossibility of explaining, by the classical deterministic propositions, the entire diversity of terrestrial phenomena. It was impossible not to see that Newtonian mechanics is capable of describing terrestrial phenomena only if it is combined with an extraordinary diversity of force laws; salvation from this was sought in the attempt to construct the small world surrounding us out of bricks, each of which resembled the great celestial world.

There is no doubt that such a conception increased our physical knowledge to an entirely extraordinary degree and promoted its development down to the most recent times. Only one thing it did not do: it did not save determinism, and indeed in general it did not become any support for it. All the actual successes of physical atomistics, beginning with Boltzmann, were achieved solely thanks to the fact that statistical considerations were added to the propositions of classical physics.

Of course, from the differential equations of classical physics one can derive certain assertions concerning a multitude of systems, concerning an aggregate, without all the properties of the individual systems being known or included in the calculations. Such assertions are, for example, the theorems on the motion of the center of gravity and of areas, the energy equation, and so forth. They follow directly from the differential equations valid for individual systems and, of course, belong wholly to deterministic physics. The original idea of atomistics was also that the desired conclusions could be reached by means of considerations of this kind, without addi-

tion of further hypotheses. This is shown, for example, by Boltzmann’s attempt to derive Maxwell’s distribution law of velocities from the laws of collision of elastic spheres. However, it has long been known that such a mode of inference is impossible: without the addition of typically statistical concepts, such as molecular disorder, the uniform distribution of known features, and also purely theoretical-probabilistic reasoning, one cannot arrive at any conclusions other than the previously mentioned propositions of mechanics. The transition from the physics of an individual elementary body—an atom, proton, electron, etc.—to a macroscopic phenomenon is accomplished only through the mediation of statistics.

2. Basic concepts of statistics.

Here I must digress somewhat in order to develop certain basic statistical concepts necessary for understanding the questions under discussion. Everyone knows that wherever statistics is involved, two assumptions must be made. One is that the matter must concern a set, a large number of individual phenomena or individual processes—the application of the word “probable” to a single, non-repeating event lies outside the rational theory of probabilities; the second assumption is what a physicist usually characterizes by the expression “molecular disorder” and what can be spoken of, rather, as the absence of order. Both basic circumstances require greater clarification, and this will best be achieved with a concrete example. Let us therefore recall, first of all, the statistics of the results of a game of dice. The individual event here consists in a single throw of a die, which ends with counting the spots that have landed face up; the totality of the game will be characterized by a sequence of integers lying between 1 and 6. Now the first idealization that we must accept, in order to obtain a concept suitable for theoretical justification, is that we imagine this sequence to be infinite, never-

breaking off. If among the first \(n\) throws the one appeared \(n_1\) times, the two \(n_2\) times, and so on, then we form the quotients \(n_1:n\), \(n_2:n\), and then assume that these quotients, the so-called relative frequencies, tend to definite limits if the number of observations \(n\) becomes larger and larger. In such a case the limits are called, for short, the probabilities of the appearance of the one, the two, and so on, in the series of games under consideration.

The assumption of an indefinitely long series of experiments is necessary for the construction of the theory of probabilities. One must not be troubled by the thought that in very many, even in the majority of practical problems, we are dealing with series of quite definite extent, for example, with a given number of molecules, and the like. If, while remaining with the game of dice, we ask what is the probability that in a sequence of 100 throws there will appear once the sequence of the numbers from 1 to 6 in their natural order, then this will be only a somewhat more complicated case than the one considered earlier. Here the individual event is not a single, but a hundredfold throw of the die, and the object of observation is the appearance or nonappearance of the indicated sequence of numbers in one such hundredfold throw. But we can speak of probability only in the case when we imagine the whole process of hundredfold throwing and observation as capable of being repeated indefinitely long. The assertion that we make on the basis of probability theory can relate only to what will happen under an infinite repetition of the process. That it is nevertheless possible to obtain conclusions of great practical value depends on the special form of the assertion.

The second proposition concerns the absence of order, or molecular disorder. It is self-evident that the latter cuts fundamentally into the structure of statistical theory; to express this precisely is also not difficult if we first again turn to the example of a simple sequence of a game of dice. Let us imagine a player who has staked on the number 6 and, as they say, has a chance or

... view of a gain equal to the probability of the number 6 and equal, consequently, to the limit of the ratio \(n_6:n\), which, as is known, for a fair die is

\[ \frac{1}{6}. \]

If a series of throws revealed some regularity, for example, such that every tenth throw regularly did not give a 6, or that the six never appeared more than three times in a row, etc., then the player would be able to improve his chances by choosing those games on which he stakes. He could, for example, make it a rule to skip every tenth throw of the die, or to skip the one that follows after a threefold occurrence of the six, and so on. The sign of the absence of a rule consists in the fact that no such improvement of chances by means of selection occurs. We give the following formulation of this principle, which is called the principle of the absence of a system of play and which in some respect may stand alongside the principle of the impossibility of a perpetuum mobile: whatever method we use to establish a choice between games, or in a more general sense—between the elements of an aggregate—if only the play is continued sufficiently long, the chance, and consequently also the limit of the relative frequency of the occurrence of the six in the selected series, remains strictly equal to the limit for the entire series as a whole.

Of course, it is impossible to verify the realization of this assumption in any concrete case. It is impossible in principle to continue a series of experiments to infinity, and it is likewise impossible to investigate “all” rules of selection, to see whether they change the chances. An aggregate satisfying both requirements—the existence of limits and with respect to rules of selection—we shall briefly call a “collective.” As is evident, this is not an empirical object, but the same kind of idealized abstract concept as, for example, the concept of a sphere in geometry or of a rigid body in mechanics. In exactly the same way, it is impossible to establish by a finite number of measurements whether a given body is a sphere or not. It is well known that we can construct a reasonable theory only on the basis of such idealized

concepts and that, with respect to representatives of exact natural science, no stricter justification is required here.

It is now no longer difficult to formulate more precisely what the tasks consist of whose solution is accessible to probability theory, or, in other words, to the statistical theory of a series of phenomena. We call a distribution in a collective the totality of the probabilities existing within it; for example, the six fractions which, as limits of relative frequencies, correspond to the faces of an arbitrary die, whether fair or unfair.

Next one can establish that there are collectives, connected with one another, whose distributions determine one another. For example, if the distributions relating to two separate dice are known, then one can compute the probabilities that generally exist for a joint game with both dice. Thus one may say: the exclusive task of probability theory consists in computing, from given distributions within a known initial collective, the distributions in collectives derived from it. This proposition, which completely delineates the nature of probability theory, substantially restricts it in two directions: first, one can draw conclusions about an unknown probability only on the basis of a given probability; and second, the result of the computation can always be only again a probability; one can assert something only about the limit under an unlimited repetition of the experiment. In other branches of exact natural science similar definitions have long been known; it is known that Newtonian mechanics permits one to find the sought final velocities only from given initial velocities,—in probability theory, unfortunately, one still often encounters a formulation of questions that reveals complete ambiguity with regard to the necessary data of the problem and the same ambiguity with regard to the possible conclusions to be obtained.

Another ambiguity, always arising when the discussion concerns statistics or probability, relates to the confusion of the inexactness of a statement with a statistical statement. The starting point of all statistics we may

to see only in the presence of a collective consisting, according to this definition, of an infinite sequence of observations that are exact in themselves. After each game with dice it is quite exact and unambiguous to observe which of the 6 numbers from 1 to 6 has appeared, and only the aggregate of observations gives, to a certain degree, a “blurred” picture with a mean value or mathematical expectation of 3.5 (for a “proper” die) and a “dispersion” which, according to the usual definition, is here equal to 2.92 when the six sides have equal probability. As we shall see, matters are no different in those cases that lie closer to physics. In any case, statistics can begin only when there are unambiguous and definite individual observations at hand. Of course, this does not mean that, for example, the position of a pointer must be read with an arbitrary number of decimal places—we shall return to this question in detail.

Now one must say something more about one definite difficulty—this is the last point of the purely mathematical preparation—about a difficulty that for a long time introduced great obscurity into physical statistics; a difficulty whose overcoming is a necessary condition for a completely irreproachable theory. The question is whether, and in what way, one can subject to the statistical mode of consideration the temporal course of phenomena in a system left to itself. What is meant by this is easily understood from the following example from the theory of Brownian motion. This is a well-known experimental problem, consisting in establishing the number of Brownian particles in an optically isolated region of space over definite intervals of time. For this purpose, by the well-known Poisson formula, one computes the probability \(w_x\) of the number of particles \(x\), for a mean value \(a\), equal to \(a^x e^{-a} : x!\), and compares it with the relative frequency of the occurrence of \(x\). However, on the other hand it is clear that the sequence of these observations certainly does not form a collective, and that, strictly speaking, there can be no question here of probability, because the assumption of the absence of regularity has not been fulfilled. If, for example,

the observed number of particles fluctuates about the mean value 1.5 between 0 and 7, then one should never have, immediately after 1 or 2, the number 6 or 7; by the nature of things, changes do not occur so suddenly. Consequently, if the player wished to bet on the appearance of the number 7, it would be reasonable on his part to do so not before a 5 or 6 appears, and in no case after a 0 or 1; with such a choice the chances of winning would noticeably increase. A similar case, although not so obvious, always occurs when one attempts to apply probability theory to a naturally occurring process in a closed system that is not distorted as a result of observation. For example, although the distribution of velocities in a gas will change as a result of collisions discontinuously in the mathematical sense, nevertheless one can hardly expect arbitrarily large jumps without intermediate stages. In the case of Brownian motion one usually speaks of the “consequences of probability” (Smoluchowski), without posing the problem in its full breadth. For here we encounter nothing other than the ergodic problem, i.e. the question of how far we are entitled to extend to temporal processes, to so-called temporal ensembles, the values of probabilities obtained from combinatorial considerations. Here I shall only briefly point out that this problem is purely mathematical and completely solvable.

It is precisely from a simple game of dice that one can derive a sequence of numbers corresponding to the type of the sequence of particle numbers in Brownian motion. Here I shall write down several numbers that may be encountered in successive throws of one die: 2, 1, 6, 3, 3, 1, 5, 2, 4, 5, 4, 1, 6, 6, etc. They possess the property of the absence of a rule, i.e. by whatever rule we select numbers from the entire sequence, if the sequence is continued sufficiently, then for any numbers in the selected subsequence there will be the very same frequency that exists in the whole sequence. Now from the written sequence I form another sequence by adding together each two adjacent numbers of the first sequence: \(2+1=3\), \(1+6=7\), \(6+3=9\), etc. 6, 4, 6, 7, 6, 9, 9,

5, 7, 12. The numbers of this series lie between 2 and 12, exactly as does the number of points when playing with two dice. But the absence of regularity is violated in a quite definite way: for example, 12 never follows 4, 10 follows three, etc.; and in this the new sequence of numbers differs essentially from the collective formed by observations of the sum of the numbers of two dice.

It is self-evident that the regularities which interest us in the newly formed series of numbers can be derived mathematically from the properties of the original collective. The idea is that in the new numerical series one takes a group of elements, say \(n\) elements, and considers them as an element of some new collective. In this collective, as the attributes whose probabilities must be found, we take the frequency of occurrence in the group of some number—for example, seven. Applying the usual rule of calculation, we shall then find the probability that, for example, in a group of 1000 summed numbers there are exactly 150 sevens, i.e. that the sevens have relative frequency 0.15. A remarkable result of the calculation is the following. For large \(n\) it turns out to be “almost certain” (i.e. it has probability close to 1) that the relative frequency of sevens within the group is very close to the value calculated in playing with two dice as the probability of the number of points seven. Here, in an explicit form, we have the complete mathematical solution of the ergodic problem, namely, the law of transferring combinatorially computed probabilities to temporal processes which, taken by themselves, are not collectives. With some inaccuracy one may say that the law of large numbers can be transferred, to a certain extent, to a temporal process—although not the original concept of probability. The connection with the physical problem will become clearer if one takes into account that the probabilities of change of the differences of two successive numbers are directly given. (The differences lie between 0 and 5, form a collective, and have, respectively, probabilities \(\frac{6}{36}\), \(\frac{10}{36}\), \(\frac{8}{36}\), etc.) In essence the situation is the same for two physical problems, men-

mentioned above: the probabilities of change or transition are given, or are assumed to be given, and from them one must determine the statistics of the course of the process. One very general proposition reads as follows, in the formulation adopted in statistical mechanics: if the instantaneous state of a mechanical system is represented by a point in phase space, and the coordinates are chosen so that the energy surface is represented by a sphere, and if it is assumed that the system undergoes random changes by jumps, so that for every magnitude of the jump, independently of direction, there is a definite probability characteristic of it, then, with probability close to 1, one should expect that after some time the entire energy surface will be covered, with uniform density, by phase points traversed by the system. This proposition, mathematically provable, completely replaces the so-called ergodic hypothesis, which asserts, as is known, that the curve of the system’s trajectory—regarded continuously—covers the entire energy surface. Instead of this definite but highly hypothetical assertion, there appears within the framework of a rational theory of probabilities a provable assertion, which, to be sure, formulates only a statistical datum.

3. Application to Questions of Theoretical Physics

I now turn to applying to the interpretation of physical questions certain consequences that follow from the clarification of the fundamental statistical concepts. I shall place first the proposition that every assertion of physics must express a factual event accessible to observation and to actual experimental verification. In this connection, the decisive property of observations is their repeatability. An experiment must be described in such a way that it can be repeated as often as desired, at different times, in different places, and by different persons. Difficulties are already contained in this concept of repeatability. For, undoubtedly, the complete totality of all physical

of processes is just as unique as the history of mankind. The temperature curve of some locality does not present pure periodicity; no thunderstorm proceeds exactly as the preceding one,—and in the strict sense no experiment can be independent of what is taking place in the more remote surroundings. But the very meaning of the scientific mode of consideration is that, within the single aggregate process, small events limited in time and space are singled out which, with a known approximation, repeat themselves. Only such approximately repeatable events are the subject of physics.

Here the concept of approximation should be developed somewhat. In general, the result of a certain physical observation is obtained by reading the indication of a pointer, and so on. Of course, the reading can be made only with a limited degree of accuracy, and, for example, it can never be decided experimentally whether a certain quantity is expressed by an irrational number or not. One may assert, without any restriction of generality, that the result of an observation is always only an integer, namely the number of the smallest units still accessible to measurement. The numbers may be large if it is possible to count or estimate many points; but since these are finite integers, in form the results of pointer readings do not differ from the indications of the number of pips in throwing dice, or from the number of particles in Brownian motion. What, then, does it mean that a process is approximately repeated? Evidently only this: that if a certain prescribed system of observations is carried out many times, the results will be integers more or less close to one another; only with a very crude choice of units of measurement can it happen that the numbers will not differ.

The circumstance that ever more exact observations are associated with such fluctuations, of course, did not escape the attention of physicists even in the heyday of determinism; however, it was regarded as something secondary, as a contingent circumstance having nothing in common with the foundations of the knowledge of nature. Nevertheless, a mathemati-

mathematical treatment of “inaccuracies,” the so-called theory of errors of Gauss, which we can directly apply. The first and most essential premise of the theory of errors—going much further than the quantitatively specialized assumptions leading to the Gaussian function—may be expressed as follows, using the concepts introduced by us: the results of repeated measurements of a definite physical quantity, with an unchanged experimental setup, form a collective. After what has been said earlier, these words contain two precise assertions. First, it is asserted that, as the series of experiments is continued, the result of each separate measurement possesses a definite limiting frequency; second, that the results follow one another without any regularity; in other words, that an arbitrary selection of them (in the sense explained above) will not change these limiting frequencies. The totality of limiting values corresponding to the separate possible events, or probabilities, forms a distribution within the collective. (As is known, in the theory of errors it is customary to introduce, in one way or another, hypotheses giving the distribution a quite definite, unambiguous form—the form of the Gaussian function; however, in our reasoning we do not consider this.) Only one thing is important: to each distribution there corresponds a certain mean value, computed by known methods, which we also designate as the value of the mathematical expectation. If, moreover, one recalls that by means of the object of measurement and the measuring setup there is determined not a single number, but a collective, an infinite sequence of numbers of a definite kind, then the following assertion becomes clear: the true value of a measurement is called the mathematical expectation of the corresponding collective. This number depends on the object of measurement connected with the measuring instrument.

Now we are already in a position to discuss the claims of the theory of determinism and its distinctive features. From the differential equation of the motion of a point one may, for example, de—

to include that the point under consideration must reach a certain place after one or another time \(t\). This number \(t\) can be computed from the equation with arbitrary precision, always to however many decimal places—if it is assumed that this has been done independently of the numerical values of the physical constants, by means of the appropriate exclusion and choice of units of measurement. The naive perception of determinism consists in ascribing reality to the number calculated in this way, a reality directly provable by a definite experiment. We already know that such unambiguity does not exist, that the prescription of a certain repeatable experiment determines not one number, but an infinite sequence of numbers that are in general different. In order to give any reasonable meaning at all to the concept of a deterministic event, we must employ the concepts: ensemble, distribution, mathematical expectation of a distribution, and then we shall be able to say: a theory is confirmed by experiment if the calculated value and the “true value of the measurement” coincide, i.e. the mathematical expectation determined by the object of measurement and the measuring apparatus and obtained, strictly speaking, only after an infinitely large number of measurements. Of course, in this case there must already have been taken into account what in the theory of errors is called a “systematic” error—by means of the physical theory of the process, including also the measuring process.

Convinced adherents of the idea of causality will probably not agree with such an interpretation of the role of deterministic events. In fact, more or less consciously, they add to the state of affairs just set forth the following arbitrary assumption, which says approximately this: to every theoretical event there corresponds an infinite sequence of different experimental arrangements of increasing precision, such that, when measured in constant units, the dispersions of the corresponding distributions become smaller and smaller and ultimately tend to zero. It is clear that this assumption is a far-reaching extrapolation beyond the limits

actually accessible observation; however, this assumption must be rejected, not only because everywhere that we have spoken of limiting values, we in a similar fashion pass beyond the domain of the observable. But the assumption of an unlimited increase in precision directly contradicts every atomic hypothesis. If there does not exist unlimited divisibility of all physical quantities, then the fineness of measurement also cannot be arbitrarily increased. In the macroscopic domain one could at most say that the scatter of observations can in principle be reduced to atomic dimensions. But what sense is there in speaking of the exact fulfillment of equations, of an exact agreement between theory and observation, when a system of differential equations is applied to atoms, electrons, and so forth? One would have to imagine that there exist measuring instruments whose sensitivity surpasses even atomic dimensions, and, obviously, all physical content would thereby be rejected. Heisenberg has recently been especially insistent on the need to describe atomic experiment concretely, thereby reviving the discussion of causality and statistics.

I shall dwell for a moment longer in the macroscopic domain, in order to establish the connection between what has just been said and earlier remarks on determinism. We have established that deterministic mechanics of the point is fully applicable, for example, to the motion of celestial bodies, but that in other cases—for example, in the case of the Galton board or the motion of small particles in air—it proves powerless, if not formally, then in essence. This assertion, or rather its first part, must now be limited, taking into account that a theory can be judged only in connection with the observations that serve to test it. Only so long as the measuring instrument is so crude that it can serve as a suitable object for deterministic physics could it lead to an unambiguous event, and one may speak of the applicability of determinism. However, as soon as the measuring instrument is improved, it will encounter the same difficulties that we earlier established for

of the motion of small particles, and the result of observation will be an ensemble; the assertion of the original deterministic theory will have only a statistical meaning, as a statement about the mathematical expectation of certain ensembles. Thus there emerges a double limitation on the applicability of the classical physics of differential equations: first, it is actually applicable only to a definite domain of macroscopic events, and second, within this domain it has a genuinely deterministic character only when limited to sufficiently coarse measurements. Returning to the image of Laplace’s demon, we would have to say: he can solve only a small part of his problem, and the result should then be regarded only as an approximation; if the problem goes beyond the limits of a certain specified domain, it becomes meaningless; if one strives to increase the accuracy, then the result of the calculations will yield only a mathematical expectation within an ensemble.

Thus the statistical method of consideration proves to be inevitable in two respects. However, it is necessary to emphasize: we never speak of a contradiction between a series of observations and classical theory; we are never compelled to say that in some particular process a proposition of deterministic physics is violated. Such an assumption was made only once, in recent years, in the well-known work of Bohr, Kramers, and Slater, and was soon again rejected because of its lack of foundation. The systematic theory, which I have been following for more than ten years, despite all the freedom it has granted to indeterminism, has never known any other form of untenability of deterministic physics than the fact that in certain known cases it acts in vain, i.e. becomes insufficient for solving problems.

The recent development of microphysics has brought about a new turn in the general appraisal of the statistical method of consideration. Above I said that all atomistics arose from the effort to save, at least in a small measure, determinism, which was slipping away from...

the domain of the large. This striving was based on the false assumption that from individual, deterministically described events one can derive, in their essential features, the laws governing their totality, without constructing any hypotheses. However, if one does not treat the word “probability” too lightly, it must be admitted that this is impossible, if only because probabilities are calculated exclusively from given (or assumed) probabilities: there is not a single smallest proposition of the kinetic theory of gases which would follow from classical physics alone, without assumptions of a statistical character. For considerations of another kind, and not for these, attempts to give determinism a foundation in atomic physics have now been abandoned: it has been acknowledged that elementary processes themselves are wholly incapable of causal description. This is a direct consequence of the requirement that theory be considered only in connection with experiment, which serves to test it.

Modern radiation physics proceeds from the fact that to every process of radiation there is assigned a definite conservative mechanical system; for example, to monochromatic radiation—a point elastically connected with a center (Planck’s oscillator), to hydrogen radiation—a point in a Coulomb force field (Bohr’s atomic model), and so on. From the Hamiltonian function of the given mechanical system there are derived, on the one hand, the frequencies of oscillations, and on the other hand the amplitudes of oscillations, i.e. the distributions of intensities,—according to known formal-mathematical prescriptions. Up to this point everything proceeds in the sense of the physics of differential equations. Now, however, one must take into account that this mechanical system is present in the radiation process not singly, but in billions of copies. Born was the first to establish here a reasonable connection with his explanation: the intensity calculated from Schrödinger’s wave equation, as a function of the state parameter, gives the probability that one of the specimens of the system will be found at the corresponding point of phase space. This proposition retains a definite content even if one abandons the idea of a multitude of individual particles. For ve-

the probability of a state, multiplied by the total number of particles, gives nothing other than the mathematical expectation, or mean value, of the particles located at the given place. Thus, according to Born’s interpretation, Schrödinger’s differential equation gives only the mathematical expectation of a physical event, only the mean value of what is expected at a given place. This is the same interpretation that we ought to have given to every event of the physics of differential equations in the macroscopic domain, when we limited ourselves to statements only about what is actually observed—as was done above. And the justification given by Heisenberg for this intrusion of indeterminism into atomic physics coincides exactly with what was said above about nondeterminacy in the macroscopic domain.

As is well known, Heisenberg shows that an experiment is inconceivable in which both the position and the velocity of a very small body could be measured exactly at the same time. More precisely, since the product of both “uncertainties” is a finite number of the order of magnitude of Planck’s quantum of action, neither of the two quantities can be determined quite exactly. As the accuracy of one measurement increases, the uncertainty of the other increases. But uncertainty should not be confused with statistics, as I have already indicated. Likewise, what is essential in the new point of view is not that the process of observation affects the observed object. Such an effect is often encountered, for example when the velocity of a flow is measured with a Pitot tube or the strength of a field is investigated with a test body. Such influences constitute “systematic errors” of observation, and they can be eliminated by corrections, i.e., one must formulate differential equations that are in principle suitable for the aggregate system consisting of the object of observation and the measuring instruments. The matter is fundamentally different when, together with Heisenberg, we consider an atomic particle in a microscope for γ-rays, in order to determine its position as accurately as possible, i.e., when we subject it to the action of very short radiation. The high-energy light quanta now falling on the particle and affecting the magnitude of its momentum in the sense

the Compton effect, are subject to an undetermined process, knowable only statistically; although collisions do not directly form an ensemble, nevertheless (like the number of particles of Brownian motion) they represent a temporal event of the ergodic type considered above, with a limited absence of regularity. In any case, the values of momentum obtained under the influence of collisions with light quanta indirectly lead to an ensemble, and the so-called uncertainty is the measure of dispersion of the corresponding distribution. This may also be described as follows: a concrete process of measurement in microphysics is not an elementary process, but a statistical event. Thus statistics enters into atomic measurements inevitably, just as it does into macroscopic ones.

In conclusion, one may draw the following inference from these hastily sketched thoughts: strict determinism, usually ascribed to the classical physics of differential equations, is only apparent; it does not hold if the theory is accepted in principle only in connection with the experiment that serves to test it, i.e., if one restricts oneself only to what is sensuously perceived, or “in principle” observable. In the macroscopic realm, the indeterminate is hidden partly in the objects of observation, and partly penetrates through the measurement processes; every microphysics, however, carries a statistical element within it, since only this element ensures the transition to mass phenomena, and every measurement is precisely such a phenomenon.

I believe that such a conception, based on a clarification of the basic concepts of statistics, will help to smooth out the rift between causality and statistics that permeates modern physics.

  1. Report at the Fifth Congress of Physicists and Mathematicians in Prague, September 16, 1929. Naturwissenschaften, 1930. Trans. by E. L. Starokodamskaya. 

Submission history

On Causal and Statistical Regularity in Physics[^1]