Full Text
STATISTICAL THEORIES OF MATTER, RADIATION, AND ELECTRICITY1
Karl Darrow, New York
INTRODUCTION
The chief subject of the present article is the advances of the atomic theory in two areas that until now had lain outside its domain. As is well known, atomic theory seeks to explain the observed properties of material bodies of microscopic dimensions by regarding the latter as an aggregate of the smallest particles, each of which is endowed with only a few of the simplest qualities. Any gas possesses, for example, elasticity, viscosity, entropy, and temperature. Theoretically speaking, all these properties might also be ascribed to each of its atoms. Modern atomic theory of gases, however, approaches the question in an entirely different way. It seeks to explain the four properties mentioned above (and many others) as characteristic features of an aggregate of an enormous number of identical particles, each of which individually possesses none of these properties and is characterized exclusively by its position, velocity, mass (sometimes also by its moment of inertia), and capacity to undergo elastic collisions with other particles. Such an approach has, in general, fully justified itself. We do not ascribe viscosity, entropy, and temperature to each individual atom, but, on the contrary, derive
(qualitatively and quantitatively) their existence from the properties of an aggregate of atoms. The theory on the basis of which these conclusions are drawn is called statistical. It is based on certain fundamental propositions which, in their original form, constitute what is called classical statistics. The successes of classical statistics are one of the testimonies to the validity of the corpuscular theory of matter. Earlier—before the brilliant experiments with individual atoms appeared, thanks to which atomic theory gained universal recognition—this testimony was almost the only one.
Radiation in certain respects resembles a gas. For example, radiation enclosed in a vessel whose walls are maintained at a definite temperature possesses, like a gas, entropy, temperature, and elasticity. It seems natural to try to explain these properties by proceeding from ideas analogous to the atomic theory of gases, i.e. by considering radiation as an aggregate of an enormous number of particles—photons or light corpuscles. Now this idea seems quite natural; but at the time when everyone regarded light as something wave-like, it would probably have seemed extremely strange. Even after Einstein destroyed these old conceptions, it took about twenty years for the concept of a “light gas” to grow out of quantum theory. The historical development here proceeded not as in the atomic theory of matter, but in the reverse direction: the existence of light corpuscles was proved (in experiments on the Compton effect and the photoelectric effect) long before the statistical theory of these corpuscles had developed. As we now know, the chief difficulty here consisted in the fact that, even starting from the concept of light particles and assigning to them the correct values of energy and momentum, one could not use for them the same statistics that had given such good results for material atoms. Bose showed how the foundations of statistics had to be changed in order to construct a consistent atomic theory of radiation in thermal equilibrium with the walls of a vessel.
The second of the new achievements of atomic theory amounts, in part, to the revival of a theory first proposed 30 years ago, according to which negative electricity in metals is a collection of moving particles that, in their motion, collide with the atoms of the metal. In its time the development of this theory was, of course, carried out by the methods of classical statistics. Owing to a number of serious discrepancies with experiment, it was subsequently abandoned and only recently revived again by Pauli and Sommerfeld on the basis of a modified statistics. This modified statistics, first given by Fermi and, independently of him, by Dirac, does not coincide with the radiation statistics introduced by Bose, although it is very reminiscent of it. It cannot yet be said that the renewed theory of the “electron gas” gives a satisfactory explanation of all the numerous phenomena accompanying the motion of electricity and heat in metals and at their boundary surfaces. However, its initial successes are so great that further development will apparently proceed not along the path of abandoning it (as had seemed almost inevitable before the introduction of the new statistics), but along the path of correcting certain details.
Do these new ideas affect the atomic theory of material gases? It turns out that, in application to material gases, all three kinds of statistics—the classical one and the two new ones—lead practically to identical results. Discrepancies appear only at very low temperatures and very high densities; and even here it is difficult to say in whose favor the experimental data speak. But let us even suppose that these data turn out to be on the side of one of the new statistics. In that case it would only be necessary to abandon one of the fundamental propositions of the kinetic theory of gases and replace it with another; the whole superstructure of formulas and equations in which kinetic theory comes into contact with experiment would remain practically untouched. In theoretical physics this is, fortunately, much easier to do than in architecture.
Recently, by atomic theory one usually
...is understood as the theory of the structure of the atom. In reality, however, this latter theory encompasses a domain into which its predecessor did not enter. No one ever thought that all the properties of a gas could be explained as statistical properties of an aggregate of particles. Some of these properties the original atomic theory ascribed to individual atoms, thereby in essence declining to explain them. Among these were all phenomena connected with spectra. At the point where statistical theory stopped, the work of the creators of atomic models began. In particular, Bohr proposed a model of the hydrogen atom capable, at least within considerable limits, of explaining the Balmer series and the rest of the line spectrum of “atomic hydrogen.” Following Rutherford, he constructed this model out of two particles. Thus he and his followers began to develop an atomic theory of the atom, standing one step deeper—or, rather, higher—than the atomic theory of matter, which had served as its starting point.
How, then, does the new “atomic theory of the atom” differ from its predecessor? The basic difference in the method and aims of the two theories comes down to the fact that in the first we are dealing with a comparatively small number, while in the second—with an enormous number of elementary particles.
Bohr’s model of the hydrogen atom contains only two particles; models of other atoms contain no more than a few dozen particles. In a system consisting of two particles, one can determine the position and velocity of each particle with any desired degree of accuracy and completely describe its motion in its orbit. Even in systems with a nucleus and several dozen electrons one can attain a certain accuracy; it is enough to recall the diagrams of electron orbits in heavy atoms that were so current some 6–7 years ago. Perhaps such definiteness was unjustified, but in any case it was possible. The situation is quite different in statistical theory.
A model of a cubic centimeter of gas under ordinary conditions of temperature and pressure contains about \(10^{23}\) particles.
A single mention of this enormous figure is enough to understand how pointless it is to ascribe to each particle a definite initial position and velocity. Merely to write down the initial conditions, not to speak of drawing conclusions from them, would take longer than the lifetime of the entire human race.
At first glance this difficulty seems insurmountable. But in reality the matter is by no means so terrible. For the application of the statistical method it is quite unnecessary to know exactly the position and velocity of every particle. It is enough to know only the function which indicates what fraction of the total number of particles is contained in a given small (but not too small!) volume, and what fraction of the total number of particles has momenta lying in a given small (but not too small!) interval. These definitions are considerably more modest and vague, but are quite sufficient for our purpose—that is, for explaining the origin of entropy, temperature, viscosity, thermal conductivity, diffusion, and so on.
Moreover, if atomic theory could give a more definite picture, i.e. indicate with absolute accuracy the positions and velocities of all atoms, then the concepts of entropy and temperature would lose all meaning altogether. The existence of these quantities in our theory is due exclusively to the vagueness of the description. They seem to us physical realities only because our sense organs and measuring instruments are not capable of reacting directly to the actions of individual atoms. Therefore the brain, too, imposes upon itself a certain vagueness, refusing an overly exact description of the motion of particles; the fact that, owing to the enormous number of particles, this refusal is made “sincere” changes nothing in principle. Exact information about every atom is not only impossible, but also unnecessary, even undesirable, to obtain. Let us recall Aesop’s fable of the fox and the inaccessible grapes: in the present case the grapes really would set one’s teeth on edge.
It is possible, however, that the grapes do not exist at all. One of the most striking assertions of theoretical
of physics in recent years is the idea that even in an atom with several particles, even when speaking of a single particle, it is impossible to specify exactly its position and velocity. Applying statistical methods to an aggregate of a large number of particles, we indicate only those approximate limits within which their coordinates and momenta lie. It is possible that this uncertainty and diffuseness is rooted in the very nature of things. In that case, the distinction I have emphasized between the theories of the structure of the atom and the statistical theory of matter and radiation would turn out, in essence, to be the distinction between an incorrect approach to a known domain of phenomena and a correct approach to all phenomena.
The ultimate aim of any statistical theory is the determination of the above-mentioned function, the so-called distribution function. As we have already said, this function shows for what fraction of the total number of particles we assume the coordinates and momenta to be contained in given small intervals. The task of statistical theory is to derive this assumption from certain more fundamental assumptions. Of course, one may say that this derivation amounts merely to choosing fundamental assumptions according to one’s taste. In any case, from whatever point of view we proceed, only the distribution function itself is subjected to experimental verification: indirectly, because it gives numerical values for thermal conductivity, viscosity, specific heat; and directly, since in recent times direct methods for studying it have been found.
The distribution function is usually introduced by an equation of the form:
\[ dN=f(x,\ y,\ z,\ p_x,\ p_y,\ p_z)\cdot dx\,dy\,dz\,dp_x\,dp_y\,dp_z. \tag{1} \]
This equation refers to an aggregate of a definite number \(N\) of particles occupying a known bounded volume: for example, to a gas in a vessel, to radiation in a cavity, to electrons in a wire. \(dN\) represents the number of particles whose coordinates lie in the interval \(dx\) in \(x\), in the inter-
in an interval \(dy\) in \(y\), in an interval \(dz\) in \(z\), and the components of momentum—in an interval \(dp_x\) in \(p_x\), in an interval \(dp_y\) in \(p_y\), in an interval \(dp_z\) in \(p_z\). The expression “in an interval \(dx\) in \(x\)” means “in the interval between \(x\) and \(x+dx\).” The function \(f\) is the distribution function with respect to the given variables, i.e. with respect to the coordinates of the particles referred to the axes of a Cartesian system in ordinary three-dimensional space and with respect to the momenta referred to the same axes. In the present case it is more convenient to deal with the components of momentum, and not of velocity, since 1) the components of momentum satisfy canonical equations, 2) passing to the study of ensembles of photons, we shall see that momenta play there the same role as in ensembles of material atoms, whereas the velocities of all photons are identical. Knowing the form of the distribution function in one set of variables, we can easily express it also in other variables.1 Further, in order to pass from the distribution function in all independent variables to the distribution function in some of them, one must, as is known, integrate it over the entire range of variation of the remaining variables. For example, in order, in the case represented by equation (1), to find the distribution with respect to the component \(f_z\), one must integrate \(f\) over the entire range of variation of the first five variables.
The product \(dx\,dy\,dz\) is an element of volume in ordinary, or coordinate, space; the product \(dp_x\,dp_y\,dp_z\) is an element of volume in momentum space, in which each particle is represented by a point with coordinates equal to the components of the particle’s momentum; finally, the product \(dx\,dy\,dz\,dp_x\,dp_y\,dp_z\) is an element of volume in the so-called phase space. The function \(f\) expresses the distribution of our collection of particles over this six-dimensional phase space. In some cases—for example.
\[ F(v_1,v_2,\ldots)=f(u_1,u_2,\ldots)\frac{\partial(u_1,u_2,\ldots)}{\partial(v_1,v_2,\ldots)} \]
considering an ensemble of electrons in a non-uniformly heated metal or an ensemble of oscillators—one must continually deal with all 6 dimensions. In other cases, when photons are in question, and likewise atoms or electrons in a volume where the temperature and potential are everywhere the same, one may assume that the distribution over coordinates is uniform (i.e., that \(f\) does not depend on \(x, y, z\)) and omit it altogether from consideration. Then the problem reduces to finding the distribution in the three-dimensional space of momenta. But even in these simplest cases it would, of course, be more consistent always to operate in phase space. Unfortunately, the human brain is so constructed that, however much we may reason about a space with six or with six trillion dimensions, we cannot imagine any space other than three-dimensional.
In an equation of type (1), the product of differentials appearing on the right-hand side must be neither too large nor too small. If this volume element is so large that \(f\) changes appreciably from one of its points to another, then the factor multiplying it—which by definition is the mean value of \(f\) in the given element—must be computed by methods of integral calculus. On the other hand, if this element is so small that it contains only a few molecules, then the product of \(f\) by its volume may be several times greater or smaller than the number of molecules actually contained in it. This is easiest to see from the following example: suppose that space is divided into elements whose number is 10 times greater than the total number of particles. Then in at least \(9/10\) of all elements the number of particles will be zero, with \(f\) greater than zero; in the remaining elements the number of particles will surely be greater than the product of \(f\) by the volume of the element. Dividing space into such small cells makes the description too exact and therefore unsuitable for our purposes.
Already here it is appropriate to emphasize that the product of differentials appearing in equations of type (1) is not
one must not confuse them with those elementary cells of phase space which will be discussed below and which play such a great role in both the old and the new statistics. Each volume element sufficiently large to be used in equation (1) contains a very large number of such cells. In other words, the subdivision of phase space into elementary cells, which will be discussed below, is too fine to be used in connection with the distribution function. Failure to understand this fact often leads to great confusion.1
Speaking of the distribution function, we tacitly assume that there exists some stable, constant, unchanging distribution of atoms in a gas, of photons in a closed cavity, of electrons in a wire. This assumption requires justification, since it is by no means self-evident. On the contrary, at first glance it seems that the more numerous the particles, the more sharply the distribution should change from one moment to the next. For example, an assembly of \(10^{20}\) particles appears at first sight to be such a continuous chaos that it is meaningless to seek in it any stable distribution. Experience, however, shows the opposite. The density of a gas in a vessel remains constant and uniform all the time; the gas does not accumulate in one corner and does not suddenly become hot at one end and cold at the other. In radiation enclosed in a closed vessel with heated walls, the intensity of any portion of the spectrum remains constant so long as the temperature of the walls is constant. The distribution of velocities in a stream of electrons emitted by an incandescent filament also, apparently, does not change. Moreover, if by some artificial method one makes the density or temperature of a gas different at different points and then leaves it to itself,
then all nonuniformities very quickly even out and the gas will again return to a stable state.
As is well known, such a transition of a gas from an unstable state to a normal, stable one is characterized by the increase of a certain function \(S\), which is called entropy. In some of the simplest cases this change in entropy can be calculated. In a stable state the entropy attains a maximum value, which can be determined (up to an additive constant) by knowing the temperature and other constants characterizing the gas. It is further known that the entropy \(S\) of a gas and its energy \(E\) are related to the absolute temperature by the equation
\[ \frac{dS}{dE}=\frac{1}{T}, \tag{2} \]
which, properly speaking, also serves as the definition of absolute temperature.
Thus, if in some way we obtain a statistical value of the entropy, i.e., knowing the distribution function we can at any moment calculate the entropy \(S\), then the stable distribution must be recognized as that for which \(S\) reaches its greatest value (for the given number of particles and the given total energy). Here there can no longer be any choice. Conversely, any hypothesis made concerning the form of the function \(S\) can be tested experimentally by whether the distribution corresponding to the maximum \(S\) is indeed stable. As will be shown below, the distribution function corresponding to the maximum of the entropy includes the derivative \(\frac{dS}{dE}\), i.e. the absolute temperature. Thus the temperature appears in the course of the derivation itself, and not by virtue of some special assumption.
This method, first proposed by Boltzmann and subsequently developed by Planck, plays a very substantial role. One choice of the function \(S\) leads to the classical Maxwell–Boltzmann distribution; another—to the Bose or Fermi distribution (the distinction between the latter two appears in another point).
Each of these functions has a logarithmic form, i.e. is proportional to some other function which is called the probability. When theoretical physics speaks of probability, this means that all hope has been lost of finding a causal connection among the observed phenomena. This is the weak side of Boltzmann’s method. According to Boltzmann, the “distribution corresponding to the maximum of entropy” is the “most probable distribution”; we can even find the numerical value of its “probability,” which so exceeds the numerical value of the probabilities of other distributions that the given distribution may safely be regarded as “normal.” But it remains completely unproved why one “most probable” distribution must always be followed by another, and why in place of any “improbable” distribution there must always (or at least in the majority of cases) come a more probable one. Furthermore, the very mechanism of the transition from one distribution to another remains entirely unexplained, i.e. the mechanism of mutual collisions in which particles exchange energies and velocities, as a result of which the ensemble as a whole approaches a stable state. True, there exist such statistical methods (they will be mentioned briefly in the last paragraph) which take all these aspects into account. But in the method by which the distribution laws of Maxwell–Boltzmann, Bose, and Fermi–Dirac will be derived below, the concept of causality is absent.
We shall consider the application of these laws to ensembles of freely flying particles, first in the case of the absence of an external field, and then in the case of the presence of an external field (electrostatic or gravitational) derivable from a potential.
The reader will probably be surprised that we say so little about oscillators, whose statistics, developed by Planck, served as the first modification of classical statistics and as the starting point of the whole of quantum theory, i.e. the most significant achievement of physics
over the last 25 years. The history of this period of development is extremely curious, but I can dwell only on some of its main points.
Planck oscillators served two purposes; they: 1) helped Planck derive the law of the steady distribution of radiant energy in a closed cavity (Planck regarded radiation as a system of waves in thermal equilibrium with myriads of oscillators forming the walls of the cavity), and 2) enabled a number of investigators to develop the theory of the specific heats of solids. Bose statistics made the first purpose entirely unnecessary, since, using these statistics and regarding radiation as a system of corpuscles, one can derive the Planck distribution without introducing oscillators at all. As for the second, as early as 1912 (now this seems especially remarkable!) Debye, instead of the usual picture of a solid as a lattice consisting of vibrating atoms, introduced a new conception of a solid as a system of standing waves in a continuous medium. Debye’s theory, of course, did not deny the existence of atoms, but in its statistical calculations the place of each atom was taken by a group of standing waves. We have now become accustomed to the idea that in certain theoretical arguments, instead of a free electron, a quantum, or even an atom enclosed in a certain bounded volume, one may consider a group of standing waves in that volume. Thus it turns out that a free particle, an oscillator, and a group of standing waves are closely connected with one another; it is possible that they merely express different points of view on one and the same object. Therefore, although in this article I shall use almost exclusively strictly corpuscular terminology, it is possible that all the results set forth in it can be translated into the language of oscillators or into the language of waves.
Certain assumptions must be made concerning the particles themselves. Each of them is characterized by its position, momentum, and energy. Momentum may be expressed in terms of mass and velocity, but it is usually more convenient to have
the matter directly with ourselves. As is known, the basic difference between photons, on the one hand, and electrons and atoms, on the other—i.e., between those particles from which we shall construct radiation, and those from which we shall construct gases and electricity in metals—reduces to different relations between momentum and energy.
Experiments with material bodies of macroscopic dimensions have led to the establishment of the well-known equations relating the kinetic energy \(K\) and the momentum \(p\) to the mass \(m\) and the velocity \(v\):
\[ K=\frac{1}{2}mv^{2}; \qquad p=mv, \tag{3} \]
which are usually assumed to be valid also for the smallest particles of matter and electricity. On the other hand, at the time when the corpuscular theory of light was only just gaining acceptance—let us recall that the wave theory held unlimited sway even after Planck had arrived at quanta on the basis of oscillator statistics—Einstein proposed, at different times (1905 and 1917), the following two formulas expressing the relation of the energy and momentum of a photon to its wavelength:
\[ E=\frac{hc}{\lambda}, \qquad p=\frac{h}{\lambda}. \tag{4} \]
From the historical point of view it is interesting that Einstein arrived at the latter formula on the basis of statistical investigations of thermal equilibrium between photons and atoms. The first of formulas (4) received experimental confirmation in the photoelectric effect and in a number of other phenomena; the second—in the Compton effect. These experiments are sufficiently widely known, so we shall not dwell on them here.
From equations (3) it follows that, for particles of matter and electricity,
\[ E=\frac{1}{2m}p^{2}=\frac{1}{2m}(p_x^{2}+p_y^{2}+p_z^{2}). \tag{5} \]
For light particles, however, equations (4) give
\[ E=pc. \tag{6} \]
The difference between a light gas and an electronic or material gas is to a certain extent explained by the different form of formulas (5) and (6). But its basis, as we shall see below, is rooted not in the latter, but in the difference between statistical methods.
Classical Statistics
In what follows we shall consider material gases, radiation in a closed cavity, and negative electricity in metals as collections of particles, each of which has a definite position and momentum. Such a collection can be represented by a set of points in ordinary space, along whose axes the coordinates of the particles \(x, y, z\) are laid off, and by a set of points in momentum space, along whose axes the momenta of the particles \(p_x, p_y, p_z\) are laid off.
The simplest illustration of the method of classical statistics is furnished by the problem of the most probable distribution of particles in ordinary space. In this case the solution of the problem can be foreseen in advance. Indeed, it is clear that if in a given bounded volume there is nothing that would distinguish one part of it from another (i.e. there is no variable external field in space), then the particles must tend to be distributed uniformly throughout the entire volume. Statistical theory must lead precisely to such a result. The uniform distribution must turn out to be the most probable. How, then, must “probability” be defined in order that it should prove greatest in the case of a uniform distribution?
But first of all, what is a uniform distribution? Let us divide space—mentally, of course—into equal cells. A uniform distribution is one in which the number of particles in each cell is on the average the same. In order for such a definition to make sense, the cells must evidently have finite, not too small, dimensions. For example, their linear dimensions must be no smaller than the mean distance between particles; otherwise, “uniform dis-
“determination” would be impossible. It is also just as inexpedient to subdivide space too finely as it is to view the picture under a microscope. The quantity that we wish to determine loses its meaning if the accuracy is too great. Both for this reason and in order to make legitimate the familiar mathematical approximations, it is usually assumed that each cell contains a large number of particles.
Let \(N\) be the total number of particles; \(m\) the number of cells into which the volume \(V\) has been subdivided; finally, \(N_i\) the number of particles in the \(i\)-th cell. The given distribution is characterized by the numbers \(N_1, N_2, \ldots N_i, \ldots N_m\).
At the foundation of classical statistics lies the fact that, if each particle is endowed with some distinguishing mark—for example, if it is assigned a definite number—then the given distribution can be obtained in various ways. Indeed, suppose that the particles are arranged in a certain order according to the distribution \(N_1, N_2, \ldots N_m\). By rearranging the particles in any way from one cell to another, while keeping the total number of particles in each cell unchanged, we obviously do not disturb this distribution. The total number of such arrangements, i.e. the number of combinations of \(N\) into \(N_1, N_2, \ldots N_m\), is given by the well-known formula:¹
\[ W=\frac{N!}{N_1!N_2!\ldots N_m!}. \tag{7} \]
¹ Let us imagine that we have \(m\) boxes and an urn with \(N\) numbered, but otherwise perfectly identical, balls, which are drawn in succession and placed in the boxes according to the following rules of the game: the first \(N_1\) balls that come to hand are placed in the 1st box; the next \(N_2\) balls in the 2nd box, and so on. Having distributed all the balls in this way, we determine the contents of each box and then repeat the same process indefinitely. Obviously, the sequence in which the balls appear from the urn may be different—the number of such different sequences is \(N!\). If, as the result of two trials, the contents of the boxes turn out to be different, then it is clear that in the first case the balls came out in one order, and in the second case in another order. However, the converse assertion is not true. Take some definite sequence of balls; it can be shown that
This function attains its minimal value, equal to unity, in the case when all particles are gathered in one cell, i.e., under the most nonuniform of all possible distributions. Its maximal value, as we shall now show, is attained for a uniform distribution.1
Instead of the function \(W\) itself, we shall consider its logarithm. It is obvious that \(\lg W\) (which in what follows will be identified with entropy) reaches its maximum simultaneously with \(W\).
We have:
\[ \lg W=\lg N!-\sum_i \lg N_i! \tag{8} \]
Let us use Stirling’s formula for factorials of large numbers—one of the formulas most frequently used in mathematical statistics. This formula states:
\[ x!=(2\pi x)^{\frac{1}{2}}\left(\frac{x}{e}\right)^x \]
\[ \lg x!=x\lg x-x+\frac{1}{2}\lg(2\pi x). \tag{9} \]
The first two terms of the last expression serve as a sufficient approximation already for \(x=10\). In this approximation we have
\[ \lg W=\mathrm{const}-\sum N_i\lg N_i. \tag{10} \]
There exist \(Q-1=(N_1!N_2!\cdots N_m!-1)\) other sequences that lead to exactly the same result. Indeed, there are \(N_1!\) ways in which the first \(N_1\) balls can leave the urn without losing their place in the first box; \(N_2!\) ways in which the next \(N_2\) balls can leave without losing their place in the second box, and so on. Thus, among the \(N!\) sequences, each \(Q\) lead to the same result, so that in all we have \(\dfrac{N!}{Q}\) different results.
Let \(W^0\) be the value of \(W\) for some given distribution \(N_1^0, N_2^0 \ldots N_m^0\); and let \(W = W^0 + \delta W\) be its value for a slightly different distribution \(N_1^0 + \delta N_1, N_2^0 + \delta N_2, \ldots, N_m^0 + \delta N_m\). The change in \(\lg W\) in passing from the first distribution to the second is, to the first approximation,
\[ \delta \lg W = \frac{\delta W}{W} = - \sum_i (1 + \lg N_i^0)\,\delta N_i . \tag{11} \]
If \(W^0\) is the maximum value of \(W\), then the difference between \(\lg W^0\) and the value of \(\lg W\) for another, slightly different distribution must, to the first approximation, be equal to zero. In other words, the expression standing on the right-hand side of equality (11), i.e., the “first variation,” must be zero for all possible values \(\delta N_1 \ldots \delta N_m\) (since the total number of particles \(N\) is assumed unchanged, only those values \(\delta N_1 \ldots \delta N_m\) are “possible” whose sum is equal to zero).
But from equation (11) it is immediately clear that its right-hand side is equal to zero in the case when all \(N_i^0\) have the same value, equal, say, to \(a\). Indeed, in this case
\[ \delta \lg W = -\sum_i (1 + \lg a)\,\delta N_i = \mathrm{const.}\sum_i \delta N_i, \tag{12} \]
and for all possible variations \(\sum_i \delta N_i = 0\), for, as was said earlier, the total number of particles must remain constant.
Thus we have obtained the desired result: the most probable distribution will be that in which the number of particles in all cells is the same. If we had not succeeded in doing this, then our method of “counting the possible ways of realizing the given distribution” would have proved unsuitable. Since the functions \(W\) and \(\lg W\) attain their greatest value for the uniform distribution, which intuitively seems the most probable and in fact prevails in gases, and their smallest value—for the most nonuniform of all distributions, it is есте-
to take \(W\) essentially as the measure of the “probability” of the given distribution.
Let us note that the preceding derivation is valid for any value of the constant \(a\). The magnitude of this constant is determined by the total number of particles in the ensemble: \(a=\dfrac{N}{m}\).
Let us now apply the same method to finding the distribution with respect to momenta.
For this purpose we divide momentum space, similarly to coordinate space, into equal-volume cells, each of which would contain a large number of particles. The distribution, as in the preceding case, is determined by the number of particles in each cell; consequently, its “probability” \(W\) is still given by formula (7). The quantities \(\lg W\) and \(\delta \lg W\) also have the same form as before. The essentially new point, however, is the following. Since the energy of a particle depends on its position in momentum space, different distributions with respect to momenta correspond, generally speaking, to different values of the total energy of the ensemble. Therefore a small change in the distribution, connected with a variation \(\delta \lg W\), generally also causes a variation of the total energy \(\delta E\).
The central point of all subsequent reasoning is the identification of the quantity \(\lg W\) with entropy.
More precisely, we assume
\[ S=k\lg W, \tag{13} \]
where \(k\) is a constant factor of proportionality, whose value must be determined from experiment.
Consider a gas in a state of thermal equilibrium at temperature \(T\). If an infinitely small additional energy \(dE\) is imparted to it and one waits until thermal equilibrium is again established, then as a result its entropy will increase by an amount \(dS\), determined by the equation
\[ \frac{dS}{dE}=\frac{1}{T}. \tag{14} \]
Thus, if our model of the gas and the expression for the entropy correspond to reality, then, in passing from the most probable distribution corresponding to the total energy \(E\) to the most probable distribution corresponding to the total energy \(E+dE\) (with the total number of particles unchanged), the variation of the function \(\lg W\) must be equal to the quantity \(\frac{1}{kT}\,dE\).
Let us suppose that, for our initial distribution, the total energy is equal to \(E\), and consider an arbitrary small change of it, connected with an increase of the energy by \(dE\). The new distribution will, obviously, differ only slightly from the most probable distribution corresponding to the energy \(E+dE\). Consequently, one may say that the most probable distribution at energy \(E\) is that for which the first variation is equal to \(\frac{dE}{kT}\). This expression, as was to be expected, vanishes if we consider distributions with the same \(E\).
It is easy to show that the distribution satisfying the above condition has the form:
\[ N_i=\alpha e^{-\frac{\varepsilon_i}{kT}}, \tag{15} \]
where \(\alpha\) is a constant, and \(\varepsilon_i\) is the mean energy of the particles in the \(i\)-th cell, connected with the mean momentum of these particles by the equation:
\[ \varepsilon=\frac{1}{2m}(p_x^2+p_y^2+p_z^2). \tag{16} \]
Indeed, using equation (11) and substituting into it the value of \(\lg N_i\) from equation (15), we have:
\[ \delta S=-k\sum(1+\lg N_i)\delta N_i =-k(1+\lg\alpha)\sum\delta N_i +\sum \varepsilon_i\frac{\delta N_i}{T} =\frac{\delta E}{T}, \tag{17} \]
which was what had to be proved.
The value of the constant \(\alpha\), as in the preceding case, is determined by the total number of particles. Let us denote this number by \(N\) and consider the cells as small
cubes of volume \(H\), so that their number per unit volume of momentum space is \(\frac{1}{H}\). The density \(\rho\) of particles in momentum space, i.e., in other words, the momentum distribution function is, obviously, \(\frac{N_i}{H}\):
\[ \rho=\frac{N_i}{H}=\frac{\alpha}{H}e^{-\frac{\varepsilon_i}{kT}} =\frac{\alpha}{H}e^{-\frac{p_x^2}{2mkT}}e^{-\frac{p_y^2}{2mkT}}e^{-\frac{p_z^2}{2mkT}}. \tag{18} \]
The integral of this expression over the whole momentum space is the total number of particles \(N\):
\[ N=\int_{-\infty}^{\infty}\int_{-\infty}^{\infty}\int_{-\infty}^{\infty} \rho\,dp_x\,dp_y\,dp_z. \tag{19} \]
The integration is carried out easily, since the triple integral is equal to the product of three ordinary integrals:
\[ N=\frac{\alpha}{H} \left[ (2mkT)^{\frac{1}{2}} \int_{0}^{\infty} e^{-w^2}\,dw \right]^3, \tag{20} \]
where \(w\) is a symbolic notation for each of the three momenta. Hence
\[ \alpha=\frac{NH}{(2\pi mkT)^{\frac{3}{2}}}. \tag{21} \]
Finally, for the number of particles in the \(i\)-th cell we obtain the expression:
\[ N_i=\frac{NH}{(2\pi mkT)^{\frac{3}{2}}}e^{-\frac{\varepsilon_i}{kT}}, \tag{22} \]
which contains 4 constants: \(m\), \(N\), \(k\), and \(H\). The first three can be determined on the basis of experimental data (I shall discuss the methods of determination below); in particular, the third is a universal constant named after Boltzmann in his honor, although Boltzmann himself did not give its numerical value. The fourth constant—the volume of a cell \(H\)—does not enter into the expressions for the distribution functions: neither into the function \(\rho\), nor into the distribution function with respect to
energy (which will be discussed below), nor, finally, the basic distribution function \(f\) with respect to coordinates and momenta, determined by equation (1), which I can now present in the form:
\[ f=\frac{N}{V(2\pi mkT)^{\frac{3}{2}}}\, e^{-\frac{p_x^2+p_y^2+p_z^2}{2mkT}}. \tag{23} \]
(here \(V\) denotes the volume of that part of ordinary space which is occupied by our ensemble). The fact that \(H\) drops out is extremely deceptive; for it suggests that not only the volume of the cells has no essential significance, but also that the introduction of the cells themselves serves only as an intermediate step for calculating the distribution function, and that they may therefore be taken arbitrarily small, like differentials in differential calculus. However, precisely the latter is incorrect. From the very course of the proof it follows that the cells must possess a finite volume. As will become clear below, the subdivision of space into cells must be regarded as a quantum postulate; even in the case where it serves to derive the Maxwell–Boltzmann law, which at first sight seems diametrically opposed to everything asserted by quantum theory.
Our next step will be the derivation of the distribution with respect to energy. As I already noted at the beginning, the distribution we are considering is isotropic with respect to all directions of motion of the particles in ordinary space. Mathematically this is expressed by the fact that \(p_x\), \(p_y\), and \(p_z\) enter into the distribution function symmetrically; the physical reason for the isotropy lies in the fact that in our initial assumptions we did not single out any direction over the others (the situation is different in the presence of an external field—there all the results obtained below lose their validity). An isotropic distribution can be completely characterized by the distribution function with respect to energy. This latter can be obtained directly from the distribution function with respect to momenta, if one passes from
Cartesian coordinate system in momentum space and a polar one.^1
We shall carry out this calculation in a somewhat different way. Let us divide momentum space into spherical “layers” by means of a series of concentric spheres with their center at the origin of coordinates. To each sphere there corresponds a definite value of $\varepsilon$, and to each layer an interval from $\varepsilon$ to $\varepsilon+d\varepsilon$. Consider any one of these layers, say the $s$-th. Let the values of the energy on the spheres bounding it be $\varepsilon_s$ and $\varepsilon_s+d\varepsilon$, the radii of these spheres $r_s$ and $r_s+dr$, and, finally, the volume of the layer $dV$. Then, taking into account formula (5), p. 237, we obtain
\[ r_s=(2m\varepsilon_s)^{1/2},\qquad dr=\frac{m}{(2m\varepsilon_s)^{1/2}}\,d\varepsilon \]
\[ dV=4\pi r^2\,dr,\qquad dn=\frac{2\pi}{H^3}(2m)^{3/2}\varepsilon_s^{1/2}\,d\varepsilon . \tag{24} \]
Let us suppose at first that the volume of each layer is many times greater than the volume of an elementary cell. In that case the number of cells $Q_s$ in the $s$-th layer is a large number equal to
\[ Q_s=\frac{dV}{H^3} =\frac{2\pi}{H^3}(2m)^{3/2}\varepsilon_s^{1/2}\,d\varepsilon . \tag{25} \]
In each of these cells there is, on the average,
\[ N_s=\alpha e^{-\varepsilon_s/kT} \tag{26} \]
particles. Consequently the total number of particles in our layer is equal to
\[ M_s=Q_sN_s =\frac{2N}{\sqrt{\pi}(kT)^{3/2}}\, \varepsilon_s^{1/2}e^{-\varepsilon_s/kT}\,d\varepsilon =F(\varepsilon)\,d\varepsilon . \tag{27} \]
Let us denote the absolute value of the momentum $\sqrt{p_x^2+p_y^2+p_z^2}$ by $p$, and the polar angles of the momentum vector by $\theta$ and $\varphi$. Then we shall have
\[ \Phi_p\,dp\,d\theta\,d\varphi = \frac{\alpha}{H^3} e^{-p^2/2mkT}\, p^2\sin\theta\,d\theta\,d\varphi\,dp . \]
Hence, integrating over $\theta$ and $\varphi$, we obtain the distribution function with respect to $p$, from which, by means of formula (8), one may derive a certain distribution function with respect to energy:
Obviously, the coefficient multiplying \(d\varepsilon\) in formula (27) is a distribution function with respect to energy. (In what follows we, writing it as \(\varphi_s\), shall discard the subscript \(s\).) The value of \(c_s\) I took from equation (21), but it could have been obtained directly by integrating \(F\) from \(\varepsilon=0\) to \(\varepsilon=\infty\) and equating the integral to the total number of particles \(N\).
The fact that the quantity \(M_s=F(\varepsilon_s)d\varepsilon_s\) splits into two factors—the number of cells in the \(s\)-th layer, \(Q_s\), and the mean number of particles in each cell, \(N_s\)—considerably facilitates the comparative study of different kinds of statistics. As we shall see below, in passing from one statistic to another sometimes the first factor changes, sometimes the second, and sometimes both together.
In particular, from the Maxwell–Boltzmann law one can pass to that distribution which Planck introduced for oscillators simply by changing the factor \(Q_s\). Until now we divided momentum space into equal cells, the number of which in the layer \(S\), enclosed between the spheres \(\varepsilon_s\) and \(\varepsilon_s+d\varepsilon_s\), was proportional to \(\varepsilon_s^{1/2}d\varepsilon_s\). Let us now consider cells whose volume continuously increases with distance from the origin of coordinates, so that their number in the layer \(S\) is proportional simply to \(d\varepsilon_s\), without the factor \(\varepsilon_s^{1/2}\).
This form of expressing Planck’s postulate is somewhat unusual, although essentially this was the route taken by Planck himself. It is usually said that Planck singled out certain “allowed” values of the energies of the particles of the ensemble, separated from one another by equal intervals \(b\), and consequently forming a series of the form \(a, a+b, a+2b\ldots\), where \(a\) and \(b\) are constants. To each of these allowed values there corresponds a sphere in momentum space. In the layer \(s\) there are approximately \(\dfrac{d\varepsilon_s}{b}\) such “allowed spheres”; this approximate value is the closer to the true one the larger \(\varepsilon_s\) and \(d\varepsilon_s\) are in comparison with \(b\). For statistical conclusions it is immaterial how one represents to oneself the disposition of the particles of this layer: whether one considers the surfaces—*
of the very “allowed” spheres or the cells bounded by them. This difference in methods of description is important for other questions, but in the statistical calculations it plays no role. I shall therefore use sometimes the first picture and sometimes the second. In the present case it is more convenient to consider the subdivision of momentum space into cells having the form of thin concentric spherical layers with center at the origin, whose volume increases continuously with distance from the latter.
If the layer \(s\) is so large that it contains a large number of such cells or allowed spheres, then we may confine ourselves to the first approximation and put:
\[ Q_s=\frac{d\varepsilon}{b}. \tag{28} \]
Hence, using expression (26) for the number of particles in one cell, we obtain:
\[ M_s=Q_sN_s=\frac{\alpha}{b}e^{-\frac{\varepsilon_s}{kT}}\,d\varepsilon_s =F(\varepsilon_s)d\varepsilon_s, \tag{29} \]
where \(M_s\), as before, is the number of particles whose energy lies between \(\varepsilon_s\) and \(\varepsilon_s+d\varepsilon_s\). The value of the constant is still fixed by the condition that the integral of \(F\) over the entire range of variation of the energy, from \(0\) to \(\infty\), must be equal to \(N\):
\[ \int_{0}^{\infty}F(\varepsilon)d\varepsilon =\frac{\alpha kT}{b}\int_{0}^{\infty}e^{-w}\,dw=N, \tag{30} \]
whence we obtain the following expression for the distribution function with respect to energy:
\[ F(\varepsilon)=\frac{N}{kT}e^{-\frac{\varepsilon}{kT}}. \tag{31} \]
This function, of course, contains none of the results obtained by Planck. It has the same smooth and continuous form as the Maxwell–Boltzmann function itself, and is completely independent of the constant \(b\), i.e., of the magnitude of the interval between the succes-
relative allowed energy values. This constant dropped out for the very same reason for which the constant \(H\) dropped out of function (23)—i.e., because of the artificially introduced continuity. Indeed, in performing the integration (30) to determine \(\alpha\), we tacitly assumed that all the allowed energy values lying in the interval \(d\varepsilon_s\) are so close to one another that they may be identified with \(\varepsilon_s\) itself. In other words, we leveled out the discreteness that had originally been introduced by the idea of allowed spheres. Therefore there is nothing surprising in the fact that in function (31), as in the Maxwell–Boltzmann law, not a trace of this discreteness remained!
In order to avoid this leveling, it is clearly necessary, in determining \(\alpha\), to use not integration but summation. Planck’s postulate makes this summation mathematically extremely easy. Indeed, since the number of particles in the \(i\)-th cell is
\[ N_i=\alpha e^{-\frac{\varepsilon_i}{kT}}, \tag{32} \]
the total number of particles is equal to the sum
\[ N=\sum N_i=\alpha e^{-\frac{a}{kT}}\sum_{i=0}^{\infty} e^{-\frac{ib}{kT}} =\frac{\alpha e^{-\frac{a}{kT}}}{1-e^{-\frac{b}{kT}}} \tag{33} \]
[since by Newton’s binomial formula \(1+x+x^2=(1-x)^{-1}\)] whence for \(\alpha\) we have exactly:
\[ \alpha=N e^{\frac{a-b}{kT}}\left(e^{\frac{b}{kT}}-1\right) \tag{34} \]
and for the number of particles in the \(i\)-th cell we obtain the formula
\[ N_i=N e^{\frac{b}{kT}}e^{-\frac{ib}{kT}}. \tag{35} \]
Here, in contrast to the Maxwell–Boltzmann law, the discreteness of momentum space, entering into the classical picture, is accepted and legitimized. Not
Although Planck did not introduce discreteness into classical statistics where it had not been before. He merely was the first to draw attention to it. Not confining himself to the investigation of those problems in which discreteness can be neglected, Planck considered also those cases in which it must be taken into account, and he took it into account.
As we have already indicated, the distribution (35) was proposed by Planck not for freely moving particles, but for oscillators. A “Planck oscillator” is a particle performing simple harmonic oscillations along a straight line about a position of equilibrium, to which it is attracted by a force proportional to the displacement. Like a free particle, its state at each instant of time is characterized by the values of the coordinate \(q\) and the momentum \(p\) (with the position of equilibrium usually being taken as the origin of coordinates). But, in contrast to a free particle, the energy of an oscillator depends not only on \(p\), but also on \(q\), and is a function of the form \(Ap^{2}+Bq^{2}\). Therefore here we cannot consider the space of momenta separately, but must deal with the phase space of both variables \(p\) and \(q\). In principle, we ought always to proceed in this way; but since for an assembly of freely moving particles the phase space has six dimensions, it was more expedient to consider one momentum space. Such a compromise was possible because the energy depended only on the momenta (it was harmful only in one respect, which will be discussed below); in the present case it is impossible, but also unnecessary, since the number of dimensions of the phase space is only two.
Let us imagine this two-dimensional phase space as a plane, each point of which is referred to a Cartesian coordinate system \(p, q\). Suppose that the masses and frequencies of all the oscillators are the same, i.e. that they all possess identical values of the constants \(A\) and \(B\), while the amplitudes may be different. The point representing a given oscillator in phase space describes an elliptical orbit with its center at the origin of coordinates. Different ampli-
different ellipses correspond to these amplitudes. Since the energy of an oscillator depends on its amplitude, for different ellipses we have different values of the energy, and conversely. Therefore, if we divide the phase space into cells by a series of concentric ellipses with their center at the origin of the coordinates, then each cell will correspond to a definite region of values of the energy. In this, the present case differs essentially from the preceding one in that equal intervals of energy correspond to unequal cells (i.e., unequal volumes of phase space).
Thus, if our ellipses divide the phase space into equal cells, then these ellipses themselves correspond to values of the energy forming an arithmetic progression of the form: \(a,\ a+b,\ a+2b,\ldots,\ a+ib,\ldots\). Whether we regard these values of the energy as the only “allowed” ones and allow the oscillators to choose among these values, or whether we assume that the oscillators are distributed uniformly over all the cells, is already a secondary question. In the present case it is easy to say in what the difference will consist here. Independently of whether the oscillators are distributed among the cells or whether they are located on the ellipses themselves, classical statistics leads to the distribution (35). But if one wishes to calculate the mean energy of the oscillators, then in the first case, as the energy of the particles of the \(i\)-th cell, one will have to take the arithmetic mean between \(a+ib\) and \(a+(i+1)b\), while in the second case simply the energy of the particles on the \(i\)-th ellipse, \(a+ib\). In other words, the transition from the case of a discrete series of possible energies to the case of a uniform distribution over cells reduces, in essence, to replacing the original discrete series by another series, the terms of which are the arithmetic means of the terms of the original series. I dwell on this question in order to emphasize that already in the very subdivision of phase space into cells there is implicite contained the quantum postulate.
As is known, Planck assumed that the quantity \(b\)—i.e., the interval between two successive possible-
energies or, with another method of description, between two extreme values of the energy in each cell—is the product of a certain universal constant \((h)\) and the frequency of the oscillator \((\nu)\). Under this assumption the area of each cell is equal to \(h\),1 whatever the frequency of the oscillator may be. From this follows the following general principle of quantum theory. Suppose that each element of a given ensemble is characterized by \(n\) independent coordinates (i.e., has \(n\) degrees of freedom) and by the same number of momenta. Then Planck’s generalized principle says: for an ensemble of particles, each of which has \(n\) degrees of freedom, phase space must be divided into cells of volume \(h^n\).
Let us now see what classical statistics gives, under this assumption, for an ensemble of such particles in which energy and momentum are related by the relation \(\varepsilon = cp\) (as for light corpuscles), and not \(\varepsilon = \dfrac{p^2}{2m}\) (as for material particles). In this case, as in the preceding one, to each value of the energy there corresponds in momentum space a definite sphere with its center at the origin. But the numerical relations obtained are different. Instead of formulas (24), we now have
\[ \varepsilon_s = c r_s, \qquad d\varepsilon_s = c\,dr_s \]
\[ dV = 4\pi r_s^2\,dr_s = \frac{4\pi}{c^3}\,\varepsilon_s^2\,d\varepsilon_s \tag{36} \]
where \(dV\) is the volume of the layer \(S\) enclosed between the spheres \(\varepsilon_s\) and \(\varepsilon_s + d\varepsilon_s\). Let us divide momentum space into equal cells of volume \(H\) and suppose that \(N_i\) is so insignifi-
varies considerably in passing from one cell to another, so that even in the \(S\)-th layer, containing a large number of cells, \(N_i\) may be set equal to a certain mean value \(N_s\). In this case, the number of cells in the layer \(S\) is equal to
\[ Q_s=\frac{dV}{H}=\frac{4\pi}{c^3H}\varepsilon_s^2\,d\varepsilon_s \tag{37} \]
Using for \(N_i\) the classical value (15) and identifying it with \(N_s\), we obtain for the total number of particles in the \(S\)-th layer the expression:
\[ M_s=Q_sN_s=a\,\frac{4\pi}{c^3H}\,\varepsilon_s^2 e^{-\frac{\varepsilon_s}{kT}}\,d_s\varepsilon_s =F(\varepsilon_s)\,d_s\varepsilon . \tag{38} \]
The constant \(a\) is calculated by the same method as before. As a result, for the distribution function we have
\[ F(\varepsilon)=\frac{N}{3k^3T^3}\, \frac{\varepsilon^2}{e^{\frac{\varepsilon}{kT}}} \tag{39} \]
Such, according to classical statistics, is the “uniform” distribution by energy for a light gas under the assumption that momentum space is divided into equal cells. Experiment, however, leads to an entirely different distribution:
\[ F(\varepsilon)=\frac{8\pi V}{c^3h^3}\cdot \frac{\varepsilon^2\,d\varepsilon}{e^{\frac{\varepsilon}{kT}}-1} \tag{40} \]
At first glance it may seem that (39) serves as a limiting case of (40), i.e., that the actual distribution by energy, like Planck’s distribution for oscillators, can be found simply by carrying out the same derivation more accurately. In reality, however, this is not so. It is true that the second factor of expression (39) is indeed the limit of the second factor of expression (40) at very low temperatures. But the first factor in (40) contains only the volume of the gas and certain universal constants, whereas the first factor in (39) contains the temperature and one arbitrary constant—the total number of particles in the ensemble. Therefore it is by no means possible to say that one of these factors serves
previous one. Further, it must be noted that our assumption that the volume of each cell of phase space is equal to \(h^3\) was in no way reflected in the form of formula (39). Bose showed that, in order to obtain (40) instead of expression (39), it is necessary to rebuild the entire foundation of classical statistics.
Bose Statistics
Let us divide the momentum space of the photons, as before, into equal-volume cells, and consider the various distributions of the particles among these cells. Our task, as before, is to find the “most probable” distribution associated with the entropy of the collection. But, in contrast to the preceding case, we shall introduce an entirely new definition of “distribution” and hence a new method of counting the possible ways of realizing it, i.e. of finding its probability.
Suppose that the distribution of particles in the former sense of the word is known to us; that is, it is known that in one cell there are \(N_0\) particles, in another cell—\(N_1\) particles, and in general in the \(i\)-th cell—\(N_i\) particles.
Let us count the number of cells containing not a single particle—let it be equal to \(Z_0\). Let us count the number of cells containing one particle—let it be equal to \(Z_1\). In general, let \(Z_i\) be the number of cells containing \(i\) particles. Let us write down all the numbers \(Z_i\) thus obtained.
We shall now give a new term to the quantity which we earlier called “distribution”: some new name—for example, “placement.” We shall characterize a “distribution” not by a system of numbers \(N_0, N_1, N_2,\ldots\), but by the quantities \(Z_0, Z_1, Z_2,\ldots\) and by the value of the total energy of the collection.
Obviously, every distribution in the new sense of the word will correspond to an enormous number of distributions in the old sense, and [[unclear: several words]].
Such a change of terminology may seem inconvenient, but the slightest confusion leads, in my view, to a greater inconvenience than the use of a different term than “distribution” for the quantity always called by that name in the new statistics.
encompasses a whole series of distributions in the old sense of the word. These latter we shall regard as different ways of realizing the first. Passing finally to the new terminology, this may be expressed as follows: as the probability of a given distribution \(W^*\) we shall take the number of arrangements covered by it; earlier, as the probability of each arrangement \(W\), we took the number of combinations covered by it (see equation (9)). Now, however, all arrangements are considered equiprobable.
We proceed to the definition of the number \(W^*\). In computing it, it must be borne in mind that not all the arrangements covered by a given sequence of values \(Z_i\) interest us, but only that part of them which corresponds to the prescribed value of the total energy of the ensemble.
As before, in addition to dividing the momentum space into equal cells of volume \(H\), we shall consider its subdivision into spherical layers, each of which is large enough to contain a large number of cells and, at the same time, small enough that within its limits one and the same value of the energy \(\varepsilon\) may be assigned to its cells. As a result we shall obviously arrive at a “leveled” formula. It is remarkable that, although Planck introduced quanta into physics precisely by abandoning that “leveling” which was usual in classical statistics, the quantum formula for radiation is now derived by a method in which this leveling is legitimized!
Let us consider any one of our layers, for example the \(s\)-th. Let \(Q_s\), or \(Z\), be the total number of cells in this layer; \(Z_{is}\) the number of those cells in which there are \(i\) particles; and, finally, \(M_s\) the total number of particles in the layer. According to our scheme, the number of ways of realizing the given distribution characterized by the numbers \(Z_{is}\) is given by the formula
\[ W_s^*=\frac{Q_s!}{Z_{0s}!\,Z_{1s}!\,Z_{2s}!\cdots},\qquad Q_s=\sum_i Z_{is}. \tag{41} \]
Since, by assumption, the energy values for all cells of our layer are approximately the same, all these different...
ways of realizing the distribution \(Z_{is}\) correspond, approximately, to one and the same total energy (and also to one and the same total number of cells and total number of particles).
Let us carry out the same count for all the remaining layers. The total number of ways of realizing the given distribution that satisfy the condition of constancy of the total energy, the total number of particles, and the total number of cells is evidently equal to the product of all the quantities \(W_s^*\),
\[ W^*=\prod_s W_s^* . \tag{42} \]
We form, for the same purpose as in the preceding case, the expression \(\lg W^*\) and use Stirling’s formula (assuming that none of the \(Z_i\) is less than, approximately, 10)
\[ \lg W^*=\sum_s \lg W_s^* =\sum_s \left[ Q_s \lg Q_s-\sum_i Z_{is}\lg Z_{is}\right]. \tag{43} \]
Let us suppose further that in the present case the entropy is defined not by the logarithm of the function \(W\) from formula (10), but by the logarithm of the function \(W^*\)
\[ S=k\lg W^* . \tag{44} \]
Under this assumption, a small change \(\delta E\) of the total energy of the ensemble \(E\), caused by the transition from the given distribution \(Z_{is}\) to a slightly different one \(Z_{is}+\delta Z_{is}\), must be connected with the corresponding change of the expression \(k\lg W^*\) by the formula:
\[ \delta S=\delta(k\lg W^*)=\frac{\delta E}{T}. \tag{14} \]
But the first variation of \((k\lg W^*)\) is equal to:
\[ \delta(k\lg W^*)=-k\sum_s\sum_i (1+\lg Z_{is})\,\delta Z_{is}. \tag{45} \]
In order to satisfy condition (14), it is necessary to put
\[ Z_{is}=a_s e^{-\frac{\varepsilon_s}{kT}} . \tag{46} \]
Indeed, substituting this expression into formula (45), we obtain:
\[ \delta(k \lg W^*) = -k \sum_s \sum_i \left(1+\lg a_s-\frac{i\varepsilon_s}{kT}\right)\delta Z_{is} = \]
\[ = -k \sum_s (1+\lg a_s)\sum_i \delta Z_{is} +\frac{1}{T}\sum_s \varepsilon_s \sum_i i\,\delta Z_{is} = \frac{\delta \varepsilon}{T} \tag{47} \]
since \(\sum_i \delta Z_{is}=0\), owing to the constancy of the number of cells in each layer, while the quantity \(\varepsilon_s \sum_s i Z_{is}\) is equal to the total energy of the particles in layer \(s\).
Thus, Bose statistics, by introducing a new definition of entropy expressed by equation (44), leads to a new equilibrium distribution. This distribution is given by equation (46), which fixes the values of the numbers \(Z_{is}\). In order to be able to compare it with the results of the old statistics, it is necessary to obtain an expression for the new distribution function in energy.
Let us first find the values of the constants \(a_s\). For this it is necessary to take the sum of the quantities \(Z_{is}\) over all \(i\) for each layer and equate it to the total number of cells in the given layer \(Q_s\):
\[ Q_s=\sum_i Z_{is} = a_s \sum_i e^{-\frac{i\varepsilon_s}{kT}} = a_s\left(1-e^{-\frac{\varepsilon_s}{kT}}\right)^{-1}. \tag{48} \]
Substituting this value of \(a_s\) into (46), we obtain something already rather similar to familiar results.
For the total number of particles \(M_s\) in layer \(S\) we have
\[ M_s=\sum_i i Z_{is} = \frac{a_s e^{-\frac{\varepsilon_s}{kT}}} {\left(1-e^{-\frac{\varepsilon_s}{kT}}\right)^2} = Q_s\cdot \frac{1}{e^{\frac{\varepsilon_s}{kT}}-1}, \tag{49} \]
and this already looks quite familiar.
The number of cells \(Q_s\), for the case when they all have the same volume \(H_s\), we have already calculated [see eq. (37)]; consequently:
\[ M_s= \frac{4\pi}{c^3H} \frac{\varepsilon_s^2}{e^{\frac{\varepsilon_s}{kT}}-1} \,d\varepsilon_s = F(\varepsilon_s). \tag{50} \]
Here, \(F(\varepsilon)\) is the “smoothed” distribution function with respect to energy in the new statistics. In contrast to the functions with which we dealt in the old statistics, it depends directly on the volume of the elementary cell \(H\). Further, whereas in the old statistics we could obtain quantum formulae only by abandoning approximations, here we arrive at these formulae even while admitting approximations. I consider it necessary to emphasize this distinction.
Let us finally introduce the quantum postulate, which is a generalization of Planck’s principle for oscillators, i.e. let us assume that the volume of an elementary cell of phase space is equal to \(h^3\), and let us further suppose that this volume is the product of the volume of an elementary cell of momentum space, \(H\), by the total volume \(V\) occupied by the gas. This assumption gives:
\[ H=\frac{h^3}{V}. \tag{51} \]
Let us note that the latter relation assumes an especially elegant form if, instead of cells, one considers the possible values of the energy and, instead of the corpuscular picture, uses the wave picture. Indeed, in such a picture formula (51) determines a discrete series of possible wavelengths, and the latter coincide with the lengths of those standing waves which can be formed in a cube of volume \(V\).
Thus, the new statistics leads to the following distribution:
\[ F(\varepsilon)\,d\varepsilon = \frac{4\pi V}{h^3}\, \frac{\varepsilon^2\,d\varepsilon}{e^{\varepsilon/kT}-1}. \tag{52} \]
Discarding the factor \(V\), we obtain the number of particles in unit volume whose energy lies between \(\varepsilon\) and \(\varepsilon+d\varepsilon\), under an equilibrium distribution characterized by the temperature \(T\). The product of this quantity by \(\varepsilon\) gives the total energy of these particles in unit volume. If the particles under consideration are photons, then their wavelengths lie
within the limits from [[unclear: formula]] to [[unclear: formula]], and the frequencies within the limits from [[unclear: formula]] to [[unclear: formula]]. Passing from the variable \(\varepsilon\) to the new variables \(\lambda\) and \(\nu\), we obtain a distribution function expressing the dependence of the density of radiant energy on wavelength and frequency. The density of energy computed in this way agrees with the observed one with only one difference: it lacks a factor of two. The appearance of this extra factor of two is explained by the polarization of light. As a result we arrive at the formula for black radiation:
\[ \rho(\nu)\,d\nu = \frac{8\pi h}{c^{3}}\, \frac{\nu^{3}\,d\nu}{e^{h\nu/kT}-1} \tag{53} \]
the validity of which thus serves as a justification for the introduction of the new statistics.
Let us note that the new statistics makes it possible to determine exactly the number of photons in a unit volume at a given temperature, whereas in the old statistics the total number of atoms entered as an arbitrary constant. This is explained by the profound physical difference between a light gas and a material gas. Knowing the temperature and the volume of a box containing helium, I still cannot say anything about the number of helium atoms in it; the density of the latter may vary in any manner at constant volume and constant temperature. But knowing the temperature and the volume of a closed cavity containing radiation, I can determine exactly the magnitude of the radiant energy enclosed in it, i.e. the total number of light quanta. The new statistics is in agreement with this experimental fact. The question arises: can one try to apply the new statistics to a material gas and thereby avoid the absurd conclusion that the total number of atoms of this gas is determined by its temperature and volume? It turns out that this is possible. For this it is necessary to replace the distribution (46) by a more general distribution containing an arbitrary constant \(B\):
\[ Z_{is}=\sum_{r} g_{ir} e^{-iB-\varepsilon_{ir}/kT} \tag{54} \]
Substituting this into the expression \((k \lg W)\), we obtain instead of (47) the equation:
\[ \delta(k \lg W^{*}) = - k \sum_s (1+\lg a_s)\sum_i \delta Z_{is} + kB\sum_s \sum_i i\,\delta Z_{is} + \frac{1}{T}\sum_s \varepsilon_s \sum_i i\,\delta Z_{is}. \tag{55} \]
If the distribution (54) corresponds to reality, then the right-hand side of the last equation must be equal to \(\frac{\delta E}{T}\), i.e. not only \(\sum_i \delta Z_{is}\), but also \(\sum_s \sum_i i\,\delta Z_{is}\) must be zero. But \(\sum_s \sum_i i\,\delta Z_{is}\) is indeed \(=0\) for all variations under which the total number of particles remains constant. Consequently, the distribution (51) does indeed possess the greatest entropy among all distributions corresponding to the same total energy and total number of particles. The distribution (46) is still more “exceptional”: the value of its entropy is the greatest in comparison with all other distributions corresponding to the same total energy, including also those which possess a somewhat different total number of particles. For a material gas, the distribution (54) may be recognized as the most probable. The fact that for a light gas the distribution (46) is required apparently expresses one of the deepest differences between matter and radiation.
Carrying out the same calculations as above, we arrive at the following expression for the number of particles in the \(s\)-th layer:
\[ M_s = Q_s \frac{1}{e^{B+\frac{\varepsilon_s}{kT}}-1}. \tag{56} \]
In the case of a light gas, we put \(B=0\) and used the value of \(Q_s\) from equation (37). In the case of a material gas, \(Q_s\) is determined by (25), which, for \(H=\frac{h^3}{U}\), gives:
\[ M_s = \frac{2v}{h^3(\pi)^{\frac12}} (2\pi m)^{\frac32} \frac{(\varepsilon_s)^{\frac12}\,d\varepsilon_s} {e^{B+\frac{\varepsilon_s}{kT}}-1} = F(\varepsilon_s)\,d\varepsilon_s . \tag{57} \]
To calculate the constant \(B\), it is necessary to integrate the expression obtained over all possible values of the energy from \(0\) to \(\infty\) and to set the result equal to the total number of particles, \(N\).
Einstein proposed the distribution (57) as a possible replacement for the classical Maxwell–Boltzmann distribution. It is difficult to say which of these distributions corresponds more closely to the experimental data, since with increasing temperature formula (57) approaches the classical one more and more, and under ordinary conditions of temperature and pressure they are completely indistinguishable. Resolving this question would be of great importance, since it would show which definition of probability, and consequently of entropy, is applicable to a material gas. The reader has perhaps already noticed that applying the Bose method to the problem of the most probable distribution of particles in ordinary space would lead to discrepancies with classical statistics, i.e., with experimental facts. Therefore we, unfortunately, must always operate in six-dimensional phase space.
Fermi Statistics
The statistics proposed by Fermi and, independently of him, by Dirac is based on the same fundamental principles as Bose statistics, i.e., it proceeds from the same method of counting the possible ways of realizing a given distribution and from the same definitions of probability and entropy. But one additional assumption of a restrictive character is introduced into it, namely: it is postulated that each cell can contain no more than a certain limiting number of particles. In particular, for a gas in the absence of an external field it is assumed that each cell is either empty or contains at most one particle.
The impetus for the emergence of the Fermi–Dirac theory was Pauli’s “exclusion principle.” This principle may be expressed in the following way. According to Bohr’s theory…
of the structure of the atom, intra-atomic electrons can revolve only in certain “allowed” orbits, each of which is characterized by several so-called “quantum numbers.” (Recently the concepts of “allowed orbits” and “quantum numbers” have acquired a certain vagueness; but as a first approximation one may use the original scheme.) To this restrictive principle of Bohr, Pauli added yet another restrictive principle, stating that in each orbit characterized by a given set of quantum numbers there may be at most one electron (perhaps it would be more accurate to say not “one electron,” but “a certain small number of electrons”). The connection of this principle with one of Fermi’s basic assumptions will soon become clear to us. This is not the place to enumerate all the experimental facts that speak in favor of Pauli’s principle; in any case they are so numerous that they can serve to a certain extent as confirmation of the admissibility of Fermi’s idea—or rather, can make this idea admissible (which, in the absence of these facts, it might not have seemed to be).
All the reasoning proceeds in the present case quite analogously to the reasoning that led to the derivation of the distribution (56), with the sole difference that the sum over \(i\) reduces to two terms: \(i=0\) and \(i=1\). Indeed, according to our fundamental assumption, to characterize the distribution in the given layer it is necessary to specify only two numbers \(Z_{is}\): the number of empty cells \(Z_{0s}\) and the number of cells containing one particle each \(Z_{1s}\). It is easy to show that, for a given total energy and total number of particles, the distribution (54) still has the greatest probability:
\[ Z_{0s}=\alpha_s,\qquad Z_{1s}=\alpha_s e^{-B-\frac{\varepsilon_s}{kT}} \tag{58} \]
Hence, for the number of particles in the layer \(S\), we obtain:
\[ M_s=Q_s\,\frac{1}{e^{B+\frac{\varepsilon_s}{kT}}+1} \tag{59} \]
In the present case, it is obviously meaningless to use for \(Q_s\) the value (37), since, as we have already seen, for a photon gas Bose’s formula is valid. Therefore Fermi took the value of \(Q_s\) valid for material gases, and obtained:
\[ M_s=(2\pi m)^{\frac{3}{2}} \frac{\varepsilon_s^{\frac{1}{2}}\,d\varepsilon_s} {e^{\frac{\varepsilon_s-\mu}{kT}}+1} =F(\varepsilon_s)\,d\varepsilon_s \tag{60} \]
This formula serves as the starting point for the modern theory of the electron gas in metals, revived and reworked by Pauli and Sommerfeld. The second part of the present article will be devoted to an exposition of this theory.
Application of Fermi Statistics to Free Electrons in Metals
We shall consider a given piece of metal as a certain volume bounded by walls, filled with “free” electrons, whose distribution is determined by Fermi’s formula.
The Fermi distribution function contains, as an arbitrary constant, the total number of electrons \(N\). It also contains the volume occupied by these \(N\) particles, \(V\).* We shall assume the latter to be equal to the total volume of the given piece of metal. Such an assumption, as though completely ignoring the existence of atoms and regarding the metal as a vacuum filled only with free electrons, seems at first sight extremely strange, even if it is regarded only as a first approximation. However, the following two facts speak in its favor: 1) slowly moving electrons can pass through atoms (at least through the atoms of some elements) as freely as through empty space, and 2) a plane wave can pass without scattering through an assemblage of particles, each of which is capable of causing scattering, provided only that these particles are arranged in the form of a regular lattice and the distance between them
*
less than the wavelength. The velocities ascribed to electrons in metals are so insignificant, and their wavelengths so large, that such behavior is in principle quite possible.
The “walls” play the role of an agent preventing the electrons from leaving the volume occupied by the metal. This agent is usually represented in the form of a sudden and sharp jump of potential at the boundary surface of the metal. Let us denote the normal component of the electron’s velocity by \(u\). Then the action of the walls may be expressed as follows: so long as the “normal component of the kinetic energy” of the electron, \(\frac{1}{2}mu^2\), remains less than some limiting value \(W_a\), the electron cannot pass through the boundary surface and is thrown back into the metal; when, however, \(\frac{1}{2}mu^2 > W_a\), the electron emerges outside, with its kinetic energy reduced by \(W_a\). According to the newest views, the electron can sometimes emerge outside even when \(\frac{1}{2}mu^2 < W_a\), and conversely can remain in the metal when \(\frac{1}{2}mu^2 > W_a\); for the present, however, we shall use the original, simpler conceptions. The constant \(W_a\) is called the work function.
Like the constant \(N\), the work function is an arbitrary constant of our theory. Modern physics strives to reduce all differences between metals to different values of these two constants. Later we shall see that, in order to achieve this aim, it is necessary to introduce still other constants—first of all the one which in the older theory played the role of the electron’s “mean free path.” But even with only these two constants a number of valuable results can be obtained.
Let us again rewrite Fermi’s formula, expressing the distribution by energy in an ensemble of \(N\) particles contained in a volume \(V\) at temperature \(T\):
\[ F(\varepsilon) = G \frac{2}{\sqrt{\pi}}\cdot \frac{V}{h^3}(2\pi m)^{3/2}(\varepsilon)^{1/2} \frac{1}{A e^{\varepsilon/kT}+1}. \tag{61} \]
In order that our notation coincide with Sommerfeld’s, we have put \(e^{v}=\frac{1}{A}\) and have introduced the additional factor \(G\) (to which subsequently the value 2 will be assigned), for reasons that will be indicated below. The corresponding distribution function with respect to coordinates and momenta is
\[ f(x,y,z,p_x,p_y,p_z)=\frac{G}{h^3}\frac{1}{A e^{\varepsilon/kT}+1}. \tag{62} \]
The first step, as in classical statistics, must consist in determining the constant \(A\) by integrating the function \(F(\varepsilon)\) over all values of the energy from 0 to \(\infty\) and equating the resulting expression to the total number of particles \(N\):
\[ \int_0^\infty F(\varepsilon)d\varepsilon=N. \tag{63} \]
But this step, so easy in classical statistics, is here extremely difficult. Indeed, the integral of \(F(\varepsilon)\) does not belong to the number of functions well known in mathematics and cannot even be expressed by any combination of such functions; thus the relation between \(A\) and \(N\) cannot be expressed by any simple equation. However, Sommerfeld succeeded in deriving for the integral (63) two series expansions, one of which is valid for \(A<1\), and the other for very large \(A\). Thanks to a fortunate coincidence—which may even seem too fortunate to be correct—in all real cases it is quite sufficient to use the first one or two terms of one of these series.
Let us first consider an approximation which, as will become clear later, is unsuitable for the electron theory of metals; namely, let us suppose \(A\) to be so small that in the denominator of the expression \(F(\varepsilon)\) the second term may be neglected in comparison with the first. In this case the distribution law is close to the Maxwell–Boltzmann one, and therefore the constant \(A\) must have such a
value for which \(F(\varepsilon)\) would in the limit coincide with the classical expression (27), i.e. the value
\[ A=\frac{Nh^{3}}{GV}(2\pi mkT)^{-\frac{3}{2}}. \tag{64} \]
It would, however, be a great error to regard this value of \(A\) as correct under all circumstances. It can be used only if \(A\) is considerably less than 1; in other words, if the right-hand side of equation (64) is extremely small. The question arises: in what concrete physical cases does this occur?
It turns out that for all material gases under ordinary laboratory conditions \(A\) is very small, i.e. the new statistics leads practically to the same results as the old one. Discrepancies between them appear only at such high densities and such low temperatures at which most gases are already in the liquid state, or in any case very far from that “ideal” state with which statistics alone is concerned. The only gases for which it may be possible to detect these discrepancies are helium and hydrogen.
At this point Fermi apparently stopped. But Pauli observed that if the new statistics is applied to an electron gas of the density which it—according to Rikke and Drude—must possess inside metals, then the deviations from the classical distribution must be expressed considerably more sharply. Indeed, first, the mass of each electron \(m\) is many times smaller than the mass of an atom or molecule of any material gas. Secondly, \(N\), the number of electrons in a given piece of metal—if one assumes it equal to the number of metal atoms contained in it—is thousands of times greater than the number of particles contained in an equal volume of material gas. The expression (64) for \(A\) contains \(N\) in the numerator and \(m^{\frac{3}{2}}\) in the denominator, and therefore for the hypothetical electron gas in metals it is already not small. Conse-
Consequently, in the present case, expression (64) for \(A\) cannot be regarded as correct.
Let us note that, whereas the quantity \(m\) is known to us from experiment, we know nothing definite about the quantity \(N\). We cannot directly determine how many free electrons are contained in a given piece, say in a unit volume, of a metal. Therefore, as I have already said above, the density of the electron gas
\[ n\left(=\frac{N}{V}\right) \]
is an arbitrary constant of the theory. The problem is, starting from some definite value of \(n\), to obtain correct numerical values for at least 5–6 quantities characterizing the given metal. So long as the methods of classical statistics were applied to the electron gas, this proved impossible. Taking \(n\) equal to the number of atoms of the metal in unit volume, we obtained inordinately high values for the specific heat (and, as has recently become clear, for the magnetic susceptibility); while, on passing to smaller \(n\), we came into disagreement with experiment in other respects. The general opinion inclined, however, to the view that \(n\) could be considered close to the number of atoms, or an integral multiple of this number. At the basis of it lay, in my opinion, the conviction that, since free electrons are detached from atoms and since all atoms are alike, each of them on the average supplies as many electrons as any other. In any case, for Pauli and Sommerfeld it was natural to try to apply Fermi statistics under the assumption that the number of electrons is equal to the number of atoms, and to see what the combination of these two assumptions would yield.
Substituting for \(m\) the value of the electron mass, for \(\frac{N}{V}\) the number of atoms in unit volume of any metal, for \(T\) any temperature from absolute zero to several thousand degrees, and finally for \(G\) any small integer, we shall see that the right-hand side of equality (64) assumes an extraordinarily large value. For example, for \(\frac{N}{V} = 5.9 \cdot 10^{22}\)
(number of atoms in \(1\ \mathrm{cm}^3\) of silver), \(T=300^\circ K\), \(G=2\), Sommerfeld obtained
\[ \frac{n h^3}{G}(2\pi m k T)^{-\frac{3}{2}}\simeq 2400, \tag{65} \]
i.e. a result that invalidates equality (64).
Hence it is clear that in the present case one must use that one of the above-mentioned expansions of the integral \(\int F(\varepsilon)\,d\varepsilon\) which is valid for large \(A\). The first two terms of this expansion have the form:
\[ \int_0^\infty F(\varepsilon)d\varepsilon = N = \frac{GV}{h^3}\frac{4\pi}{3}(2mkT\lg A)^{\frac{3}{2}} \left[1+\frac{\pi^2}{8}(\lg A)^{-2}+\cdots\right]. \tag{66} \]
Restricting ourselves to the first term of this expression and substituting into it the above-indicated values of \(N/V\) and \(T\), we do indeed obtain for \(\lg A\) a very large value \((\lg_e A=325)\). Hence one may conclude that for the electron gas in metals it is sufficient to take the first two terms, and in some cases even only the first term, of the series written above.
Thus, in the first and second approximations for \(A\), the following formulas are valid:
\[ 2mkT\lg A = h^2\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}} \]
first approximation;
\[ 2mkT\lg A = h^2\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}} \left[ 1- \frac{(2\pi m k T)^2}{12h^4} \left(\frac{3n}{4\pi G}\right)^{-\frac{4}{3}} \right] \tag{67} \]
second approximation.
(In calculating the second approximation, the value of \(\lg A\) from the first approximation is substituted into the second term of the series.)
Substituting one of these approximations into the functions (61) and (62), we obtain, with the desired degree of accuracy, the final expression for the postulated distribution of electrons, into which there enters only one arbitrary constant \(n\). We shall now consider the application of these formulas in particular concrete cases.
Specific Heat
As we have already mentioned above, the principal difficulty for the old theory of the electron gas, based on classical statistics and on the assumption that the number of free electrons is equal to the number of atoms, was the problem of specific heats. Let us see how this difficulty is resolved in the new theory.
Knowing the law of distribution with respect to energy, we can now determine the total energy of the ensemble as a function of temperature:
\[ E=\int_{0}^{\infty}\varepsilon F'(\varepsilon)\,d\varepsilon . \tag{68} \]
Substituting into this formula the classical expression (28), we find
\[ E=\frac{3}{2}NkT. \tag{69} \]
Using instead the Fermi distribution, in the first approximation we have:
\[ E=E_0+\frac{1}{2}\gamma VT^2, \]
where
\[ E_0=\frac{2\pi}{5}\frac{VGh^2}{2m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}}, \]
\[ \gamma=\frac{\pi}{3}nG\frac{(2\pi k)^2}{h^2}\left(\frac{3n}{4\pi G}\right)^{\frac{1}{3}}, \tag{70} \]
i.e. a formula of a completely different form.
As is known, the specific heat is expressed by the derivative \(\frac{dE}{dT}\). Thus, in the classical theory it has a constant value, whereas in Fermi statistics it is proportional to the absolute temperature—which agrees with Nernst’s heat theorem—and even at room temperatures constitutes only a small fraction of the classical value.
Experimentally, the specific heat of the electron gas cannot be measured separately from the specific heat of the lattice.
atoms forming the metal. At first glance this makes any experimental verification of formulas (69) and (70) impossible. However, one may say with complete certainty that the classical formula is unacceptable one way or another. Indeed, the observed values of the specific heats of metals agree so well with those values which statistical theory (both old and new) assigns to atoms alone, that for electrons there simply “is no room left”—more precisely, there is no room left for the additional heat capacity \(\frac{3nk}{2}\) required by the classical theory. If, however, \(n\) is taken so small that the quantity \(\frac{3nk}{2}\) could be neglected, then this, as I have already said, leads to difficulties in other questions. The quantity \(\gamma VT\), which the new statistics gives, may be neglected even in the case when \(n\) is equal to the number of atoms and \(T\) reaches several hundreds of degrees. Anyone who in his time did not experience the defeat suffered by the old theory of the electron gas on the problem of specific heats will hardly understand the joy which the victory of the new statistics in this question brought to physicists. By this victory the new statistics redeemed all its shortcomings in other domains.
Fundamental Properties of the Fermi Distribution
Let us now dwell on some fundamental properties of the Fermi distribution, which has so brilliantly withstood its first test.
For this purpose let us substitute in function (62), instead of the constant \(A\), its first approximation, given by formula (67):
\[ f=\frac{G}{h^{3}}\frac{1}{e^{\frac{\varepsilon-W_i}{kT}}+1}, \quad \text{where } W_i=\frac{h^{2}}{2m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}}. \tag{71} \]
At absolute zero, the exponential term is equal either to infinity or to zero, depending on whether we have \(\varepsilon > W_i\) or \(\varepsilon < W_i\). Thus, at absolute zero the density of electrons in phase space is equal to—
constant value \(\dfrac{G}{h^3}\) for all values of the energy less than \(W_i\), and is equal to zero for all values of the energy greater than \(W_i\).
This striking result is easily obtained without any calculations, as a consequence of the basic postulates of Fermi’s theory. By definition, absolute zero is that temperature at which the system completely ceases to give up any energy to the external space. Therefore, at absolute zero we must have the most “dense” of all possible distributions. If one proceeds from the fundamental assumption that each cell of phase space can contain at most one electron, then the densest distribution will obviously be that in which all electrons are situated within a certain sphere centered at the origin of coordinates, one in each cell. The volume of this sphere must be such that all the electrons can be accommodated in it. With such a distribution, the number of electrons per unit volume of phase space inside the sphere is equal to the number of cells per unit volume, i.e. to the reciprocal of the volume of an elementary cell, while outside the sphere it is equal to zero. Thus the process of cooling the system is accompanied by the formation of a kind of “crystal lattice” in phase space.
All the preceding arguments are valid not only for the phase space of momenta. In momentum space, a sphere of radius \(p\), i.e. of volume \(\dfrac{4\pi p^3}{3}\), contains \(\dfrac{4\pi p^3}{3} \big/ \dfrac{h^3}{V}\) elementary cells. Equating this expression to the total number of electrons \(N\) and solving the resulting equation for \(p\), we find the radius of the sphere in which all the electrons would be placed one in each cell. If, however, the volume of the sphere is set equal to \(\dfrac{N}{G}\), then the value of \(p\) gives us the radius \(p_m\) of that sphere in which the electrons would be accommodated \(g\) in each cell. This quantity obviously represents the maximum
the value of the electron momentum at absolute zero. The velocity corresponding to this maximum momentum \(v_m\) is equal to \(\dfrac{p_m}{m}\) and consequently is given by the expression
\[ v_m=\frac{h}{m}\left(\frac{3n}{4\pi}\right)^{\frac{1}{3}}, \tag{71a} \]
while the corresponding kinetic energy is \(\frac{1}{2}mv_m^2=W_i\).
Let us find the mean value of the electron velocity \(v\), and also of its square, cube, etc., at absolute zero in a gas obeying the Fermi distribution. The general formula for the mean value of any expression of the type \(v^s\) is:
\[ \overline{v^s}=\frac{1}{n}\int v^s f(v)\,dv =\frac{1}{n}\frac{4\pi G}{m^3h^3}\int_0^{v_m} v^{s+2}\,dv. \]
In particular,
\[ \overline{v^{-1}}=\frac{3}{2v_m};\quad \overline{v}=\frac{3v_m}{4};\quad \overline{v^2}=\frac{3v_m^2}{5};\quad \overline{v^3}=\frac{1}{2}v_m^3. \tag{72a} \]
The corresponding means for the Maxwellian distribution are
\[ \overline{v^{-1}}=2\left(\frac{m}{2\pi kT}\right)^{\frac{1}{2}};\quad \overline{v}=2\left(\frac{2kT}{\pi m}\right)^{\frac{1}{2}};\quad \overline{v^2}=\frac{2kT}{m};\quad \overline{v^3}=4\pi\left(\frac{2kT}{\pi m}\right)^{\frac{3}{2}}. \tag{72b} \]
If the energy values \(\varepsilon\) are plotted along the abscissae, then the curve of the distribution function with respect to coordinates and momenta \(f\) will at first have the form of a horizontal straight line situated at a distance \(\dfrac{G}{h^3}\) from the axis of abscissae; while the curve of the distribution function with respect to energy \(F\) will have the form of a concave parabola. Beginning with the abscissa \(\varepsilon=W_i\), both curves will coincide with the axis of abscissae.
This is the situation at absolute zero; what happens when the temperature rises? As Sommerfeld showed, when the temperature is raised, the right angles on the distribution curve are gradually and very slowly rounded off, and the curve continues all the time to pass through the middle of the vertical segment
\(BC\) (Fig. 1). The end of the curve, as in the classical case, approaches the abscissa axis. But at room temperatures, and even at considerably higher ones, the distribution differs so little from that which occurs at absolute zero that many phenomena can be excellently explained on the basis of the above-described “compressed”—so-called “ideally degenerate”—distribution. In particular, when calculating resistance and certain other characteristic constants of metals, one may use the mean values (72a). Alongside them, however, there also exist such constants—for example the coefficient of thermal conductivity—in calculating which one cannot use the mean values derived for the case of absolute zero, and more exact formulas have to be taken. On these questions we refer the reader to the cited works of Sommerfeld.
Fig. 1
Graphical representation of the Fermi distribution function \(f\), plotted with respect to \(\dfrac{\varepsilon}{W_i}\) as the independent variable, for an electron gas containing \(6.5 \cdot 10^{22}\) particles per \(\text{cm}^3\) at temperature \(0\) (the graph consisting of straight-line segments) and \(1500^\circ K\). The quantity \(W_i\) is equal to 6 equivalent volts.
The calculation gives unexpectedly large values for \(W_i\): for example, 5.6 volts for silver, 5.7 for tungsten, 6.0 for platinum. Obviously, the quantity \(W_i\) depends on the compactness of the lattice; it is the greater, the more closely the atoms are arranged. In potassium and sodium the distance between atoms is comparatively large, and therefore for them \(W_i\) is equal, respectively, to 2.1 and 2.3 volts.
The difference between this and the classical distribution is enormous. Previously we assumed that at ordinary temperatures the energy of electrons in metals fluctuates, according to Maxwell’s law, around a mean value equal to approximately \(0.02\) volts. According to present ideas, the range of possible energies extends from \(0\) to
6 volts, and the greatest number of electrons falls precisely at the top of the curve (although at the very top the curve breaks off sharply).
The pressure of the electron gas is related to the energy density by the same equation as in the classical theory:
\[ p=\frac{2}{3}\cdot\frac{E}{V}, \]
and, consequently, it varies proportionally to the total energy, i.e., it has an immeasurably large value (of the order of hundreds of thousands of atmospheres) at absolute zero, and then slowly increases proportionally to \(T^2\). I do not know whether it will be possible to construct a manometer for measuring these electron pressures, but in any case it would have to be made extremely strong.
At first glance, all that has been set forth above is in sharp contradiction with experiment. In fact, the measurement of thermionic currents seems to show that the work function which prevents electrons from escaping from a metal is far from reaching 6 volts (and in some metals is even 2 volts). How, then, do the rapidly moving electrons remain in the metal? It turns out that the new theory, while increasing the kinetic energy of the electrons, at the same time increases the obstacle which they have to overcome. Here we encounter the first of the recent experiments confirming the new statistics.
EMISSION OF THERMOELECTRONS
According to the simplest ideas, the thermionic current is produced by those electrons for which \(\frac{1}{2}mu^2>W_a\), where \(u\) is the component of the velocity normal to the bounding surface of the metal, and \(W_a\) is a constant interpreted as the work function, i.e., as a retarding jump of potential at the boundary of the metal. In other words, the kinetic energy of the thermoelectrons must be so great that they can “jump over” the wall surrounding them.
It is obvious that any emission of thermoelectrons disturbs the distribution of the electron gas inside the metal, since
it represents an uncompensated loss of electrons. From the standpoint of thermodynamics, it is more natural to consider the case in which the loss of electrons is compensated by their simultaneous influx into the metal. But in that case the current force, like the force of a heat flow between two bodies of the same temperature, could not be measured. In any event, by assigning to the internal electrons a Maxwell or Fermi distribution, we may be committing an error (unless the emission is infinitely small). In real cases this error is apparently insignificant.
Thus the simplest theory of the thermoelectronic current can be expressed by the equation
\[ i = e\left(\frac{m}{h}\right)^3 G \int_{u=u_0}^{\infty} du \int_{-\infty}^{\infty} dv \int_{-\infty}^{\infty} dw\, \frac{u}{A e^{\frac{1}{2}m\left(u^2+v^2+w^2\right)/kT}+1}. \tag{73} \]
This expression serves as the mathematical formulation of what was said at the beginning of this paragraph. In writing it, we assumed that the Fermi distribution remains valid even for electrons that have already flown out of the metal. The factor \(e\) denotes the charge of the electron; \(i\) denotes the density of the thermoelectronic current in electrostatic units. The factor \(m^3\) entered because, instead of momenta, we took the components of velocity \(u, v, w\) as the independent variables. The quantities \(u_0\) and \(W_a\) are connected by the equation
\[ W_a=\frac{1}{2}mu_0^2. \tag{74} \]
In carrying out the integration, \(s\) must of course be replaced by \(\frac{1}{2}m\left(u^2+v^2+w^2\right)\). Substituting for \(A\) the first approximation from equation (67) and using the symbol \(W_i\), introduced in equality (71), we obtain:
\[ i = e\left(\frac{m}{h}\right)^3 G \int_{0}^{\infty} \int_{0}^{\infty} \int_{0}^{\infty} \frac{u}{e^{\frac{1}{kT}\left(\varepsilon-W_i\right)}+1} \,du\,dv\,dw. \tag{75} \]
The integration is carried out very easily if the second term of the denominator is neglected in comparison with the first. For all
values of \(\varepsilon\) smaller than \(W_i\), such a rejection would be illegitimate, since for them the second term is many times larger than the first. But if \(W_a\) is considerably greater than \(W_i\), i.e. if \(\dfrac{W_a-W_i}{kT}\) is positive and considerably greater than unity, then all thermoelectrons belong to the very upper part of the curve, in which the dominant role is played by the first term. Replacing the integrand by \(u e^{\frac{E}{kT}-\frac{W_i}{kT}}\), we obtain, from well-known formulas,
\[ i=\frac{2\pi e m a}{h^3}(kT)^2 e^{-\frac{W_a-W_i}{kT}} . \tag{76} \]
It follows from this that if one plots along the abscissae the quantity \(\dfrac{1}{T}\), and along the ordinates the expression \(\lg i-2\lg T\), then one should obtain a straight line (provided only that \(W_a\) does not depend on the temperature). Experiment confirms this conclusion of the theory. The slope of the straight line varies from metal to metal, and its magnitude—according to our theory equal to \(\dfrac{W_a-W_i}{kT}\)—shows that the difference \(W_a-W_i\) fluctuates within the limits from one to 5–6 volts. Thus our approximate method proves to be quite justified even at white-heat temperature. It is of interest to compare the result obtained with what the classical theory gives.
Application of the Maxwellian distribution (with substitution, instead of \(A\), of the value (64)) leads to the formula
\[ i=\frac{en}{\sqrt{(2\pi m)}}(kT)^{\frac{1}{2}} e^{-\frac{W_a}{kT}} . \tag{77} \]
On checking this formula, it turns out that the experimental points fall well on the straight line \(\lg i-\frac{1}{2}\lg T\), just as on the straight line \(\lg i-2\lg T\). Indeed—as is known to every physicist who has dealt with thermoelectrons—a function of the form \(e^{-\frac{c}{T}}\) changes so rapidly with a change of \(\dfrac{1}{T}\) that the form of the curve practically does not depend on whether some constant stands before it as a factor.
or a small degree \(T\). Therefore, even if one is convinced that \(W_a\) does not depend on \(T\), the experimental curves do not directly make it possible to choose between the classical and the new theories. But classical theory assumed that the slope of the straight line is determined by the expression \(\dfrac{W_a}{k}\). Consequently, if the Fermi distribution is correct, this means that physicists had all along been taking values of the work function that were too low. In experimentally studying curves that agree equally well with formulas (76) and (77), and in determining the constant appearing in the exponent, usually denoted by \(-\dfrac{b}{k}\), they took \(b = W_a\), whereas according to the new statistics \(b = W_a - W_i\). The difference between these two equalities is extremely substantial, since, if the number of free electrons in a metal is equal to the number of atoms, then \(W_i\) is equal to approximately 6 volts, i.e. greater than \(b\) itself.
Is there another method for determining the work function, besides investigating the curve of the thermoelectronic current? If such a method existed, it would not only make it possible to choose between the old and the new statistics, but also—in the case where it turned out that \(W_a > b\), i.e. that the new statistics is correct—would allow one to experimentally determine the difference \(W_a - b = W_i\), and with it the arbitrary constant \(n\), the only unknown quantity in the expression for \(W_i\). In other words, we would obtain the possibility of experimentally determining the number of electrons in a unit volume of the electron gas.
It turns out that such a direct and immediate method of determining the work function is provided by the phenomena of electron diffraction in crystals. In those phenomena where negative electrons behave like waves, the work function enters into the expression for the refractive index. The refractive index of a metal can be determined from the diffraction pattern produced by electrons incident upon it. A detailed investigation of this diffraction pattern was carried out by Davisson and Ger—
by measuring on a nickel crystal. The value of the work function obtained by them greatly exceeds the value of the thermoelectronic constant \(b\)—so much so that for \(n\) one obtains a value approximately twice as large as the number of atoms of the metal.^1 To be sure, much in their results is still unclear; in particular, it is not clear why the values of the work function derived from the refractive index depend on the velocity of the electrons.
The coefficient in the expression \(T^2 e^{-\frac{W_a-W_i}{kT}}\) on the right-hand side of equation (76) contains only universal constants (including \(G\)) and therefore has the same value for all metals. (This fact, as Richardson already showed 20 years ago, can be derived directly from the fundamental principles of thermodynamics, without making any assumptions concerning the distribution of electrons by velocities.) Its value
\[ \frac{2\pi k^2 e m}{h^3}\,G \]
differs only by the factor \(G\) from the value obtained by Dushman,^2 whose numerical magnitude, in ordinary units, is \(60.2\) amperes per \(\text{cm}^2\,\text{deg}^2\). For a number of metals this value agrees well with the experimental data. Hence one might conclude that \(G=1\), but such a conclusion would destroy the above-presented theory of paramagnetism.
It remains to assume that half of those electrons which approach the boundary surface with a velocity sufficient for escape nevertheless are reflected. Under such an assumption, our formula contains the factor \(1-r\), where \(r\) is the reflection coefficient, in the present case equal to \(\frac{1}{2}\), compensating for the influence of the factor \(G\), to which the value 2 is usually assigned. By assigning other values to \(r\), one can explain why experiment sometimes gives, for the constant entering equation (76), values smaller
^1 L. Rosenfeld, E. E. Witmer, Z. Physik 49, 534, 1928 (l. c. infra). For 6 other metals (Al, Cr, Cu, Ag, Au, Pb), they use values of the refractive index given by Rupp and calculate the corresponding \(n\).
^2 This factor is usually denoted by the letter \(A\); however, in the present article this symbol is used in another meaning.
than 60.2 amperes, lying between 60.2 and 120.4. True, sometimes this constant is also considerably greater than 120, so that here, apparently, not everything has yet been clarified.^1 Fowler and Nordheim tried to calculate the reflection coefficient by the methods of wave mechanics. They arrived at a number of interesting results in the field of thermoelectronic currents, cold discharge, and the photoelectric effect. To obtain these results, however, the new statistics alone is not sufficient; one must also make use of the theory of refraction and reflection of electron waves. Since the latter goes beyond the scope of the present article, I shall indicate here only some of the main points.
Up to now we have regarded a metal as an equipotential region surrounded by a surface on which the potential undergoes a sharp jump, followed again by an equipotential region of external space. According to Fowler and Nordheim, and also according to Schottky and certain other investigators, a metal is an equipotential region surrounded by a surface beyond which there is a force field; moreover, the force of the latter is a function of the distance from the surface. From this point of view, the quantity \(W_i\) is the integral of the field force from the surface of the metal to infinity. The electrons arriving from within fall upon this surface. The fraction of electrons that passes through the field depends on their kinetic energy and on the form of the curve expressing the dependence of the field force on the distance (and not only on its integral). The mean value of this quantity for electrons of different velocities—it depends, obviously, on the distribution of the electrons by velocities—is the above-mentioned coefficient \(r\). For some of the simplest curves it can be calculated. Such, in general outline, is the path for establishing—
^1 If one uses an empirical equation of the form \(i = aT^n e^{-b/T}\), then, by choosing in a suitable way the dependences of \(W_i\) and \(r\) on the temperature, one can, for any \(a\) and \(b\), make it agree with the preceding theory. But in doing so, as Fowler observed, “the number of hypotheses not amenable to verification becomes too great for fruitful discussion.”
K. DARROW
Determination of the influence of surface conditions on the emission of thermoelectrons and the cold discharge.¹
Cold Discharge²
Let us now suppose that between the given metal and an oppositely charged electrode there is a difference of potentials, and that near the metal the field strength is very large. Let us further assume that this field penetrates right up to the surface of the metal. Then throughout the entire external space we shall have, besides the above-mentioned “internal” field, also an “external” field; and the form of the curve expressing the dependence of the field strength on distance will change somewhat. Knowing the initial form of this curve, it is easy to determine what change is introduced into it by the appearance of the external field. In some of the simplest cases it is possible to calculate also the corresponding change in the coefficient \(r\), i.e. the magnitude of that additional electron flux which is caused by the external field. This current is called the “cold discharge.”
Nordheim assumed that, in the absence of an external potential difference, the field strength obeys the usual inverse-square law (i.e. that the “internal field” serves, as it were, as the “electrical image” field of a certain unit charge in our conductor). Under this assumption, for the current strength of the cold discharge \(i\) the following approximate formula is obtained:
\[ i = c' S F^2 e^{-c''/F}, \tag{78} \]
where \(F\) is the external field strength, \(S\) is the surface area of the metal, and \(c'\) and \(c''\) are constants whose values are determined by the constant \(W_i\). The experimental data at our disposal (a detailed account of them is given in Nordheim’s article in Physikalische Zeitschrift)² show that the dependence between the current and the field strength is indeed given by a formula of the form (78), but that the values predicted by the theory—
¹ On the theory of Fowler and Nordheim, see the article by R. Fowler, Uspekhi fizicheskikh nauk 10, 135, 1930. — Ed.
² On the cold discharge, see Uspekhi fizicheskikh nauk 8, 533, 1928. — Ed.
values of \(c'\) and \(c''\) exceed their actually observed values many times over (for \(c'\) the discrepancy is even by an order of magnitude, and for \(c''\)—by several tens of times). One may try to explain this discrepancy by saying that, in the discharge, only small areas of the metal surface actually take part, constituting a negligible fraction of its total area, and that on these small areas the field strength has a considerably greater value than it would have if the metal were homogeneous everywhere. Under this assumption, the ratio of the observed value to that predicted for \(c'\) is equal to the fraction of the total area covered by “active” regions; and for \(c''\) it is equal to the reciprocal of the factor by which the field strength must be multiplied.
The shortcoming of such an explanation is that it is too simple. Until we have direct data on the magnitude of the active regions and of the field strength prevailing on them, equation (78) must be regarded as an equation with two arbitrary constants, not very convenient for testing those basic assumptions on which it is based. An investigation of the experimental data carried out by Nordheim showed, however, that the ratio between the predicted and the observed value for \(c''\) always fluctuates between 10 and 20, while for \(c'\) it is close to \(10^{-10}\) (at least for tungsten). This regularity, to a certain extent, may serve as confirmation of our theory.
Photoelectric Effect
According to the usual ideas, the elementary photoelectric effect occurs in the following way: a quantum of light, falling upon a metal, gives up all its energy \(E_0\) to some electron, initially at rest, and knocks it out of the metal. On emission, the kinetic energy of the electron is decreased by \(W_a\) (or by an even larger amount, since part of it may still be lost while the electron is moving inside the metal). A more exact consideration, according to which the electron initially is not at ...
is at rest, but belongs to an electron gas governed by the laws of classical statistics, leads practically to the same results, since in this case too the initial energy of the electron is negligibly small in comparison with the energy which it receives from the quantum. If, however, the electron gas obeys the new statistics, then it contains electrons whose kinetic energy is close to \(W_i\). But, on the other hand, if the new statistics is correct, then the loss of kinetic energy by \(W_i\) is greater than had previously been assumed. As a result, both with the old and with the new statistics we obtain, for the maximum kinetic energy of the electrons knocked out by quanta of frequency \(\nu\), Einstein’s equation:
\[ E_{\max}=E_0+W_i-W_a=h\nu+\mathrm{const} \tag{79} \]
with only this difference: in the second case the additive constant has the form \(W_i-W_a\), while in the first case it is simply \((-W_a)\). In both cases, this additive constant must coincide with the thermoelectronic constant \(b\).
The new theory has one essential difference (which, perhaps, is an advantage) from the old: it presupposes the existence of a sharply bounded maximum of kinetic energy. More precisely, according to the new theory, the curve of the energy-distribution function at \(E=E_{\max}\) undergoes a sudden jump from some finite value to zero and forms at this point a sharp angle with the axis of abscissas. According to classical statistics, however, based on the Maxwellian distribution, the approach of the curve to the axis of abscissas must be asymptotic in character. Moreover, the new statistics makes it possible to draw definite conclusions concerning the form of the curve for values of \(E\) less than \(E_{\max}\). A large amount of experimental data on this question was obtained by Ives1 and his collaborators, but their interpretation is complicated by the fact that the electron partially loses energy before leaving the metal, the magnitude of which is unknown.
PARAMAGNETISM OF THE ELECTRON GAS
The magnetic susceptibility of the electron gas was calculated by Pauli even before the appearance of Sommerfeld’s work on specific heats. We did not, however, consider it advisable to follow the chronological order, since one additional complicating circumstance enters into Pauli’s calculations.
This circumstance is connected with that basic assumption which explains why an electron gas can in general be a magnet—that is, with the assumption of “electron magnetism.” It is possible that, in speaking here of an “assumption,” I am being excessively scrupulous, since the existence of electron magnetism may be regarded as proved by experiments with the gyromagnetic effect and by the successes of the “rotating electron” in the explanation of spectra. These experiments lead to the following value of the magnetic moment of the electron \(\mu_0\):
\[ \mu_0=\frac{eh}{8\pi m_0 c}, \tag{80} \]
where \(m_0\) is the mass of the electron in the state of rest. Further, they show that when the electron is in a magnetic field, the vector of its magnetic moment must be parallel or antiparallel to the field. Let us denote the angle between the moment of the electron and the field by \(\theta\); according to the foregoing, \(\theta\) is equal to \(0\) or \(\pi\).1 As is known, when a magnet \(M\) makes an angle \(\theta\) with a magnetic field \(H\), its “additional magnetic energy” is \(-MH\cos\theta\).2 Thus in an electron gas, when a magnetic field \(H\) is switched on, the energy of each electron is increased or decreased by
\[ \lambda=\frac{ehH}{8\pi m_0 c}. \tag{81} \]
As we have already said above, instead of dealing with cells of phase space, one can, at least in some cases, make use of the representation in terms of “possible values of the energy.” In the present problem it is precisely this latter representation that proves most convenient. Using it, one may say that the appearance of a magnetic field in an electron gas causes each possible value of the energy to split into two new ones. In fact, to each allowed energy \(E\) there correspond two possible energies: \(E+\Delta\) and \(E-\Delta\), or, more briefly: \(l+m\Delta\), where \(m\) may be equal to \(+1\) or \(-1\).
Pauli assumed that the most probable distribution of electrons over this double series of possible energies is determined according to the new statistics, including Fermi’s postulate, which in the present case says that to each of the allowed values of the energy there corresponds at most one electron.
In the preceding paragraph we denoted the number of those cells, or allowed values of the energy in the layer \(S\), which contain \(i\) electrons, by \(Z_{is}\), and the total number of allowed values of the energy in the layer \(S\) by \(Q_s\). According to what has been said above, in a magnetic field we have \(Q_s\) allowed energies shifted upward by \(\Delta\) in comparison with the original ones, and the same number of allowed energies shifted downward by \(\Delta\) in comparison with the original ones. Let us denote the number of those cells in the “shifted upward” group which contain \(i\) electrons by \(Z_{is,+1}\), and the corresponding number for the “shifted downward” group by \(Z_{is,-1}\). We now consider the distribution:
\[ Z_{ism}=a_{sm} e^{-iB-i(\varepsilon_s+m\Delta)/kT} \qquad \begin{matrix} m=+1,-1\\ i=0,1 \end{matrix} \tag{82} \]
It is easy to show that this distribution is an equilibrium one. Indeed, if formula (82) is substituted into the expression for the probability \(W^*\) according to Bose’s formula, and if, as usual, \(S=k\lg W^*\) is put, then it turns out that when the energy of the assembly undergoes a variation \(\delta E\)—under the condition that the total number of cells and the total number of particles remain constant—the first variation of the entropy is equal to \(\dfrac{\delta E}{T}\), which serves as the characteristic criterion ...
sign of the equilibrium distribution. The proof of this assertion, which can be carried out quite analogously to the corresponding arguments in the preceding cases [see, for example, formula (47)], as well as the determination of the constants \(a_{sm}\), we leave to the reader. Finally, our distribution takes the form:
\[ Z_{is+1}=\frac{1}{e^{-B+\frac{\varepsilon_s+\Delta}{kT}}+1}, \]
\[ Z_{is-1}=\frac{1}{e^{-B+\frac{\varepsilon_s-\Delta}{kT}}+1}, \]
\[ Q_s=\frac{2}{(\pi)^{1/2}}\frac{V}{h^3}(2\pi m)^{3/2}(\varepsilon_s)^{1/2}. \tag{83} \]
These formulas may be regarded as descriptions of two electron gases, one of which consists only of parallel poles, and the other—only of magnetic poles antiparallel to it. In reality we always have a mixture of these two gases. Thus, for example, the quantity \(Z_{is+1}\) represents the number of parallel magnets whose energies are nothing other than the initial allowed values of the energy of the layer \(S\), shifted downward by \(\Delta\). To be sure, owing to the shift, these magnets themselves no longer lie in the layer \(S\); in other words, their energies no longer lie between \(\varepsilon_s\) and \(\varepsilon_s+d\varepsilon_s\), but between \((\varepsilon_s-\Delta)\) and \((\varepsilon_s-\Delta+d\varepsilon_s)\). But in order to simplify the calculations it is more convenient to take the integral over the initial—unshifted—values of the energy. In this case the total number of magnets in the “parallel” gas will be given by the equation:
\[ N_1=\frac{2}{(\pi)^{1/2}}\frac{V}{h^3}(2\pi m)^{3/2} \int_0^\infty \frac{\varepsilon^{1/2}\,d\varepsilon} {e^{-B+\frac{\varepsilon-\Delta}{kT}}+1}. \tag{84} \]
For brevity let us denote the constant before the integral by \(L\) and introduce the following abbreviation
\[ \varphi(u)=\int_{\varepsilon=0}^{\infty} \frac{\varepsilon^{1/2}\,d\varepsilon} {e^{u+\frac{\varepsilon}{kT}}+1}. \tag{85} \]
Then, expanding \(N_1\) in powers of the variable \(\dfrac{\Delta}{kT}=\dfrac{\mu_0H}{kT}\), we shall have:
\[ N_1=L\left[\varphi(B)-\frac{\mu_0H}{kT}\varphi'(B)+\text{terms of higher order in }H\right]. \tag{86} \]
Similarly, for the total number \(N_2\) of electron magnets in the “antiparallel” gas, we obtain the formula:
\[ N_2=L\left[\varphi(B)+\frac{\mu_0H}{kT}\varphi'(B)+\text{terms of higher order in }H\right]. \tag{87} \]
The total magnetic moment of the “parallel” gas is equal to \(N_1\mu_0\), the total magnetic moment of the “antiparallel” gas to \(-N_2\mu_0\), the first of them being directed along the field, and the second against the field. Consequently the magnetic moment of the entire collection of electrons is \((N_1-N_2)\mu_0\). We shall consider only such values of \(H\) for which, in the expansions of \(N_1\) and \(N_2\), it is possible to restrict ourselves to the first two terms. In this approximation, the resulting magnetic moment is proportional to \(H\), i.e. the quotient of its division by \(H\)—the magnetic susceptibility of the electron gas \(\chi\)—is a constant. Experiment shows that in almost all paramagnetic substances the susceptibility in fact does not depend on \(H\), even for very large values of the latter. Therefore, the error connected with the omission of higher-order terms in the expansions (86) and (87) is in all probability insignificant. Our approximate formula for \(\chi\) has the form
\[ \chi=\frac{(N_1-N_2)\mu_0}{H}=-\frac{2L\mu_0^2}{kT}\varphi'(B)=-\frac{N\mu_0^2}{kT}\frac{\varphi'(B)}{\varphi(B)}. \tag{88} \]
It now remains only to make the last step—to express the hitherto undetermined constant \(B\) through the total number of particles of the collection \(N\). This number is \(N=N_1+N_2\). Neglecting terms of higher order in \(H\), we have:
\[ N=2L\varphi(B). \tag{89} \]
Equation (89) essentially coincides with equation (63), which relates the Sommerfeld constant \(A\) to \(N\), since, as before, \(e^B=\dfrac{1}{A}\). In order that this coincidence be complete, it is necessary to put \(G=2\) in formula (63),
as we did above (for this purpose we introduced the factor $G$).
The further calculations proceed as follows: in the right-hand side of equation (66) we set $\lg A = -B$, differentiate it with respect to $B$; then into the derivative we substitute the value of $B$ obtained by equating the right-hand side of equation (56) to the quantity $N$, i.e. the value given in (67), and finally we substitute the resulting result into (88). As a result, for $\chi$ one obtains the following expression:
\[ \chi = 12 \left(\frac{\pi}{3}\right)^{2/3} \mu_0^{\,2} n^{1/3} m_0 h^{-2}. \tag{90} \]
Let us now turn to the experimental results. Can it be assumed that the magnetic susceptibility of some metal is determined exclusively by its free electrons? Here we encounter the same uncertainty as in the problem of specific heats. The process of magnetization of an ordinary paramagnetic metal is made up of processes of threefold character: the orientation of electrons, the orientation of atoms, and, finally, that distortion of the orbits of intra-atomic electrons which gives rise to diamagnetism. Apparently, no theory can determine exactly what fraction of the total magnetic moment is supplied by each of these three effects separately. It can, however, be stated with certainty that in the alkali metals the second effect is absent. Indeed, spectroscopic data show with complete obviousness that the magnetic moment of an ion of any alkali metal—that is, of an atom without one outer electron—is equal to zero. Therefore, if each atom of an alkali metal were to lose its outer electron, then in a magnetic field there would be no orientation of the ions, and the number of electrons in the electron gas would be equal to the number of atoms. The values of $\chi$ calculated on this assumption should somewhat exceed the actual values of the susceptibility, since the diamagnetic effect, which we have not taken into account, is opposite in sign to the paramagnetic one and therefore partially neutralizes it. The order of magnitude of this diamagnetic effect we can
determine; moreover, it turns out that it is the larger, the greater the atomic number of the metal.
The value of \(Z\) given above is valid for absolute zero; for higher temperatures one must use not one, but the first two terms of the expansion (66). It turns out, however, that the correction thereby obtained is insignificant. Like the mean energy and the pressure, the magnetic susceptibility of the electron gas is practically the same throughout the entire temperature range extending from absolute zero to room temperature and considerably above. Experiment shows that the susceptibility of the alkali metals indeed does not depend on temperature. This fact seemed so surprising that, in essence, it served as the first impetus for the investigations of Pauli set forth above. In fact, if the electron gas obeyed the laws of classical statistics, then (with the number of electrons equal to the number of atoms) the susceptibility of the metal would increase as the temperature decreased and would attain enormous values at absolute zero.
According to the experimental data available to Pauli when he published his theory, the observed values of the permeability for potassium and sodium were somewhat lower, and for rubidium and cesium appreciably higher, than the calculated ones. This discrepancy could be explained either by a diamagnetic effect or by some flaw in the theory. A recent paper by Canadian investigators, which appeared quite à propos, gave considerably better agreement with the theory. This can be seen at least from the following table (taken from the article by E. C. Stoner, in which the original sources are also indicated):
| Na | K | Rb | Cs | |
|---|---|---|---|---|
| Pauli theory . . . | 0.66 | 0.52 | 0.49 | 0.45 |
| Experiment: | ||||
| Mac-Lennan et al. . | 0.61 | 0.42 | 0.31 | 0.42 |
| Lön . . . . . . . . | 0.65 | 0.54 |
(All numerical values must be multiplied by \(10^6\).)
It is remarkable that the agreement turns out to be too go-
simpler. The expected diamagnetic effect should have been greater than the small discrepancies between the experimental and theoretical values that are observed for sodium, potassium, and cesium. It is possible that, in connection with this, a more detailed theoretical development of the diamagnetic effect will be required. Returning once more to the question of the coefficient \(G\), let us recall that the introduction of the coefficient \(G=2\) into equality (61) is, in essence, equivalent to the assumption that the electron gas is a mixture of two assemblies of particles completely independent of one another, each of which separately obeys Fermi statistics. Such a representation is somewhat unusual, but nevertheless one cannot do without it.
Theory of Electrical Conductivity
The new theory of electrical conductivity, developed by Houston and Bloch, is based on the wave theory of electricity, according to which the interior of a metal is filled not with moving particles, but with standing waves, the number of nodes and antinodes of the latter being equal to the number of free electrons. In other words, its foundation is not Fermi statistics alone, but Fermi statistics plus wave theory. Of course, if one considers that Fermi statistics implicite includes wave mechanics and conversely, then the latter reservation becomes inessential. But the corpuscular theory of negative electricity is 1) the most customary, 2) the most convenient for the description of most of those phenomena in which electrons figure. Therefore, in expounding the new theory of electrical conductivity, I shall use mainly the language of corpuscles, although later I shall have to introduce an assumption which is, in essence, equivalent to the transformation of corpuscles into waves.
Let us imagine a metallic rod whose length, measured along the \(x\)-axis, is equal to \(d\). Suppose that a potential difference \(V\) is applied to the ends of this rod, so that along the \(x\)-axis there acts a field \(E=\frac{V}{d}\). If the electrons moved in the metal perfectly freely, then
each of them, arriving at the negative end of the rod with no initial velocity, would immediately begin to move in the direction of the positive end and would reach it with kinetic energy \(eV\) and the corresponding velocity \(\left(2\frac{eV}{m}\right)^{1/2}\). Of course, in reality this is not so. When along some wire we have a potential gradient, then at all its points the same heat evolution is observed, and no signs are seen that at one end the electrons move faster than at the other.
It is therefore necessary to assume that the free motion of an electron is interrupted at definite intervals of time and that at each such interruption the electron loses that kinetic energy and velocity which it had acquired under the influence of the field since the preceding interruption. More precisely, on the average the increase in kinetic energy and velocity during the time elapsing between interruptions must be compensated by their loss at the interruptions.
In the corpuscular theory the role of these interruptions is played by the impacts or collisions of electrons with atoms. The loss of all the acquired velocity at each collision can be explained by saying that our electron, striking an atom, rebounds in the direction perpendicular to the field. True, such an explanation seems somewhat artificial. But the same result can be reached if electrons and atoms are regarded as elastic spheres and it is assumed that in collisions the atoms remain motionless. Indeed, it is easy to show that under this assumption the angle between the direction of motion of the electron before and after the collision is, on the average, equal to \(90^\circ\). In other words, on the average the electron rebounds just as often forward as backward, quite independently of the original direction of its motion. A rigorous proof of this assertion will be given below.
Here we encounter one difficulty, which I wish now to mention, although I cannot yet indicate by what means it can be removed. In the subsequent calcu-
collisions we shall assume that at the end of each free path the electron loses not only its velocity, but also all the kinetic energy acquired by it under the influence of the field during this path. But, in a collision with an infinitely heavy sphere, there is no loss of kinetic energy at all. In a collision with a sphere whose mass is equal to the mass of an atom, the electron loses kinetic energy, but does not wholly lose its velocity. The theory of this latter case was developed by Compton and Gerth in connection with the problem of the electrical conductivity of gases. It can also be applied to the problem of the electrical conductivity of metals, although here it gives no especially good results in comparison with other hypotheses.
Thus, at each collision the electron loses the velocity acquired by it since the preceding collision and returns to its initial state. Denote the mean interval of time between collisions by \(t_0\). Since the acceleration of the electron is equal to \(\frac{eE}{m}\), the velocity acquired by it at the end of the period \(t_0\) is \(\frac{eEt_0}{m}\), and its mean value is equal to half this quantity. Let us note that this acquired velocity by no means represents the entire velocity of the electron. On the contrary, the mean velocity of thermal motion, which we shall denote by \(v\), is many times greater than that additional velocity which the electron can acquire in an ordinary (not very large) electric field during the time of a free path. The whole action of the field reduces to the fact that it slightly bends the rectilinear trajectories of the electrons in the intervals between collisions. This is so even in classical statistics, not to mention the new one. Denote the mean free path of the electron by \(l\). Then \(t_0=\frac{l}{v}\) and the mean acquired velocity is \(\frac{1}{2}\frac{eEl}{mv}\). Correspondingly1
the corresponding current density is equal to the product of this quantity by the number of electrons in a unit volume and by the charge of each electron. Thus, for the current density at \(E=1\), which, by definition, is nothing other than the conductivity \(\sigma\), we obtain the formula:
\[ \sigma=\frac{1}{2}\frac{ne^{2}l}{mv}. \tag{91} \]
The constant \(l\)—the mean free path—is the third arbitrary constant of the electron theory of metals.
I fear that all the preceding arguments may seem somewhat out of date; nevertheless, in essence they contain the whole corpuscular theory of electrical conductivity. The hypothesis of elastic spheres is of an exclusively auxiliary character and serves only as a concrete expression of the fundamental fact that the electrons in a metal, being under the influence of an external field, alternately acquire and lose velocity. The mean free path is that average distance over which a continuous increase of velocity takes place.
The first test of formula (91) is provided by the dependence of resistance on temperature. At one time it was thought that here, as in the problem of specific heats, the old theory of the electron gas comes into conflict with experiment.
Indeed, experiment shows that the resistance
\[ \rho=\frac{1}{\sigma} \]
of every metal changes rapidly with temperature. For most metals this change is proportional to \(T\), and at very low temperatures it is even faster (not to mention the remarkable phenomenon of superconductivity). According to equation (91), the cause of this regularity must be sought in the nature of the dependence of the quantities \(v\), \(n\), and \(l\) on temperature (since \(e\) and \(m\) are universal constants).
As is known, according to the laws of classical statistics \(v\) is proportional to \(T^{1/2}\). Thus, in order that \(\rho\) be proportional to \(T\), the product \(nl\) must be propor-
tional to \(\dfrac{l}{T^{1/2}}\). Fermi statistics imposes still stricter requirements, since in it the mean velocity of thermal motion practically does not depend at all on temperature, and therefore all the “responsibility” for the behavior of \(\rho\) falls on the quantity \(nl\). In this respect the first step of the new statistics is a step backward.
Can the temperature dependence sought be reduced to a change in the quantity \(n\)? Under this assumption, \(n\) would have to decrease as the temperature rises. It seems natural that \(n\) should increase with temperature, since free electrons may appear upon ionization of atoms, and the cause of ionization is thermal motion. The hypothesis that \(n\) decreases with temperature appears very strange, although with its aid Utermann achieved some success.
Thus, part of the “responsibility”—even all the responsibility, if one uses the new statistics and, in addition, regards \(n\) as constant—must fall on the mean free path. But here the hypothesis of elastic collisions collapses. It gives for the mean free path a value that is independent of temperature,\(^1\) if the thermal expansion of the metal is neglected. At first sight the matter appears hopeless.
In order to escape this difficulty, Wien put forward the following idea. If one assumes that an electron sometimes passes freely through an atom and that the probability of such passage is the smaller the more energetic the oscillations of the atom, then the mean free path should decrease with increasing temperature. In other words, we would obtain the desired result.
Modern physics approaches the question in almost the same way. We still consider that the probability of—
\(^1\) According to the usual formula of the kinetic theory of gases
\[ l=\frac{1}{N\pi R^{2}}, \]
where \(N\) is the number of resting spheres per unit volume, and \(R\) is the sum of the radii of the resting and the moving sphere.
the repulsion—or, more precisely, the scattering—of an electron by an atom increases with increasing amplitude of the latter’s oscillation. We now see the reason for this in the fact that, during oscillation, the atom moves away from its equilibrium position in the crystal lattice, i.e. changes its position relative to neighboring atoms. The probability of scattering depends not only on the presence of an atom in the electron’s path, but also on its position relative to the other atoms of the crystal.
Such an assumption is, in essence, a compromise between the corpuscular and wave theories. Indeed, on what grounds do we speak of the wave nature of a beam of light incident on a grating, or, say, of a beam of X-rays incident on a crystal? The chief evidence in favor of this wave hypothesis is the fact that the scattering (diffraction) of a beam by an aggregate of regularly arranged slits in a grating, or of atoms in a crystal, differs essentially from scattering by a single slit or a single atom. Thus, for example, in the case of a regularly arranged aggregate of slits, light is not scattered at all in certain directions, whereas in the case of a single slit these directions are in no way distinguished from others. To explain these facts one has to resort to the idea of interference of waves. In particular, it is assumed that in the “dark bands” of the diffraction pattern, the waves scattered by different slits or particles are in opposite phases and therefore mutually annihilate one another. In corpuscular language this may be expressed by saying that a beam of light is a stream of particles scattered by atoms, and that the laws of this scattering are determined by the arrangement of the atoms in a regular lattice; in particular, for example, the probability that a particle will be thrown off in the direction of one of the shadow bands is equal to zero.
I do not think that such a compromise constitutes a complete victory for the wave theory, although the development of modern theoretical physics apparently proceeds precisely in this direction. In any case, if we wish to express the phenomena of diffraction in crystals in the language of the corpuscular theory
(regardless of whether we are dealing with light waves or with negative electricity), it is necessary to introduce the concept of the probability of scattering, i.e., to assume that the mean free path depends on irregularities in the arrangement of the atoms.
This principle takes on an especially simple and remarkable form if one is dealing with a beam whose wavelength is greater than the distance between the lattice points. In this case, when the scattering particles are arranged regularly, scattering does not occur at all. Although each particle individually scatters waves, they pass through the entire lattice without change. In corpuscular language this means that, as the structure of the atoms becomes ideally regular, the probability of deflection of each particle tends to zero, and the mean free path tends to infinity.
It follows from this that the resistance of an ideal crystal is equal to zero when all the atoms are at rest at the lattice points, and gradually increases as the temperature rises. The law of this increase can be determined if two things are known: 1) the relation between the scattering of waves by the particles of the lattice and the amplitudes of oscillation of these particles about their equilibrium positions, and 2) the relation of these amplitudes to temperature. The second of these questions is the subject of the theory of the specific heats of solids, developed chiefly by Debye. The first question was investigated in detail by Debye and other physicists in connection with the problem of the scattering of X-rays by crystals. Transferring the results they obtained to the theory of the diffraction of electron waves, Houston showed that, over a wide range of temperatures, the resistance of an ideal crystal varies in proportion to the absolute temperature.
In order to determine not only the dependence of resistance on temperature, but also the numerical value of the resistance itself, it is necessary to calculate the mean free path of the electron. According to the old corpuscular theory, this path is determined by the sizes of those atoms with which the electrons collide; according to the wave theory, however,
it depends on the scattering power of each individual atom. (The determination of this scattering power is a problem of the new mechanics. Proceeding in this way, Houston obtained, for some metals, good agreement with experiment.
Another possibility for the appearance of irregularities in the crystal lattice consists in replacing some part of the atoms of the given element, scattered disorderly through the lattice, by atoms of some other element. We encounter such a phenomenon in certain alloys, the so-called “solid solutions.” As is known, the resistance of an alloy exceeds the resistance of the element predominating in it. Nordheim showed that the observed dependence of the resistance on the percentage composition of the alloy can be explained on the basis of the diffraction theory of electrical conductivity.¹
Thus we have established that the concept of the mean free path can be expressed in terms of wave theory and that, on the basis of this theory, one can determine the dependence of the free path on temperature. Let us now show how, on the basis of the corpuscular picture, one can construct a theory of thermal and electrical conductivity and of the thermoelectric effect in crystals.
We shall use the so-called method of the perturbed distribution function, first proposed by Lorentz. Those Maxwell and Fermi functions with which we have dealt up to now are isotropic in character, i.e., the components of the velocity \(\xi, \eta, \zeta\) enter into them only in the form of the combination \(\xi^{2}+\eta^{2}+\zeta^{2}(=v^{2})\).² These “standard” functions are applicable, obviously, only in the case when the temperature and the potential are everywhere constant. For the same metal in which there is an electric field
¹ The idea that the “free path of the electron” is equal to the distance from one irregularity in the crystal lattice to the next was advanced even before the appearance of the wave theory of negative electricity.
² In what follows I shall, as is customary, use the components of velocity instead of the components of momentum.
or the temperature gradient—and also for an alloy whose chemical composition changes from point to point—are unsuitable. If the direction of the gradient of the variable quantity—whether the electric potential, temperature, or something else—is taken as the \(x\)-axis, then one must expect that \(\xi\) enters into the distribution function differently from \(\eta\) and \(\zeta\).
It is easy, however, to show that, as a rule, these deviations from the standard function are comparatively insignificant. Therefore Lorentz assumed that in the presence of a potential gradient parallel to the \(x\)-axis the distribution function has the form:
\[ f = f_0(v) + \xi g(v), \tag{92} \]
where \(f_0(v)\) is the standard distribution function (according to Lorentz, coinciding with the Maxwellian), and \(\xi g(v)\) is a small additive term. The function \(g\) must be such that \(f\) remains constant, despite collisions of electrons with atoms. Lorentz showed that in each of the three cases considered below there exists its own \(g\).
We shall begin with the simplest case of a homogeneous metal situated in a constant electric field at constant temperature. In this case the distribution function must be the same at all points. We shall dwell on its calculation in detail, although the final formula will differ only slightly from formula (91). As has just been indicated, the problem consists in finding such a function \(g\) of the combination \((\xi^2+\eta^2+\zeta^2)^{1/2}\) that the distribution \(f_0+\xi g\) should be stationary, i.e., that the number of particles in any cell of velocity space should remain constant. Let us consider in particular the cell enclosed between the planes \(\xi\) and \(\xi+d\xi\), \(\eta\) and \(\eta+d\eta\), and \(\zeta\) and \(\zeta+d\zeta\). Let us call it the cell \(C\). In order not to have to deal with a large number of Greek letters, we shall denote the volume of this cell \(d\xi\,d\eta\,d\zeta\) by \(d\tau\). The number of particles in it is equal to \(f\,d\tau\). This number must remain unchanged, despite the fact that particles continuously enter and leave \(C\), both under the influence of the field and under the influence of collisions.
The electric field \(E\) imparts to each particle a constant acceleration
\[
\alpha=\frac{eE}{m},
\]
owing to which a continuous displacement of particles from one cell to another occurs. It is easy to see that the number of particles that in this way leave \(C\) per unit time (as the “unit of time” in the present case it is more convenient to take not a second, but some interval of time small in comparison with the interval of time between two successive collisions) is, in the first approximation, equal to
\[
\alpha f(\xi+d\xi,\eta,\zeta)\,d\eta\,d\zeta.
\]
This loss is partly compensated by the influx of those particles which originally lay behind the plane \(\xi\); however, the compensation is incomplete, since the number of the latter is equal to
\[
\alpha f(\xi,\eta,\zeta)\,d\eta\,d\zeta,
\]
and as a result there is a net loss of
\[
\alpha \frac{\partial f}{\partial \xi}\,dt
\]
particles.2 It must, in its turn, be compensated by an additional influx of particles under the influence of collisions.
Let \(\alpha\,dt\) be the number of those particles of the cell \(C\) which during the time \(dt\) underwent collisions that forced them to leave this cell, and let \(b\,dt\) be the number of those particles which originally were in other cells and, under the influence of collisions during the time \(dt\), entered the cell \(C\). According to the preceding, the function \(g\) must be chosen so that the net loss of electrons under the influence of the field is exactly compensated by an increase in their number under the influence of collisions, i.e. so that the relation holds:
\[
\frac{eE}{m}\frac{\partial f}{\partial \xi}=b-a.
\tag{93}
\]
Hence it is clear that first of all we must express \(b-a\) through the distribution function.
The number of those particles of velocity \(v\) which undergo collis-
collisions per unit time is already known to us. Indeed, for each particle the number of collisions is equal to \(l/\tau\), so that
\[ a=\frac{v}{\tau} f\, d\xi\, d\eta\, d\zeta . \tag{94} \]
In order to calculate \(b\), it is necessary to classify all collisions according to their results, i.e. according to which cells of velocity space the colliding particles fall into. In velocity space a particle of speed \(v\) lies on a sphere of radius \(v\) with center at the origin. In a collision with an immobile atom only the direction of the velocity changes, but not its magnitude, i.e. in velocity space the particle passes from one point of this sphere to another. Consequently, only those electrons which lie in the same spherical layer as \(C\) can fall into the cell \(C\). We shall derive an expression for the number of particles entering the cell \(C\) from any cell \(C'\) of the same layer and conversely—from \(C\) into \(C'\). The integral of the difference between these numbers over all \(C'\) is the desired quantity \(b-a\).
First let us determine for how many particles, after collisions, the trajectory (in coordinate space, of course) is deflected by an angle lying between \(\theta\) and \(\theta+d\theta\). In order to be deflected through the angle \(\theta\), an electron must strike an atom in such a direction that the angle between the normal (i.e. the line of centers at the moment of collision and the initial trajectory of the electron) is equal to \(\psi=\frac{1}{2}(\pi-\theta)\). Denote the radius of the atom by \(R\) and assume that, in comparison with it, the radius of the electron may be neglected.\(^1\) At some given instant of time our \(f\, d\xi\, d\eta\, d\zeta\) electrons are located in a unit volume of the metal and belong to a definite cell \(C\) of velocity space. Imagine that each of these electrons is, in coordinate space, the center of two circles lying in the plane—
\(^1\) The following formulas are also valid without this restriction, if only by \(R\) one understands the sum of the radii of the atom and the electron; but such a generalization, as far as I know, has no practical significance.
plane perpendicular to its trajectory, whose radii are \(R\sin\psi\) and \(R(\sin\psi+d\sin\psi)=R(\sin\psi+\cos\psi\,d\psi)\). As the electron moves, these circles move along the normals to their plane with velocity \(v\). In unit time each pair of circles describes a pair of cylinders \(v\). The volume of the layer enclosed between these cylinders is
\[ v\cdot 2\pi R^2\sin\psi\cos\psi\,d\varphi . \]
Multiplying this quantity by the number of electrons \(f\,d\xi\,d\eta\,d\zeta\), we obtain the total volume of all such layers. In order to obtain the number of electrons of interest to us, falling on atoms at angles lying between \(\varphi\) and \(\varphi+d\varphi\), i.e. being deflected through angles lying between \(\theta\) and \(\theta+d\theta\), it is evidently necessary to multiply this volume by the number of atoms in unit volume, \(N\).
\[ N f\,d\xi\,d\eta\,d\zeta\cdot 2\pi v R^2\sin\psi\cos\psi\,d\psi = f\,d\xi\,d\eta\,d\zeta\cdot N\pi R^2v\cdot \frac12\sin\theta\,d\theta \tag{95} \]
The total number of collisions per unit time experienced by the electrons of interest to us is obtained by integrating expression (95):
\[ Z\,dt=\pi N R^2 v f\,d\tau . \tag{96} \]
Substituting this expression in (95), we obtain:
\[ Z\cdot 2\sin\psi\cos\psi\,d\psi = Z\cdot \frac12\sin\theta\,d\theta . \tag{97} \]
Let us note that deflections less than \(90^\circ\) and greater than \(90^\circ\) occur with equal frequency, so that on average the electrons, as we have already mentioned, lose every trace of their initial direction of motion. Further, comparing (96) with (94), we obtain the following expression for the mean free path:
\[ l=\frac{1}{N\pi R^2}, \tag{98} \]
the same as that indicated in the note. In order to bring out the most characteristic features of expression (97), however, it is necessary to return to velocity space.
In this space, those electrons whose trajectories are deflected through angles lying between \(\theta\) and \(\theta+d\theta\) jump from the cell \(C\) into other cells of the above-mentioned
of the spherical layer. These latter fill a spherical zone bounded by two cones with common vertex at the origin, whose axes pass through \(C\), and whose solid-angle halves are respectively \(\theta\) and \(\theta+d\theta\). As is known, the area of this zone is proportional to \(\sin\theta\,d\theta\). Hence there follows the important consequence that the electrons flying out of \(C\) as a result of collisions are distributed uniformly over the whole sphere. In this case the cell \(C\) may be any cell of the given layer.
Let us now consider the process of exchange of electrons between two arbitrary cells of the layer: the cell \(C\) \((\xi,\eta,\zeta)\) of volume \(d\tau\) and the cell \(C'\) \((\xi',\eta',\zeta')\) of volume \(d\tau'\). The number of particles jumping from \(C\) to \(C'\) is equal to the total number of collisions in \(C\), multiplied by the ratio of the volume \(C'\) to the total volume \(V\) of the layer. Conversely, the number of particles jumping from \(C'\) to \(C\) is equal to the total number of collisions in \(C'\), multiplied by the ratio of the volume \(C\) to the total volume of the layer. As a result, in \(C\) there is obtained the remainder:
\[ Z(\xi',\eta',\zeta')\,d\tau'\left(\frac{d\tau}{V}\right) - Z(\xi,\eta,\zeta)\,d\tau\left(\frac{d\tau'}{V}\right). \tag{99} \]
With the aid of formulas (96) and (98), this expression may be written as:
\[ \frac{v\,d\tau}{l}\,[f(\xi',\eta',\zeta')-f(\xi,\eta,\zeta)]\,d\tau'. \tag{100} \]
The integral of this quantity over \(\xi',\eta',\zeta'\), extended over the entire spherical layer, is the total number of particles obtained by the cell \(C\) as a result of collisions, which above we denoted by \((b-a)\,d\tau\). Using Lorentz’s expression for the distribution function \(f\), and taking into account that throughout the whole region of integration the quantity \((\xi'^2+\eta'^2+\zeta'^2)^{1/2}\) has values close to \(v\), we obtain:
\[ b-a=\frac{v}{l}\,g(v)\int(\xi'-\xi)\,\frac{d\tau'}{V}. \tag{101} \]
To facilitate the integration, it is convenient to introduce in velocity space a polar coordinate system. If the origin is left in the former place, and the polar axis is drawn through the cell \(C\), then the radius vector will be a величи-
\(b\), and the polar angle by \(\theta\). We shall subdivide our layer into cells by parallels and meridians. Then for each cell we shall have:
\[ \frac{dz'}{r}=\frac{1}{4\pi}\sin\theta\,d\theta\,d\varphi, \tag{102} \]
where \(\varphi\) is the azimuth. Hence we obtain: \({}^{1}\)
\[ b-a=\frac{r}{4\pi}\,g(v)\int_{0}^{\pi}d\theta\,\sin\theta \int_{0}^{2\pi}d\varphi\,(\xi'-v). \tag{103} \]
Thus our problem reduces to the calculation of \(\xi'-\xi\), i.e. of the change which the \(X\)-component of the velocity undergoes when the electron passes from \(C'\) to \(C\). It is not difficult to show \({}^{2}\) that
\[ \xi'-\xi=-2v\cos\psi\cos\omega =-2v\sin\frac{\theta}{2}\cos\omega, \tag{104} \]
where \(\psi\), as before, denotes the angle between the initial trajectory of the electron and the line of centers at the moment of collision, and \(\omega\) is the angle between the line of centers and the \(X\)-axis. As is known, \({}^{3}\) the angles \(\omega,\psi,\varphi\) and \(\beta=\arccos\frac{\xi}{v}\) (= the angle between the per-
\({}^{1}\) We note that our formulas differ somewhat from Lorentz’s, since by \(\theta\) we denote an angle twice as large as Lorentz’s.
\({}^{2}\) Let \(v\) and \(v'\) be the velocity vectors of the electron before and after the collision, and let \(c_{1}\) be the unit vector in the direction of the line of centers at the moment of collision. The components of \(v\) and \(v'\) along the line of centers are equal in magnitude and opposite in direction; while the components perpendicular to the line of centers are equal both in magnitude and in direction. In vector notation this may be written as:
\[
v\cdot c_{1}=-v'\cdot c_{1};\quad
v-(v\cdot c_{1})=v'-(v'\cdot c_{1})c_{1}=v'+(v\cdot c_{1})c_{1},
\]
whence
\[
v'-v=-2(v\cdot c_{1})c_{1}=-2v\cos\psi\,c_{1};\quad
\xi'-\xi=-2v\cos\psi\cos\omega.
\]
\({}^{3}\) Consider two planes \(P_{1}\) and \(P_{2}\) intersecting along a vertical axis, the dihedral angle between them being equal to \(\varphi\). Through some point \(O\) on this axis draw a horizontal plane \(H\), and consider two unit segments \(OR_{1}\) and \(OR_{2}\), the first of which lies in the plane \(P_{1}\) and makes with the vertical axis an angle \(\beta\), while the second lies in the plane \(P_{2}\) and makes with the vertical axis an angle \(\psi\). Denote the angle between these segments by \(\omega\). The heights of the points \(R_{1}\) and \(R_{2}\) above the plane \(H\) are respectively \(\cos\beta\) and \(\cos\psi\). Finally, note on the vertical-
with the initial trajectory of the electron and the axis \(X\)) are connected by the following relation:
\[ \cos \omega=\cos \beta \cos \psi-\sin \beta \sin \psi \cos \varphi . \tag{105} \]
Substituting this into equation (103) and integrating, we obtain:
\[ b-a=\frac{v\xi}{l}\,g(v), \tag{106} \]
Finally, equation (93), expressing the condition of stationarity of the distribution function \(f_0+\xi g\), takes the form:
\[ \frac{eE}{m}\frac{\partial}{\partial \xi}(f_0+\xi g)=\frac{v\xi}{l}\,g . \tag{107} \]
If the quantity \(\xi g(v)\) is small in comparison with \(f_0(v)\), then the second term on the left-hand side may be neglected; whence, using the equality
\[ \frac{\partial f_0}{\partial \xi} = \frac{df_0}{dv}\frac{\partial v}{\partial \xi} = \frac{\xi}{v}\frac{df_0}{dv}, \]
we arrive, finally, at the formula
\[ \xi g(v)=\frac{\xi l}{v^2}\frac{eE}{m}\frac{df_0}{dv}. \tag{108} \]
Such is the correction introduced into the distribution of electrons by the electric field. Let us note that, in accordance with Lorentz’s original assumption, \(g\) contains \(\xi,\eta,\zeta\) only in combination with \(v\).
Each electron, passing in a given unit of time through any surface drawn inside the metal, gives an electric current of strength \(e\). At the same time electrons moving in opposite directions produce opposite effects, so that the actually observed current is the excess of the flux of electricity moving in one direction over the flux moving in the other direction. Let us imagine a plane surface element \(da\), perpendicular to the field, i.e. to the axis \(x\). We shall classify all electrons passing through it according to the values of \(\xi\). Suppose that the number of those electrons,
line passing through the point \(R_2\), the third point \(R_3\) at a distance \(\cos \beta\) from the plane \(H\). Calculating the sides of the right triangle \(R_1R_2R_3\) and applying the Pythagorean theorem to them, we obtain formula (105).
which in unit time pass through this element and for which, at the moment of arrival, the \(x\)-component of the velocity lies in the interval from \(\xi\) to \(\xi+d\xi\), is equal to \(H(\xi)\,d\xi\,da\). These electrons are situated in a right prism with base \(da\) and height \(\xi\):1
\[ H(\xi)\,d\xi\,da=\xi\,da\,\wp \int_{-\infty}^{\infty} d\eta \int_{-\infty}^{\infty} d\zeta\, f(\xi \eta \zeta). \tag{109} \]
Since, for electrons moving in opposite directions, \(\xi\) has different signs, the integral of the expression just written over all values of \(\xi\), when multiplied by \(C\), gives the total current through the area \(da\). It is obvious that if \(f\) is isotropic, then this same integral is equal to zero, i.e. enormous quantities of electricity pass through \(da\) in both directions, mutually compensating one another. This compensation ceases only when a non-isotropic perturbing term appears in the distribution function. Using Lorentz’s expression, we obtain for the current \(I\) through a unit area perpendicular to the field the expression:
\[ I=e\int H(\xi)\,d\xi = e\int_{-\infty}^{\infty}\int_{-\infty}^{\infty}\int_{-\infty}^{\infty} d\xi\,d\eta\,d\zeta'\,g(v). \tag{110} \]
Substituting formula (108) for \(g(v)\), transforming the integral to polar coordinates with the polar axis along \(Ox\), and integrating over the angles, we have:
\[ I=\frac{Ce^{2}E}{m}\,\frac{4\pi}{3}\int_{0}^{\infty} v^{2}\left(-\frac{df_{0}}{dv}\right)\,dv. \tag{111} \]
The final integration is carried out easily if one substitutes the Maxwell function for \(f_{0}\), and less easily if one uses the Fermi function. Integrating by parts and noting—
assuming that \(v\) vanishes at the lower limit and \(f_0\) at the upper limit, formula (111) may be written in the following form:
\[ I=-\frac{8\pi l e^2 E}{3m}\int_0^\infty f_0 v\,dv . \tag{112} \]
The quotient obtained by dividing the last integral by the number of electrons per unit volume \(h\) is nothing other than the mean value of the reciprocal of the electron velocity in the absence of a field \(\overline{v^{-1}}\), but multiplied by \(\frac{1}{4\pi}\). Denoting this mean value by \(\overline{v^{-1}}\), we obtain for the electrical conductivity the following formula
\[ \sigma=\frac{I}{E}=\frac{2}{3}\frac{e^2 h}{m}\,\overline{v^{-1}}, \tag{113} \]
which in essence coincides with the formula derived earlier (91).1 The value of the quantity \(\overline{v^{-1}}\) is given by formulas (72 b) and (72 a). Putting \(G=2\) in the latter, we finally obtain
\[ \sigma=\frac{4}{3}\frac{e^2 l h}{(2\pi m kT)^{1/2}} \quad \text{(according to the old statistics),} \tag{114 a} \]
\[ \sigma=\frac{8\pi}{5}\frac{e^2 l}{h}\left(\frac{3h}{8\pi}\right)^{2/3} \quad \text{(according to the new statistics).} \tag{114 b} \]
As Sommerfeld showed, all the arguments leading to the derivation of formula (108) remain valid also in the case when \(l\) is a function of \(v\). Then, in equality (111), \(l\) remains under the integral sign, and the integral itself becomes equal to the quantity
\[ \frac{h}{4\pi}\,v^{-2}\frac{d(lv^2)}{dv}. \]
Such a generalization is sometimes useful.
Homogeneous metal with a temperature gradient; thermal conductivity
Let us now proceed to finding such a function \(g\) that the distribution \((f_0+\xi g)\) remains stable in the presence
in the metal of a constant temperature gradient directed along \(Ox\). Having found this function, we shall be able to calculate the integral:
\[ W = \frac{1}{2m}\iiint d\xi\,d\eta\,d\zeta\, g(v)v^2\xi^2, \tag{125} \]
which is completely analogous to integral (110), with the sole differences that in it 1) \(g\) has another form; 2) in place of the charge \(e\) there stands the kinetic energy \(\frac{1}{2}mv^2\). This integral is, evidently, equal to the total flux of kinetic energy carried by the electrons through a unit surface perpendicular to the direction of the temperature gradient, i.e. it represents that fraction of the total heat flux which is carried by the electrons. The standard distribution function \(f_0\) includes both the temperature \(T\) and the coordinate \(x\). However, an isotropic distribution varying from point to point cannot, under our conditions, be stable, since with it the temperature gradient would quickly be leveled out. In other words, the stable distribution must be non-isotropic. In order to find it, it is necessary, as in the preceding case, to obtain the condition of equilibrium between particles entering and leaving each cell of phase space under the influence of 1) collisions, 2) the temperature gradient. The present case is complicated by the fact that here one has to deal not with velocity space, but with the whole phase space, i.e. to consider the motion of electrons through six-dimensional cells \(d\xi\,d\eta\,d\zeta\,dx\,dy\,dz\), and not through three-dimensional cells \(d\xi\,d\eta\,d\zeta\). Let us consider any one of such cells
\[ d\xi\,d\eta\,d\zeta\,dx\,dy\,dz = ds, \]
containing
\[ f(x,y,z,\xi,\eta,\zeta)\,ds \]
electrons. The first 3 factors in \(ds\) show within what limits the velocities of these electrons lie; the second 3 factors, within what limits their coordinates lie. Let us call those values of velocities and coordinates which correspond to the cell \(ds\) “allowed.” Electrons possessing allowed coordinates, thanks to collisions, continuously enter and leave the allowed region of velocities. The net
the gain obtained by this cell \(ds\) is given, as before, by expression (106). On the other hand, electrons possessing allowed velocities continuously enter the allowed region of coordinates \(dx\,dy\,dz\)—from the side of smaller or larger \(x\)’s, depending on whether we are dealing with positive or negative \(\xi\). By exactly the same method by which we obtained formula (93), it is easy to show that the net loss of electrons caused by this is equal to \(\xi \dfrac{df}{dx}\). The condition of equilibrium reads
\[ \xi \frac{df}{dx}=b-a=-\frac{\xi}{c}\,g(v). \tag{116} \]
Putting here \(f=f_0\) (in view of the smallness of the additive term \(\xi g\)), we obtain everything necessary for determining the function \(g\), and with it also the expression \(W\) (for the prescribed form of the standard function \(f_0\)).
At this point, however, we encounter one difficulty. If the \(g\) calculated in this way is substituted into expression (110) for the density of the electric current \(I\), then it turns out that this expression is not equal to zero. In other words, from formulas (116) there follows the conclusion that every transfer of heat from one part of the metal to another is accompanied by a corresponding transfer of electricity. Such a conclusion is in clear contradiction with experiment. In order to avoid it, it remains to suppose that whenever a temperature gradient appears in a metal, an internal electric field arises in it, cancelling that electric current which arises in the transfer of heat. The temperature gradient causes the appearance of a potential gradient; and the distribution function must be stable with respect to both of these gradients. With respect to the cell \(ds\), this means that the net gain \(b-a\) must be balanced by the sum of the net losses in coordinate space \(\left(\xi \dfrac{df}{dx}\right)\) and in velocity space \(\left(\alpha \dfrac{\partial f}{\partial \xi}\right)\). Hence, denoting our hypothetical electric field by \(E\), and the acceleration imparted by it to each ...
an acceleration to the electron by means of \(\alpha\left(=\dfrac{eE}{m}\right)\), we obtain two equations:
\[ \alpha \frac{\partial f}{\partial \xi}+\xi \frac{\partial l}{\partial x}=b-a=\frac{\xi v}{l}\,g \tag{117} \]
and
\[ \frac{I}{e}=\iiint d\xi\,d\eta\,d\zeta\,\xi^{2}g=0 \tag{118} \]
for determining the quantities \(E\) and \(g\).
Lorentz solved these equations starting from the Maxwell–Boltzmann distribution, and obtained:
\[ W=\frac{8}{3}\frac{n l k^{2}T}{(2\pi m kT)^{1/2}}\,\frac{dT}{dx}; \qquad \sigma=\frac{1}{2}\frac{k}{m}\frac{dT}{dx}. \tag{119} \]
Sommerfeld, however, used the Fermi function and obtained for the degenerate case:^[In deriving formula (120) one has to use the second approximation for the Fermi distribution function (equation 107 a), since the first approximation (equation 107 b) simply gives \(w=0\).]
\[ W=\frac{8\pi^{3}}{9}\frac{l k^{2}T}{h}\left(\frac{3n}{8\pi}\right)^{2/3}\frac{dT}{dx}. \tag{120} \]
The coefficient of \(\dfrac{dT}{dx}\) in these expressions is usually called the coefficient of thermal conductivity and is denoted by \(\chi\). Let us note that \(\chi\), like \(\sigma\), contains the constants \(n\) and \(l\). The ratio \(\dfrac{\chi}{\sigma}\), however, does not depend on these constants.
Wiedemann–Franz Ratio
For the quantity \(\dfrac{\chi}{\sigma}\), the old and the new statistics give almost identical values, differing only by a numerical factor. Both of these values depend only on \(T\) and on the universal constants \(k\) and \(e\):
\[ \frac{\chi}{\sigma}=2\left(\frac{k}{e}\right)^{2}T\;(=4.2\cdot 10^{-11},\ \text{for }T=291^\circ K) \tag{121} \]
according to the old statistics, and
\[ \frac{\chi}{\sigma}=\frac{1}{3}\pi^{2}\left(\frac{k}{e}\right)^{2}T\;(=7.1\cdot 10^{-11},\ \text{for } T=291^\circ K) \tag{122} \]
according to the new statistics.
This “Wiedemann–Franz relation” serves as strong support for the proponents of the theory of the electron gas. All the other formulas of this theory contain the constants \(n\) and \(l\) (or at least one of them) and therefore cannot serve as a decisive test of the theory, since, by using relative arbitrariness in the choice of \(n\) and \(l\), one can in each individual case obtain agreement with experiment. True, so long as we remained on the ground of classical statistics, the interpretation of one experiment came into contradiction with the interpretation of another, so that in general the result turned out to be unfavorable for our theory; but it was impossible to indicate at precisely which point its predictions were certainly erroneous. If, however, it were found that the observed values of \(\frac{\chi}{\sigma}\) differed greatly from the quantity \(2\left(\frac{k}{e}\right)^{2}T\), then the theory of the electron gas would have received a mortal blow. But precisely at the point where the discrepancy would have been fatal, we obtain agreement—not perfect, but too good to be accidental. For most metals the ratio \(\frac{\chi}{\sigma}\) at room temperature is close to \(6\) or \(7\cdot 10^{-11}\). This circumstance, more than any other, supports confidence in the validity of the fundamental propositions of the electron theory of metals.
The mean value of the quantity \(\frac{\chi}{\sigma}\) at \(291^\circ K\) for twelve metals (Al, Cu, Ag, Au, Ni, Zn, Cd, Pb, Sn, Pt, Pd, and Te) is equal to \(7.11\cdot 10^{-11}\). As we see, the agreement with the new statistics is more than good. It is possible that in part it is explained by chance that crept in when the average was calculated. It is difficult to say whether this agreement can be counted among the pieces of evidence speaking in favor of the new statistics. Let us recall that Drude, starting from the very crude idea that all electrons move
with the same velocity1 arrived at the value \(\frac{\chi}{\sigma}=6.3\cdot 10^{-11}\). It is remarkable that the lengthy calculations of Lorentz only worsened the agreement which Drude had obtained naively, by a simple route.
According to both statistics, \(\frac{\chi}{\sigma}\) should be proportional to \(T\). Over a fairly wide range of temperatures experiment confirms this conclusion of the theory. Discrepancies appear only at very low temperatures. To explain them one may suppose that the transfer of electricity involves not free electrons, but also electrons passing directly from atom to atom, or that elastic vibrations of the atoms of the metal also participate in the transfer of heat. Indeed, since insulators, in which free electrons are absent, are somehow capable of transmitting heat, it is clear that heat transfer can occur in the same way in metals as well, and consequently the observed values of \(\chi\) and \(\frac{\chi}{\sigma}\) may exceed the values calculated on the basis of the electron-gas theory.
Internal potentials
As we have just seen, in a metal in which heat transfer takes place and there is no electric current, it is necessary to introduce a hypothetical internal electric field. It can be shown that the same kind of field must exist also in the case when there is no current, but the number of electrons per unit volume varies from point to point. Indeed, in the absence of a field the electrons would evidently have to diffuse from regions of higher density into
regions of lowest density. Formulas (117) and (118) give us the possibility of calculating the strength of this field.
Returning to these equations, let us introduce in velocity space the polar system of coordinates \(v,\theta,\varphi\); then we shall have \(\xi=v\cos\theta\). Multiply both sides of equation (117) by \(\cos\theta\), so that its right-hand side is proportional to \(\xi^{2}g\). Integrating the resulting equality over the whole velocity space, we obtain:
\[ \alpha \int d\tau \,\frac{\partial f_0}{\partial \xi}\cos\theta +\int d\tau \xi\,\frac{\partial f_0}{\partial x}\cos\theta=0, \tag{123} \]
since the integral on the right-hand side, by virtue of formula (118), is equal to zero.
Carrying out the usual transformations, we have:
\[ \alpha \int d\tau \,\frac{\partial f_0}{\partial v}\cos^2\theta +\frac{\partial}{\partial x}\int d\tau \, f_0 v\cos^2\theta=0. \tag{124} \]
Integration over the angles gives
\[ \frac{4\pi}{3}\alpha \int \frac{\partial f_0}{\partial v}v^2\,dv +\frac{4\pi}{3}\frac{\partial}{\partial x}\int f_0 v^3\,dv=0. \tag{125} \]
Finally, integrating the first term by parts and using the fact that, for any statistics, \(f_0\) vanishes at one limit of integration and \(v\) at the other, we obtain:
\[ -\frac{2}{3}\alpha \int 4\pi f_0 v^{-1}\cdot v^2\,dv +\frac{1}{3}\frac{\partial}{\partial x}\int 4\pi f_0 v\cdot v^2\,dv=0. \tag{126} \]
The integrals appearing on the left-hand side of this formula are, obviously, proportional to the mean values of the functions \(v^{-1}\) and \(v\), taken over all electrons of the cell \(dx\,dy\,dz\) of coordinate space. More precisely, they are equal respectively to the mean reciprocal velocity and the mean velocity, multiplied by the number of electrons in the cell under consideration, \(n\,dx\,dy\,dz\). Consequently, formula (126), reducing by the factor \(\frac{1}{3}\,dx\,dy\,dz\), may be rewritten as:
\[ -2n\overline{v^{-1}}+\frac{d(nv)}{dx}=0. \tag{127} \]
This equation must be satisfied by the acceleration \(\alpha\)—more precisely, by the accelerating field \(E=\frac{m\alpha}{e}\)—in order that, notwithstanding
the presence of the gradient \(\dfrac{d(nv)}{dx}\), the strength of the electric current was equal to zero. The existence of the hypothetical field \(E\) is conditioned precisely by this gradient, of which the temperature and concentration gradients are particular cases.
In the case of classical statistics, the further calculations are extremely simple, since \(\bar v\) depends only on the temperature, while \(n\) may vary quite arbitrarily. Up to now we have assumed that \(n\) is constant, while \(T\) varies along the \(x\)-axis, so that
\[ \frac{d(nv)}{dx}=n\frac{dv}{dx}=n\frac{dv}{dT}\cdot\frac{dT}{dx}. \tag{128} \]
Using this formula, the reader may verify the validity of the expressions for \(a\) and \(W\) given in equation (119). We are now interested in the case when \(T\) and \(v\) are constant, while \(n\) varies along the \(x\)-axis. In this case:
\[ \frac{\partial(nv)}{\partial x}=v\frac{\partial n}{\partial x}. \tag{129} \]
Substituting this into formula (127), remembering that for the Maxwell distribution
\[ v=\frac{2kT}{m}\,\bar v^{-1} \tag{130} \]
and, finally, canceling both sides of the resulting equality by \(v\), we have
\[ an=\frac{eE}{m}n=\frac{kT}{m}\frac{\partial n}{\partial x}. \tag{131} \]
Equation (131) also determines the electric field sought. Integrating, one may give it the following more usual form
\[ n=n_0 e^{\frac{eE}{kT}(x-x_0)}=n_0 e^{\frac{V-V_0}{kT}}. \tag{132} \]
This is the famous Boltzmann equation, according to which variations of the density in any assembly of particles whose temperature is the same at all points are always accompanied by the presence of a force field impeding the motion of these particles—and conversely. More precisely, if in
at two points \(P\) and \(O\) the concentration of particles is respectively \(n\) and \(n_0\), then—according to Boltzmann—there must be present in the assembly such a force field that, when a particle is displaced from \(O\) to \(P\), its potential energy increases by \(-kT\lg \dfrac{n}{n_0}\). On the other hand, if the particles under consideration are electrons and if we have an electric field derivable from a potential, whose values at \(P\) and \(O\) are, respectively, \(V\) and \(V_0\), then this change in potential energy is given by the expression \(e(v-v_0)\).
Boltzmann’s equation has become so deeply rooted in modern physics that any replacement of it by some other equation appears to be something very strange and undesirable. Nevertheless, it is easy to see that, substituting into equation (127) Fermi’s distribution for absolute zero, we arrive at an equation different from Boltzmann’s. Indeed, since according to the new statistics \(v\) depends on \(n\)—i.e. the mean velocity of the electrons depends on their concentration—the second term in equation (127) is no longer proportional to \(\dfrac{\partial n}{\partial x}\). Formulae (72a) give
\[ nv=\frac{3}{4}nv_m=\frac{3hn}{4m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}};\qquad \overline{v^{-1}}=\frac{3}{2v_m}. \tag{133} \]
Substituting these values in (127), we obtain:
\[ \frac{eE}{m}=a=\frac{h^2}{3m^2}\left(\frac{3}{4\pi G}\right)^{\frac{2}{3}}n^{-\frac{1}{3}}\frac{\partial n}{\partial x}, \tag{134} \]
whence, after integration, it follows:
\[ -E(x-x_0)=+(V-V_0)=\frac{h^2}{2ml}\left[\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}}-\left(\frac{3n_0}{4\pi G}\right)^{\frac{2}{3}}\right]. \tag{135} \]
This formula replaces Boltzmann’s equation in the new statistics.
Using equality (71), equation (135) can be rewritten as follows:
\[ e(V_1-V_0)=(W_i)_1-(W_i)_0. \tag{136} \]
Thus, in order that between two specimens of an electron gas obeying Fermi statistics—whose temperatures are equal to absolute zero, and the maximum kinetic energy of whose electrons is, respectively, $W_{a1}$ and $W_{a0}$—there should exist stable equilibrium, there must be present such a potential difference that the kinetic energy of the fastest electron of the first group, having passed into the second group, is equal to the kinetic energy of the fastest electron of the second group. Expressed in this form, the statement is easy to remember and even seems almost self-evident.
Let us now consider two different metals in contact with one another. In order to avoid a discontinuity, one may imagine that these metals are soldered to one another by means of an alloy in which both metals are mixed in a definite proportion, continuously varying from 0 to 100%. The density of electrons in each of our metals is different, since the number of electrons per unit volume is equal to the number of atoms, which changes in passing from one metal to the other. If we assume that the soldering process itself does not affect the magnitude of the electron concentration at points sufficiently removed from the place of soldering, then it follows from this that between our metals there must exist a definite potential difference, whose magnitude is given by formula (132) or (135), depending on whether we use the old or the new statistics. For example, in the example of potassium and silver considered by Sommerfeld, this difference reaches 4.2 volts,^1 with potassium serving as the negative pole (in the calculation, in accordance with what was said above, it is assumed that the electron concentration is equal to the atomic concentration). The classical formula, however, gives a considerably lower figure—about 0.04 volts. The discrepancy between the two theories is striking. Both statistics explain the existence of an “internal differ-
^1 Sommerfeld himself, taking $G=1$, obtained a value equal to 5.7 volts; the value of 4.2 volts corresponds to $G=2$.
of potentials” for different electron concentrations, but the magnitude of this difference at a given concentration is different in the two theories; moreover, in all practical cases the new statistics gives values considerably larger than the old.
Although direct measurement of the electric potential inside a metal is impossible, there are phenomena which show that between two metals in contact, and also between two parts of one metal maintained at different temperatures, there is indeed a difference of potentials.
These phenomena are called thermoelectric. Among them are the Peltier and Thomson effects, the occurrence of a thermo-electromotive force, etc. The existence of an “internal potential gradient” manifests itself here in the fact that the heat evolution caused by the passage of a current no longer obeys Joule’s law alone. In order to obtain a theory of thermoelectric phenomena, it is evidently necessary to consider the problem of heat transfer by an electron gas whose distribution function, in addition to the perturbing action of the electric field, is also subject to the perturbing action of one of two other factors—the temperature gradient or variable concentration—which we have so far considered separately. This problem will be discussed in the next section.
Fig. 2.
In conclusion, in order to avoid misunderstandings, I wish to note that the “internal” difference of potentials, the expression for which has just been derived, has nothing in common with the so-called “contact” difference of potentials between two metals. In order to determine what this latter is equal to, let us consider the arrangement shown in Fig. 2, where \(A\) and \(B\) are two metals in contact with one another along surface 1 and ending at boundary surfaces 2 and 3. Let us consider some-
some electron in the metal \(A\) and compare the difference of potentials which it has to overcome in order to get to the point \(P_2\), adjacent directly to surface 2, with the difference of potentials which it must overcome in order to pass through the metal \(B\) and get to the point \(P_3\), directly adjacent to surface 3. Recalling the notation which we used in the chapter on thermionic currents, it is easy to see that the difference of potentials between \(P_2\) and \(P_3\) is given by the expression:
\[ (W_{aA}-W_{\omega B})-(W_{iA}-W_{iB})=b_A-b_B. \tag{137} \]
This is precisely the contact difference of potentials. We see that, according to the new statistics, it (at the temperature of absolute zero) is equal to the difference of the values of Richardson’s constant \(b\), usually considered as a function of the work. According to the old statistics, however, it differs from \(b_A-b_B\) by an amount equal to the internal difference of potentials between our metals at the surface of contact 1. It is possible that it will be possible to establish which of these conclusions agrees better with experimental data.
Theory of Thermoelectric Phenomena
Let us now proceed to the determination of the amount of heat liberated in a metal through which an electric current passes and in which, in addition, there is an internal gradient of potential, caused either by a temperature gradient or by a concentration gradient, or by both together. These calculations lead to formulas which can be tested experimentally and thereby provide a new test for the fundamental propositions of Fermi statistics.
The amount of heat liberated per unit time in a unit volume of a conductor, with simultaneous passage along the \(x\)-axis of a flow of electricity of density \(I\) and a flow of heat of density \(W\), is equal to \(IE-\dfrac{\partial w}{\partial x}\). Here \(E\) is the total force of the electric field, i.e., generally speaking, the sum of the external and hypothetical internal fields. Denote
knowing the corresponding acceleration through \(\alpha\), and the amount of heat liberated per unit volume through \(r\), we have
\[ -r=\frac{m\alpha}{e}I-\frac{\partial N}{\partial x}. \tag{138} \]
The density of the heat flux is given by formula (115), which we rewrite here:
\[ W=\frac{1}{2}m\int v^{4}\cos^{2}\theta\,g\,d\tau, \tag{139} \]
where the function \(g\), as before, represents the non-isotropic perturbing term in the distribution function. This function and the acceleration \(\alpha\) are determined by two equations:
\[ \alpha\frac{\partial f_{0}}{\partial \xi}+\xi\frac{\partial f_{0}}{\partial x} =\frac{v_{\xi}}{l}g=\frac{v^{2}\cos\theta}{l}g, \tag{140} \]
\[ \frac{I}{e}=\int \xi^{2}g\,d\tau=\int v^{2}\cos^{2}\theta\,g\,d\tau, \tag{141} \]
which coincide with equations (117) and (118), with the sole difference that the electric current \(I\) is no longer assumed to be zero.
Multiply both sides of equation (140) by \(\frac{1}{2}mv^{2}\cos\theta\) and integrate over the entire velocity space. Then the integral on the right-hand side will be equal to \(\frac{W}{l}\); expanding the integrals on the left-hand side, we find:
\[ W=\frac{1}{6}ml\left(-4\alpha nv+\frac{\partial(nv^{3})}{\partial x}\right). \tag{142} \]
Next multiply the parts of equation (140) by \(\cos\theta\) and again integrate over the entire velocity space. In this case on the right-hand side we obtain \(\frac{I}{le}\) and arrive at a formula which is a generalization of formula (127):
\[ -2\alpha nv^{-1}+\frac{\partial(nv)}{\partial x}=\frac{3I}{el}. \tag{143} \]
Using these equations, one can, with the aid of equality (138), express \(r\) through the mean free path, universal constants, and the mean values of various powers of \(v\). Choosing for \(f_{0}\) the Maxwell–Boltzmann function,
Calculating the expressions \(\alpha\) and \(\dfrac{\partial W}{\partial x}\), and recalling the formula for \(\sigma\) (equation 114a) and \(z\) (equation 119), we obtain:
\[ I\frac{m\alpha}{e} = I\frac{1}{2}\frac{k}{e}\frac{\partial T}{\partial x} = \frac{I^2}{\sigma}; \qquad \frac{\partial W}{\partial x} = \frac{\partial}{\partial x}\left(\varkappa\frac{\partial T}{\partial x}\right) + 2I\frac{k}{e}\frac{\partial T}{\partial x}. \tag{144} \]
It follows from this that the amount of heat released per unit time per unit volume is
\[ r = \frac{I^2}{\sigma} + \frac{\partial}{\partial x}\left(\varkappa\frac{\partial T}{\partial x}\right) + \frac{3}{2}\frac{k}{e}\frac{\partial T}{\partial x}I. \tag{145} \]
The first term of this expression is, evidently, Joule heat; the second term is not directly connected with the electric current and in general does not depend on the agent by which the appearance of the heat flux is caused. We are therefore interested only in the third term. This term is proportional to the first power of \(I\), and therefore for one direction of the current it expresses absorption, and for the other—emission of heat. The first case, evidently, occurs when the electrons move from lower temperatures to higher ones. In this motion the electrons acquire just such energy as is sufficient to raise their temperature to the temperature of the region into which they enter. Thus, the coefficient \(\dfrac{3}{2}\dfrac{k}{e}I\) represents, as it were, the specific heat of the electron gas, which is equal to the specific heat of any other monatomic gas with the same number of particles.
If, instead of the Maxwell distribution, one takes the Fermi distribution and substitutes in (142) and (143) the values \(\bar v\), \(v^1\), and \(\overline{v^3}\) for absolute zero, it turns out that in the right-hand side of equation (138) the terms with the first power of \(I\) cancel each other. This could have been expected, since, according to the preceding, the sum of these terms represents the specific heat of the electron gas, which for the Fermi distribution at absolute zero is equal to zero. Passing, however, to the second approximation, Sommerfeld showed that the coefficient of \(I\) in the expression for \(r\) is indeed pro-
proportional to the specific heat of the electron gas (and consequently to the temperature \(T\)) and is given by the formula:
\[ \frac{2\pi^2}{3}\,\frac{mk^2}{eh^2}\left(\frac{2\pi G}{3n}\right)^{\frac{2}{3}}T. \tag{146} \]
Experiment shows that, when a current passes through a nonuniformly heated homogeneous wire, heat is indeed evolved, and the expression for the rate of this evolution contains a term proportional to the current, changing sign when the direction of the current is changed. This “Thomson heat” (and the specific heat of the electron gas associated with it) has always proved to be smaller than the value which classical statistics gave for it under the assumption that the number of free electrons is equal to the number of atoms.
The new statistics, under this assumption, gives, both for the Thomson heat and for the specific heat, considerably smaller values, amounting to approximately \(1\%\) of the classical values (at room temperature), and agreeing (at least in order of magnitude) with the experimental data. Furthermore, experiment gives some indication that, over a fairly wide range of temperatures, the Thomson heat is indeed proportional to \(T\).
Finally, let us consider the Peltier heat—i.e., that term proportional to the current and changing its sign with a change in the direction of the current, which is observed when electricity passes through the junction (or contact surface) of two metals. Obviously, this heat is nothing other than the term with the first power of \(I\) in the expression \(I^2\frac{ma}{e}\), which, when \(\frac{dT}{dx}=0\), is equal to \(r\) with, for \(\alpha\), the value (143) substituted, under the assumption that \(n\) changes from the value corresponding to one metal to the value corresponding to the other metal. Using classical statistics and denoting the concentration of electrons in the 1st and in the 2nd metal respectively by \(n\) and \(n_1\), we obtain for this term the expression
\[ \frac{kT}{e}\lg\frac{n}{n_0}\,I. \tag{147} \]
Its value for \(J=1\) is, obviously, the internal difference
potentials between our metals. Using the new statistics, however, for absolute zero we obtain zero, and in the second approximation—the expression
\[ \frac{2\pi^2}{3}\frac{m(kT)^2}{eh^2} \left[ \left(\frac{4\pi G}{3n}\right)^{\frac{2}{3}} - \left(\frac{4\pi G'}{3n^0}\right)^{\frac{2}{3}} \right]. \tag{148} \]
Putting \(T=1\), we obtain a value considerably smaller than the internal potential difference between metals—a value of the order of fractions of a millivolt. This is also approximately the order of magnitude of the actually observed Peltier heat. It is curious that this fact, in itself, agrees equally well with both statistics. According to classical statistics, the internal potential difference between two metals is, generally speaking, small and is equal in magnitude to the Peltier heat at unit current. According to Fermi statistics, on the other hand, the internal potential difference is, generally speaking, large, but the Peltier heat constitutes only a small fraction of it.
Addendum
In the present article a whole series of questions of undoubted interest has been omitted; information about them the reader will find in the appended bibliography. Among them, in particular, are:
Sommerfeld’s theory of the Hall effect.
Houston’s generalization of Sommerfeld’s theory of the internal potential difference, containing, among other things, the theory of the Peltier heat for the case in which the current passes between two differently oriented crystals of one and the same substance;
the application of the methods of the new quantum mechanics to the problem of the interaction between atoms and free electrons, proposed by Bloch;
the determination of the arrangement of intra-atomic electrons by statistical methods, carried out by Fermi;
Fowler’s work on fluctuations in the new statistics.
In addition, the bibliography gives references to a number of other interesting problems.
LITERATURE.
J. M. Adams. The value of \(W_a\) for silver. Z. Physik. 52, p. 882 (1929). W. Anderson. Z. Physik. 50, pp. 874—877 (1928). Astrophysical applications. H. M. Barlow. Phil. Mag. (7) 7, pp. 458—470 (1929). Critique of the electron theory of metals. R. S. Bartlett. Proc. Roy. Soc. A121, pp. 456—464 (1928). Thermionic and cold discharge. E. S. Bieler. Journ. Franklin Inst. 206, pp. 65—82 (1928). General considerations concerning the statistics of the electron gas. H. F. Biggs. Nature, 121, pp. 503—506 (1928). Review of general works. F. Bloch. Z. Physik. 52, pp. 555—559 (1928). Quantum mechanics of electrons in crystal lattices. F. Bloch. Susceptibility and change of resistance in a magnetic field. Z. Physik. 53, pp. 216—227 (1929). Bose. Z. Physik. 26, pp. 178—181 (1924). The first article on Bose statistics. Z. Brillouin. C. R. 184, pp. 589—591 (1927). P. A. M. Dirac. Proc. Roy. Soc. A112, pp. 661—677 (1926). J. W. M. Du Mond. Proc. Nat. Acad. Sc. 14, pp. 875—878 (1928); Phys. Rev. (2) 33, pp. 643—658 (1929). Scattering of X-rays by free electrons. C. Eckart. Z. Physik 47, pp. 38—42 (1928). Contact differences of potentials. A. Einstein, Berliner Sitzungsberichte, 1927, pp. 261—267; 1925, pp. 3—17. Applications of Bose statistics with modifications for a material gas. E. Fermi. Z. Physik. 36, pp. 902—912 (1926). Fermi statistics. E. Fermi. Z. Physik. 48, pp. 73—79 (1928). Application of statistics to electrons inside the atom. E. Fermi. Z. Physik. 49, pp. 550—554 (1928). Statistical estimate of \(S\)-levels. E. Fermi, Quantentheorie und Chemie (Hirzel. 1928), pp. 95—111. Application of statistics to atomic problems. V. Fock. Z. Physik. 49, pp. 339—357 (1928). Dirac’s statistical equation. R. H. Fowler. Proc. Roy. Soc. A117, pp. 549—552 (1928). Thermions. R. H. Fowler. Z. Nordheim, Proc. Roy. Soc. A119, pp. 173—181 (1928). R. H. Fowler. Proc. Roy. Soc. A120, pp. 229—232 (1928). Photoelectric effect. R. H. Fowler. Proc. Roy. Soc. A122, pp. 36—49 (1929). Thermionic constant A.—R. H. Fowler. Monthly Notices Royal Astron. Soc. 87, pp. 114—122 (1926). Astrophysical applications. J. Frenkel. Z. Physik. 49, pp. 31—45 (1928). Theory of the magnetic and electric properties of metals at absolute zero. J. Frenkel, N. Mirolubow. Z. Physik. 49, pp. 885—893 (1928). Theory of metallic electrical conductivity on the basis of wave mechanics. J. Frenkel. Z. Physik. 50, pp. 234—248 (1928). Cohesion. R. Fürth. Z. Physik. 48, pp. 323—339 (1928) and 50, pp. 310—318 (1929). Fluctuations in the new statistics. G. E. Gibson, W. Heitler. Z. Physik. 49, pp. 465—472 (1928). The chemical constant in the new statistics. E. H. Hall. Proc. Nat. Acad. Sc. 14, pp. 366—380, 802—811 (1928). Arguments against the theory of the electron gas. W. V. Houston. Z. Physik. 47, pp. 33—37 (1928). Cold discharge. W. V. Houston. Z. Physik. 48, pp. 449—468 (1928). Conductivity and thermionic effect. W. V. Houston. Phys. Rev. (2) 33, pp. 361—363 (1929). Cold discharge. O. Klein.
Z. Physik. 53, pp. 157–165 (1929). Reflection of electrons from a potential barrier on the basis of Dirac dynamics. E. Kretschmann. Z. Physik. 48, pp. 739–744 (1928). Continuation of Sommerfeld’s work. J. E. Lennard-Jones. Proc. Phys. Soc. Lond. 40, pp. 320–537 (1928). Review. J. E. Lennard-Jones. H. J. Woods. Proc. Roy. Soc. A120, pp. 727–735 (1928). Theory of the electron gas with modifications for atoms in metals. A. C. G. Mitchell. L. Physik. 50, pp. 570–576 (1928). Entropy of the electron gas. L. Nordheim. Physik. Zeits. 30, pp. 177–196 (1929). General review of the new theory of thermionic emission, cold discharge, and the photoelectric effect. Earlier works: Z. Physik. 46, pp. 833–855 (1928); Proc. Roy. Soc. A121, pp. 626–639 (1928); A119, pp. 173–181 (1928). L. Nordheim. Proc. Roy. Soc. A119, pp. 689–698 (1928). Another way of deriving the distribution law. L. Nordheim. Naturwiss. 16, pp. 1042–1043 (1928). Electrical conductivity of alloys. J. R. Oppenheimer. Proc. Nat. Acad. Sc. 14, pp. 363–365 (1928). Cold discharge. W. Pauli. Z. Physik. 41, pp. 81–102 (1927). Paramagnetism. W. Pauli. Probleme der modernen Physik. (Hirzel, 1928) pp. 30–45. H-theorem. F. Rasetti. Z. Physik. 49, 536–549 (1928). Statistical estimate of M-levels. O. K. Rice. Phys. Rev. (2) 31, pp. 1051–1059 (1928). Electrocappillarity. O. W. Richardson. Proc. Roy. Soc. A117, pp. 719–730 (1928). Cold discharge. L. Rosenfeld. E. E. Witmer. Z. Physik. 47, pp. 517–521 (1928). “Gas” of radiation. L. Rosenfeld. E. E. Witmer. Z. Physik. 49, pp. 534–541 (1928). Refractive index of metals for electron waves, the magnitude \(Wa\). R. Ruedy. Phys. Rev. (2) 32, pp. 974–978 (1928). Dispersion of metals. E. Schroedinger, Phys. Zts. 27, pp. 95–101 (1927). A. Sommerfeld. Z. Physik. 47, pp. 1–32, 43–60 (1928). Fundamental work on the application of Fermi statistics to electrons in metals. Preliminary note. Naturwiss. 15, pp. 825–832 (1927). A. Sommerfeld. Naturwiss. 16, pp. 374–381 (1928). Extension of the article just cited. A. T. Watermann. Phil. Mag. (7) 6, pp. 965–970 (1928). Dependence of electrical conductivity on pressure. A. T. Watermann. Proc. Roy. Soc. A121, pp. 28–41 (1928). Cold discharge. G. Wentzel. Probleme der modernen Physik (Hirzel, 1928), pp. 79–87. Photoelectric effect.
-
Drude calculated the ratio \(\frac{k}{e}\), knowing neither \(k\) nor \(e\) separately. He made use of the fact that this ratio is equal to \(\frac{N_{0}k}{N_{0}e}\) (\(N_{0}\) is the number of molecules in a gram-molecule, the so-called Loschmidt number), where \(N_{0}k\) is the gas constant \(R\), and \(N_{0}e\) is Faraday’s constant for the electron. ↩↩↩↩↩↩↩↩↩↩↩
-
In fact, here we are dealing not with a loss, but with an increase in the number of particles, since \(\dfrac{dt}{d\xi}<0\). ↩↩