THEORY OF THE METALLIC STATE \*
L. Nordheim
Submitted 1935 | SovietRxiv: ru-193501.88107 | Translated from Russian

Full Text

THEORY OF THE METALLIC STATE *

L. Nordheim

Introduction

The phenomenological theory of electricity regards electrical conductivity as a certain material constant, depending in a definite way on the state of the substance (for example, on temperature). In thermodynamics it proves possible to relate electrical conductivity to other phenomenological constants and to establish a number of general relations describing the phenomena of electrical conduction, thermal conduction, and thermoelectricity.

However, contemporary physics cannot be satisfied with such a theory, which has the character of a purely descriptive one. In particular, the task of atomistics is to reduce the whole diversity of empirical constants to a small number of universal constants, and in atomic theory we already have a quite satisfactory solution of this problem. In the theory of electrical phenomena, the first attempts in this direction belong to Riecke, Drude, and Lorentz.^1 Yet, despite the considerable successes of the first works, the further development of the theory proved to be connected with very considerable difficulties and ultimately led to a number of contradictions. Only in recent years has the successful development of quantum theory and statistics made it possible to give, in general, a satisfactory picture of the phenomena. The present article is devoted to an exposition of the results achieved.

It should be noted from the very beginning that the final goal has not yet been attained here. The theory of the electrical, magnetic, and thermal properties of a metal ought, strictly speaking, to consist in the solution of the more general problem of the structure of the crystal. At present there can hardly be any doubt that we already know the basic physical laws underlying the solution of this general problem. Nevertheless, we are only at the very beginning of the path that should lead to the full development of the theory. At present we are still compelled to make various hypotheses and assumptions, which ought to be substantiated more precisely in the future.

* Müller-Pouillet, IV, 4, 11th ed., translated by S. G. Kalashnikov.

The development of all modern physics has shown that electricity has an atomistic nature and is associated with definite material carriers: electrons, nuclei, ions. Electrical conductivity therefore denotes the transport of carriers of electricity. Tolman’s experiments showed that in metals these carriers are electrons. Consequently, in a metal there must be a certain number of electrons that are less bound in their motion than the electrons entering into the composition of atoms and molecules. Thus we arrive at the notion of conduction electrons as a special kind of mobile gas, which can flow under the action of an external electric field and thereby accounts for the electrical conductivity of a metal. We shall call these electrons “free electrons.” This concept must at first denote only the highest degree of mobility.

Completely free motion of the electrons would, obviously, also give rise to infinite electrical conductivity of the metal, since under the action of an external field such electrons could be accelerated arbitrarily. Therefore it proves necessary further to assume that the electrons continually undergo a braking, for example as a result of collisions, so that the kinetic energy accumulated by the electrons is transferred in the collision processes to the atoms of the metal.

§ 1. Elementary theory of Drude. Difficulties of classical conceptions

We thus arrive at a certain approximate picture of the phenomenon of electrical conductivity, which is the starting point of Drude’s theory. Since all other classical theories are only a generalization and supplement to Drude’s theory, we shall dwell on it in somewhat greater detail.

Let \(N\) be the number of free electrons in one cubic centimeter of a metal. This quantity may also be a function of the state of the substance, varying, for example, with temperature, pressure, and so on. Let the mean velocity of the thermal motion of the electrons be \(v_0\). When an external field \(F\) is applied (for example, in the direction \(X\)), the electrons acquire an additional component of velocity in the direction of the field. With respect to the collision process we shall make, formally, the simplest assumption: that the additional momentum is, on the average, completely lost by the electron after a mean free path \(l\) in a collision with an atom of the metal and passes into thermal motion. The mean velocity of the electrons in the direction of the field determines both the transport of electricity and the transport of kinetic energy (heat).

Let us calculate all these quantities quantitatively. The acceleration of an electron is \(\dot{x} = \dfrac{eF}{m}\), and the additional velocity acquired by an electron during the time \(t\) will be \(\Delta v = t\,\dfrac{eF}{m}\). By \(t\) one should understand the time spent

traversed by an electron over the path \(l\), equal (if \(\Delta v\) is considerably smaller than \(v_0\)) to \(t=\dfrac{l}{v_0}\), and, consequently:

\[ \Delta v=\frac{el}{mv_0}\,F. \]

The mean value of the additional velocity is equal to half the maximum, and we obtain:

\[ \overline{\Delta v}=\frac{el}{2mv_0}\,F. \tag{1} \]

The magnitude of the current produced by the electrons in the direction of the field we shall find to be:

\[ i_x=\sum e\overline{v}_x=\sum e(\overline{v}_{0x}+\overline{\Delta v})=\sum e\,\overline{\Delta v}, \]

since the mean value of the velocity component of the chaotic thermal motion vanishes. Taking (1) into account, we find the current produced by the field \(F\):

\[ i=\frac{Ne^2l}{2mv_0}\,F \tag{2} \]

and the electrical conductivity

\[ \varkappa=\frac{Ne^2l}{2mv_0}. \tag{3} \]

We obtain Ohm’s law. This Drude formula follows directly (to within a numerical factor) from dimensional considerations as well; \(\varkappa\) must be proportional to: the number of electrons \(N\), the acceleration \(\dfrac{e}{mv_0}\) in a field of unit intensity, the charge of the electron, and the mean free path \(l\).

Owing to the loss of energy by the electrons in collisions, heat must be liberated. This heat must be equal to the kinetic energy transmitted in collisions; per unit time it will be expressed by the collision term \(N\dfrac{v_0}{l}\), multiplied by the kinetic energy corresponding to the velocity \(\Delta v=2\overline{\Delta v}\). Therefore, for the amount of heat liberated in each cubic centimeter of metal per unit time, we obtain:

\[ Q=N\frac{v_0}{l}\frac{m}{2}\left(\frac{elF}{mv_0}\right)^2 =\frac{Ne^2lF^2}{2mv_0} =\frac{i^2}{\varkappa}, \tag{4} \]

i.e. Joule’s law.

Up to now we have not at all touched upon the question of the magnitude of the velocity of the chaotic motion of the electrons in the absence of a field. Let us make the plausible assumption that the laws of this motion are the same as for an ordinary gas. Then

\[ \frac{1}{2}mv_0^2=\frac{3}{2}kT. \tag{5} \]

Let us note that this assumption is in any case necessary if we do not wish to go beyond the framework of classical statistics. Then we can describe the phenomenon of thermal conduction in me-

metal just as is done in the case of an ideal monatomic gas. For the coefficient of thermal conductivity we obtain in this case:

\[ \lambda=\frac{1}{3} C_v v_0 l=\frac{1}{2} N v_0 l k \tag{6} \]

(since \(C_v=\frac{3}{2}kN\)) and, on the basis of (3), (5), and (6), we find:

\[ \frac{\lambda}{\varkappa}=\frac{m v_0^2 k}{e^2}=3\left(\frac{k}{e}\right)^2 T. \tag{7} \]

Thus the ratio of the coefficients of electrical conductivity and thermal conductivity turns out to be a simple universal function of temperature, containing no material constants of the metal whatever. Relation (7) is the expression of the Wiedemann–Franz law.

The primitive theory set forth above requires refinement in many respects. The subsequent classical theories, first and foremost the theories of Lorentz and Bohr, had the aim of introducing these refinements by a more accurate calculation of mean values and a more detailed analysis of the collision processes. The results obtained in this way turned out to coincide in the main with the results of the simple theory, and the whole difference was reduced merely to the appearance of a somewhat different numerical factor. Since Lorentz’s theory will enter as a special case into the further exposition, we shall not dwell in greater detail on the development of the classical theory.

The results obtained above seem at first quite satisfactory. However, upon more detailed examination the entire state of the theory proves questionable. On the basis of the classical theory it is impossible to make unambiguous statements about those physical quantities which enter into its formulas; among such quantities are the number of free electrons, their mean velocity, and the mean free path. We shall give here only the most important considerations.

Since the electrons must take part in thermal motion on an equal footing with the atoms, the presence of free electrons must necessarily affect the value of the heat capacity of a metal. The latter is determined, according to classical ideas, only by the number of degrees of freedom, and therefore for each free electron, just as for a gaseous atom, there should on average be a heat capacity of \(\frac{3}{2}k\). The heat capacity of a metal should therefore, in comparison with an insulator, be greater by some amount determined by the number of free electrons. This, however, is not the case. The Dulong–Petit law, according to which only the vibrational degrees of freedom of the atoms of a solid are essential for the magnitude of the heat capacity, is also well justified for metals. We must, consequently, assume that the number of free electrons is small in comparison with the number of atoms (of the order of \(1\%\)). But then it becomes necessary to ascribe to the electrons an abnormally large mean free path, of the order of hundreds of interatomic distances, in order that formula (3) should give the correct value of the electrical conductiv-

ties. The admission of such a mean free path is associated, in the classical theory, with great difficulties (which, as will be seen below, disappear in wave mechanics).

However, even after making such an admission, we still do not remove all the difficulties. There are a number of grounds for believing that the number of conduction electrons is considerably larger. Thus, for example, it is difficult to understand how such a small number of electrons is compatible with the considerable temperature dependence of the electrical conductivity. Since the mean velocity \(v_0\) increases with temperature as \(T^{\frac{1}{2}}\), while at the same time experiment shows that, in the region of ordinary temperatures (at low temperatures the contradictions only increase), the electrical conductivity decreases as \(\frac{1}{T}\), the product \(Nl\), according to (3), should be proportional to \(T^{-\frac{1}{2}}\). The free path, according to all classical conceptions (for example, those considering the process of elastic collisions on atoms as on solid spheres), should, to a first approximation, be independent of temperature; consequently, we come to the conclusion that the number of conduction electrons \(N\) decreases with increasing temperature. We may think that the process leading to the formation of free electrons must be similar to dissociation or ionization. But any such process would always give a strong increase in the number of electrons with temperature, a result directly opposite to the experimental data. An approximate constancy of the number of electrons \(N\) could be expected if this number of electrons were a multiple of the number of atoms; in that case one could expect that the further removal of electrons would require a significantly greater expenditure of energy. Such a conception, leading to a number of conduction electrons of the same order as the number of atoms, should be preferred to all others.*

One could cite a whole series of other phenomena, such as, for example, electro-optical effects, thermoelectric phenomena, etc., which likewise lead to values of \(N\), \(e\), and \(v\) that, within the framework of the classical theory, do not agree with one another.

Despite a whole series of very ingenious attempts, the prospects for creating a satisfactory theory remained hopeless until recent years, when our conceptions of the mechanics and statistics of electrons underwent a fundamental revision. The merit of Pauli³ and Sommerfeld⁴ consisted in transferring the new

* Direct experimental proof that the number of conduction electrons is comparable with the number of atoms was given recently by I. Kikoin and I. Fakidov² in studying the Hall effect on molten alkali metals. For molten alkali metals one may expect the best agreement with theory (here both the classical and the new theories lead to almost identical results), since the alkali metals are typical metals. In addition, the influence of the anomalies discussed in Part IV of the present article should be minimal for them.

…notions into the theory of the metal. Then, in the works of a number of authors, above all Houston, Bloch, and Peierls,* the theory was developed and deepened to such an extent that at the present time we already have a quite satisfactory picture of a large number of phenomena connected with the electrical conductivity of metals.

§ 2. Classical and quantum-mechanical description of the system

The significance of quantum mechanics for the theory of metals is determined by two circumstances. First, the laws of motion of the electron are changed in a fundamental way as compared with classical mechanics; second, the statistics is changed. By the latter we mean the formation of mean values for a large number of systems or, correspondingly, the determination of the most probable state of such an ensemble for given external parameters—total energy, volume, etc. Since the study of statistics requires only quite negligible knowledge from the field of quantum mechanics, while at the same time making it possible to obtain many important results, we shall base the exposition on statistics. Later (Part IV) we shall also consider those questions for which the alteration of the laws of mechanics themselves proves to be essential.

From the logical point of view, of course, the reverse course would be more expedient. However, with the method of exposition we have chosen, we gain the advantage that at the outset we do not bind ourselves to any special model representations and can first examine in detail those questions which are in no way connected with them.

The quantum-mechanical change in the laws of motion manifests itself in statistics in the fact that only possible stationary states with discrete values of the energy are considered.** Such an approach is not new; it originates in Bohr’s classical theory, and here the new quantum mechanics introduces no additional complications. For the statistical method it is completely immaterial what objects are to be studied and, consequently,

* Houston’s work⁵ is very important historically, since in it it was first clearly shown that, for the electrical resistance of a metal, the decisive factor is the disturbance of the crystal lattice by thermal vibrations. In detail, however, the theory does not correspond to the present state of quantum mechanics. The foundations of the modern theory were given by Bloch⁶ and were substantially supplemented and developed by Peierls⁷ (see also Nordheim’s article⁸). A survey of the exposition of the theory in elementary form is given by Darrow,⁹ and quite thoroughly by Brillouin.¹⁰ The latter works also serve as introductions to quantum statistics. For quantum statistics specifically one should mention the monographs of Fowler, Uhlenbeck, and Peierls.¹¹

** Here it will be sufficient for us to consider only the case of a discrete energy spectrum, which obtains whenever the system under consideration is confined to a finite volume. This assumption proves unavoidable in all statistical calculations. Therefore we shall not enter here into an analysis of the difficulties that arise for systems with a continuous energy spectrum (aperiodic motions).

what variables we use to describe the state of a system; the only essential point is that these variables determine the state unambiguously. In each individual case we shall indicate which parameters are chosen for the description.

In classical theory one uses the phase space of canonical coordinates and momenta. The state of a system (at a definite instant of time) is determined by specifying the volume element (cell) of phase space in which the representative point is to be found; moreover, for an exact description of the system it proves necessary to pass to a division of phase space into infinitely small cells.

In the old Bohr theory the cells were chosen of a quite definite size and form, and to each cell there corresponded a strictly definite value of the energy. To such a division of phase space into cells there corresponds, in the new quantum theory, the enumeration of Schrödinger wave functions; here, too, by specifying the number (the quantum number) the energy is also determined.* We may therefore retain the term “cell” for this case as well. The circumstance that the coordinates are then not exactly determined and that only a certain probability distribution is given for them is in no way an obstacle. (The change in the method of describing position in space must, of course, be taken into account in those questions where this spatial distribution enters directly.) In § 6 we shall show in more detail that, for the systems of interest to us, the subdivision of phase space into cells is indeed equivalent to the enumeration of eigenfunctions.

Quantum mechanics possesses one feature which leads to very substantial consequences. Namely, for an assembly of identical particles, quantum mechanics gives new integrals of the equations of motion. Thus, for example, when we have the simplest case of two identical particles not interacting with one another, then in classical physics the state of such a system is determined by the state of each of the particles separately. Suppose, for example, particle 1 is in state \(a\), and particle 2 in state \(b\). This state will be macroscopically identical with the state in which particle 1 is in \(b\), and particle 2 in \(a\). Meanwhile, in counting the different states in classical statistics, the state \(a + b\) would have to be counted twice. By observations we could, of course, only establish that some one (1 or 2) of the particles is in state \(a\), and some other one (2 or 1) is in state \(b\). Einstein expressed this state of affairs in the following words: one cannot paint one electron red and the other green; we cannot introduce such distinguishing marks that would not affect the other properties of the system.

* We shall not consider the analysis of a much more general method of describing the state of a system, which is provided by the transformation theory of quantum mechanics.

However, there is one logical requirement that must be imposed on any theory, provided it is to be more than a mere description: it must contain no elements that are in principle unobservable. Therefore there is nothing surprising in the fact that quantum theory, which in fact seriously operates with this requirement, assumes that the two states discussed above, indistinguishable from one another, must in reality be regarded as one and the same.

This leads to a substantial change in all statistical calculations. At the same time, however, it should be especially emphasized that the new method of counting is, of course, not applicable to such systems as in reality consist of distinguishable elements. In these cases classical statistics remains valid. The latter applies, for example, to a game of dice, since dice can perfectly well be painted in different colors without any alteration of their properties essential for the game. The considerations set forth are readily generalized to the case of an arbitrarily large number of identical particles. In quantum statistics, states differing only by a permutation of identical particles are not regarded as distinct and therefore are not counted separately.

In the statistics of electrons it is necessary to take into account one further circumstance. The study of atomic spectra leads to the conclusion that a special restriction is imposed on the state of electrons in an atom (the Pauli principle), according to which in each stationary state there can be at most one electron.* This principle, as is known, makes it possible to explain the construction of the periodic system of the elements and the structure of spectra and may therefore be regarded as established with the highest degree of reliability. Such an additional restriction must evidently have very substantial consequences also in statistical calculations.

Mathematically the considerations set forth may be represented as follows. Suppose, for example, that there are two identical particles, and let us first assume that there is no interaction between them; then the Schrödinger equation for the whole system will be:

\[ \{H(q_1)+H(q_2)\}\psi(q_1q_2)=E\psi(q_1q_2). \]

Here \(H(q_i)\) are the energy operators for the separate systems: they have the same form and differ only in their arguments. \(E\) is the energy parameter, \(\psi(q_1q_2)\) the wave function of the complete system. We shall also take into account the spin of the electron, i.e., we shall regard \(q\) as denoting symbolically both the spatial coordinates and the spin. The equation separates if we put:

\[ \psi_{nm}(q_1q_2)=\psi_n(q_1)\cdot\psi_m(q_2), \tag{1a} \]

\[ E_{nm}=\varepsilon_n+\varepsilon_m, \tag{1b} \]

* This principle was established by Pauli \(^{12a}\) at first purely empirically in the study of atomic spectra. The quantum-mechanical classification into symmetry classes was discovered by Heisenberg \(^{13}\) and Dirac. \(^{14}\)

i.e., the eigenfunction of the complete system can be represented as the product of the eigenfunctions of the component systems, and the total energy is simply equal to the sum of the separate energies. Then for each of the component systems we obtain the equation:

\[ H(q)\psi_n(q)=\varepsilon_n\psi_n(q). \tag{2} \]

In addition to (1a), the same energy value will also correspond to the function \(\psi_{nm}=\psi_m(q_1)\psi_n(q_2)\), which differs from the first only by a permutation of the arguments. We therefore have here a degenerate system. In the presence of some perturbation, for example in the case of interaction of the particles, as the solution one would have to take not the functions \(\psi_{nm}\) themselves, but a suitable linear combination of them. As calculation shows, in this case the eigenfunctions of the zeroth approximation are obtained with different symmetry properties. Namely, upon permutation of both particles the eigenfunction either does not change at all (the symmetric case), or changes sign (the antisymmetric case). This symmetry property is preserved also in the presence of any perturbation that acts only identically on both particles. Thus all solutions are split into non-combining series, each of which corresponds to its own symmetry.

For the case of many particles the result is analogous. We again first considered the case of absence of interaction. Here too one obtains a splitting of the solutions into non-combining series with definite symmetry properties. Of these, two types should be especially noted, one of which is completely symmetric and has the eigenfunction:

\[ \psi^s_{n_1,n_2,\ldots}=C\sum_P \psi_{n_1}(q_1)\psi_{n_2}(q_2)\ldots, \tag{3} \]

and the other is completely antisymmetric:

\[ \psi^s_{n_1,n_2}=C \left| \begin{array}{cccc} \psi_{n_1}(q_1) & \psi_{n_2}(q_1) & \ldots & \psi_{n_n}(q_1)\\ \psi_{n_1}(q_2) & \psi_{n_2}(q_2) & \ldots & \psi_{n_n}(q_2)\\ \cdot & \cdot & \cdot & \cdot\\ \psi_{n_1}(q_n) & \psi_{n_2}(q_n) & \ldots & \psi_{n_n}(q_n) \end{array} \right|. \tag{4} \]

In expression (3) the summation is carried out over all permutations \(q\), and the resulting function is therefore symmetric with respect to all particles, whereas the determinant (4), when any two rows (i.e. particles) are interchanged, changes, as is known, its sign, i.e. is antisymmetric with respect to them. \(C\) denotes a constant normalizing factor.

Each of the eigenfunctions written down (and along with it the state belonging to one of the two classes of symmetry considered) is uniquely determined by specifying a series of numbers \(n_1,\ldots,n_z\), without regard to the order of the series. This series of indices of the eigenfunctions indicates how many particles are in the given-

in the same state (cell).* At the same time one cannot ask which particular particles are in one state or another; this is precisely what leads to the new method of counting mentioned above. In the antisymmetric case the determinant becomes zero when any two indices (the states of individual particles) are equal to one another, since then the corresponding rows of the determinant become identical. For such a state there is no antisymmetric eigenfunction. We thus obtain an analytic formulation of the Pauli exclusion principle.

For the many-body problem, just as for the case of two particles, the symmetry properties are preserved also in the presence of any perturbation, provided only that it acts in the same way on all particles; the latter occurs, for example, in the interaction of particles (in this case, of course, the special forms (3) and (4) are not preserved). Thus, if a system at some moment in time (for example, at the formation of the world) was in a state with a definite symmetry, then this character of the symmetry will be preserved for all subsequent times. Quantum mechanics here gives no grounds for selecting a priori any one of these symmetry classes; it guarantees only its preservation. We thus have here a far-reaching analogy with the integrals of classical mechanics.

If the Pauli principle holds for electrons in atoms (i.e., if even two electrons cannot be in the same quantum state in atoms), then it must hold for all electrons in the universe. Otherwise, upon replacing an atomic electron (for example by ionization and the subsequent capture of some other electron), the symmetry character of the atom itself could change, which is not observed. This unambiguously determines that, with respect to all electrons, the world is antisymmetric. We shall have to touch later on the question of the character of the symmetry and the statistics of other particles (protons, etc.) (§ 7).

§ 3. New statistics

The fundamental problem of statistics is the following. There is an assembly of \(N\) identical particles, whose interaction may, in a first approximation, be neglected. Each such particle may occupy a series of stationary states in which it has the energy \(\varepsilon_1, \varepsilon_2, \ldots, \varepsilon_k\), and the total number of particles \(N\) is very large. The question is posed of the behavior of such a system as a whole.

For describing the entire system, various possibilities present themselves, depending on the required degree of accuracy.

I. Microscopic distributions, or complexions. For each individual particle it must be specified

* Such an unambiguous definition by means of numbering is possible only for these two cases and is impossible for a large number of intermediate symmetry classes. Thus these two classes are the only ones satisfying the logical requirement of the indistinguishability of identical particles.

its state. Such a method of description, both in classical physics and in quantum physics, far surpasses all possibilities of observation.

II. The distribution function. The possible energy states are divided into intervals (numbered by the index \(s\)), each interval comprising values of the energy so close together that, for the latter within a single interval, only their mean value \(\varepsilon_s\) can be specified. Let \(A(s)\) be the number of stationary states (cells) within the interval \(s\), and \(N_s\) the number of particles. The distribution function is determined by specifying these quantities. For a sufficiently large number of particles and a sufficiently fine subdivision, the distribution function may be represented by a smooth curve \(f(\varepsilon)\), if the abscissae are taken to be \(\varepsilon_s\), and the ordinates the number of particles corresponding to this value of the energy (Fig. 1). By also marking the divisions of the intervals on the axis of abscissae, we obtain \(N_s\) as the mean height of the shaded portion, and

\[ N_s = f(\varepsilon) A_s . \]

In general, for the number of particles corresponding to the energy interval \(d\varepsilon\), we obtain \(f(\varepsilon)\, d\varepsilon\). Such a subdivision into intervals is a purely formal device and in no case has any physical meaning.

Fig. 1. Distribution function; cells and intervals.

Fig. 1. Distribution function; cells and intervals.

III. The macroscopic state. Only quantities pertaining to the whole system as a whole are specified: the total energy (temperature), the total force referred to some external parameter (pressure), etc.; in other words, thermodynamic quantities.

Some definite distribution function may be realized by means of a large number of different complexes, and each macroscopic state by a large number of different distribution functions. The fundamental premise of all statistics, including quantum statistics, is contained in the following assertion.

For given external parameters there is one distribution function which will be realized predominantly often, namely that for which the number of complexes is the greatest. This distribution dominates over all the others to such an extent that we may be certain that in reality we shall almost always encounter precisely this distribution.* For each macroscopic state there exists such a definite distribution function. The mean value of any quantity is always

\[ \text{* The probability of a deviation of the state from the most probable one can also be calculated statistically, as, for example, is done in the theory of fluctuations. Let us note that the latter, in quantum statistics, can be carried out with the same generality as in classical statistics.} \]

will correspond precisely to this most probable distribution.

The root of this situation lies in the fact that all complexes belonging to identical macroscopic parameters are assumed to occur equally often, provided only that the system is left to itself for a sufficiently long time. In classical statistics this is justified by Liouville’s theorem, which asserts that equal elements of the volume of phase space correspond to equal a priori probabilities, and by a hypothesis of the ergodic type, according to which any one of the states under consideration can actually be realized in the course of time. In quantum mechanics, Liouville’s theorem is replaced by the assertion that all discrete, nondegenerate states possess equal statistical weights.* The special ergodic hypothesis thereby proves superfluous, since quantum mechanics permits the corresponding law to be derived directly.^16 We shall not, however, deal here in greater detail with these difficult questions of principle.

These general foundations are valid both for classical and for quantum statistics. In the latter only the method of counting complexes is changed, in accordance with the necessity of taking into account the results set forth in § 2. Thus a complex is specified as follows:

I. Classical theory (Boltzmann): for each particle a definite cell is specified in which it is located. Complexes obtained by permuting particles among different cells are considered distinct.

II. Quantum theory: for each cell the number of particles contained in it is specified. Complexes obtained by permuting particles among different cells are considered identical. Here it is necessary further to distinguish the following cases:

a) any number of particles may be found in each cell: Bose—Einstein statistics** (the symmetric case);

b) in each cell there may be no more than one particle: Fermi—Dirac^19 statistics (Pauli principle, antisymmetric case).

The following example explains the different methods of counting the number of complexes. There are two balls (particles), which we distribute among three urns (cells). The possible complexes are shown in the diagram given. In Boltzmann statistics we may number the particles; in the other two cases this cannot be done, and the particles are shown by crosses.

Of the first six states of Boltzmann statistics, each pair is reproduced as one state in Einstein—Bose (and Fermi—Dirac) statistics; the last three (both balls in one and the same urn) are considered the same in both statistics. In Fermi—Dirac statistics, however, the latter are not allowed at all.

* That this follows necessarily from the principles of quantum mechanics was shown by Dirac.^15

** On the original application (before the development of quantum mechanics) to light quanta, see Bose’s paper.^17 The physical significance was indicated by Einstein.^18

DIAGRAM 1

Boltzmann

\(a\) \(b\) \(c\)
1 2 0
2 1 0
1 0 2
2 0 1
0 1 2
0 2 1
1,2 0 0
0 1,2 0
0 0 1,2

9 complexes

Einstein—Bose

\(a\) \(b\) \(c\)
× × 0
× 0 ×
0 × ×
×× 0 0
0 ×× 0
0 0 ××

6 complexes

Fermi—Dirac

\(a\) \(b\) \(c\)
× × 0
× 0 ×
0 × ×

impossible

3 complexes

Let us further clarify, by this example, typical questions arising in statistics.

  1. The probability that a definite particle (for example 1) is in a definite urn (for example \(a\)). This question has meaning only in Boltzmann statistics. We compute according to the rule: the probability is equal to the ratio of the number of favorable cases to the total number of possibilities, i.e. \(3:9, 1:3\).

  2. The probability that some one particle is in urn \(a\) (by the same rule):

Boltzmann (B) \(4:9\); Einstein—Bose (E.—B.) \(2:6\); Fermi—Dirac (F.—D.) \(2:3\).

  1. The mean number of particles in urn \(a\), i.e. the sum of the particles in \(a\) over all states, divided by the number of states:

B. \(6:9\); E.—B. \(4:6\); F.—D. \(2:3\).

  1. The probability of finding not a single particle in \(a\):

B. \(4:9\); E.—B. \(3:6\); F.—D. \(1:3\).

  1. The probability of finding two particles at once in \(a\):

B. \(1:9\); E.—B. \(1:6\); F.—D. \(0\).

This example shows how much the results of the three statistics differ from one another.

It should be emphasized once more that the new statistics are applicable only when we are dealing with identical particles. If, on the contrary, there is any possibility whatever of distinguishing the latter from one another, then Boltzmann statistics must always be applied; the proper vibrations of an elastic body may serve as an example, since they are all distinct from one another.

Let us proceed to the counting of the number of complexes of a definite distribution function. The latter is specified, as already mentioned, by the numbers of particles \(N_s\) belonging to the various intervals; the number of cells in the interval is \(A_s\). It is required to determine in how many different ways such a distribution can be realized.

I. Boltzmann statistics. First let us distribute the \(N_s\)-particles in the interval \(s\) among its \(A_s\)-cells. This is the problem of distrib-

of distributing \(N_s\) balls among \(A_s\) urns, with an arbitrary number of balls possibly being in each urn. Such a distribution can be carried out in \(A_s^{N_s}\) ways, since each particle may be in any cell, independently of how many particles are already there. In this case distributions that differ in the order in which the particles follow one another (definite particles in definite cells) are counted as distinct. Each distribution of the interval \(s\) can be combined with any distribution of another interval, and the total number of such possibilities is represented by the product

\[ \prod_s A_s^{N_s}. \]

At the same time, however, we have not yet taken into account the possibility of permuting particles between different intervals, whereas the distributions resulting from this in Boltzmann statistics must be regarded as distinct. The number of such permutations, as is known, is equal to

\[ \frac{N!}{\prod_s N_s!}, \]

where \(N=\sum_s N_s\) is the total number of particles. Finally the number of complexes will be:

\[ K_{\mathrm{B}}=N!\prod_s \frac{A_s^{N_s}}{N_s!}. \tag{1} \]

II. Bose—Einstein Statistics. Here the problem consists in distributing \(N_s\)-particles among \(A_s\)-cells without taking account of the sequence of particles. For this purpose one proceeds as follows. All the elements (cells and particles) are arranged in a row in an arbitrary sequence, but in such a way that at one end, for example on the right, there is a cell. Gathering all particles into the nearest cells to their right, we obtain a definite distribution of particles among cells. The total number of possible combinations will be given by the number of permutations of \(N_s+A_s-1\) elements (one cell drops out), i.e. \((N_s+A_s-1)!\). This number must, however, be reduced, since all combinations that differ by a permutation of particles (or cells) give one and the same distribution. Consequently, the written number must also be divided by \(N_s!\) and by \((A_s-1)!\) (one cell was fixed). We thus obtain for the number of different combinations within an interval:

\[ \frac{(N_s+A_s-1)!}{N_s!(A_s-1)!}. \]

The total number of complexes is again obtained by multiplying the numbers for the separate intervals:

\[ K_{\mathrm{E.-B.}}=\prod_s \frac{(N_s+A_s-1)!}{N_s!(A_s-1)!}, \tag{2} \]

where the written expression is already final, since permutations of particles between different intervals introduce nothing new (multiplication by the factor \(\frac{N}{\prod_s N_s!}\) is not required).

III. Fermi—Dirac statistics. The problem is the same as in case II, except that in each cell there can be no more than one particle. The number of cells must therefore be greater than the number of particles. The problem may also be formulated as follows: from the number \(A_s\) of cells, \(N_s\) contain one particle each, the remaining ones are empty. It is necessary, in this way, to divide \(A_s\) objects into two groups, of sizes \(N_s\) and \(A_s - N_s\). The number of ways, as is known, is equal to:

\[ \frac{A_s}{N_s!(A_s-N_s)!}. \]

The total number of complexions will therefore be

\[ K_{\Phi-\mathrm{D}.}=\prod_s \frac{A_s}{N_s!(A_s-N_s)!}. \tag{3} \]

It is again unnecessary to take into account the permutation of particles among the various intervals.

§ 4. The Most Probable Distribution

We have calculated the number of complexions for a given distribution of the numbers \(N_s\). According to the general program of statistics, one must now, among all possible distributions, seek the one for which the number of complexions is the greatest.* In doing so, one should consider only those states which, under the given external conditions, could transform into one another; for such states the total number of particles and the total energy must be prescribed (the first condition, however, is not obligatory for light quanta). The problem reduces to finding the maximum of \(K\) as a function of \(N_s\) under the additional conditions:

\[ \sum_s N_s=N, \tag{1} \]

\[ \sum_s \varepsilon_s N_s=E. \tag{2} \]

In practice it is convenient to seek the maximum of \(\lg K\), which can always be done, since \(\lg\) is a monotone function of its argument. Further we shall make use of Stirling’s formula:

\[ \lg N! = N \lg N - N, \tag{3} \]

the application of which here is quite legitimate, since the division into intervals can always be carried out so that only factorials of very large quantities enter.

* The method under consideration leads most quickly to the goal in simple cases. However, the method connected with canonical ensembles of Gibbs is logically and mathematically more satisfactory. Some more complicated problems (for example, fluctuations) are solved very elegantly by the latter method. Here, however, a somewhat more extensive mathematical apparatus is required. We therefore confine ourselves only to an indication of the literature.²⁰

The resulting simple problem for a relative maximum is solved in the usual way. Multiplying the auxiliary conditions by the Lagrange indeterminate multipliers \((-\alpha,\ -\beta)\) and adding them to the original function, we require that the derivatives of the resulting new function with respect to the varied variables vanish. The values of the multipliers are obtained, as usual, by again using the auxiliary conditions. We then obtain:

I. Boltzmann statistics. According to § 3, (1) and § 4, (3):

\[ \lg K_{\mathrm{B}}=N\lg N+\sum_s \left(N_s\lg A_s-N_s\lg N_s\right), \]

the function to be differentiated (by Lagrange’s rule) is:

\[ L_{\mathrm{B}}=N\lg N+\sum_s \left(N_s\lg A_s-N_s\lg N_s\right)-\alpha\sum_s N_s-\beta\sum_s \varepsilon_s N_s, \]

and hence the extremal conditions are:

\[ \frac{\partial L_{\mathrm{B}}}{\partial N_s}=\lg A_s-\lg N_s-1-\alpha-\beta\varepsilon_s=0. \]

This gives:

\[ N_s=A_s e^{-1-\alpha-\beta\varepsilon_s}=\frac{A_s}{e^{1+\alpha+\beta\varepsilon_s}} . \tag{4} \]

— the Maxwell–Boltzmann distribution law. The constants \(\alpha\) and \(\beta\) can be determined by substitution into (1) and (2). Their physical meaning will be discussed below.*

II. Bose–Einstein statistics. The calculations are carried out according to exactly the same scheme [cf. § 3, (2)]:

\[ \lg K_{\mathrm{E.-B.}}=\sum_s\left\{(N_s+A_s)\lg (N_s+A_s)-N_s\lg N_s-A_s\lg A_s\right\}. \]

(We neglect unity in comparison with the large numbers \(N_s\) and \(A_s\).)

\[ \left. \begin{aligned} L_{\mathrm{E.-B.}}&=\sum_s\left\{(N_s+A_s)\lg (N_s+A_s)-N_s\lg N_s-A_s\lg A_s\right\}\\ &\quad-\alpha\sum_s N_s-\beta\sum_s \varepsilon_s N_s;\\[4pt] \frac{\partial L_{\mathrm{E.-B.}}}{\partial N_s}&=\lg\left(\frac{A_s}{N_s}+1\right)-\alpha-\beta\varepsilon_s=0,\\[4pt] N_s&=\frac{A_s}{e^{\alpha+\beta\varepsilon_s}-1}. \end{aligned} \right\} \tag{5} \]

The last expression gives the Bose–Einstein distribution law.

* Usually, as the Lagrange factor, instead of \(\alpha\) one introduces \(\alpha^*=\alpha+1\), in order to eliminate the 1 in the exponent (4). However, such a device violates the symmetry in the formulas of the three statistics and creates difficulties in thermodynamic interpretation. With the notation chosen by us, this does not occur.

III. Fermi—Dirac Statistics [cf. § 3, (3)]:

\[ \lg K_{\mathrm{F.-D.}}=\sum_s\left\{-N_s\lg N_s-(A_s-N_s)\lg(A_s-N_s)+A_s\lg A_s\right\}, \]

\[ \left. \begin{aligned} L_{\mathrm{F.-D.}}&=\sum_s\left\{-N_s\lg N_s-(A_s-N_s)\lg(A_s-N_s)\right.\\ &\qquad\left.+A_s\lg A_s-\alpha N_s-\beta\varepsilon_s N_s\right\},\\[4pt] \frac{\partial L_{\mathrm{F.-D.}}}{\partial N_s} &=\lg\left(\frac{A_s}{N_s}-1\right)-\alpha-\beta\varepsilon_s=0,\\[4pt] N_s&=\frac{A_s}{e^{\alpha+\beta\varepsilon_s}+1}. \end{aligned} \right\} \tag{6} \]

The last expression gives the Fermi—Dirac distribution law. That the distribution laws obtained correspond to a maximum of the number of complexes could be easily verified by calculating the second derivatives.

Introducing the symbol

\[ \gamma= \begin{cases} 0 & \text{for B.}\\ -1 & \text{for E.—B.}\\ +1 & \text{for F.—D.} \end{cases} \tag{7} \]

we can combine the formulas of the various statistics. We obtain:

\[ N_s=\frac{A_s}{e^{1-\gamma^2+\alpha+\beta\varepsilon_s}+\gamma} \tag{8} \]

(up to the constant term \(N\lg N\) in Boltzmann statistics, which plays no role);

\[ \lg K=\sum_s\left\{N_s\lg\frac{A_s-\gamma N_s}{N_s}-\gamma A_s\lg\frac{A_s-\gamma N_s}{A_s}\right\}= \]

\[ =\sum_s\left\{(N_s-\gamma A_s)\lg(A_s-\gamma N_s)-N_s\lg N_s+\gamma A_s\lg A_s\right\}. \tag{9} \]

Substituting (8) into (9) and noting that

\[ \frac{A_s-\gamma N_s}{N_s}=e^{1-\gamma^2+\alpha+\beta\varepsilon_s}, \qquad \frac{A_s-\gamma N_s}{A_s} =\frac{1}{1+\gamma e^{-1+\gamma^2-\alpha-\beta\varepsilon_s}}, \]

we find for the number of complexes \(\overline K\) in the most probable state:

\[ \lg\overline K =\sum_s\left\{N_s(1-\gamma^2+\alpha+\beta\varepsilon_s) +\gamma A_s\lg(1+\gamma e^{-\alpha-\beta\varepsilon_s})\right\} \tag{10} \]

(in calculating the second term we neglect the influence of \(-1+\gamma^2\) in the exponent in comparison with the factor \(\gamma\)).

The introduction of intervals \(s\) and of the numbers \(A_s\) was, in the preceding arguments, only an artificial device that made it possible to operate throughout with large numbers and to apply Stirling’s formula. In order to reduce the influence of this arbitrary choice on the final result, one may introduce, instead of the quantities \(N_s\), the numbers of particles in an interval, the most probable number of particles in a cell \(n_k\). Then we obtain:

\[ n_k=\frac{N_s}{A_s} =\frac{1}{e^{1-\gamma^2+\alpha+\beta\varepsilon_k}+\gamma}. \tag{11} \]

The numbers \(n_k\) are now no longer large and, moreover, for example in F.—D. statistics, always \(n_k\leq 1\). The number of complexes can be expres-

expressed through the number \(n_k\); moreover, here, instead of summing over intervals, one will already have to sum over cells \(k\):

\[ \lg K=\sum_k \left\{ n_k \lg \frac{1-\gamma n_k}{n_k}-\gamma \lg(1-\gamma n_k)\right\}= \]

\[ =\sum_k \left\{(n_k-\gamma)\lg(1-\gamma n_k)-n_k\lg n_k\right\}, \tag{12a} \]

\[ \lg \overline{K}=\sum_k \left\{ n_k(1-\gamma^2+\alpha+\beta \varepsilon_k)+\gamma \lg(1+\gamma e^{-\alpha-\beta\varepsilon_k})\right\}. \tag{12b} \]

§ 5. Thermodynamics of Statistics

It can be shown that not only the classical Boltzmann statistics, but also the new statistics, could likewise serve as a basis for thermodynamics and would make it possible to give a statistical interpretation of thermodynamic quantities (first of all, temperature \(T\) and entropy \(S\)). The natural and only logically satisfactory method for this is a consistent reproduction of the course of the ideas of thermodynamics.

For the theoretical determination of \(T\) and \(S\) in thermodynamics, one shows that for all reversible, i.e. quasi-static,* processes the amount of heat \(\delta Q\) possesses an integrating divisor; then, by definition:

\[ \delta Q=T\delta S. \tag{1} \]

If it were possible to establish this relation on the basis of statistical considerations, then thereby we would also statistically define both \(T\) and \(S\). We shall show briefly how this can be done within the framework of the statistics of reversible processes.

The total energy of the system is:

\[ E=\sum_s \varepsilon_s N_s=\sum_k \varepsilon_k n_k. \tag{2} \]

Let us consider processes of two types. In processes of the first type we shall change the energy without changing the structure of the body (influx of heat); in this case, of course, the distribution of individual systems over energy states changes, i.e. the numbers \(N_s\) change. In processes of the second type we change the very structure of the body, varying some parameter \(a\) (this includes, for example, a change in the volume of the body, which we can carry out at least with the aid of a movable piston; then \(a\) will determine the position of the piston). Such a process will cause a change in the energy values \(\varepsilon_s(a)\), which we may regard as functions of this variable parameter. If, in doing so, the numbers \(N_s\) do not change, then we call such a process adiabatic. In the general case we may represent an ongoing process as the superposition of two processes of the kind just analyzed.

* By quasi-static we shall here, just as in thermodynamics, mean such a process for which at any instant of time an equilibrium distribution law may be assumed (corresponding to the given instantaneous values of the macroscopic parameters).

For an infinitely small change of state, the energy changes by

\[ \delta E=\sum_s N_s \delta \varepsilon_s+\sum_s \varepsilon_s \delta N_s=\delta A+\delta Q, \tag{3} \]

where the first term gives the work performed, and the second the quantity of heat supplied. Thus

\[ \delta A=\sum_s N_s \delta \varepsilon_s, \tag{4a} \]

\[ \delta Q=\sum_s \varepsilon_s \delta N_s, \tag{4b} \]

\[ \delta Q=\delta E-\delta A. \tag{4c} \]

The force \(F\), referred to any one of the parameters \(a\) (for example, pressure, if \(a\) denotes volume), can be obtained by summing the separate actions of all particles, i.e.,

\[ F=-\sum_s N_s \frac{\partial \varepsilon_s}{\partial a}, \qquad \delta A=-F\delta a. \tag{5} \]

We shall now show that the statistical expression for the heat supplied (4b) does indeed have an integrating factor, i.e. for all three statistics:

\[ \delta Q=\frac{1}{\beta}\,\delta \lg \overline{K}, \tag{6} \]

where \(\overline{K}\) still denotes the number of complexes of the most probable state. Hence it necessarily follows that:

\[ \beta=\frac{1}{kT}, \tag{7} \]

\[ k\lg \overline{K}=S. \tag{8} \]

The proportionality coefficient \(k\) entering here, introduced from dimensional considerations, is not yet determined; its value must be chosen so as to give agreement with the empirical temperature scale. The factor \(k\) coincides with the Boltzmann constant entering the ordinary Maxwell distribution.

Thus we automatically obtain an entropy proportional to the logarithm of the number of complexes (the thermodynamic probability), avoiding all purely speculative arguments.

We turn to the proof of the statement just made, which we can carry out simultaneously for all three statistics. According to § 4 (10) we have:

\[ \lg \overline{K}=\sum_s \left\{ N_s(1-\gamma^2+\alpha+\beta\varepsilon_s)+\gamma A_s \lg \left(1+\gamma e^{-\alpha-\beta\varepsilon_s}\right)\right\}. \]

Let us now imagine the process discussed above, in which the total number of particles \(N=\sum_s N_s\) remains constant; for such a process the additive term \(N\lg N\) of Boltzmann statistics plays no role, and we shall therefore not take it into account at all.

Then the change in \(\lg \overline{K}\) will be:*

\[ \delta \lg \overline{K} = \sum_s N_s \delta(1-\gamma^2+\alpha+\beta \varepsilon_s) + (1-\gamma^2+\alpha)\sum_s \delta N_s + \beta \sum_s \varepsilon_s \delta N_s + \gamma \sum_s A_s \delta \lg(1+\gamma e^{-\alpha-\beta\varepsilon_s}) . \]

Since \(\sum_s \delta N_s=0\), owing to the fact that \(N=\mathrm{const}\), and since \(\gamma^2=+1\) for E.—B. and F.—D. statistics, on the basis of § 4 (8) we have:

\[ \gamma A_s \delta \lg(1+\gamma e^{-\alpha-\beta\varepsilon_s}) = \frac{\gamma A_s \gamma e^{-\alpha-\beta\varepsilon_s}} {1+\gamma e^{-\alpha-\beta\varepsilon_s}} \,\delta(-\alpha-\beta\varepsilon_s) = -\gamma^2 N_s \delta(\alpha+\beta\varepsilon_s). \]

Therefore, for the two latter statistics we obtain:

\[ \delta \lg \overline{K}=\beta \sum_s \varepsilon_s \delta N_s, \tag{9} \]

which is what was required to be proved [the relations (4b) and (6) should be taken into account]. For Boltzmann statistics \((\gamma=0)\), taking into account that

\[ N_s=A_s e^{-1-\alpha-\beta\varepsilon_s}, \]

we have:

\[ \sum_s N_s \delta(1+\alpha+\beta\varepsilon_s) = \]

\[ = \sum_s A_s e^{-1-\alpha-\beta\varepsilon_s} \delta(1+\alpha+\beta\varepsilon_s) = -\sum_s \delta N_s=0, \]

which again leads to relation (9).

For the entropy we obtain:

\[ \frac{1}{k}S = N(1-\gamma^2+\alpha) +\frac{1}{kT}E +\gamma \sum_s A_s \lg(1+\gamma e^{-\alpha-\beta\varepsilon_s}) \tag{10} \]

or, if instead of summation over intervals one introduces summation over cells:

\[ \frac{1}{k}S = N(1-\gamma^2+\alpha) +\frac{1}{kT}E +\gamma \sum_k \lg(1+\gamma e^{-\alpha-\beta\varepsilon_k}). \tag{11} \]

By relation (11) the entropy is determined only for the most probable state corresponding to thermodynamic equilibrium. Meanwhile, the concept of entropy is more general and has meaning for all states, not only equilibrium ones. We can also obtain the general expression for the entropy from formula § 4 (12a):

\[ S=k\lg K = k\sum_k \left\{ n_k \lg \frac{1-\gamma n_k}{n_k} -\gamma \lg(1-\gamma n_k) \right\}, \tag{12} \]

which is determined by specifying only the numbers \(n_k\).

* The intervals for the varied states must be chosen so that they are obtained adiabatically from the initial ones. Then \(A_s\) need not be varied.

It is often more convenient, instead of the energy \(E\), to operate with the free energy \(F\), which in thermodynamics is defined as

\[ F=E-TS. \tag{13} \]

According to (10) we have:

\[ F=-kT\{\alpha N+Z\}, \tag{14} \]

where

\[ Z=-(\gamma^2-1)N+\gamma\sum_s A_s \lg(1+\gamma e^{-\alpha-\beta\varepsilon_s})= \]

\[ =-\sum_s A_s\left\{(\gamma^2-1)e^{-1-\alpha-\beta\varepsilon_s} -\gamma\lg(1+\gamma e^{-\alpha-\beta\varepsilon_s})\right\} \tag{15} \]

is customarily called the “sum of states” (Zustandssumme). In classical statistics

\[ Z=\sum_s A_s e^{-1-\alpha-\beta\varepsilon_s} =\sum_k e^{-1-\alpha-\frac{\varepsilon_k}{kT}}; \tag{16a} \]

in quantum statistics

\[ Z=\gamma\sum_s A_s \lg(1+\gamma e^{-\alpha-\beta\varepsilon_s}) =\gamma\sum_k \lg\left(1+\gamma e^{-\alpha-\frac{\varepsilon_k}{kT}}\right). \tag{16b} \]

All the most important quantities can be simply calculated with the help of this sum. Thus, for example, always

\[ n_k=-kT\frac{\partial Z}{\partial \varepsilon_k}; \qquad N=-\frac{\partial Z}{\partial \alpha}; \tag{17} \]

the force referred to any one of the parameters \(a\):

\[ P=-\sum_k n_k\frac{\partial\varepsilon_k}{\partial a} =kT\sum_k\frac{\partial Z}{\partial\varepsilon_k}\frac{\partial\varepsilon_k}{\partial a} =kT\frac{\partial Z}{\partial a}. \tag{18} \]

The quantity \(\alpha\) can be determined from the additional condition:

\[ \sum_k n_k =\sum_k \frac{1}{e^{1-\gamma^2+\alpha+\beta\varepsilon_k}+\gamma} =N. \tag{19} \]

Its thermodynamic meaning can be clarified by differentiating the free energy with respect to the number of particles \(N\). According to (14)

\[ \left(\frac{\partial F}{\partial N}\right)_{T=\mathrm{const}} =-kT\left\{\alpha+N\frac{\partial \alpha}{\partial N}+\frac{\partial Z}{\partial N}\right\} \]

and, substituting the expression for \(N\) from (17):

\[ \alpha=-\frac{1}{kT}\left(\frac{\partial F}{\partial N}\right)_{T=\mathrm{const}} =-\frac{\varepsilon_0}{kT}. \tag{20} \]

The expression obtained coincides with the definition of the chemical potential \(\varepsilon_0\) for an individual particle.*

* In classical statistics, since \(\gamma=0\), the sum of states reduces to the familiar expression (16a). At the same time this shows the expediency of the notation chosen by us for \(\alpha\), since only with this method of notation can the chemical potential be represented by relation (20); the latter would have been impossible, for example, if we had introduced, as is usually done, \(\alpha^*=\alpha+1\). Of course the difference disappears if \(\alpha\) is a large number (classical statistics).

THEORY OF THE METALLIC STATE

From the thermodynamic interpretation of $\alpha$ given above there follows a way of solving more complicated problems. If we have a system whose energy depends on several quantum numbers (for example, particles which, besides the kinetic energy of translational motion, also possess internal energy; an example may be an electron in a magnetic field with its two possible spin orientations, rotating atoms and molecules with their systems of definite excitation levels), then for all particles with the same internal state one may specify a definite state function $Z_i$. The state function $Z$ of the whole system will be:

\[ Z=\sum_i Z_i, \]

where $\alpha$ has one and the same value in all $Z_i$. Relation (18) will remain valid. Instead of (17) we shall have:

\[ n_k^{(i)}=-kT\frac{\partial Z_i}{\partial \varepsilon_k^i}, \tag{21} \]

and $\alpha$ will now be determined from the relation:

\[ N=\sum_{k,i} n_k^{(i)}. \tag{22} \]

All values of the energy $\varepsilon_k^{(i)}$ must, of course, be referred to one and the same zero level.

§ 6. Quantum-mechanical description of the translational motion of a free particle

Up to this point we have left completely open the question of what the objects of our statistical investigation are. We shall now consider the simplest case, the case of a gas of material particles, without taking account of interaction or internal energy.* In what follows we shall extend the reasoning to the case of the presence of an external field, and in Part IV we shall consider the case of particles in a periodic field, which for electrons in a metal is much closer to the true picture than the notion of free electrons.

As is well known, quantum mechanics gives a dualistic description of phenomena and thus makes it possible to interpret processes both in terms of corpuscular concepts and in terms of wave concepts. Both methods of description are equivalent, provided only that the uncertainty relation given by quantum theory is taken into account.

For a free particle of mass $m$ the Schrödinger equation will have the form:

\[ -\frac{h^2}{8\pi^2 m}\Delta\psi-U\psi=\frac{h}{2\pi i}\frac{\partial\psi}{\partial t}, \tag{1} \]

* We cannot enter here into an analysis of the problem of a photon gas (a space filled with radiation), which is especially characteristic for the new statistics.

where \(U\) is the potential energy. In the case of a free particle (absence of an external field) the potential energy is constant, and the solutions of equation (1) will be plane waves:

\[ \psi = C e^{-\frac{2\pi i}{h}(\varepsilon t-\mathbf{r}\mathbf{p})} = C e^{-\frac{2\pi i}{h}(\varepsilon t-xp_x-yp_y-zp_z)} . \tag{2} \]

Here \(C\) is a normalizing factor, and \(p_x, p_y, p_z\) are constants, for which, by substituting the solution into the original equation (1), we obtain:

\[ p^2 = |\mathbf{p}|^2 = p_x^2 + p_y^2 + p_z^2 = 2m(\varepsilon - U). \]

Solution (2) thus corresponds to the state of an electron with momentum \(\mathbf{p}\) (components \(p_x, p_y, p_z\)). On displacement along the normal to the plane of the wave front \(xp_x + yp_y + zp_z = \mathrm{const}\) by a distance

\[ \lambda = \frac{h}{p} = \frac{h}{\sqrt{2m(\varepsilon-U)}} = \frac{h}{mv} \tag{3} \]

we find identical values of \(\psi\). This gives, for the wavelength \(\lambda\), the well-known de Broglie relation. The particle velocity \(v\), introduced here formally, coincides with the group velocity of the waves (2). The position of the particle in space is not determined by the solution (2), in full agreement with the uncertainty relation, since we suppose the momentum to be exactly specified. Localization in space can be effected only by forming a wave packet, i.e. by a superposition of waves of type (2) with different values of the energy (or, correspondingly, of the momentum). The transport of mass (and for charged particles also the transport of charge) may be obtained from the wave function (a plane wave or a packet) by calculating the expression

\[ \mathbf{s} = -\frac{h}{4\pi i m}\left(\psi \operatorname{grad}\bar{\psi} - \bar{\psi}\operatorname{grad}\psi\right), \tag{4} \]

which gives the current density.

For the case of unbounded space all values of \(\varepsilon\) and \(\mathbf{p}\) are possible, and we obtain a continuous spectrum of the characteristic numbers of equation (1). In real problems, however, one always has to deal with a system bounded by some finite volume. Physically this means that in passing through the wall the potential energy increases very strongly. We shall come, therefore, to a picture of a kind of potential well bounded by very high walls. In the limiting case of an infinitely high potential barrier the problem can be solved by introducing definite boundary conditions. As such conditions one may require, for example, that the function \(\psi\) vanish at the boundary. Then, in the region under consideration, only standing waves will be possible. For a cube with edge \(K\), the boundary conditions will be: \(U=\mathrm{const}\) for

\[ -\frac{K}{2} < x, y, z \le +\frac{K}{2} \]

and \(U=\infty\) outside the cube; the solutions of equation (1) satisfying such boundary conditions will be:

\[ \psi = C e^{-\frac{2\pi i}{h}\varepsilon t} \sin \frac{\pi x k_x}{K} \sin \frac{\pi y k_y}{K} \sin \frac{\pi z k_z}{K}, \]

where here \(k_x, k_y, k_z\) must already be integers. We obtain, for a finite volume, a discrete series of possible states, as we assumed in quantum statistics. These standing waves give, according to (4), a flux equal to zero, because the spatial coordinates enter into them as real variables. In the corpuscular picture (for example in the old quantum mechanics of Bohr) they correspond to the motion of particles alternately reflecting from the walls of the vessel.

If we wish also to take into account those cases in which a transition of electrons from one metal into another is possible, then the boundary condition \(\psi=0\) must be replaced by another one. In this case we may choose cyclic boundary conditions:

\[ \psi(x+K)=\psi(x)\ \text{etc.} \tag{5} \]

As the solution of the equation we then obtain:

\[ \psi=e^{-\frac{2\pi i}{h}\varepsilon_{\mathbf{k}}t}\,\psi_{\mathbf{k}}, \tag{6a} \]

where

\[ \psi_{\mathbf{k}}=\frac{1}{\sqrt{K^{3}}}\,e^{\frac{2\pi i}{K}(\mathbf{kr})} = \frac{1}{\sqrt{K^{3}}}\,e^{\frac{2\pi i}{K}(xk_x+yk_y+zk_z)}; \tag{6b} \]

\(k_x,k_y,k_z\) are integers.

The constant factor here is chosen so that the function is normalized; indeed, on integrating over the whole volume \(V\):

\[ \int \int \int_{-\frac{K}{2}}^{+\frac{K}{2}} \psi_{\mathbf{k}}\bar{\psi}_{\mathbf{k}'}\,dx\,dy\,dz = \int \psi_{\mathbf{k}}\bar{\psi}_{\mathbf{k}'}\,dV = \delta_{\mathbf{k}\mathbf{k}'} = \begin{cases} 1 & \text{for } \mathbf{k}=\mathbf{k}',\\ 0 & \text{for } \mathbf{k}\ne\mathbf{k}' . \end{cases} \tag{7} \]

since the volume \(V=K^3\). \(\mathbf{k}\) is the wave vector; its magnitude, in the notation chosen, gives the number of nodes of the eigenfunction. Substituting, as before, the solution into Schrödinger’s equation, we find the relation between the wave vector and momentum and energy:

\[ \mathbf{p}=\frac{h\mathbf{k}}{K};\qquad k^2=|\mathbf{k}|^2=\frac{2mK^2}{h^2}(\varepsilon_{\mathbf{k}}-U). \tag{8} \]

Let us note that in the preceding arguments we could, without loss of generality, have taken the potential energy inside the box to be equal to zero.

In the case under consideration, each eigenfunction will already correspond to a definite current which, the value of which we find by integrating the current-density equation (4) over the whole volume. Carrying out this integration we find:

\[ \mathbf{S}=\int \mathbf{s}\,dV=\frac{h}{mK}\mathbf{k}=\frac{\mathbf{p}}{m}. \tag{9} \]

For statistics the most important question is that of the distribution of the characteristic numbers. From (8) we find for the possible values of the energy:

\[ \varepsilon_{\mathbf{k}}=\frac{h^2}{2mK^2}\,|\mathbf{k}|^2, \tag{10} \]

where \(\mathbf{k}\) corresponds to three integers, which may be either positive or negative. Hence it follows that the number of characteristic numbers \(i(\varepsilon)\,d\varepsilon\) corresponding to the energy interval \(\varepsilon,\ \varepsilon+d\varepsilon\) is given by the number of points of the integer lattice in \(k\)-space contained in the spherical layer of radius

\[ \sqrt{2m\varepsilon}\,\frac{K}{h} \]

and the thickness

\[ \frac{1}{2}\sqrt{\frac{2m}{\varepsilon}}\,\frac{K}{h}\,d\varepsilon . \]

Thus

\[ A(\varepsilon)\,d\varepsilon = 2\pi (2m)^{\frac{3}{2}} \frac{K^3}{h^3}\, \varepsilon^{\frac{1}{2}}\,d\varepsilon . \]

(It should be noted that this result can also be obtained by using the boundary condition \(\psi=0\).)

We have assumed that to each state of translational motion (defined by the triple of numbers \(k\)) there corresponds one single characteristic number, or, in other words, that a particle at rest can be in only one single state. This assumption, however, will not be valid if the particles have additional degrees of freedom. Thus, for example, if a particle possesses some angular momentum \(j\frac{h}{2\pi}\) (\(j\) an integer or half-integer), then, as is known, it can have \(2j+1\) different orientations relative to a given direction, and in the absence of an external field these different states are energetically completely identical. In this case, in reality, \(2j+1\) quantum states will correspond to each translational motion. In order to take this circumstance into account in the calculation, we introduce a “weight” factor of the state \(G\), giving the degree of degeneracy of the translational motion under consideration. In the case where there is a moment, \(G=2j+1\). The electrons, which interest us here first of all, possess spin; their mechanical moment is \(s=\frac{1}{2}\frac{h}{2\pi}\), and for them only two orientations are possible. For electrons, therefore, \(G=2\). Taking further into account that \(K^3\) gives the total volume \(V\), we finally obtain:

\[ A(\varepsilon)\,d\varepsilon = VG\,\frac{2\pi(2m)^{\frac{3}{2}}}{h^3}\, \varepsilon^{\frac{1}{2}}\,d\varepsilon . \tag{11} \]

As is known, the number of natural oscillations of any region in the asymptotic case, i.e. for large values of the energy, does not depend at all on the special form of the region (box). This means that, when the geometrical form of the region is changed, only the distribution of the very lowest energy values can change somewhat, whereas for the higher ones we may be sure of the validity of the distribution law (11). This circumstance makes it possible to take as the basis of the calculations the simple form of a cube, without thereby introducing any special restrictions.

The result obtained in (11) can also be interpreted in a somewhat different way. Namely, instead of the number of states with a definite energy, we could seek the number of states corresponding to a definite interval of momenta \(dp_x, dp_y, dp_z\). According to

(8) this number is equal to the number of points of the integer lattice contained in the parallelepiped (in \(k\)-space) with edges

\[ \frac{K}{h}\,dp_x,\quad \frac{K}{h}\,dp_y,\quad \frac{K}{h}\,dp_z; \]

it is equal, if the weight \(G\) is taken into account and \(K^3=V\) is put, to

\[ A(p_x,p_y,p_z)\,dp_x\,dp_y\,dp_z = V\,\frac{G}{h^3}\,dp_x\,dp_y\,dp_z. \tag{12} \]

Expression (11) is completely equivalent to (12) and can be obtained from it by integration over the layer:

\[ \varepsilon < \frac{p_x^2+p_y^2+p_z^2}{2m} < \varepsilon+d\varepsilon . \]

The various volume elements inside \(V\) are completely equivalent. We can therefore divide the volume \(V\) into cells as well and pose the question of the number of states corresponding to the element \(dx\,dy\,dz\,dp_x\,dp_y\,dp_z\) of the six-dimensional phase space of coordinates and momenta. This number is equal to

\[ A(\mathbf p,\mathbf r)\,dp_x\,dp_y\,dp_z\,dV = \frac{G}{h^3}\,dp_x\,dp_y\,dp_z\,dV. \tag{13} \]

Thus the quantization of normal vibrations must be interpreted in terms of the corpuscular picture as a division of phase space into finite cells of volume \(h^3\) (if the weight \(G\) is not taken into account). This requirement is not new; it was already put forward in Bohr’s quantum theory. We can therefore also use the representations of the corpuscular picture. The assertion that both methods of description lead to identical results is only one of the formulations of the general correspondence principle. By adopting such a division of phase space into cells of magnitude \(h^3\), we no longer need, for carrying out statistical calculations, any other consequences of quantum mechanics. It should be noted, however, that this proposition itself needs a more serious justification.

§ 7. The material gas. Degeneracy criterion

In applying the results obtained to various concrete cases, one must first of all decide which statistics the system under study is to obey. In the case of electrons we know from atomic theory that the Pauli principle is valid for them. Hence it follows unambiguously that electrons must obey Fermi–Dirac statistics. As has already been emphasized, this property does not follow from any other laws of nature known to us, but is simply an empirically established fact.

For protons, too, Fermi–Dirac statistics must hold. Such a conclusion follows unambiguously from the experimental study

intensities of the band spectrum of \(H_2\) and from the fact that ortho- and parahydrogen exist.*

The question proves, however, more difficult for complex particles, such as, for example, more complex nuclei, atoms, and molecules. If these particles form complexes that do not change during the processes under consideration (for example, atoms and molecules under the condition of absence of ionization and dissociation), then quite definite theoretical statements become possible, following directly from the symmetry properties of the eigenfunctions. Thus, for example, considering hydrogen atoms \(H\), consisting of a proton and an electron, we find that the eigenfunction of a system of any two elementary particles will exhibit the symmetry proper to these elementary particles. In interchanging two atoms as a whole, we perform both an interchange of electrons and an interchange of protons. Each separate interchange corresponds to a change of sign of the eigenfunction, and consequently, under a double interchange the sign will not change. Hence it follows that hydrogen atoms must obey Bose—Einstein statistics.

Obviously, one can establish the corresponding rule also for the case of arbitrarily complex particles. If a complex particle consists of an odd number of elementary particles, each of which individually obeys Fermi—Dirac statistics, then the whole particle as a whole will likewise obey Fermi—Dirac statistics; this will be true quite independently of whether elements obeying Bose—Einstein statistics enter into the composition of our particle or not, since in both cases the sign of the eigenfunction of the system will change to the opposite one under each interchange. If, on the contrary, a complex particle consists of an even number of antisymmetric particles (in the particular case contains none of them at all), then such particles must obey Bose—Einstein statistics.\(^{26}\)

On the basis of the available empirical material one may assert that the application of Fermi—Dirac statistics to protons and electrons does not give rise to any doubts. For atoms and molecules as a whole, the question of the applicability of different statistics could not be checked experimentally, since the difference in the results of applying different statistics in this case, as we shall see below, is too small. As for atomic nuclei, the study of band spectra leads to the conclusion that the rule indicated here is not always fulfilled. Thus, for example, \(\alpha\)-particles and \(O\) nuclei obey Bose—Einstein statistics, as they should, because they contain an even number of Fermi—Dirac particles. Nitrogen nuclei, however, also obey Bose—Einstein statistics, whereas theoretically one would expect the applicability of Fermi—Dirac statistics\(^{27}\) [the nitrogen nucleus should contain 14 protons (the atomic weight is 14) and 7 electrons, since the nuclear charge is \(Z = 14 - 7\); hence, 21 elementary particles]. The impression is created that nuclear electrons are immaterial for statistical calculations.**

These questions, belonging to the most interesting questions of contemporary physics, are connected in the closest way with the laws of the structure of nuclei and with the nature of elementary particles; however, we cannot enter here into a more detailed discussion of them.

* The theoretical interpretation of the intensities was given by Heisenberg\(^{21}\) and Hund.\(^{22}\) The corresponding experimental data are given in the work of Rasetti.\(^{23}\) On ortho- and parahydrogen, see the works of Bonhoeffer and Harteck\(^{24}\) and of Eucken and Hiller.\(^{25}\)

** If it is assumed that nuclear electrons combine with protons into neutrons, then the rule indicated above will mean that neutrons obey Fermi—Dirac statistics. At present, however, we still have no relevant experimental data.

For a material gas the distribution law, according to § 6 (11) and § 4 (8), will be:

\[ N(\varepsilon)d\varepsilon = -\frac{A(\varepsilon)d\varepsilon}{e^{\alpha+\frac{\varepsilon}{kT}}+\gamma} = V\frac{2\pi G(2m)^{\frac{3}{2}}}{h^3} \frac{\varepsilon^{\frac{1}{2}}d\varepsilon}{e^{\alpha+\frac{\varepsilon}{kT}}+\gamma} \tag{1} \]

or, for individual states:

\[ n(\varepsilon)=\frac{1}{e^{\alpha+\frac{\varepsilon}{kT}}+\gamma}. \tag{2} \]

The constant \(\alpha\) is determined from the additional condition

\[ N=\sum_s N_s=\int_0^\infty N(\varepsilon)d(\varepsilon) = V\frac{2\pi G}{h^3}(2m)^{\frac{3}{2}} \int_0^\infty \frac{\varepsilon^{\frac{1}{2}}d\varepsilon}{e^{\alpha+\frac{\varepsilon}{kT}}+\gamma}. \tag{3} \]

This constant turns out to depend both on the temperature and on the concentration of particles \(N/V\). From the expressions given it is already clear that the distribution law, and together with it all the properties of the gas, approach the classical one when \(\gamma(=\pm 1)\) is small in comparison with the exponential function, i.e. when \(\alpha\) assumes a large positive value. This means that in this case one must have \(N(\varepsilon)\ll A(\varepsilon)\), i.e. the number of particles for any energy interval \(d\varepsilon\) must be small in comparison with the number of available cells. Then, on the one hand, the difference between the results of the Fermi—Dirac and Bose—Einstein statistics disappears, since the number of particles per cell becomes small compared with unity and the Pauli exclusion principle no longer plays any role. On the other hand, the results of both quantum statistics begin to coincide with the results of classical statistics, since in this case the number of permutations for the different complexes becomes the same and therefore ceases to play a role in determining the most probable distribution.*

It is therefore very important to have a definite criterion which would make it possible to decide under what conditions (density, temperature) degeneracy may be expected, i.e. a noticeable deviation from the classical laws. Introduce into relation (3) the new variable \(x=\varepsilon/kT\). Then

\[ N= V\frac{2\pi G(2m)^{\frac{3}{2}}}{h^3} (kT)^{\frac{3}{2}} \int_0^\infty \frac{x^{\frac{1}{2}}dx}{e^{\alpha+x}+\gamma}, \]

* This number is given by the factor \(\dfrac{N!}{\prod_s N_s!}\) in § 3, (1). When, in the overwhelming number of distributions, the particles are in different cells, this number will practically always coincide with \(N!\).

which can also be written in the following form:

\[ J_{\frac12}(\alpha)=\frac{N}{V}\frac{h^3}{2\pi G}(2mkT)^{-\frac32}; \quad J_{\frac12}(\alpha)=\int_0^\infty \frac{x^{\frac12}\,dx}{e^{\alpha+x}+\gamma}. \tag{4} \]

When \(\alpha\) is positive and very large, \(J_{\frac12}(\alpha)\) will be small in comparison with unity, and conversely. Thus we obtain the desired criterion:

\[ \frac{N}{V}\frac{h^3}{2\pi G}(2mkT)^{-\frac32} \begin{cases} \ll 1 & \text{there is no degeneracy},\\ \to 1 \text{ or } \gg 1 & \text{there is degeneracy}. \end{cases} \tag{5} \]

This expression contains, in addition to the universal constants \(h\) and \(k\), the concentration of particles, the mass of the particles, and the temperature. If all these quantities are given, then one can immediately decide whether degeneracy occurs or not. The onset of degeneracy is favored by a large concentration of particles, a low temperature, and a small particle mass. The temperature \(T_e\) at which expression (5) is exactly equal to unity may be called the degeneracy temperature; from (5) we obtain:

\[ kT_e=\frac{h^2}{2m}\left(\frac{n}{2\pi G}\right)^{\frac32}. \tag{5a} \]

Let us consider the example of a real gas, for instance He, for which the conditions for the onset of degeneracy are especially favorable. Put \(n=10^{22}\), \(m=6.6\cdot10^{-24}\), \(G=1\). Then for the degeneracy temperature we find \(T_e=6.5^\circ\). We therefore cannot hope to show the degeneracy of a gas by deviations from the classical laws, since degeneracy must occur at such a low temperature and at such a high concentration that the influence of van der Waals corrections will greatly exceed the influence of degeneracy.\(^{28}\) The situation is different for electrons, whose mass is considerably smaller (\(m=0.9\cdot10^{-27}\), \(G=2\)) and which in a metal have a very high concentration (of the order of one electron per atom, i.e. \(10^{23}\) per \(\mathrm{cm}^3\)), a concentration very difficult to realize for molecules (compression). Taking this concentration, we find for the degeneracy temperature of the electron gas a value of the order of \(7\cdot10^4\). Such an electron gas, even at the highest attainable temperatures, will deviate very strongly from the laws of classical statistics. In what follows we shall consider only electrons and shall confine ourselves to the Fermi—Dirac statistics (\(\gamma=+1\)), which alone is of practical importance (we shall not discuss the problem of a photon gas, which obeys Bose—Einstein statistics).

Besides the distribution function, it is useful to have an expression also for other thermodynamic quantities. The energy of the gas as a function of temperature is obtained as equal to

$$ E=\sum_s \varepsilon_s N_s=\int_0^\infty \varepsilon N(\varepsilon)\,d\varepsilon =V\frac{2\pi G(2m)^{\frac32}}{h^3}\int_0^\infty \frac{\varepsilon^{\frac32}\,d\varepsilon}{e^{\alpha+\frac{\varepsilon}{kT}}+1} = $$

$$ =V\frac{2\pi G}{h^3}(2m)^{\frac32}(kT)^2J_{\frac32}; \quad J_{\frac32}=\int_0^\infty \frac{x^{\frac32}\,dx}{e^{\alpha+x}+1}, \tag{6} $$

and the magnitude of the gas pressure we find according to § 5 (18):

$$ p=kT\frac{\partial Z}{\partial V} =kT\frac{\partial}{\partial V} \left\{ V\frac{2\pi G}{h^3}(2m)^{\frac32} \int_0^\infty \lg\left(1+e^{-\alpha-\frac{\varepsilon}{kT}}\right)\varepsilon^{\frac12}\,d\varepsilon \right\} = $$

$$ =\frac{2\pi G}{h^3}(2m)^{\frac32}kT \int_0^\infty \lg\left(1+e^{-\alpha-\frac{\varepsilon}{kT}}\right)\varepsilon^{\frac12}\,d\varepsilon. \tag{7} $$

Carrying out integration by parts [putting \(u=\lg\left(1+e^{-\alpha-\frac{\varepsilon}{kT}}\right)\); \(v=\varepsilon^{\frac12}\), \(uv=0\) at the limits] and taking (6) into account, we find:

$$ \frac{pV}{kT} =V\frac{2\pi G}{h^3}(2m)^{\frac32}\frac{2}{3} \int \frac{\varepsilon^{\frac32}}{kT}\, \frac{d\varepsilon}{e^{\alpha+\frac{\varepsilon}{kT}}+1} =\frac{2}{3}\frac{E}{kT}. \tag{8} $$

The well-known relation of the classical theory of gases proves to be valid here as well.

The expressions for the entropy and the free energy are obtained directly, according to § 5 (10) and (14).

For further operations with the relations obtained, it is necessary to examine in more detail the integrals entering into (4)—(8).

§ 8. Integral formulas of the Fermi—Dirac statistics

The integrals occurring in formulas (4) and (8), which by the substitution \(\frac{\varepsilon}{kT}=x\) are reduced to the type:

$$ J_k(\alpha)=\int_0^\infty \frac{x^k}{e^{\alpha+x}+1}\,dx \tag{1} $$

must be calculated as functions of the parameter \(\alpha\). Since these integrals are not evaluated elementarily, we shall seek approximate expressions for the most important special cases.

I. Weak degeneracy. Here \(\alpha\gg 1\) and is positive (in the limiting case we must obtain the classical theory). The exponential function in the denominator is always considerably greater than unity, and we may put

$$ \frac{1}{e^{\alpha+x}+1} =\frac{e^{-(\alpha+x)}}{1+e^{-(\alpha+x)}} =e^{-(\alpha+x)}\sum_{n=0}^{\infty}(-1)^n e^{-(\alpha+x)n} = $$

\[ = \sum_1^\infty (-1)^{n-1} e^{-(a+x)n}, \]

\[ \left. \begin{aligned} J_k &= \sum_1^\infty (-1)^{n-1} e^{-an} \int_0^\infty x^k e^{-nx}\,dx = \\[3pt] &= \sum_1^\infty (-1)^{n-1} e^{-an} \frac{1}{n^{k+1}}\int z^k e^{-z}\,dz, \\[3pt] J_k &= \sum_1^\infty (-1)^{n-1} \frac{\Gamma(k+1)}{n^{k+1}} e^{-an}, \end{aligned} \right\} \tag{2} \]

where \(\Gamma\) is the gamma function. In particular, for

\[ \Gamma\!\left(\frac{3}{2}\right)=\frac{1}{2}\sqrt{\pi};\qquad \Gamma\!\left(\frac{5}{2}\right)=\frac{3}{4}\sqrt{\pi} \]

we obtain:

\[ J_{\frac12}(a)=\frac{1}{2}\sqrt{\pi}\left(e^{-a}-\frac{e^{-2a}}{2^{\frac32}}+\frac{e^{-3a}}{3^{\frac32}}+\ldots\right) \to \frac{1}{3}\sqrt{\pi}\,e^{-a}, \tag{3a} \]

\[ J_{\frac32}(a)=\frac{3}{4}\sqrt{\pi}\left(e^{-a}-\frac{e^{-2a}}{2^{\frac52}}+\frac{e^{-3a}}{3^{\frac52}}+\ldots\right) \to \frac{3}{4}\sqrt{\pi}\,e^{-a}. \tag{3b} \]

II. Strong degeneracy. \(\alpha\) must be very large and negative. Put \(-\alpha=a\gg 1\). Then

\[ J_k(a)=\int_0^\infty \frac{x^k dx}{e^{-a+x}+1}. \tag{4} \]

For this integral at large values of \(a\), Sommerfeld proposed an asymptotic series. It is not difficult to give an approximate expression also for integrals of a more general form, which we shall need below:

\[ J=\int_0^\infty \frac{\psi(x)\,dx}{e^{-a+x}+1} =\int_0^\infty \psi(x) f_0(x)\,dx;\quad f_0=\frac{1}{e^{-a+x}+1}, \tag{5a} \]

\[ K=\int_0^\infty \frac{\varphi(x)\,dx}{(e^{-a+x}+1)(e^{a-x}+1)} = -\int_0^\infty \varphi(x)\frac{\partial f_0(x)}{\partial x}\,dx. \tag{5b} \]

Here \(\psi\) and \(\varphi\) are completely arbitrary functions, subject only to certain restrictions, which will be discussed below. \(J\) and \(K\) are related to one another by the relation:

\[ K(\varphi)=-\varphi f_0\Big|_0^\infty+\int \frac{d\varphi}{dx} f_0\,dx = -\varphi f_0\Big|_0^\infty+J\!\left(\frac{d\varphi}{dx}\right). \tag{6} \]

We can obtain an approximate expression for \(K\) by taking into account that \(\dfrac{\partial f_0}{\partial x}\) has, at \(x=a\), an extremely steep maximum, so that only the form of the function \(\varphi\) near this point has a noticeable influence on the value of the integral. Introduce

\[ z=x-a, \]

\[ \varphi(z)=\varphi(0)+z\varphi'(0)+\frac{z^2}{2!}\varphi''(0)+\ldots \]

This expansion of \(\varphi\) in a Taylor series is suitable for computing the integral only in the case where the series converges in the region \(|x|\leq a\) and if outside this interval \(\varphi\) tends to infinity more slowly than an exponential function. Then

\[ K=\int_{-a}^{+\infty} \frac{\varphi(0)+z\varphi'(0)+\dfrac{z^2}{2!}\varphi''(0)+\ldots} {(e^z+1)(e^{-z}+1)}\,dz. \]

If, moreover, \(a\) is very large, then with great accuracy (the order of the error is \(e^{-a}\)) the lower limit may be replaced by \(-\infty\). Taking into account that the odd powers from our series drop out, we obtain the asymptotic expression

\[ K=\sum_0^\infty \varphi^{(2n)}(0)\frac{K_{2n}}{(2n)!}, \]

where

\[ K_{2n}=\int_{-\infty}^{+\infty} \frac{z^{2n}\,dz}{(e^z+1)(e^{-z}+1)} = 2\int_0^\infty \frac{z^{2n}e^{-z}}{(1+e^{-z})^2}\,dz. \]

Using the binomial series:

\[ \frac{1}{(1+e^{-z})^2} = 1-2e^{-z}+3e^{-2z}-4e^{-3z}+\ldots, \]

we obtain:

\[ K_0=2\int_0^\infty \frac{dz}{(e^z+1)(e^{-z}+1)} = 2\int_1^\infty \frac{dt}{(1+t)^2}=1, \]

\[ K_{2n}=2\sum_1^\infty (-1)^{l-1}l\int_0^\infty z^{2n}e^{-lz}\,dz; \]

and since

\[ \int_0^\infty e^{-lz}z^{2n}\,dz=\frac{(2n)!}{l^{2n+1}}, \]

we finally find:

\[ K=\{\varphi+2(c_2\varphi^{\mathrm{II}}+c_4\varphi^{\mathrm{IV}}+\ldots)\}_{z=0,\ \text{i.e. }x=a}, \tag{7} \]

where

\[ c_{2n}=\sum_1^\infty \frac{(-1)^{l-1}}{l^{2n}} = 1-\frac{1}{2^{2n}}+\frac{1}{3^{2n}}-\frac{1}{4^{2n}}+\ldots \tag{8a} \]

In particular, for \(n=1\) we obtain the well-known series:

\[ c_2=1-\frac{1}{2^2}+\frac{1}{3^2}-\frac{1}{4^2}+\ldots=\frac{\pi^2}{12}. \tag{8b} \]

Expression (7) gives the desired asymptotic representation. To form it, it proves necessary to know the values of \(\varphi\) and of its derivatives only at the point \(x=a\). From (7) we also easily find the corresponding expression for \(J_k(\alpha)\). Setting

\[ \varphi(x)=\frac{1}{k+1}x^{k+1};\quad \psi(x)=\frac{\partial\varphi}{\partial x}=x^k, \]

we find from (6) and (7):

\[ J_k = K(\varphi)=\frac{a^{k+1}}{k+1}+2\{c_2 k a^{k-1}+c_4 k(k-1)(k-2)a^{k-3}+\ldots\} = \]

\[ =\frac{a^{k+1}}{k+1} \left\{1+2\left(c_2\frac{(k+1)k}{a^2} +c_4\frac{(k+1)\ldots(k-2)}{a^4}+\ldots\right)\right\}. \tag{9} \]

As a particular case we obtain hence:

\[ J_{\frac12}(a)=\frac{2}{3}a^{\frac32} \left(1+\frac{3c_2}{2a^2}+\ldots\right) =\frac{2}{3}a^{\frac32} \left(1+\frac{\pi^2}{8a^2}+\ldots\right), \tag{10a} \]

\[ J_{\frac32}(a)=\frac{2}{5}a^{\frac52} \left(1+\frac{15c_2}{2a^2}+\ldots\right) =\frac{2}{5}a^{\frac52} \left(1+\frac{5\pi^2}{8a^2}+\ldots\right) \tag{10b} \]

(for the case of large positive \(a=-\alpha\)).

§ 9. Fermi—Dirac Gas

Relation 7 (4), on the basis of § 8 (3) and (10), becomes, for the two limiting cases, the following:

\[ \frac{N}{V}\frac{h^3}{2\pi G}(2mkT)^{-\frac32} = \begin{cases} \frac12\sqrt{\pi}\left(e^{-\alpha}-\dfrac{e^{-2\alpha}}{2^{3/2}}+\ldots\right) & \text{for } \alpha\gg 1, \quad (1a)\\[1.2em] \dfrac23(-\alpha)^{\frac32} \left(1+\dfrac{\pi^2}{8\alpha^2}+\ldots\right) & \text{for } -\alpha\gg 1. \quad (1b) \end{cases} \]

The first case (1a) obtains when the left-hand side of the equation is \(\ll 1\) (weak degeneracy), the second case (1b)—when the left-hand side is \(\gg 1\) (strong degeneracy). Solving the equation with respect to \(\alpha\) and retaining only the first term of the expansion, we find:

\[ e^{-\alpha}=\frac{nh^3}{G}(2\pi mkT)^{-\frac32} \quad \text{for } \alpha\gg 1,\ \left(n=\frac{N}{V}\right), \tag{2a} \]

\[ -\alpha=\frac{h^2}{2kT m}\left(\frac{3n}{4\pi G}\right)^{\frac23} \quad \text{for } -\alpha\gg 1,\ \left(n=\frac{N}{V}\right). \tag{2b} \]

For the case (2b) we shall find also the second approximation. We obtain it most simply if we substitute the value (2b) into the correction term (1b) and use the expansion:

\[ \left(1+\frac{\pi^2}{8\alpha^2}\right)^{-\frac23} = 1-\frac{2}{3}\frac{\pi^2}{8\alpha^2}+\ldots = 1-\frac{c_2}{\alpha^2}+\ldots \]

We obtain:

\[ -\alpha= \frac{h^2}{2mkT}\left(\frac{3n}{4\pi G}\right)^{\frac23} \left\{1-\frac{(2\pi mkT)^2}{12h^4} \left(\frac{4\pi G}{3n}\right)^{\frac43}\right\} = \]

\[ =\frac{\mu}{kT} \left\{1-c_2\left(\frac{kT}{\mu}\right)^2\right\}, \tag{3} \]

where

\[ \mu=\frac{h^{2}}{2m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}} . \tag{4} \]

I. Weak degeneracy—\(\alpha \gg 1\). In this case one may neglect unity in the denominator of the expression § 7 (1), and then, using (2a), we obtain the Maxwell distribution law:

\[ \frac{N(\varepsilon)}{N}\,d\varepsilon = \frac{2}{\sqrt{\pi}}(kT)^{-\frac{3}{2}} e^{-\frac{\varepsilon}{kT}}\varepsilon^{\frac{1}{2}}\,d\varepsilon . \tag{5a} \]

From § 7 (4) and (6) and § 8 (3a) and (3b) we find the well-known relation of classical kinetic theory,

\[ E=\frac{3}{2}NkT. \]

Using further § 7 (8), we obtain the equation of state of an ideal gas in the usual form:

\[ pV=NkT=RT. \tag{5b} \]

Let us also find, for this limiting case, the expression for the entropy from § 5 (10) or (11). Neglecting unity in comparison with \(e^{-\alpha}\) (but not in comparison with \(\alpha\)), we may put:

\[ \sum \lg\left(1+e^{-\alpha-\frac{\varepsilon}{kT}}\right) = \sum e^{-\alpha-\frac{\varepsilon}{kT}} = \sum n_k = N \]

and then we find the expression for the entropy in the form:

\[ S=kN\left(1+\alpha+\frac{E}{NkT}\right) = kN\left(\frac{5}{2}+\alpha\right). \]

Substituting the value of \(\alpha\) from (2a), introducing the gas constant \(R=kN\) and the particle concentration (5b) \(n=\frac{p}{kT}\), we finally obtain

\[ S=R\left\{\frac{5}{2}\lg T-\lg p+C\right\}, \tag{6a} \]

where

\[ C=\lg\frac{G(2\pi m)^{\frac{3}{2}}(ek)^{\frac{5}{2}}}{h^{3}} . \tag{6b} \]

(\(e\) is the base of natural logarithms).

Thus we arrive at the well-known Stern—Tetrode expression with the correct expression for the entropy constant \(C\). It should be noted that Bose—Einstein statistics (\(\gamma=-1\)) would have led to exactly the same result. Thus the value of \(C\) given here should be regarded rather not as a consequence of the new statistics, but as a consequence of the quantization of motions (the division of phase space into cells \(h^{3}\)). In classical statistics, only an inconvenient, confusing (and incorrect) term \(\lg N\) would enter as the entropy constant; here it is eliminated altogether.

II. Strong degeneracy—\(a \gg 1\). Let us now consider the case important for us, that of strong degeneracy, when \(a=-\alpha\) is a large positive number. In this case the correction term in (3) will be small. This correction term is the beginning of an expansion in positive powers of \(\dfrac{kT}{\mu}\), which starts with a quadratic term that plays a role only at very high temperatures. We may therefore put:

\[ \alpha=-\frac{\varepsilon_0}{kT};\quad \varepsilon_0=\mu\left\{1-\frac{\pi^2}{12}\left(\frac{kT}{\mu}\right)^2+\cdots\right\}, \tag{7} \]

where \(\varepsilon_0\) may be regarded as constant up to so high a temperature at which \(kT\) becomes comparable with \(\mu\), i.e. up to the degeneracy temperature [§ 7 (5)]. The distribution function, according to § 7 (1), in this case will be:

\[ N(\varepsilon)\,d\varepsilon = V\frac{2\pi G(2m)^{\frac{3}{2}}}{h^3} \frac{\varepsilon^{\frac{1}{2}}\,d\varepsilon} {e^{\frac{-\varepsilon_0+\varepsilon}{kT}}+1}, \tag{8a} \]

or, if it is referred to momenta [cf. § 6 (12)]:

\[ N(\mathbf{p})\,dp_x\,dp_y\,dp_z = \frac{VG}{h^3} \frac{dp_x\,dp_y\,dp_z} {e^{\frac{-\varepsilon_0+\varepsilon}{kT}}+1} = \frac{VG}{h^3}f_0\,dp_x\,dp_y\,dp_z, \tag{8b} \]

where

\[ f_0=\frac{1}{e^{\frac{-\varepsilon_0+\varepsilon}{kT}}+1}. \tag{9} \]

The distribution function found differs in a characteristic way from the Maxwell law (5a); it is shown graphically in Fig. 2. Below the critical value

\[ \varepsilon=\mu=\frac{h^2}{2m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}} \tag{10} \]

at not too high temperatures (i.e. \(kT \ll \mu\)) the magnitude of the exponential function in the denominator of (9) is small in comparison with unity. Therefore, for \(\varepsilon<\mu\), an approximately constant value \(f_0\) is obtained. Conversely, for values of \(\varepsilon\) above the critical one the exponential function is considerably greater than unity, and we obtain an exponential course of the distribution function in agreement with the Maxwell law. The fall of the curve \(f_0\) near the critical value of \(\varepsilon\) is the steeper the lower the temperature \(T\), and in the limit as \(T\to0\) we obtain a rectangular form of the curve with a sharp break. Only at very high temperatures do we arrive at the pure exponential function, as required by classical statistics.

THEORY OF THE METALLIC STATE

Thus, even at absolute zero temperature, the particles are not at rest, but possess the most varied energies up to the maximum value \(\varepsilon=\mu\); the gas has a definite energy of absolute zero. A visual justification of the energy of absolute zero is contained in the Pauli principle, according to which in each cell \(h^3\) there may be no more than 1 (or, respectively, \(G\)) particles. The distribution at \(T=0\) will be such that all \(N\) of the lowest cells are occupied; in this case no energy can any longer be taken away from the gas. As the temperature is raised, the distribution over the cells “loosens,” and in the end the density of the distribution in phase space becomes so small that there is a negligible probability of finding an electron in any one of the cells \(h^3\). Then (9) passes into the usual Maxwellian distribution (5a).

Fig. 2. The Fermi–Dirac distribution function.

The value of the maximum energy at \(T=0\) is given by expression (10). Assuming that for each atom of the metal there are \(z\) free electrons and denoting by \(n_0\) the number of atoms in \(\mathrm{cm}^3\) (so that \(n=zn_0\)), we find, for various metals, the following values of the quantity \(\mu z^{-2/3}\), independent of \(z_0\), in volts:

TABLE 1

Metal Li Na K R Cs Cu Ag Au Mg Ca Hg
\(n_0\cdot 10^{-22}\) 4,65 2,54 1,33 1,08 0,85 8,48 5,88 5,90 4,22 2,29 4,19
\(\mu z^{-2/3}\) (V) 4,7 3,1 2,1 1,8 1,5 7,1 5,5 5,5 4,4 3,0 4,4
Metal Al Zr Pb Th Ta Mo W Fe Ni Pd Pt
\(n_0\cdot 10^{-22}\) 6,00 4,26 3,31 2,86 5,50 6,40 6,25 8,49 9,00 6,54 6,62
\(\mu z^{-2/3}\) (V) 5,6 4,4 3,8 3,3 5,3 5,8 4,8 7,1 7,3 5,9 5,9

The maximum energy of absolute zero is obtained by multiplying the figures given by \(z^{2/3}\), if \(z\) is the number of electrons split off per atom. If even only one electron per atom conducts

itself as free, we obtain for the energy of the absolute zero in various metals a value from 2 to 10 V. This value turns out to be extraordinarily large in comparison with \(kT\) \((=8.55\cdot 10^{-5}\cdot T\,V)\) at all attainable temperatures.

For the heat capacity and pressure of the electron gas we find substantial deviations from the classical laws. The total store of energy is given by the expression § 7 (6). According to § 7 (4) we obtain:

\[ E = NkT\frac{J_{\frac{3}{2}}}{J_{\frac{1}{2}}}. \]

Using further the expansion § 8 (10) and § 9 (3), we find the expression for the energy in the form of a series:

\[ E = N\frac{3}{5}\mu\left\{1+\frac{5\pi^2}{12}\left(\frac{kT}{\mu}\right)^2+\cdots\right\} = \]

\[ = V\left\{\frac{3nh^2}{10m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}} +\frac{\pi^2 m}{2h^2}\left(\frac{4\pi G}{3n}\right)^{\frac{2}{3}}(kT)^2\right\}. \tag{11} \]

The total energy is thus made up of a constant term, the energy of the absolute zero, and of a variable part increasing with temperature, which is precisely what determines the heat capacity. On the basis of (11) we obtain for the heat capacity, referred to one electron, the value:

\[ c_v=\frac{\partial E}{\partial T}\frac{1}{N} =\frac{3}{2}k\frac{\pi^2 kT}{3\mu} =\frac{3}{2}k\frac{kT\,2\pi^2m}{h^2}\left(\frac{4\pi G}{3n}\right)^{\frac{2}{3}}. \tag{12} \]

The heat capacity increases linearly with temperature; moreover, at \(T=0\), \(c_v=0\). But even at comparatively high temperatures \(c_v\) is still considerably smaller than the classical value \(\frac{3}{2}k\), given by the law of equipartition of energy. Putting \(\mu=x\,V\) (see the table given above), we find the value of the coefficient in (12):

\[ \frac{\pi^2}{3}\frac{kT}{\mu}=2.81\cdot 10^{-4}\frac{T}{x}. \]

We thus obtain, as the first great success of the theory, the removal of the principal difficulty of the classical conceptions, indicated in § 1. It now proves entirely possible to assume a very large number of free electrons (of the order of the number of atoms) without thereby noticeably changing the magnitude of the heat capacity of the metal. The value of the heat capacity given by (12) cannot be detected by present-day methods.

In accordance with the presence of a large energy of the absolute zero, there also arises a pressure at absolute zero, which we can find from § 7 (8) and § 9 (11):

\[ p=\frac{2}{3}\frac{E}{V}=\frac{2}{5}n\mu =\frac{nh^2}{5m}\left(\frac{3n}{4\pi G}\right)^{\frac{2}{3}}. \tag{13} \]

For the electron gas in a metal, the pressure is found to be of the order of \(10^5\) atm. This pressure does not manifest itself in any way in various processes, because it is compensated by electrostatic forces that prevent the electrons from flying out of the metal, and thus its role is reduced only to maintaining equilibrium in the metal.

The expression for the entropy is also of interest. From § 5 (10), § 6 (11) one obtains:

\[ S = kN \left\{ \alpha + \frac{E}{NkT} + \frac{V}{N} \frac{2\pi G(2m)^{\frac{3}{2}}}{h^3} \int_0^\infty \varepsilon^{\frac{1}{2}} \lg \left( 1 + e^{-\alpha - \frac{\varepsilon}{kT}} \right) d\varepsilon \right\}, \]

and according to § 7 (7) and (8):

\[ S = kN \left\{ \alpha + \frac{E}{NkT} - \frac{pV}{NkT} \right\} = kN \left\{ \alpha + \frac{1}{3}\frac{E}{NkT} \right\}. \]

Here, by § 7 (4) and (6):

\[ \frac{E}{NkT} = \frac{J_{\frac{3}{2}}}{J_{\frac{1}{2}}}, \]

so that, calculating the entropy per mole of substance \((kN = R)\), we find

\[ S = R \left( \frac{1}{3} \frac{J_{\frac{3}{2}}}{J_{\frac{1}{2}}} + \alpha \right). \tag{14} \]

Taking into account (2a) and the expansion § 8 (3), for large \(T\) we obtain the classical expression (6); conversely, for low temperatures we find from (14), § 8 (10a) and (10b), that the entropy tends to zero \((\alpha = -\infty)\) in full agreement with Nernst’s heat theorem.

§ 10. Theory of the Johnson Effect*

The general principles of statistics are valid for any physical system. However, as we shall see below, one and the same system can often be described in quite different ways, and a rational choice of the method of description can lead to very considerable simplifications. An instructive example, which is closely connected with the problems set forth here and, moreover, has great practical significance, is the phenomenon of the appearance in a conductor of electromotive forces caused by the random thermal motion of the carriers of electricity. The existence of this effect, which is one of the main causes of the interfering noises in tube amplifiers, can be proved by direct experiment.**

* §§ 10, 11 were written with the participation of A. Etzrod (A. Etzrod), Göttingen.
** See Johnson’s work.\(^{30}\) The phenomenon was first pointed out by Schottky.\(^{13}\)

A visual picture of the phenomenon that can be given here consists in the following: owing to the thermal motion of the electrons, fluctuations of the density of electricity are produced, and these in turn cause a change of the potential. A direct atomistic calculation of the phenomenon, however, cannot be carried out simply, since every such change of density also changes the distribution of the potential, and we would have to take into account here also the work that is performed in every change of a uniform distribution. In this case the interaction of the electrons can no longer be neglected. However, by choosing a suitable way of describing the system, namely by using the Fourier expansion in normal modes, we can here too readily attain the desired result.^32

Let us consider a circuit (Fig. 3) consisting of a conductor \(I\), closed on a contour \(II\). In what follows we shall at all times assume that the whole system is at the same temperature \(T\). The conductor \(I\) has a purely ohmic resistance \(R\); the resistance of the remaining part \(R_\nu\) is in general complex and therefore depends on the frequency (index \(\nu\)). Then, as a consequence of the thermal motion of the electrons, \(I\) acts as a generator producing an alternating voltage of all possible frequencies. The magnitude of the current that arises will be determined by the total resistance of the circuit \(R + R_\nu\).

In such a process a certain amount of energy is transferred from \(I\) to \(II\), and conversely. In the state of thermal equilibrium the two quantities of energy must be equal on the average, and this equality must hold for oscillations of any interval of frequencies. Indeed, if, for example, \(I\) produced a higher voltage at frequency \(\nu_1\), and \(II\) at frequency \(\nu_{II}\), then by inserting a lossless resonant circuit tuned to the frequency \(\nu_1\) we could reduce the transfer of energy from contour \(II\); \(II\) would be heated at the expense of the energy of \(I\); we would then arrive at a contradiction with the second law of thermodynamics. By analogous reasoning we also conclude that the voltage produced by \(I\) cannot depend on the nature of this conductor. Thus, for example, replacing \(I\) by another conductor with the same \(R\), we shall not change the transfer of energy from \(II\) to \(I\), since the electrical properties of the complete circuit have remained unchanged. Consequently, the transfer of energy from \(I\) to \(II\) also cannot change. The required electromotive force must, for any interval of frequencies, be a universal function of the resistance of the conductor and of its temperature. It cannot depend on any other quantities at all.

In order to calculate the magnitude of the electromotive force, we may, without diminishing the generality of the result, consider some special system particularly suitable for calculations. We shall choose the circuit shown in Fig. 4: two conductors with the same purely ohmic resistance \(R\) are connected by wires of length \(l\), whose self-inductance and capacitance per unit length are chosen so that precisely

\[ R=\sqrt{\frac{L}{C}}; \]

the ohmic resistance of the wires themselves is zero. With such a choice of the lines, as is known, reflection at the ends disappears,

thus the energy sent from one end is wholly transferred to the resistance at the other end. This will also be true for the oscillatory energy created by thermal motion.

In a state of thermal equilibrium, traveling waves will propagate in the wires, and they will create two equal energy fluxes sent by the resistances at the ends. If the resistances are suddenly disconnected, then the instantaneous energy will prove to be caught in the line. Its magnitude must be equal to that energy which arises at temperature \(T\), as a consequence of thermal motion in the line itself. Our line is an electrical oscillatory system which, like a string,* possesses

Figure diagrams

Fig. 3. On the theory of the Johnson effect. Diagram I.

Fig. 4. On the theory of the Johnson effect. Diagram II.

a series of independent natural oscillations, whose wavelengths are

\[ \frac{2l}{1},\quad \frac{2l}{2},\quad \frac{2l}{3}\ldots \frac{2l}{n}, \]

The corresponding frequencies will be:

\[ \nu_n=\frac{n\upsilon}{2l},\quad (n=1,2,3\ldots), \tag{1} \]

where \(\upsilon\) is the velocity of propagation of the waves in the line. Therefore the number of natural oscillations falling in the frequency interval \(d\nu\) will be equal to \(\dfrac{2l}{\upsilon}\,d\nu\). Each natural oscillation behaves like an ordinary oscillator; its mean thermal energy** will be \(kT\), and for the mean energy in a specified frequency interval we obtain:

\[ \varepsilon(\nu)d\nu=kT\frac{2l}{\upsilon}\,d\nu, \tag{2} \]

\[ \overline{\phantom{xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx}} \]

* Because for the line as well the equation of oscillation will hold:

\[ \frac{\partial^2 I}{\partial x^2}=\frac{1}{\upsilon^2}\frac{\partial^2 I}{\partial t^2} \]

with boundary conditions \(I=0\) at the ends.

** Here one may use the classical expression for the mean energy of an oscillator, since only large frequencies are of interest, for which \(h\nu \ll kT\). Only such large frequencies can pass through the macroscopic oscillatory circuits used in measurements. Therefore the question of the measure in which it would be more consistent to quantize the natural oscillations, as well as the question of the break-off of the solid-body spectrum at small frequencies, analogous to how this is done in Debye’s theory of heat capacity, need not be touched upon here.

and for the energy density one obtains

\[ kT=\frac{2}{v}\,d\nu . \]

If the resistances are now connected again, the energy in the line must remain exactly the same, and moreover for any frequency interval. But now there is no longer reflection at the ends, and an exchange of energy takes place between the resistances at the edges. The sum of the two fluxes, equal in magnitude, must produce precisely the energy density found above, and consequently (since the flux is equal to the product of the density and the velocity) each resistance must send into the line, per unit time in the frequency interval \(d\nu\), the energy

\[ kT\,d\nu . \]

Exactly the same amount of energy will also be absorbed by each resistance.

From this the magnitude of the electromotive force \(E(\nu)\) can now be obtained at once. The resultant current \(I(\nu)\) is equal to \(\frac{E(\nu)}{2R}\) (\(R\) being the total resistance of the circuit); the energy absorbed, for example, by conductor I is equal to:

\[ I^2R\,d\nu=\frac{E_\nu}{4R}\,d\nu . \]

Comparing the expression obtained with the value \(kT d\nu\), we find:

\[ E_\nu d\nu=4RkT\,d\nu, \tag{3} \]

which determines the magnitude of the mean electromotive force.

The result found, in agreement with the preceding considerations, turns out to be entirely independent of the nature of the conductor and of the particular features of the circuit. Its experimental confirmation in the experiments described below is proof that the carriers of electricity in metals participate in thermal motion. As to the actual mechanism of the phenomenon of electrical conductivity, as well as the nature of the carriers themselves, no conclusions can, of course, be drawn from this.

§ 11. Experimental Proof of the Johnson Effect. Determination of Loschmidt’s Number

In the effect under consideration one has to deal with rapid voltage oscillations of very small amplitude, and therefore very sensitive methods of measuring alternating voltage are required for its proof.

Johnson used for this purpose a six-stage vacuum-tube amplifier, to whose input terminals the conductor under investigation was connected (Fig. 5). The coupling between the individual stages was effected in the usual way and was, over wide limits, independent-

…as a function of frequency. Into one of the stages either a tuned circuit with variable damping or a band-pass filter was introduced. At the output of the amplifier a thermoelement with a galvanometer was placed; the deflections of the galvanometer, as is known, are proportional to the square of the current at the output.

For quantitative measurements the instrument at the output must be calibrated in units of voltage at the input. The relation between the mean square of the input voltage and the current at the output may be characterized by a frequency-dependent coefficient (having in the present case the dimension of electrical conductivity):

\[ \frac{\text{current at the output}}{\text{voltage at the input}} = \frac{I(\nu)}{E(\nu)} = \varkappa(\nu). \]

Substituting, instead of \(E(\nu)\), in § 10 (30) \(\dfrac{I(\nu)}{\varkappa(\nu)}\), we obtain:

\[ I^{2}(\nu)=4RkT\,\varkappa(\nu)^{2}. \tag{1} \]

In the experiment measurements cannot be made for any one definite frequency; rather, one must operate with an integral over the entire frequency region:

\[ I^{2}=4RkT\int_{0}^{\infty}\varkappa_{\nu}^{2}\,d\nu . \tag{2} \]

Fig. 5. Circuit diagram of Johnson’s experiments.

Fig. 5. Circuit diagram of Johnson’s experiments.

In practice, of course, one is limited to a finite frequency region, the width of which depends on the damping of the inserted circuit. The value of the integral can be calculated graphically from the obtained “resonance curve” and the characteristic of the amplifier.

In carrying out the experiment it turned out that even with the input terminals of the amplifier open, the galvanometer gives a certain deflection. Such a zero effect arises owing to imperfect shielding of the amplifier from external influences, and also as a result of the “shot effect” (Schroteffekt)* and of thermal motion in the other parts of the amplifier itself; all observed values are corrected by subtracting the zero effect. The results of the measurements fully confirmed the theoretical conclusions about the properties of the effect. The proportionality to temperature and the constancy of the ratio \(\dfrac{E^{2}}{R}\) were established both for electronic and for ionic conductors.

The Johnson effect gives a new possibility for an experimental determination of the Boltzmann constant [relation (2)] and

* For more detail on this, see, for example, the article by V. L. Granovskii, Advances in Physical Sciences 13, 805, 1933. (Translator’s note.)

At the same time, the Loschmidt numbers \(N\). The value of \(N\) obtained in Johnson’s own experiments turned out to be approximately \(7\%\) lower than the generally accepted one; however, the new measurements by Williams and Thatcher \(^{33}\) gave for \(k\) the value \(1.378 \cdot 10^{-16}\) erg/deg, in very good agreement with the presently accepted value \(1.371 \cdot 10^{-16}\), obtained by other methods.

The Johnson effect is of considerable interest for amplifier technology. The Johnson effect, together with the shot effect, determines the lower limit for the possibility of amplifying small alternating voltages. The shot effect can be greatly weakened by making tubes (perhaps requiring a special design) operate at full space charge. In this case, disturbances associated with thermal motion will play the dominant role. To reduce the latter, several possibilities may be indicated. First, one can make the resistance between the cathode and the grid of the first tube sufficiently small;* this path, however, is limited by a number of basic conditions of amplifier and measuring technique. Another possibility, which, to be sure, can be considered only in special cases, consists in lowering the temperature of the input resistance. And finally, for measuring circuits it is possible, by sharp tuning to a definite frequency, to greatly lower the value of the integral in (2), and thereby also reduce the magnitude \(\overline{E^2}\).

In order to give an idea of the magnitude of the effect, we give, in conclusion, a numerical example. At room temperature, in a frequency interval of 5000 hertz, the power released by the Johnson effect is about \(10^{-16}\) W. Hence we obtain for the magnitude of the voltage across a resistance of \(1\ \mathrm{M}\Omega\)

\[ E = \sqrt{10^{-16} \cdot 10^{6}} = 10^{-5}\ \mathrm{V}. \]

* This circuit element is present in the other stages as well; however, its influence in the first tube is considerably greater than in the subsequent ones.

REFERENCES

  1. E. Riecke, Ann. d. Phys. 66, 453, 545, 1898; P. Drude, ibid. 1, 566, 1900; 3, 370, 1900; 7, 687, 1902; H. A. Lorentz, Proc. Ac. Amst. 7, 438; 585, 684, 1905; The Theory of Electrons; Report Congr. Solvay 1924; N. Bohr, Metallernes Elektrontheorie. Diss. Kopenhagen, 1911.
  2. I. Kikoin and I. Fakidow, Z. Physik 71, 393, 1931.
  3. W. Pauli, Z. Physik 41, 81, 1927.
  4. A. Sommerfeld, Z. Physik 47, 1, 1928.
  5. W. V. Houston, Z. Physik 48, 449, 1928.
  6. F. Bloch, Z. Physik 52, 555, 1928; 59, 208, 1930.
  7. R. Peierls, Ann. d. Phys. 4, 121, 1930; 5, 244, 1930.
  8. L. Nordheim, Ann. d. Phys. 9, 607, 1931.
  9. K. Darrow, Elementare Einführung in die physikalische Statistik, Lpz., 1931.
  10. L. Brillouin, Die Quantenstatistik, Berlin, 1931.
  11. R. H. Fowler, Statistical Mechanics, Cambridge, 1929.
  12. G. E. Uhlenbeck, Over statistische Methoden in der Theorie der Quanta, Diss. Leiden 1927; R. Pierls, Ergeb. d. Exakt. Naturwiss. XI, 264, 1932.
    12a. Pauli, Z. Physik 31, 765, 1925.
  13. W. Heisenberg, Z. Physik 38, 411, 1926.
  14. P. Dirac, Proc. Roy. Soc. A, 112, 661, 1926.
  15. P. A. Dirac, Proc. Cambr. Phil. Soc. 25, 62, 1929.
  16. I. von Neumann, Z. Physik 57, 30, 1929; O. Klein, ibid. 72, 767, 1931.
  17. N. Bose, Z. Physik 26, 178, 1924.
  18. A. Einstein, Berl. Ber. 1924, 261; 1925, 3.
  19. E. Fermi, Z. Physik 36, 902, 1926; Dirac, Proc. Roy. Soc. A 112, 661, 1926.
  20. W. Pauli, Z. Physik 41, 81, 1927; G. E. Uhlenbeck, Diss. Leiden 1927; R. H. Fowler, Statistical Mechanics, Cambridge, 1929.
  21. W. Heisenberg, Z. Physik 41, 239, 1927.
  22. F. Hund, Z. Physik 42, 93, 1927.
  23. F. Rasetti, Proc. Nat. Ac. Sci. 15, 515, 1929.
  24. K. F. Bonhoeffer and P. Harteck, Z. physik. Chem. 4, 113, 1929.
  25. A. Eucken and K. Hiller, Z. physik. Chem. 4, 142, 1929.
  26. E. Wigner, Ung. Akad. Wiss., 1929; P. Ehrenfest and R. Oppenheimer, Phys. Rev. 37, 333, 1931.
  27. W. Heitler and G. Herzberg, Naturwiss. 17, 673, 1929.
  28. G. E. Uhlenbeck and L. Gropper, Phys. Rev. 41, 79, 1932.
  29. A. Sommerfeld, Z. Physik 47, 1, 1928; L. Nordheim, Ann. d. Phys. 9, 607, 1931, §11.
  30. J. B. Johnson, Phys. Rev. 32, 97, 1928.
  31. W. Schottky, Ann. d. Phys. 57, 571, 1918.
  32. H. Nyquist, Phys. Rev. 32, 110, 1928.
  33. Williams and Thatcher, Phys. Rev. (2) 40, 121, 1932.

Submission history

THEORY OF THE METALLIC STATE \*